A speech recognition and classification method and system based on machine learning
Through domain feature extraction, time-frequency analysis and equipment compensation, combined with the adaptive speech classification model, the problem of speech recognition classification accuracy in complex speech signal environments is solved, and accurate speech classification in noisy environments and different speaker characteristics are realized.
Patent Information
- Application Number
- CN202510281474.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing speech recognition classification methods are difficult to extract features comprehensively and accurately in complex and changeable speech signal environments, resulting in a decrease in classification accuracy, especially in noisy environments and changes in different speaker characteristics and speech speeds.
The classification task influence factor is generated through domain feature extraction, time-frequency analysis is performed to generate a standardized speech feature matrix, the target acoustic feature vector is screened based on user classification needs, and the equipment information is analyzed and generated equipment compensation feature vectors are analyzed and feature fusion is used to generate speech classification results with confidence.
It realizes accurate speech classification in complex speech signal environments, improves the accuracy and confidence of speech recognition, adapts to different signal-to-noise ratios, speaker characteristics and speech speed changes, and enhances the robustness of speech classification.
Smart Images

Figure CN119785779B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a speech recognition and classification method and system based on machine learning. Background Art
[0002] The speech signal itself has high complexity. In actual application scenarios, the signal-to-noise ratio of the speech signal fluctuates greatly. For example, in noisy environments such as streets and shopping malls, the speech will be mixed with a large amount of background noise, resulting in a significant reduction in the signal-to-noise ratio, which makes it extremely difficult to accurately extract effective features from the speech. At the same time, the speaker characteristics vary widely. The pronunciation habits, accent characteristics, and unique vocalization methods of different individuals all increase the difficulty of speech feature extraction. In addition, the speed change of the speech rate will also interfere with the accuracy of feature extraction. A fast speech rate may cause some speech information to be lost, while a slow speech rate may reduce the efficiency of feature extraction. Existing speech recognition and classification methods often have difficulty in comprehensively and accurately extracting features when dealing with these complex and variable speech signals, which will seriously affect the subsequent classification accuracy.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of this application is to provide a speech recognition and classification method and system based on machine learning, which can at least overcome the problems existing in the prior art to a certain extent. By extracting domain features from the application scenario information, a classification task influence factor including the weight of the domain-specific vocabulary is generated; by performing time-frequency analysis on the speech stream information, a standardized speech feature matrix is generated; according to the user's classification requirements, matrix features are screened to obtain a target acoustic feature vector including prosody, phoneme probability distribution, and context correlation features; the information of the acquisition device is analyzed to generate a device compensation feature vector including the sampling rate compensation coefficient, etc. The classification task influence factor, the target acoustic feature vector, and the device compensation feature vector are input into an adaptive speech classification model, and the features are fused through an attention mechanism to generate a target feature template and attention weights, and then several speech classification results and their confidence values are obtained, and finally the speech classification result information with confidence is generated to achieve accurate classification of the target speech.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of the present invention.
[0006] According to an aspect of the present application, there is provided a speech recognition and classification method based on machine learning, including: obtaining speech stream information of a target speech in a preset scenario, application scenario information of the target speech, classification requirement information of the target user, acquisition device information of the target speech, a preset speech classification model, and an acoustic training set, wherein the speech stream information of the target speech in the preset scenario includes audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions; extracting domain features from the application scenario information of the target speech to generate a classification task impact factor, wherein the classification task impact factor includes a domain-specific vocabulary weight, a real-time requirement parameter, and a fault tolerance threshold index; performing time-frequency analysis processing on the target speech stream information to generate a standardized speech feature matrix; screening the feature dimensions of the standardized speech feature matrix based on the classification requirement information of the target user to generate a target acoustic feature vector, wherein the target acoustic feature vector includes prosody features, phoneme probability distribution features, and context association features; performing signal characteristic analysis on the acquisition device information of the target speech to generate a device compensation feature vector, wherein the device compensation feature vector includes a sampling rate compensation coefficient, a frequency response curve correction parameter, and an environmental noise baseline value; performing transfer learning on the preset speech classification model based on the acoustic training set to generate an adaptive speech classification model; inputting the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and performing feature fusion through an attention mechanism to generate speech classification result information with confidence.
[0007] Another aspect of the present application is a speech recognition and classification device based on machine learning, which is characterized by including: an acquisition module, configured to acquire speech stream information of a target speech under a preset scenario, application scenario information of the target speech, classification requirement information of a target user, acquisition device information of the target speech, a preset speech classification model, and an acoustic training set, wherein the speech stream information of the target speech under the preset scenario includes audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions; a processing module, configured to extract domain features from the application scenario information of the target speech to generate a classification task influence factor, wherein the classification task influence factor includes a domain-specific vocabulary weight, a real-time requirement parameter, and a fault tolerance threshold index; perform time-frequency analysis processing on the target speech stream information to generate a standardized speech feature matrix; perform feature dimension screening on the standardized speech feature matrix based on the classification requirement information of the target user to generate a target acoustic feature vector, wherein the target acoustic feature vector includes prosody features, phoneme probability distribution features, and context association features; perform signal characteristic analysis on the acquisition device information of the target speech to generate a device compensation feature vector, wherein the device compensation feature vector includes a sampling rate compensation coefficient, a frequency response curve correction parameter, and an environmental noise baseline value; perform transfer learning on the preset speech classification model based on the acoustic training set to generate an adaptive speech classification model; input the classification task influence factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and perform feature fusion through an attention mechanism to generate speech classification result information with confidence.
[0008] According to still another aspect of the present application, an electronic device is provided, which is characterized by including: a first processor; and a memory, configured to store executable instructions of the first processor; wherein the first processor is configured to execute the above-mentioned speech recognition and classification method based on machine learning by executing the executable instructions.
[0009] According to yet another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a second processor, the above-mentioned speech recognition and classification method based on machine learning is implemented.
[0010] For a speech recognition and classification method and system based on machine learning provided by the present application, the server extracts domain features from the application scenario information to generate a classification task influence factor including a domain-specific vocabulary weight, etc.; performs time-frequency analysis processing on the speech stream information to generate a standardized speech feature matrix; screens the matrix features according to the user classification requirements to obtain a target acoustic feature vector including prosody, phoneme probability distribution, and context association features; analyzes the acquisition device information to generate a device compensation feature vector including a sampling rate compensation coefficient, etc.
[0011] Perform transfer learning on a preset speech classification model using an acoustic training set, and construct multiple groups of data sets to train the model by obtaining data features, generating sampling ratios and sampling features. According to the test results of the test sample set, if there are risk factors affecting speech recognition classification, the trained model is determined as an adaptive speech classification model using a multi-task learning framework. Input the classification task impact factor, target acoustic feature vector, and device compensation feature vector into the adaptive speech classification model, fuse the features through the attention mechanism to generate a target feature template and attention weights, and then obtain several speech classification results and their confidence values, and finally generate speech classification result information with confidence to achieve accurate classification of the target speech.
[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 The flowchart showing a method for speech recognition and classification based on machine learning provided by an embodiment of the present application;
[0014] Figure 2 The structural schematic diagram showing a device for speech recognition and classification based on machine learning provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not used to limit the present invention.
[0016] The following combines Figure 1 to describe a method for speech recognition and classification based on machine learning according to an exemplary embodiment of the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application are applicable to any applicable scenario.
[0017] In one embodiment, the present application also proposes a method and system for speech recognition and classification based on machine learning. Figure 1 Schematically shows a flowchart of a method for speech recognition and classification based on machine learning according to an embodiment of the present application. As Figure 1 shown, this method is applied to a server and includes:
[0018] S101, obtain the speech stream information of the target speech in a preset scenario, the application scenario information of the target speech, the classification requirement information of the target user, the acquisition device information of the target speech, a preset speech classification model, and an acoustic training set.
[0019] In one implementation, the preset scenario is that a customer dials a customer service hotline for consultation. The voice stream information of the target voice will include different situations. For example, Customer A dials the customer service phone on a noisy street. At this time, the signal-to-noise ratio of the voice stream information is low, and the background noise may include the sound of vehicle driving, the noise of the crowd, etc.; Customer B speaks with a strong local accent, which reflects the speaker's characteristics; Customer C speaks very fast when consulting questions, while Customer D speaks slowly. The audio data under different signal-to-noise ratios, speaker characteristics, and speaking speed conditions contained in these voice stream information will be acquired by the system. For example, Customer A says: "Hello, your product... (interfered by background noise) always has problems when in use. How to solve it?" Here, there is both unclear voice caused by low signal-to-noise ratio and the unique speaking style of Customer A.
[0020] In this intelligent customer service system, the application scenario information may be that the customer dials the customer service hotline of electronic products. This means that the voice content is likely to revolve around aspects such as the purchase, use, and after-sales of electronic products. For example, what the customer may consult about could be the battery life of a mobile phone, the software installation of a computer, etc. These information help the system initially judge the field and possible categories to which the voice belongs. The system developers set classification requirements and hope to classify customer voices into categories such as product consultation, complaint and suggestion, after-sales repair application, etc. For example, when the system detects keywords such as "the product is not easy to use" and "the experience is very poor" in the customer voice and there is a dissatisfied emotion, it tends to classify it as a complaint and suggestion; if the customer asks about the functions, usage methods, etc. of the product, it is classified as a product consultation. This is determined based on the need of the target users (here referring to the user side of the system, that is, the customer service department) to quickly classify customer voices and improve service efficiency.
[0021] When a customer dials the customer service phone, different devices may be used. Some customers use an Apple mobile phone to make the call, and its microphone has specific audio acquisition characteristics; some use an Android mobile phone, and there are also differences in audio acquisition among different brands and models of Android mobile phones. For example, some mobile phones may have better sound acquisition effect in the low frequency band, while others are more excellent in the high frequency band. There are also customers who may use a landline to make the call, and the audio acquisition frequency, noise reduction ability, etc. of the landline are also different from those of mobile phones. The system acquires this information about the acquisition devices, can understand the characteristics of the source devices of the voice data, and provides a reference for subsequent processing.
[0022] The multi-task learning framework is adopted to synchronously optimize the acoustic model and the language model as the preset speech classification model. This model as a whole belongs to the deep learning model, combining technologies such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and its variants (such as Long Short-Term Memory Network LSTM, Gated Recurrent Unit GRU), and Attention Mechanism, etc., to achieve efficient speech recognition and classification tasks.
[0023] Regarding the acoustic model part, it includes an input layer, a feature extraction layer and a sequence modeling layer, specifically as follows:
[0024] Input layer: Receives the classification task impact factor, the target acoustic feature vector and the device compensation feature vector as inputs. After these vectors are preprocessed, they are input into the model in a specific format, providing the model with information related to the scene of the speech, the acoustic characteristics, and the acquisition device.
[0025] Feature extraction layer: Adopts Convolutional Neural Network (CNN). CNN has the characteristics of local receptive fields and weight sharing, and can effectively extract acoustic features from the input speech feature vectors. For example, through convolutional operations with different-sized convolutional kernels in the time and frequency dimensions, the time-frequency features of the speech are extracted, capturing local patterns in the speech signal, such as the features of phonemes. Assuming the use of 3x3 and 5x5 convolutional kernels, the speech features are convolved multiple times, mapping the original speech feature vectors to different feature spaces to obtain more representative acoustic features.
[0026] Sequence modeling layer: Uses variants of the Recurrent Neural Network (RNN), such as Long Short-Term Memory Network (LSTM) or Gated Recurrent Unit (GRU). This layer can process the temporal information of the speech signal, solve the problem of gradient vanishing or gradient explosion existing in RNN, and better capture the long-distance dependencies in the speech. Taking LSTM as an example, each LSTM unit contains an input gate, a forget gate and an output gate, and controls the inflow and outflow of information through the gating mechanism, so as to remember long-term information. When processing the speech sequence, LSTM can effectively learn information such as the prosody features of the speech and the sequential relationship of phonemes.
[0027] Regarding the language model part, it includes an embedding layer, a recurrent neural network layer, a multi-task learning fusion layer, an attention mechanism layer and an output layer:
[0028] Embedding layer: Converts the words or characters in the speech into low-dimensional vector representations, i.e., word embeddings (WordEmbedding). For example, using pre-trained word vectors (such as Word2Vec or GloVe), each word is mapped to a vector space of a fixed dimension, so that words with similar semantics are closer in the vector space, thus providing an effective semantic representation for the language model.
[0029] Recurrent neural network layer: Also uses RNN or its variants (such as LSTM, GRU) to model the language information. This layer can learn the context relationships between words, predict the probability distribution of the next word, and thus understand the semantics and grammar structure of the speech. Through training on a large amount of text data, the language model can capture the statistical laws of the language and improve the ability to understand the speech content.
[0030] Multi-task learning fusion layer: Fuses the acoustic model and the language model by sharing some network layers and parameters. For example, in the hidden layer, some neurons receive the inputs of both acoustic features and language features simultaneously, and the features are fused through weight adjustment. At the same time, in the way of multi-task learning, different loss functions are set for the acoustic model and the language model respectively. For example, the loss function of the acoustic model can be the cross-entropy loss between the predicted phonemes and the true phonemes, and the loss function of the language model can be the cross-entropy loss between the predicted words and the true words. Through the backpropagation algorithm, the parameters of both the acoustic model and the language model are optimized simultaneously, so that the model can learn both acoustic features and language features while learning, and improve the overall performance of the model.
[0031] Attention mechanism layer: Introduces the attention mechanism in the output part of the model. The attention mechanism can dynamically allocate weights, so that when the model generates the speech classification result, it can pay more attention to the features related to the current task. For example, when fusing the classification task impact factors, the target acoustic feature vector, and the device compensation feature vector, the attention mechanism can automatically adjust the attention degree to each feature vector according to different speech contents and application scenarios, highlight the important features, and suppress the influence of irrelevant features, thus improving the classification accuracy.
[0032] Output layer: Based on the fused features, through the fully connected layer and the Softmax function, outputs the speech classification result and the confidence value of each classification result. The fully connected layer maps the output of the previous layer to the dimension of the number of classification categories, and the Softmax function converts these outputs into a probability distribution, where each probability value represents the likelihood of the speech belonging to the corresponding category.
[0033] In addition, in the CNN layer, parameters such as the size of the convolutional kernel (e.g., 3x3, 5x5, etc.), the stride, and the padding method determine the extraction effect of the convolutional operation on speech features. A smaller convolutional kernel can capture local detailed features, while a larger convolutional kernel can obtain more extensive context information. The stride determines the distance at which the convolutional operation slides on the input feature map, and the padding method affects the size of the feature map after convolution. The weight matrices (such as input weights, forget weights, output weights, etc.) and bias terms in the LSTM or GRU units are important parameters for the model to learn. These parameters determine the processing method and memory ability of the LSTM / GRU units for the input information. For example, the input weights determine the degree of influence of the input information on the current state, and the forget weights control the retention degree of the information at the previous moment. The dimension of the word embedding vector determines the representation ability of the vocabulary in the vector space. Generally speaking, a higher dimension can represent the semantic information of the vocabulary more accurately, but it will also increase the computational complexity. For example, common word embedding dimensions are 100, 200, 300, etc.
[0034] The weight matrix and bias term in the fully connected layer map the fused features to the classification category space. The size of the weight matrix is determined by the output dimension of the previous layer and the number of classification categories, and the bias term is used to adjust the reference value of the output. In the multi-task learning fusion layer, the weight parameters used to balance the losses of the acoustic model and the language model are important hyperparameters. These weight parameters determine the relative importance of the acoustic model and the language model in the overall training process. By adjusting these weights, the performance of the model on different tasks can be optimized.
[0035] The acoustic training set contains speech data of various different scenarios, speakers, and contents and their corresponding classification labels. For example, there are a large number of speech samples from the customer service scenarios of electronic products in the training set, covering various types of speech such as product consultations, complaint suggestions, and after-sales repair applications, and the accurate category information is marked. Through these training data, the preset speech classification model learns the relationship between different speech features and classifications, providing a basis for classifying new target speech in actual applications in the future.
[0036] S102, extract domain features from the application scenario information of the target speech to generate a classification task influence factor.
[0037] In one implementation, domain features are extracted from the application scenario information of the target speech to generate scenario information features described in natural language, where the scenario information features described in natural language are used to represent the scenario adaptation information of the target speech. When a customer dials the customer service hotline of an electronic product, the target speech obtained by the system is like "The tablet I bought has been charging very slowly recently. Is there a problem with the battery?" The system extracts the domain features from the application scenario information of this speech to generate scenario information features described in natural language. For example, key information such as "tablet", "slow charging", and "battery problem" are extracted from the speech, and these information constitute the scenario information features described in natural language, which are used to represent the adaptation information of this target speech to the after-sales consultation scenario of electronic products. It indicates that this speech is likely to be a consultation about after-sales problems (slow charging suspected battery failure) of electronic products (tablets), initially reflecting the association between the speech and a specific scenario.
[0038] Build a domain knowledge ontology library. In this ontology library, it includes various categories of electronic products (such as tablets, mobile phones, computers, etc.), components of each product (such as batteries, screens, motherboards, etc.), common problems (such as charging problems, abnormal screen displays, etc.), and the relationships between them. For example, it is clarified that the relationship between a tablet and a battery is a belonging relationship, and slow charging belongs to the charging problem, etc. This ontology library provides a structured knowledge framework for subsequent processing. Based on the domain knowledge ontology library, mapping processing is performed on the scenario information features described in natural language to generate structured features. "Tablet" is mapped to the tablet category under the category of electronic products in the ontology library, "slow charging" is mapped to the charging problem in the common problems, and "battery problem" is mapped to the battery-related problems of the components of the tablet. Through this mapping, the unstructured information described in natural language is transformed into structured features, which is convenient for subsequent quantification and analysis.
[0039] Process the structured features to generate a domain adaptation index, a real-time requirement parameter, and a fault tolerance threshold index.
[0040] The method includes a calculation formula for obtaining the domain adaptation index. The calculation formula is:
[0041] ;
[0042] ;
[0043] where represents the domain adaptation index, represents the number of occurrences of domain terms, represents the length of the speech segment, represents the context association strength, represents the average time span of the context window, Represents a positive number, used to prevent the denominator of the logarithmic term from being zero, Represents the i-th domain term that appears in the target speech, Represents the context information of the target speech.
[0044] Assume that within a certain time window, the number of occurrences of domain terms (such as "tablet computer", "battery") in this speech is counted is 2 times, and the length of the speech segment is calculated to be 5 seconds (assuming the total speech duration is 5 seconds). By analyzing the context information of the speech and using the formula calculate the context association strength. Here, it is assumed that through the attention mechanism, is 0.8 (the attention mechanism calculates a score based on the degree of association between each domain term and the context). The average time span of the context window is set to 3 seconds. To prevent the denominator of the logarithmic term from being zero, a positive number is taken as 0.01. Then the domain adaptation index . The higher this value, the higher the degree of adaptation of the speech to the electronics field.
[0045] Process the domain adaptation index, real-time requirement parameter, and fault tolerance threshold index to generate the classification task impact factor; since it is a customer service hotline scenario and customers hope that problems can be solved as soon as possible, the real-time requirement is relatively high. Assume that according to experience, the real-time requirement parameter is 0.8 (the value range can be determined according to actual business needs and experience, between 0 and 1, and the larger the value, the higher the real-time requirement). Considering the customer service scenario, occasional recognition errors may not have a serious impact on the overall service, but a too high error rate cannot be tolerated. Assume that the fault tolerance threshold index is 0.3. The method includes a calculation formula for obtaining the classification task impact factor, and the calculation formula is: . Among them, represents the domain adaptation index, represents the domain adaptation weight coefficient, represents the real-time requirement parameter, represents the real-time adjustment parameter, represents the fault tolerance threshold index, represents a positive number.
[0046] Assume that the domain adaptation weight coefficient is 1.2 (set according to the importance of the domain adaptation index), the real-time adjustment parameter t is 2 (set according to the specific quantification of real-time for the business), and the positive number is still 0.01. Assume that the function is the Sigmoid function. First calculate , and then through the calculation of the Sigmoid function, . This classification task impact factor comprehensively considers the domain adaptation degree, real-time requirements, and fault tolerance ability, and is used for the subsequent input adaptive speech classification model, which affects the result of speech classification. It shows that in this customer service scenario, the comprehensive matching degree of this voice with the current business scenario is relatively high, and the model should give higher weight consideration during classification.
[0047] S103. Perform time-frequency analysis processing on the target voice stream information to generate a standardized speech feature matrix.
[0048] In one implementation, perform time-frequency analysis processing on the target voice stream information to generate a number of speech frames. Among them, the speech frames are used to represent the frame information generated by the speech information in the time dimension. When the customer says the speech "The tablet computer I bought has been charging very slowly recently. Is there a problem with the battery?", the system first performs time-frequency analysis processing on the target voice stream information. Assume that the system sets every 10 milliseconds as a frame to divide the speech. This 5-second-long speech will be divided into 500 frames (5 seconds = 5000 milliseconds, 5000÷10 = 500). These speech frames discretize the speech in the time dimension, and each frame carries part of the speech information in that time period. For example, in a certain frame, it may contain the starting part of the pronunciation of the word "tablet", and another frame contains the middle or ending part of its pronunciation. Through these frames, the characteristics of the speech can be analyzed segment by segment.
[0049] Perform fast Fourier transform and conversion to Mel scale processing on a number of speech frames to generate a Mel spectrum feature sequence. Among them, the Mel spectrum feature sequence includes the frequency components and energy distribution information of the speech, and is used to represent the acoustic characteristics of the speech. Perform fast Fourier transform (FFT) on each of the generated 500 speech frames. Taking one frame as an example, FFT converts this frame of speech from the time domain to the frequency domain, and obtains the amplitude and phase information of this frame of speech at different frequencies. For example, in a certain frame, after FFT, it is found that the speech has a high energy concentration in the 200 - 500 Hz frequency band, which may be related to the pronunciation of the final sound of the character "ping" in "tablet", because different speech phonemes have different energy distribution characteristics in the frequency domain. Convert the frequency domain information after FFT processing to the Mel scale. The human ear's perception of sounds at different frequencies is not linear, and the Mel scale is more in line with the human ear's perception characteristics. Under the Mel scale, remap the speech frequencies. For example, the frequency originally at 1000 Hz may correspond to a specific Mel frequency value under the Mel scale. After performing such a conversion on each frame of speech, a Mel spectrum feature sequence is generated. This sequence contains the frequency components and energy distribution information of the speech, and can better represent the acoustic characteristics of the speech. For example, the character "chong" in "charging" will show a unique frequency component and energy distribution pattern in the Mel spectrum feature sequence, distinguishing it from other words.
[0050] Perform dynamic time warping processing on the mel-spectrum feature sequence to generate a standardized speech feature matrix. Each row of the standardized speech feature matrix represents the feature vector of one frame of speech, and each column represents the feature dimension, so that the speech features have a unified format in terms of time and feature dimensions. Since different customers speak at different speeds, even for the same content of speech, the length of their mel-spectrum feature sequences on the time axis will vary. For example, customer A speaks at a faster speed, while another customer B speaks at a slower speed. When they say the same "The tablet computer charges slowly", the number of frames of their mel-spectrum feature sequences is different. At this time, it is necessary to perform dynamic time warping (DTW) processing on the mel-spectrum feature sequence. The DTW algorithm will find the optimal alignment path between mel-spectrum feature sequences of different lengths. Suppose the mel-spectrum feature sequence of customer A's speech after processing has 400 frames, and that of customer B has 600 frames. DTW will warp them to the same length, such as 500 frames. After DTW processing, a standardized speech feature matrix is generated. Each row of this matrix represents the feature vector of one frame of speech, and each column represents the feature dimension. For example, the value in the first row and the first column may represent the energy value of the first frame of speech at a specific mel-frequency, and the value in the first row and the second column represents the energy value of the first frame of speech at another mel-frequency, and so on. In this way, regardless of whether the speech comes from a customer with a fast or slow speaking speed, it can be presented in a unified format, facilitating subsequent feature dimension screening based on the classification requirement information of the target user, and inputting it into an adaptive speech classification model for processing.
[0051] S104, perform feature dimension screening on the standardized speech feature matrix based on the classification requirement information of the target user to generate a target acoustic feature vector.
[0052] In one implementation, perform feature dimension screening on the standardized speech feature matrix based on the classification requirement information of the target user to generate prosodic features, phoneme probability distribution features, and context correlation features. The classification requirement of the target user is to classify customer speech into categories such as product consultation, complaint and suggestion, after-sales repair application, etc. The system screens the standardized speech feature matrix according to this requirement. From what the customer said, the system identifies prosodic features, phoneme probability distribution features, and context correlation features related to the classification. For example, when the customer speaks with an increased tone and a faster speed, these prosodic manifestations may imply dissatisfaction and are related to the complaint and suggestion category; while the specific phonemes contained in the speech, such as "mobile phone", "charging", "slow", etc., constitute the phoneme probability distribution features; at the same time, the context of the whole sentence revolves around the problem of slow mobile phone charging, reflecting the context correlation features.
[0053] Process the prosodic features to generate the mean fundamental frequency and the variance of the fundamental frequency. The mean fundamental frequency and the variance of the fundamental frequency are used to characterize the pitch and the variation range of the intonation, as well as the speech rate value. The system processes the recognized prosodic features. By analyzing the change of the frequency of the speech signal over time, the mean fundamental frequency and the variance of the fundamental frequency are calculated. Suppose that after calculation, the mean fundamental frequency of this segment of speech is relatively high, which means that the overall intonation of the customer's speech is relatively high, reflecting that the customer may be more excited; the larger variance of the fundamental frequency indicates a large variation range of the intonation. Considering the content, it is very likely that the customer is expressing dissatisfaction, further verifying the possibility of a complaint or suggestion. In addition, by analyzing the speech duration and the number of syllables, the speech rate value can also be calculated, and it is found that the customer's speech rate is slightly faster than the normal speech rate, which also reflects from the side that the customer's mood is relatively urgent.
[0054] Process the phoneme probability distribution features to generate the probability values of each phoneme appearing in the speech. In this sentence, phonemes such as "shou", "ji", "chong", "dian", "man" appear with a relatively high frequency. The system counts the number of times each phoneme appears in the speech and calculates the occurrence probability of each phoneme in combination with the total number of phonemes in the speech. For example, the word "shouji" appears frequently, and the phoneme probabilities of "shou" and "ji" are relatively high, indicating that the speech content is related to mobile phones. These phoneme probability distribution values help the system further determine the field and category to which the speech belongs.
[0055] Process the context association features to generate a context association score. The context association score is used to characterize the semantic association degree between the current speech segment and the context before and after. If in the customer service system, the customer also mentioned "this situation has never occurred before" before, then the current speech segment "charging is extremely slow" has a strong semantic association with the previous text, indicating that the problem has occurred recently. The system evaluates this association degree through natural language processing technology and semantic analysis algorithms and gives a context association score. Suppose this score is relatively high, indicating that the current speech segment is closely related to the context before and after, which can better assist the system in understanding the customer's intention. For example, it can more clearly determine that this is a feedback about the mobile phone charging problem rather than other irrelevant topics.
[0056] Process the mean fundamental frequency, the variance of the fundamental frequency, the probability values of each phoneme appearing in the speech, and the context association score to generate a target acoustic feature vector. Combine these values into a vector according to a certain order and weight. Suppose the vector form is [mean fundamental frequency, variance of the fundamental frequency, probability of phoneme 1, probability of phoneme 2,..., context association score]. This target acoustic feature vector contains key information in multiple aspects such as the prosody, phonemes, and context of the speech, and serves as important data for subsequent input into the adaptive speech classification model, helping the model more accurately classify the speech. In this example, this segment of speech is classified into the complaint or suggestion category.
[0057] S105, perform signal characteristic analysis on the acquisition device information of the target voice to generate a device compensation feature vector.
[0058] In one implementation, perform sampling rate analysis processing on the acquisition device information of the target voice to generate a sampling rate compensation coefficient. The system obtains the sampling rate information of the mobile phone used by customer A. Assume that the sampling rate of this mobile phone is 44.1 kHz, while the default standard sampling rate of the system is 16 kHz. To enable subsequent voice processing to be based on a unified standard, the system performs sampling rate analysis processing on the acquisition device information of the target voice. By calculating the proportional relationship between the actual sampling rate and the standard sampling rate, a sampling rate compensation coefficient is generated. In this example, the sampling rate compensation coefficient = 16 kHz / 44.1 kHz ≈ 0.363. This coefficient is used to adjust the sampling rate of the voice signal subsequently to ensure the consistency and comparability of voice features.
[0059] Obtain the frequency response curve information of the acquisition device and the noise signal of the acquisition device when there is no voice input. The system obtains the frequency response curve information of this Android mobile phone, which is a curve describing the response ability of the mobile phone microphone to sound signals at different frequencies. At the same time, the system records the noise signal of the mobile phone when there is no voice input. For example, in a quiet environment, before customer A dials the customer service phone, the signal of a short silent period collected by the mobile phone microphone is the noise signal when there is no voice input. Process the frequency response curve information to generate a frequency response curve correction parameter, where the frequency response curve correction parameter is used to characterize the gain deviation of the acquisition device in different frequency bands. The system processes the obtained frequency response curve information. Assume that through analysis, it is found that the signal gain of this mobile phone in the low frequency band (20 - 200 Hz) is 5 dB higher than that of the standard device, and the signal gain in the high frequency band (5000 - 8000 Hz) is 3 dB lower than that of the standard device. These gain deviations will affect the frequency components of the voice signal, resulting in changes in voice features. To correct these deviations, the system generates a frequency response curve correction parameter. For example, a -5 dB correction parameter is generated for the low frequency band, and a +3 dB correction parameter is generated for the high frequency band. These correction parameters are used to adjust the frequency response of the voice signal subsequently to make the frequency characteristics of the voice closer to the standard state and improve the accuracy of voice recognition.
[0060] Process the noise signal of the acquisition device when there is no voice input to generate an environmental noise baseline value. The environmental noise baseline value includes the average power and root mean square value of the noise. Determine the environmental noise baseline value by calculating the average power and root mean square value of the noise signal. Suppose that after calculation, the average power of the noise signal is 0.001W and the root mean square value is 0.032V. These values represent the noise intensity level collected by the mobile phone microphone in the current acquisition environment. The environmental noise baseline value is used as a reference for subsequent noise reduction processing of the voice signal to help the system distinguish the voice signal from the noise and improve the quality of the voice signal.
[0061] Process the sampling rate compensation coefficient, frequency response curve correction parameter, and environmental noise baseline value to generate a device compensation feature vector. Suppose these parameters are combined into a vector form: [sampling rate compensation coefficient, low-frequency band frequency response curve correction parameter, high-frequency band frequency response curve correction parameter, noise average power, noise root mean square value], that is, [0.363, -5dB, +3dB, 0.001W, 0.032V]. This device compensation feature vector contains key information about the acquisition device in terms of sampling rate, frequency response, and noise characteristics, and is input into the subsequent adaptive voice classification model to compensate for the impact of acquisition device differences on the voice signal and improve the accuracy of voice recognition classification. For example, when classifying the voice of customer A, the model will adjust the voice signal according to this device compensation feature vector to avoid classification errors caused by device factors and more accurately determine the category of the customer's voice, such as determining whether the customer is consulting product problems or making complaint suggestions.
[0062] S106, perform transfer learning on the preset voice classification model based on the acoustic training set to generate an adaptive voice classification model.
[0063] In one implementation, obtain any number of data features in the acoustic training set. Suppose the acoustic training set contains a large amount of voice data from different customers calling the electronics customer service hotline, as well as corresponding classification labels (such as product consultation, complaint suggestion, after-sales repair application, etc.). Randomly select a part of the data features from this training set, such as selecting 1000 voice data and their corresponding classification labels. These voice data cover different problem consultations, complaints, and after-sales related content of various electronic products (mobile phones, computers, tablets, etc.).
[0064] Generate a sampling ratio based on the quantity of each data feature in the acoustic training set, and process the acoustic training set based on the sampling ratio to generate a preset quantity of sampled features. Among the 1000 selected data, count the quantity of each type of data feature. Suppose there are 500 voice data of product consultation type, 300 of complaint and suggestion type, and 200 of after-sales repair application type. To ensure an appropriate proportion of each type of data in subsequent training, generate a sampling ratio according to the data quantity. For example, set the sampling ratio of product consultation type as 500 / 1000 = 0.5, complaint and suggestion type as 300 / 1000 = 0.3, and after-sales repair application type as 200 / 1000 = 0.2. These sampling ratios are used to control the proportion of each type of data extracted from the training set, avoid a certain type of data dominating in the training process, and ensure that the model can learn the features of each type of data. Suppose it is preset to generate 500 sampled features. Extract data from the training set according to the above sampling ratios. For the product consultation type, extract 500×0.5 = 250 voice data; for the complaint and suggestion type, extract 500×0.3 = 150; for the after-sales repair application type, extract 500×0.2 = 100. The extracted data constitutes the preset quantity of sampled features, which not only retains the proportional relationship of each type of data in the original training set but also reduces the data volume, facilitating more efficient model training in the follow-up.
[0065] Process any data feature with each sampled feature to generate multiple groups of data sets, where each group of data sets contains a target quantity of data samples, and at least one data sample includes identification information. Select any data feature from the originally selected 1000 data, such as a voice data of product consultation about the mobile phone battery life problem. Combine and process this data with the 500 sampled features generated previously. Take a certain quantity as a group (suppose each group contains 10 data samples) to generate multiple groups of data sets. In each group of data sets, at least one data sample has identification information (such as a classification label). For example, the first group of data sets contains 9 different voice data selected from the sampled features and 1 voice data about the mobile phone battery life problem with the classification label of "product consultation"; the second group of data sets also contains 9 different sampled feature voice data and another voice data with a classification label (which may be of complaint and suggestion type or after-sales repair application type), and so on, to generate multiple such groups of data sets.
[0066] Train a preset speech classification model based on data samples in multiple groups of data sets to generate a trained speech classification model. Use the generated multiple groups of data sets as training data and input them into the preset speech classification model. The model adopts a multi-task learning framework to synchronously optimize the acoustic model and the language model. During the training process, the acoustic model learns the acoustic features of speech (such as the frequency components, prosody, etc. of speech), and the language model learns the semantics and grammatical structures of speech (such as the relationships between words, the meanings of sentences, etc.). By continuously adjusting the parameters of the model, the model can accurately predict its classification label according to the input speech data. After multiple rounds of training, a trained speech classification model is obtained. For example, the model gradually learns to judge that the speech belongs to the battery life problem consultation in product consultation based on keywords such as "battery" and "short battery life" in the speech, as well as the speaker's tone, intonation and other features.
[0067] Process the trained speech classification model based on a test sample set to generate test results. If the data samples containing identification information in the test results are risk factors that affect speech recognition classification, then regard the trained speech classification model as an adaptive speech classification model. Among them, the adaptive speech classification model adopts a multi-task learning framework to synchronously optimize the acoustic model and the language model. Prepare a test sample set, which contains some new customer service speech data and their classification labels that have not participated in the training. Input these test data into the trained speech classification model, and the model classifies and predicts each speech data to generate test results. For example, there are 100 speech data in the test sample set, and the model predicts the category (product consultation, complaint and suggestion, after-sales repair application, etc.) to which each speech belongs and outputs the prediction results.
[0068] Check the data samples containing identification information (classification labels) in the test results. If there are many misclassified data samples in the test results, and the speech features represented by these samples are considered risk factors that affect speech recognition classification (for example, speech recognition errors caused by a specific accent, inaccurate recognition of some low-frequency words, etc.), then regard the trained speech classification model as an adaptive speech classification model. This means that the model has a certain adaptive ability when facing complex situations that may occur in actual applications, and can better handle these risk factors in subsequent speech classification tasks, improving the overall classification accuracy. For example, if it is found in the test results that many speeches with a southern accent are misclassified, and the southern accent is a common situation in the actual customer service scenario, then this trained model can be used as an adaptive speech classification model to further optimize the classification ability for such speeches with special accents in actual applications.
[0069] S107. Input the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and perform feature fusion through the attention mechanism to generate speech classification result information with confidence.
[0070] In one implementation, the adaptive speech classification model processes the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector to generate a target feature template, the attention weight for each dimension, and the attention weight related to the target frequency band frequency response curve correction parameter. In the previously set electronic product customer service hotline scenario, assume that a customer calls the customer service and says, "Your mobile phone's battery life is too poor. It runs out of power soon after being fully charged. I'm very dissatisfied!" The system inputs the previously obtained classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model. Assume that the classification task impact factor indicates that the speech is highly relevant to the electronic product complaint scenario (such as a relatively high domain adaptation index calculated as mentioned above), the target acoustic feature vector includes information such as the customer's excited tone (high average fundamental frequency, large fundamental frequency variance), the probability distribution of specific phonemes ("mobile phone", "battery", "poor battery life", etc.), and the context association information with the previous text, and the device compensation feature vector records the relevant parameters of the customer's device (such as the sampling rate compensation coefficient, the frequency response curve correction parameter, etc.).
[0071] Based on these inputs, the adaptive speech classification model first generates a target feature template. This template is an abstract representation of the current speech features. For example, it may integrate the feature combinations related to electronic product-related vocabulary in the speech, as well as the key information reflecting the customer's emotion and context. At the same time, the model calculates the attention weight for each dimension. For the acoustic feature dimensions related to "battery life", such as the energy distribution dimensions related to the pronunciation of "battery" and "life" in specific frequency bands, a higher attention weight is given because these dimensions are crucial for judging the speech category; for the parameters in the device compensation feature vector, if the response deviation of the device in certain frequency bands may affect the recognition of keywords in the speech, then the device compensation dimensions related to these frequency bands will also obtain corresponding attention weights. In addition, the attention weight related to the target frequency band frequency response curve correction parameter is also generated. For example, if the key information in the speech is concentrated in several specific frequency bands, and the device has a gain deviation in these frequency bands, then the attention weight of the frequency response curve correction parameter corresponding to these frequency bands will be higher to highlight the attention and correction of these frequency bands.
[0072] Process the target feature template, the attention weights for each dimension, and the attention weights related to the target frequency band frequency response curve correction parameters to generate several speech classification results and the confidence values attached to each classification result. Based on the generated target feature template and attention weights, the model further processes to generate several speech classification results and the confidence values attached to each classification result. The model will match and judge the speech with pre-set categories (such as product consultation, complaint and suggestion, after-sales repair application, etc.). Since the customer's speech expresses dissatisfaction with the mobile phone battery life, through the analysis of the input features, the model believes that this speech is more likely to belong to the "complaint and suggestion" category, while other categories will also be considered as possible results, but with relatively lower probabilities.
[0073] For the classification result of "complaint and suggestion", the model gives a confidence value based on the comprehensive consideration of the feature matching degree and attention weights. Suppose after complex calculations and analyses, the model believes that the confidence of this speech belonging to the "complaint and suggestion" category is 0.9. This means that the model has a high confidence in this classification result because various features in the speech, whether it is vocabulary, tone, or device-related compensation information, strongly indicate that this is a complaint and suggestion. For other possible categories, such as "product consultation", the confidence given by the model may only be 0.1 because from the overall features, the speech does not conform to the typical features of product consultation.
[0074] Process several speech classification results and the confidence values attached to each classification result to generate speech classification result information with confidence. The system will organize the classification result information, such as "Speech classification result: Complaint and suggestion, Confidence: 0.9". This information will be passed to the subsequent customer service business process, and the customer service staff can decide the processing method for this speech according to the confidence. If the confidence is high, such as 0.9 in this example, the customer service staff can directly process it according to the complaint and suggestion process to give priority to solving the customer's problem; if the confidence is low, for example, the confidence of a certain speech classification result is about 0.5, the customer service staff may further verify or adopt the method of manual assisted judgment to ensure the accuracy of classification and avoid the decline in service quality caused by misjudgment.
[0075] In this application, the server obtains the speech stream information, application scenario information, user classification requirement information, acquisition device information, preset speech classification model, and acoustic training set of the target speech. Among them, the speech stream information covers audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions. Domain feature extraction is performed on the application scenario information to generate classification task influencing factors including the weight of the domain-specific vocabulary; the speech stream information is processed through time-frequency analysis to generate a standardized speech feature matrix; matrix features are screened according to the user classification requirements to obtain a target acoustic feature vector including prosody, phoneme probability distribution, and context correlation features; the acquisition device information is analyzed to generate a device compensation feature vector including a sampling rate compensation coefficient, etc.
[0076] Transfer learning is performed on the preset speech classification model using the acoustic training set. By obtaining data features, generating sampling ratios and sampling features, multiple groups of data sets are constructed to train the model. According to the test results of the test sample set, if there are risk factors affecting speech recognition classification, the trained model is determined as an adaptive speech classification model using a multi-task learning framework. The classification task influencing factors, the target acoustic feature vector, and the device compensation feature vector are input into the adaptive speech classification model. Through the attention mechanism to fuse features, a target feature template and attention weights are generated, and then several speech classification results and their confidence values are obtained. Finally, the speech classification result information with confidence is generated to achieve accurate classification of the target speech.
[0077] In one implementation, as Figure 2 shown, this application also provides a speech recognition and classification device based on machine learning, including:
[0078] An acquisition module 201, configured to obtain the speech stream information of the target speech in a preset scenario, the application scenario information of the target speech, the classification requirement information of the target user, the acquisition device information of the target speech, the preset speech classification model, and the acoustic training set, where the speech stream information of the target speech in the preset scenario includes audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions;
[0079] The processing module 202 is configured to extract domain features from the application scenario information of the target voice to generate a classification task impact factor, where the classification task impact factor includes a domain-specific vocabulary weight, a real-time requirement parameter, and a fault tolerance threshold index; perform time-frequency analysis processing on the target voice stream information to generate a standardized voice feature matrix; perform feature dimension screening on the standardized voice feature matrix based on the classification requirement information of the target user to generate a target acoustic feature vector, where the target acoustic feature vector includes a prosody feature, a phoneme probability distribution feature, and a context association feature; perform signal characteristic analysis on the acquisition device information of the target voice to generate a device compensation feature vector, where the device compensation feature vector includes a sampling rate compensation coefficient, a frequency response curve correction parameter, and an environmental noise baseline value; perform transfer learning on the preset voice classification model based on the acoustic training set to generate an adaptive voice classification model; input the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive voice classification model, and perform feature fusion through an attention mechanism to generate voice classification result information with confidence.
[0080] The computer-readable storage medium provided in the above embodiments of the present application and the speech recognition and classification method based on machine learning provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0081] Each embodiment in the present application is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the method for evaluating speech recognition and classification based on machine learning, the electronic device, the electronic equipment, and the readable storage medium, since they are basically similar to the embodiments of the above speech recognition and classification method based on machine learning, the description is relatively simple, and the relevant parts can be referred to the partial description of the embodiments of the above speech recognition and classification method based on machine learning.
Claims
1. A speech recognition and classification method based on machine learning, characterized in that, Including: Obtain the speech stream information of the target speech in a preset scenario, the application scenario information of the target speech, the classification requirement information of the target user, the acquisition device information of the target speech, a preset speech classification model, and an acoustic training set. Among them, the speech stream information of the target speech in the preset scenario includes audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions; Extract domain features from the application scenario information of the target speech to generate a classification task impact factor, where the classification task impact factor is calculated based on a domain adaptation index, a real-time requirement parameter, and a fault tolerance threshold index; Perform time-frequency analysis processing on the target speech stream information to generate a standardized speech feature matrix; Based on the classification requirement information of the target user, perform feature dimension screening on the standardized speech feature matrix to generate a target acoustic feature vector, where the target acoustic feature vector includes prosodic features, phoneme probability distribution features, and context correlation features; Perform signal characteristic analysis on the acquisition device information of the target speech to generate a device compensation feature vector, where the device compensation feature vector includes a sampling rate compensation coefficient, a frequency response curve correction parameter, and an environmental noise baseline value; Perform transfer learning on the preset speech classification model based on the acoustic training set to generate an adaptive speech classification model; Input the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and perform feature fusion through an attention mechanism to generate speech classification result information with confidence.
2. The method according to claim 1, characterized in that Extract domain features from the application scenario information of the target speech to generate a classification task impact factor, including: Extract domain features from the application scenario information of the target speech to generate scene information features described in natural language, where the scene information features described in natural language are used to characterize the scene adaptation information of the target speech; Construct a domain knowledge ontology library; Perform mapping processing on the scene information features described in natural language based on the domain knowledge ontology library to generate structured features; Process the structured features to generate a domain adaptation index, a real-time requirement parameter, and a fault tolerance threshold index; Process the domain adaptation index, the real-time requirement parameter, and the fault tolerance threshold index to generate a classification task impact factor; The method includes a calculation formula for obtaining the domain adaptation index, and the calculation formula is: ; ; Among them, represents the domain adaptation index, represents the number of occurrences of domain terms, represents the length of the speech segment, represents the context association strength, represents the average time span of the context window, represents a positive number used to prevent the denominator of the logarithmic term from being zero, represents the i-th domain term that appears in the target speech, represents the context information of the target speech; The method includes a calculation formula for obtaining the classification task impact factor, and the calculation formula is: ; Among them, represents the domain adaptation index, represents the domain adaptation weight coefficient, represents the real-time requirement parameter, represents the real-time adjustment parameter, represents the fault tolerance threshold index, represents a positive number.
3. The method according to claim 1, characterized in that, Performing time-frequency analysis processing on the target speech stream information to generate a standardized speech feature matrix includes: Perform time-frequency analysis processing on the target speech stream information to generate a number of speech frames, where the speech frames are used to represent the frame information generated by the speech information in the time dimension; Perform fast Fourier transform and conversion to Mel scale processing on a number of speech frames to generate a Mel spectrum feature sequence, where the Mel spectrum feature sequence includes the frequency components and energy distribution information of the speech and is used to characterize the acoustic characteristics of the speech; Perform dynamic time warping processing on the Mel spectrum feature sequence to generate a standardized speech feature matrix, where each row of the standardized speech feature matrix represents the feature vector of one frame of speech, and each column represents the feature dimension, so that the speech features have a unified format in terms of time and feature dimension.
4. The method according to claim 3, characterized in that, Based on the classification requirement information of the target user, perform feature dimension screening on the standardized speech feature matrix to generate a target acoustic feature vector, including: Based on the classification requirement information of the target user, perform feature dimension screening on the standardized speech feature matrix to generate prosodic features, phoneme probability distribution features, and context correlation features; Process the prosodic features to generate a fundamental frequency mean and a fundamental frequency variance, where the fundamental frequency mean and the fundamental frequency variance are used to characterize the pitch and variation range of the intonation, as well as the speech rate value; Process the phoneme probability distribution features to generate the probability value of each phoneme appearing in the speech; Process the context correlation features to generate a context correlation score, where the context correlation score is used to characterize the semantic correlation degree of the current speech segment with the context; Process the fundamental frequency mean, the fundamental frequency variance, the probability value of each phoneme appearing in the speech, and the context correlation score to generate a target acoustic feature vector.
5. The method according to claim 1, wherein Perform signal characteristic analysis on the acquisition device information of the target speech to generate a device compensation feature vector, including: Perform sampling rate analysis processing on the acquisition device information of the target speech to generate a sampling rate compensation coefficient; Obtain the frequency response curve information of the acquisition device and the noise signal of the acquisition device when there is no speech input; Process the frequency response curve information to generate a frequency response curve correction parameter, where the frequency response curve correction parameter is used to characterize the gain deviation of the acquisition device in different frequency bands; Process the noise signal of the acquisition device when there is no speech input to generate an environmental noise baseline value, where the environmental noise baseline value includes the average power and root mean square value of the noise; Process the sampling rate compensation coefficient, the frequency response curve correction parameter, and the environmental noise baseline value to generate a device compensation feature vector.
6. The method according to claim 1, characterized in that, Perform transfer learning on the preset speech classification model based on the acoustic training set to generate an adaptive speech classification model, including: Obtain any quantity of data features in the acoustic training set; Generate a sampling ratio based on the quantity of each data feature in the acoustic training set; Process the acoustic training set based on the sampling ratio to generate a preset quantity of sampling features; Process any data feature and each sampling feature to generate multiple groups of data sets, where each group of data sets contains a target quantity of data samples, and at least one data sample includes identification information; Train the preset speech classification model based on the data samples in multiple groups of data sets to generate a trained speech classification model; Process the trained speech classification model based on the test sample set to generate a test result; If the data sample containing identification information in the test results represents a risk factor affecting speech recognition classification, then the trained speech classification model is used as an adaptive speech classification model, where the adaptive speech classification model synchronously optimizes the acoustic model and the language model using a multi-task learning framework.
7. The method according to claim 1, characterized in that Input the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and perform feature fusion through an attention mechanism to generate speech classification result information with confidence, including: Based on the adaptive speech classification model, process the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector to generate a target feature template, the attention weight of each dimension, and the attention weight related to the target frequency band frequency response curve correction parameter; Process the target feature template, the attention weight of each dimension, and the attention weight related to the target frequency band frequency response curve correction parameter to generate several speech classification results and the confidence value attached to each classification result; Process the several speech classification results and the confidence value attached to each classification result to generate speech classification result information with confidence.
8. A speech recognition and classification device based on machine learning, characterized in that, The device includes: An acquisition module, configured to acquire the speech stream information of the target speech in a preset scenario, the application scenario information of the target speech, the classification requirement information of the target user, the acquisition device information of the target speech, a preset speech classification model, and an acoustic training set, where the speech stream information of the target speech in the preset scenario includes audio data under different signal-to-noise ratios, speaker characteristics, and speech rate conditions; A processing module, configured to extract domain features from the application scenario information of the target speech to generate a classification task impact factor, where the classification task impact factor is calculated based on a domain adaptation index, a real-time requirement parameter, and a fault tolerance threshold index; perform time-frequency analysis processing on the target speech stream information to generate a standardized speech feature matrix; perform feature dimension screening on the standardized speech feature matrix based on the classification requirement information of the target user to generate a target acoustic feature vector, where the target acoustic feature vector includes prosodic features, phoneme probability distribution features, and context association features; perform signal characteristic analysis on the acquisition device information of the target speech to generate a device compensation feature vector, where the device compensation feature vector includes a sampling rate compensation coefficient, a frequency response curve correction parameter, and an environmental noise baseline value; perform transfer learning on the preset speech classification model based on the acoustic training set to generate an adaptive speech classification model; input the classification task impact factor, the target acoustic feature vector, and the device compensation feature vector into the adaptive speech classification model, and perform feature fusion through an attention mechanism to generate speech classification result information with confidence.
9. An electronic device, characterized in that, Including: A first processor; And a memory for storing the executable instructions of the first processor; Wherein, the first processor is configured to execute the machine learning-based speech recognition classification method according to any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the machine learning-based speech recognition and classification method according to any one of claims 1 to 7.