Speech recognition method, speech recognition device, electronic device and storage medium
By using spectral analysis and residual and pooling networks to process speech data, the problem of low speech recognition accuracy in Gaussian modeling is solved, and more efficient speaker identification is achieved.
Patent Information
- Application Number
- CN202210600295.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-05-30
AI Technical Summary
In existing technologies, identifying the speaker's identity based on Gaussian modeling carries a significant risk of error, affecting the accuracy of speech recognition.
The target Mel-frequency cepstral coefficients are obtained by spectral analysis. Feature extraction and processing are performed through residual networks and pooling networks to obtain the target speech vector. The vector is then decoded to identify the speaker.
It improves the accuracy of speech recognition, identifies the speaker's identity through spectral features, simplifies the model training process, and saves training time.
Smart Images

Figure CN114974219B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, a speech recognition device, an electronic device, and a storage medium. Background Art
[0002] Currently, Gaussian modeling is often used to identify the speaker identity corresponding to speech data. This method often requires relatively complex Gaussian calculations and has a high risk of error, which affects the accuracy of speech recognition. Therefore, how to improve the accuracy of speech recognition has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to propose a speech recognition method, a speech recognition device, an electronic device and a storage medium, aiming to improve the accuracy of speech recognition.
[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a speech recognition method, the method comprising:
[0005] Obtain target speech data of a target speaker;
[0006] Performing spectral analysis on the target speech data to obtain target Mel-frequency cepstral coefficients;
[0007] Performing dimension-changing processing on the target Mel-frequency cepstral coefficient to obtain a first latent state feature;
[0008] Extracting the first latent state features through a preset residual network to obtain target speech features;
[0009] Performing pooling processing on the target speech features through a preset pooling network to obtain a target speech vector;
[0010] The target speech vector is decoded to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker.
[0011] In some embodiments, the step of extracting the first latent state features through a preset residual network to obtain target speech features includes:
[0012] Performing convolution processing on the first hidden state feature through the residual network to obtain a first convolution feature vector;
[0013] Activate the first hidden state feature by using the activation function of the residual network to obtain a target activation feature vector;
[0014] The first convolution feature and the target activation feature vector are summed according to a preset weight parameter to obtain the target speech feature.
[0015] In some embodiments, the activation function includes a first function and a second function, and the step of activating the first hidden state feature using the activation function of the residual network to obtain a target activation feature vector includes:
[0016] Activate the first latent state feature using the first activation function to obtain a first activated feature vector;
[0017] Activate the first latent state feature using the second activation function to obtain a second activated feature vector;
[0018] Performing a dot product process on the first activation feature vector and the second activation feature vector to obtain the target activation feature vector.
[0019] In some embodiments, the pooling network includes a first convolutional layer, a pooling layer, and a second convolutional layer. The step of performing pooling processing on the target speech feature through the preset pooling network to obtain the target speech vector includes:
[0020] Performing dimension conversion processing on the target speech feature through the first convolutional layer to obtain a second latent state feature;
[0021] Performing pooling processing on the second latent state feature through the pooling layer to obtain an intermediate vector;
[0022] The intermediate vector is subjected to dimension change processing by the second convolutional layer to obtain the target speech vector.
[0023] In some embodiments, the step of performing pooling processing on the second latent state feature through the pooling layer to obtain an intermediate vector includes:
[0024] Performing mean calculation on the second hidden state feature through the pooling layer and the preset time series dimension to obtain a hidden state mean;
[0025] Calculating the standard deviation of the second hidden state feature through the pooling layer and the time series dimension to obtain a hidden state standard deviation;
[0026] Pooling is performed on the second latent state feature according to the latent state mean and the latent state standard deviation to obtain the intermediate vector.
[0027] In some embodiments, the step of decoding the target speech vector to obtain a target person label corresponding to the target speech data includes:
[0028] Performing label probability calculation on the target speech vector using a preset function and a preset character label to obtain a label probability vector corresponding to each preset character label;
[0029] The preset person label corresponding to the label probability vector with the largest value is selected as the target person label.
[0030] In some embodiments, the step of performing spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients includes:
[0031] Performing spectrum calculation on the target speech data to obtain a target spectrogram;
[0032] The target spectrogram is filtered to obtain target Mel-frequency cepstral coefficients.
[0033] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a speech recognition device, the device comprising:
[0034] A data acquisition module, used to acquire target speech data of a target speaker;
[0035] An analysis module is used to perform spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients;
[0036] A dimension conversion module, configured to perform dimension conversion processing on the target Mel-frequency cepstral coefficient to obtain a first latent state feature;
[0037] A feature extraction module, configured to extract the first latent state features through a preset residual network to obtain target speech features;
[0038] A pooling module, configured to perform pooling processing on the target speech features through a preset pooling network to obtain a target speech vector;
[0039] A decoding module is used to decode the target speech vector to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker.
[0040] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored on the memory and runnable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the method described in the first aspect above is implemented.
[0041] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in the first aspect above.
[0042] The speech recognition method, speech recognition device, electronic device, and storage medium proposed in this application obtain target speech data of a target speaker; perform spectral analysis on the target speech data to obtain target Mel-frequency cepstral coefficients, thereby conveniently converting the speech data into spectral features, and analyzing and processing the spectral features to obtain target Mel-frequency cepstral coefficients, so that the speaker's identity can be identified through the spectral features, thereby improving recognition accuracy. Furthermore, the target Mel-frequency cepstral coefficients are subjected to dimensionality conversion processing to obtain first latent state features; the first latent state features are subjected to feature extraction using a preset residual network to obtain target speech features; the target speech features are subjected to pooling processing using a preset pooling network to obtain a target speech vector. The latent state features corresponding to the target Mel-frequency cepstral coefficients are mapped to specific vector dimensions through dimensionality conversion processing, and the temporal state of the latent state features is changed through pooling processing to obtain a target speech vector that meets the requirements. The target speech vector is used as a feature vector for identifying the identity of the target speaker. Finally, the target speech vector is decoded to obtain a target person label corresponding to the target speech data, thereby conveniently determining the identity of the target speaker based on the corresponding target person label, thereby improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flow chart of the speech recognition method provided in an embodiment of the present application;
[0044] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.
[0045] Figure 3 yes Figure 1 Flowchart of step S104 in FIG.
[0046] Figure 4 yes Figure 3 Flowchart of step S302 in FIG.
[0047] Figure 5 yes Figure 1 Flowchart of step S105 in FIG.
[0048] Figure 6 yes Figure 5 Flowchart of step S502 in FIG.
[0049] Figure 7 yes Figure 1 Flowchart of step S106 in FIG.
[0050] Figure 8 Schematic diagram of the structure of the speech recognition device provided in the embodiment of the present application;
[0051] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0053] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0055] First, let’s analyze some of the terms used in this application:
[0056] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses theories, methods, technologies, and application systems that use digital computers or digital computer-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0057] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary subject between computer science and linguistics. It is often referred to as computational linguistics. Natural language processing includes grammatical analysis, semantic analysis, and text understanding. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing.
[0058] Information Extraction: A text processing technology that extracts specified types of entity, relationship, event, and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data is all text information extraction. Of course, the information extracted by text information extraction technology can be of various types.
[0059] Fourier transform: This function can be expressed as a linear combination of trigonometric functions (sine and / or cosine functions) or their integrals. In different research fields, the Fourier transform has many different variations, such as the continuous Fourier transform and the discrete Fourier transform.
[0060] Mel-Frequency Cipstal Coefficients (MFCCs) are a set of key coefficients used to construct the Mel-Frequency Cepstrum. A segment of a music signal can be used to generate a cepstrum that adequately represents the signal. The Mel-Frequency Cepstrum coefficients are the cepstrum (the spectrum of the spectrum) derived from this cepstrum. Unlike conventional cepstrum, the Mel-Frequency Cepstrum's most distinctive feature is that the frequency bands on the Mel-Frequency Cepstrum are evenly distributed on the Mel scale. This means that compared to the commonly seen linear cepstrum representation, these frequency bands are more closely aligned with the nonlinear human auditory system. For example, Mel-Frequency Cepstrum is often used in audio compression techniques.
[0061] Residual Networks: Residual networks can fully utilize all hierarchical features of the original image through residual dense blocks (RDBs). For very deep networks, directly extracting the output of each convolutional layer in the original image space is difficult or even impractical. RDBs serve as the building blocks of residual dense networks (RDNs). RDBs consist of densely connected layers and local feature fusion (LFF) with local residual learning (LRL). Each RDB convolutional layer has access to all subsequent layers, passing on information that needs to be preserved. LFF adaptively preserves information by connecting the previous RDB with the states of all previous layers of the current RDB to extract local dense features. Furthermore, LFF achieves extremely high growth rates by stabilizing the training of larger networks. After extracting multiple layers of local dense features, global feature fusion (GFF) is further performed to adaptively preserve hierarchical features globally. Each layer has direct access to the original image input, enabling implicit deep supervised learning.
[0062] Pooling: It is essentially a kind of sampling. It selects a certain method to reduce the dimension and compress the input feature map to speed up the operation. The most common pooling process is Max Pooling.
[0063] Decoder: It converts the previously generated fixed vector into an output sequence. The input sequence can be text, voice, image, or video; the output sequence can be text or image.
[0064] Softmax function: The Softmax function is a normalized exponential function that can "compress" a K-dimensional vector z containing any real number into another K-dimensional real vector σ(z) so that the range of each element is between (0,1) and the sum of all elements is 1. This function is often used in multi-classification problems.
[0065] Currently, Gaussian modeling is often used to identify the speaker identity corresponding to speech data. This method often requires relatively complex Gaussian calculations and has a high risk of error, which affects the accuracy of speech recognition. Therefore, how to improve the accuracy of speech recognition has become a technical problem that needs to be solved urgently.
[0066] Based on this, the embodiments of the present application provide a speech recognition method, a speech recognition device, an electronic device and a storage medium, aiming to improve the accuracy of speech recognition.
[0067] The speech recognition method, speech recognition device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the speech recognition method in the embodiments of the present application is described.
[0068] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0069] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0070] The speech recognition method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The speech recognition method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms; the software can be an application that implements the speech recognition method, etc., but is not limited to the above forms.
[0071] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0072] Figure 1 This is an optional flowchart of the speech recognition method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0073] Step S101, obtaining target speech data of a target speaker;
[0074] Step S102, performing spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients;
[0075] Step S103, performing dimension conversion processing on the target Mel-frequency cepstral coefficient to obtain a first latent state feature;
[0076] Step S104: extracting the first latent state features through a preset residual network to obtain target speech features;
[0077] Step S105, performing pooling processing on the target speech features through a preset pooling network to obtain a target speech vector;
[0078] Step S106 , decoding the target speech vector to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker.
[0079] In the embodiment of the present application, steps S101 to S106 are performed by obtaining the target speech data of the target speaker; performing spectral analysis on the target speech data to obtain the target Mel-cepstral coefficients, which can more conveniently convert the speech data into spectral features, and analyzing and processing the spectral features to obtain the target Mel-cepstral coefficients, so that the identity of the speaker can be identified through the spectral features, thereby improving the recognition accuracy. By performing dimensionality conversion processing on the target Mel-cepstral coefficients, a first latent state feature is obtained; a feature extraction of the first latent state feature is performed on the preset residual network to obtain the target speech feature; a pooling processing is performed on the target speech feature by a preset pooling network to obtain a target speech vector, which can map the latent state feature corresponding to the target Mel-cepstral coefficient to a specific vector dimension through dimensionality conversion processing, and change the temporal state of the latent state feature through pooling processing, thereby obtaining a target speech vector that meets the requirements, and using the target speech vector as a feature vector for identifying the identity of the target speaker. Finally, the target speech vector is decoded to obtain the target person label corresponding to the target speech data, thereby conveniently determining the identity of the target speaker according to the corresponding target person label, thereby improving the accuracy of speech recognition.
[0080] In step S101 of some embodiments, a web crawler can be programmed to set up a data source and then crawl data in a targeted manner to obtain target voice data. The data source can be various types of online platforms, social media, or certain specific audio databases, and the target voice data can be music materials, speech reports, chat conversations, etc. of the target speaker. Other methods can also be used to obtain the target voice data, and are not limited to these.
[0081] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0082] See also Figure 2 In some embodiments, step S102 may include but is not limited to steps S201 to S202:
[0083] Step S201, performing spectrum calculation on the target speech data to obtain a target spectrogram;
[0084] Step S202: Filter the target spectrogram to obtain target Mel-frequency cepstral coefficients.
[0085] In step S201 of some embodiments, a short-time Fourier transform (SFT) is used to calculate the sound spectrum of the target speech data to obtain a target spectrogram. Specifically, the target speech data is subjected to signal framing and windowing to obtain multiple frames of speech segments. Each frame of the speech segment is subjected to a SFT to convert the time domain features of the speech segment into frequency domain features. Finally, the frequency domain features of each frame are stacked in the time dimension to obtain a target spectrogram. Each frame of the speech segment has a frame length of 20ms, a frame shift of 10ms, and a SFT point count of 512.
[0086] In step S202 of some embodiments, the target spectrogram is filtered using a 64-dimensional Mel-cepstral filter bank. A logarithmic operation is first performed on the target spectrogram to obtain a target log spectrum. The target log spectrum is then inverse Fourier transformed to obtain a target Mel-cepstral. Furthermore, feature extraction is performed on the target Mel-cepstral to obtain target Mel-cepstral coefficients. The feature dimension of the target Mel-cepstral coefficients is T*64, where T is the number of frames of the target speech data.
[0087] Through the above steps S201 to S202, the speech data can be easily converted into spectral features, and the spectral features are filtered to obtain target Mel-frequency cepstral coefficients, so that the speaker's identity can be identified through the spectral features, thereby improving recognition accuracy.
[0088] Before step S103 in some embodiments, the speech recognition method further includes a pre-speech recognition model, where the language model includes a convolutional network, a residual network, a pooling network, and a decoding network. Among them, the convolutional network includes at least one convolutional layer, the convolution kernel size of the convolutional layer is 3, the number of channels is 512, and the convolutional network is mainly used to increase the dimension of the input vector to obtain a 512-dimensional vector, wherein the input vector can be obtained by performing spectral analysis on the sample speech data; the residual network includes at least one convolutional layer and at least one activation layer, and the residual network is mainly used to input the vector output by the convolutional network into the convolutional layer and the activation layer respectively for processing, and then perform weighted calculation on the processed vector to output the residual feature vector; the pooling network includes a first convolutional layer, a second convolutional layer and a pooling layer, which is used to increase the dimension of the residual feature vector through the first convolutional layer, and adjust the time series state of the vector after the dimensionality increase processing through the pooling layer to obtain a time-independent feature variable, and then increase the dimension of the feature variable through the second convolutional layer to obtain a sample speech vector; the decoding network is used to perform probability calculation and decoding processing on the sample speech vector output by the pooling network through a probability function, etc., so as to obtain a person label corresponding to the input vector, thereby identifying the identity of the target speaker through the person label.
[0089] Furthermore, when training the speech recognition model, the Additive Angular Margin (AAM) loss function can be used for training. The AAM loss function can be expressed as shown in formula (1):
[0090]
[0091] Where N is the number of target speakers in the sample speech dataset, θ yi,i is the intermediate variable of the sample speech vector and the weight W of the second convolutional layer yi The angle between them, m and s are fixed values, m is 0.2, s is 30, i represents the i-th speaker among N speakers, and yi represents the weight W of the second convolutional layer corresponding to the i-th person. yi , j represents the jth speaker among N speakers, yj represents the weight W of the second convolutional layer corresponding to the jth person yj .
[0092] In step S103 of some embodiments, the target Mel-cepstral coefficient is subjected to dimension-changing processing through a convolutional network to obtain a first hidden state feature, wherein the convolutional network includes at least one convolutional layer, the convolution kernel size of the convolutional layer is 3, and the number of channels is 512. The target Mel-cepstral coefficient can be subjected to dimension-changing processing through the convolutional network, and the characteristic dimension of the target Mel-cepstral coefficient is adjusted to 512, thereby filtering out irrelevant information to obtain a Mel-cepstral hidden state, i.e., the first hidden state feature.
[0093] See also Figure 3 In some embodiments, step S104 may include but is not limited to steps S301 to S303:
[0094] Step S301, performing convolution processing on the first hidden state feature through the residual network to obtain a first convolution feature vector;
[0095] Step S302: Activate the first hidden state feature using the activation function of the residual network to obtain a target activation feature vector;
[0096] Step S303: summing the first convolution feature and the target activation feature vector according to a preset weight parameter to obtain the target speech feature.
[0097] In step S301 of some embodiments, the residual network has a total of four layers, each consisting of four identical network structures. The first hidden state feature is the input of the first layer, the output of the first layer is the input of the second layer, and so on. Finally, the target speech feature is output through the fourth layer. The convolution kernels of each layer of the network structure are 3, 7, 11, and 15, respectively. Each layer of the network structure includes two convolution layers (convolution layer A and convolution layer B) and an activation layer. The first hidden state feature is convolved by convolution layer A of the residual network to capture the spatial features of the first hidden state feature and obtain a first convolution feature vector, where the size of convolution layer A is 1*1.
[0098] In step S302 of some embodiments, the activation layer of the residual network includes two activation functions, which are used to activate the first hidden state feature to obtain a target activation feature vector, wherein the target activation feature vector can be obtained by performing dot multiplication on the activation vector output by each activation function and convolving the vector obtained by the dot multiplication through the convolution layer B, wherein the size of the convolution layer A is 1*1.
[0099] In step S303 of some embodiments, the preset weight parameters can be set according to actual business needs without limitation. The specific process of summing the first convolution feature and the target activation feature vector according to the preset weight parameters to obtain the target speech feature z can be expressed as shown in formula (2):
[0100] z=w1*x+w2*(tanh(x)⊙sigmoid(x)) Formula (2)
[0101] Among them, * is a one-dimensional convolution calculation, the tanh function and the sigmoid function are activation functions, ⊙ is the dot product processing, w1 and w2 are weight parameters, and x is the first hidden state feature of the input or the output of the previous layer.
[0102] Through the above steps S301 to S303, the target Mel-cepstrum feature corresponding to the target speech data can be extracted to extract the Mel-cepstrum latent state of variable length, and the Mel-cepstrum latent state is activated to obtain the target speech feature, thereby enriching the semantic information of the target speech feature and improving the accuracy of speech recognition.
[0103] See also Figure 4 In some embodiments, the activation function includes a first function and a second function, and step S302 may include but is not limited to steps S401 to S403:
[0104] Step S401, activating the first latent state feature using a first activation function to obtain a first activated feature vector;
[0105] Step S402: Activate the first latent state feature using a second activation function to obtain a second activated feature vector;
[0106] Step S403: Perform a dot product process on the first activation feature vector and the second activation feature vector to obtain a target activation feature vector.
[0107] In step S401 of some embodiments, the first activation function is a tanh function, and the tanh function can be used to transform the element value of the input first latent state feature to between -1 and 1 to obtain a first activation feature vector.
[0108] In step S402 of some embodiments, the second activation function is a sigmoid function, and the sigmoid function can transform the element value of the input first latent state feature to between 0 and 1 to obtain a second activation feature vector.
[0109] In step S403 of some embodiments, matrix dot multiplication is performed on the first activation eigenvector and the second activation eigenvector to obtain a dot product vector, and then one-dimensional convolution is performed on the dot product vector to obtain a target activation eigenvector.
[0110] See also Figure 5 In some embodiments, the pooling network includes a first convolutional layer, a pooling layer, and a second convolutional layer. Step S105 may include, but is not limited to, steps S501 to S503:
[0111] Step S501, performing dimension conversion processing on the target speech feature through the first convolution layer to obtain a second latent state feature;
[0112] Step S502: pooling the second latent state feature through a pooling layer to obtain an intermediate vector;
[0113] Step S503: Perform dimension conversion processing on the intermediate vector through the second convolution layer to obtain the target speech vector.
[0114] In step S501 of some embodiments, the convolution kernel size of the first convolution layer is 1, and the number of channels is 1500. The first convolution layer can perform dimension-changing processing on the target speech feature, and adjust the feature dimension of the target speech feature to 1500, thereby obtaining a hidden state representing the Mel-frequency cepstral coefficient, that is, a second hidden state feature, wherein the feature dimension of the second hidden state feature is T*1500, and T is the number of frames of the target speech data.
[0115] In step S502 of some embodiments, because the length of the second latent state features is not fixed in time, the second latent state features need to be further adjusted to obtain a time-independent sentence-level vector. Specifically, the mean and standard deviation of the second latent state are calculated using a pooling layer, and then the second latent state is pooled based on the mean and standard deviation to obtain a time-independent intermediate vector, where the feature dimension of the intermediate vector is 1*3000.
[0116] In step S503 of some embodiments, the convolution kernel size of the second convolutional layer is 1, and the number of channels is 256. The second convolutional layer can perform dimension-changing processing on the intermediate vector, and adjust the feature dimension of the intermediate vector to 256, thereby obtaining a feature vector that can identify the identity of the speaker, that is, the target speech vector, wherein the feature dimension of the target speech vector is 1*256.
[0117] In the above steps S501 to S503, the target speech features are pooled through a preset pooling network to obtain a target speech vector. The latent state features corresponding to the target Mel-frequency cepstral coefficients can be mapped to a specific vector dimension through dimensionality change processing. The temporal state of the latent state features is changed through pooling processing to obtain a target speech vector that meets the requirements. The target speech vector is used as a feature vector for identifying the identity of the target speaker. The latent state can be mapped to a time-independent vector, and the target speech vector can be promoted to the sentence level, thereby improving the vector level of the target speech vector and improving the accuracy of speech recognition.
[0118] See also Figure 6 In some embodiments, step S502 includes but is not limited to steps S601 to S603:
[0119] Step S601, performing mean calculation on the second hidden state feature through a pooling layer and a preset time series dimension to obtain a hidden state mean;
[0120] Step S602, calculating the standard deviation of the second hidden state feature through the pooling layer and the time series dimension to obtain the hidden state standard deviation;
[0121] Step S603: Pooling the second latent state feature according to the latent state mean and the latent state standard deviation to obtain an intermediate vector.
[0122] In step S601 of some embodiments, the time series dimension can be set according to actual business needs, for example, the time series dimension is 3000. Feature summation processing is performed on the second hidden state features in the vector space with a time series dimension of 3000 to obtain the feature sum of the second hidden state features. The feature sum is then divided by the number of features of the second hidden state to calculate the mean of the second hidden state features and obtain the hidden state mean.
[0123] In step S602 of some embodiments, the difference between the eigenvalue of each second hidden state feature and the hidden state mean is calculated in a vector space with a time series dimension of 3000, and the standard deviation of the second hidden state feature is calculated based on a series of differences to obtain the hidden state standard deviation.
[0124] In step S603 of some embodiments, feature sampling is performed on the second hidden state feature using the hidden state mean and hidden state standard deviation of the pooling layer to obtain a time-independent intermediate vector.
[0125] Through the above steps S601 to S603, the second latent state feature with variable length can be adjusted to a time-independent feature variable, and this feature variable is promoted to the sentence level, thereby improving the vector level of the feature variable and improving the accuracy of speech recognition.
[0126] See also Figure 7 In some embodiments, step S106 may include but is not limited to steps S701 to S702:
[0127] Step S701, performing label probability calculation on the target speech vector using a preset function and preset character labels to obtain a label probability vector corresponding to each preset character label;
[0128] Step S702 : Select the preset person label corresponding to the label probability vector with the largest value as the target person label.
[0129] In step S701 of some embodiments, the preset function may be a prediction function such as a softmax function, and the preset character tags may be extracted from different data sources, for example, basic information of various characters, including identity information, personal information, and related audio and video data, may be obtained from online media and social platforms. Taking the softmax function as an example, the softmax function may be used to create a probability distribution of the target speech vector on each preset character tag. The probability distribution reflects the probability of the target speech vector belonging to each preset character tag, thereby obtaining a label probability vector corresponding to each preset character tag.
[0130] In step S702 of some embodiments, the size of the label probability vector can intuitively reflect the possibility that the target speech vector belongs to each preset character label. The larger the value of the label probability vector, the higher the degree of matching between the target speech vector and the corresponding preset character label, indicating that the target speech vector is more likely to come from the character corresponding to this preset character label. Therefore, the preset character label corresponding to the label probability vector with the largest value is selected as the target character label, thereby representing the identity of the target speaker through the target character label.
[0131] The above steps S701 to S702 can conveniently quantify the probability that the target speech vector belongs to each preset character label through a preset function to obtain a label probability vector, and then select the most appropriate preset character label as the target character label based on the size of the label probability vector, thereby confirming the identity of the target speaker based on the target character label, thereby improving the accuracy of speech recognition.
[0132] The speech recognition method of the embodiment of the present application obtains the target speech data of the target speaker; performs spectral analysis on the target speech data to obtain target Mel-frequency cepstral coefficients, which can more conveniently convert the speech data into spectral features, and analyzes and processes the spectral features to obtain target Mel-frequency cepstral coefficients, so that the identity of the speaker can be identified through the spectral features, thereby improving the recognition accuracy. Furthermore, the target Mel-frequency cepstral coefficients are subjected to dimension-changing processing to obtain first latent state features; the first latent state features are subjected to feature extraction through a preset residual network to obtain target speech features; the target speech features are subjected to pooling processing through a preset pooling network to obtain a target speech vector, which can map the latent state features corresponding to the target Mel-frequency cepstral coefficients to a specific vector dimension through dimension-changing processing, and change the temporal state of the latent state features through pooling processing, thereby obtaining a target speech vector that meets the requirements, and using the target speech vector as a feature vector for identifying the identity of the target speaker. Finally, the target speech vector is decoded to obtain a target person label corresponding to the target speech data, thereby conveniently determining the identity of the target speaker according to the corresponding target person label, thereby improving the accuracy of speech recognition.
[0133] See also Figure 8 The present invention also provides a speech recognition device that can implement the above-mentioned speech recognition method. The device includes:
[0134] The data acquisition module 801 is used to acquire target speech data of a target speaker;
[0135] An analysis module 802 is configured to perform spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients;
[0136] Dimension conversion module 803, used for performing dimension conversion processing on the target Mel-frequency cepstral coefficient to obtain the first latent state feature;
[0137] A feature extraction module 804 is configured to extract the first latent state features through a preset residual network to obtain target speech features;
[0138] The pooling module 805 is used to perform pooling processing on the target speech features through a preset pooling network to obtain a target speech vector;
[0139] The decoding module 806 is used to decode the target speech vector to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker.
[0140] In some embodiments, the analysis module 802 includes:
[0141] A calculation unit, configured to perform spectrum calculation on the target speech data to obtain a target spectrogram;
[0142] The filtering unit is used to filter the target spectrum to obtain the target Mel-frequency cepstral coefficients.
[0143] In some embodiments, the feature extraction module 804 includes:
[0144] A convolution unit, configured to perform convolution processing on the first hidden state feature through a residual network to obtain a first convolution feature vector;
[0145] An activation unit is used to activate the first hidden state feature through the activation function of the residual network to obtain a target activation feature vector;
[0146] The summing unit is used to sum the first convolution feature and the target activation feature vector according to a preset weight parameter to obtain the target speech feature.
[0147] In some embodiments, the activation function includes a first function and a second function, and the activation unit includes:
[0148] A first activation subunit is configured to activate the first latent state feature using a first activation function to obtain a first activation feature vector;
[0149] A second activation subunit is used to activate the first latent state feature through a second activation function to obtain a second activation feature vector;
[0150] The point multiplication subunit is used to perform point multiplication on the first activation feature vector and the second activation feature vector to obtain a target activation feature vector.
[0151] In some embodiments, the pooling network includes a first convolutional layer, a pooling layer, and a second convolutional layer, and the pooling module 805 includes:
[0152] A first dimension conversion unit is used to perform dimension conversion processing on the target speech feature through the first convolution layer to obtain a second latent state feature;
[0153] A pooling unit is used to perform pooling processing on the second hidden state features through a pooling layer to obtain an intermediate vector;
[0154] The second dimension conversion unit is used to perform dimension conversion processing on the intermediate vector through the second convolution layer to obtain the target speech vector.
[0155] In some embodiments, the pooling unit includes:
[0156] A mean calculation subunit is used to calculate the mean of the second hidden state feature through the pooling layer and the preset time series dimension to obtain the hidden state mean;
[0157] The standard deviation calculation subunit is used to calculate the standard deviation of the second hidden state feature through the pooling layer and the time series dimension to obtain the hidden state standard deviation;
[0158] The pooling subunit is used to perform pooling processing on the second hidden state feature according to the hidden state mean and the hidden state standard deviation to obtain an intermediate vector.
[0159] In some embodiments, the decoding module 806 includes:
[0160] A probability calculation unit is used to calculate the label probability of the target speech vector using a preset function and a preset character label to obtain a label probability vector corresponding to each preset character label;
[0161] The selection unit is used to select the preset person label corresponding to the label probability vector with the largest value as the target person label.
[0162] The specific implementation of the speech recognition device is basically the same as the specific embodiment of the above-mentioned speech recognition method, and will not be repeated here.
[0163] The present application also provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. When the program is executed by the processor, the aforementioned speech recognition method is implemented. The electronic device may be any intelligent terminal, such as a tablet computer or an in-vehicle computer.
[0164] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0165] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0166] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the speech recognition method of the embodiments of this application.
[0167] Input / output interface 903, used to implement information input and output;
[0168] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0169] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0170] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0171] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned speech recognition method.
[0172] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0173] The speech recognition method, speech recognition device, electronic device, and storage medium provided in the embodiments of the present application obtain target speech data to be recognized; perform spectral analysis on the target speech data to obtain target Mel-frequency cepstral coefficients, which can more conveniently convert the speech data into spectral features, and analyze and process the spectral features to obtain target Mel-frequency cepstral coefficients, so that the identity of the speaker can be identified through the spectral features, thereby improving recognition accuracy. Furthermore, the target Mel-frequency cepstral coefficients are subjected to dimensionality conversion processing to obtain first latent state features; the first latent state features are subjected to feature extraction through a preset residual network to obtain target speech features; the target speech features are subjected to pooling processing through a preset pooling network to obtain a time-independent target speech vector, which can map the latent state features corresponding to the target Mel-frequency cepstral coefficients to a specific vector dimension through dimensionality conversion processing, and change the temporal state of the latent state features through pooling processing to obtain a target speech vector that meets the requirements, and use the target speech vector as a feature vector for identifying the identity of the target speaker. Finally, the target speech vector is decoded to obtain a target person label corresponding to the target speech data, thereby conveniently determining the identity of the target speaker based on the corresponding target person label, thereby improving the accuracy of speech recognition. In addition, the embodiment of the present application introduces the AAM loss function to train the model parameters during the training process of the speech recognition model, so that compared with the model of traditional technology, there is no need to perform preprocessing such as silence segmentation on the input data, and the entire model training process is end-to-end model training, which can better simplify the model training process and save model training time.
[0174] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0175] It will be understood by those skilled in the art that Figure 1-7The technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or a combination of certain steps, or different steps.
[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0177] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0178] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0179] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0180] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0181] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0183] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0184] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A speech recognition method, characterized in that: The method comprises: Obtain target speech data of a target speaker; Performing spectral analysis on the target speech data to obtain target Mel-frequency cepstral coefficients; Performing dimension-changing processing on the target Mel-frequency cepstral coefficient to obtain a first latent state feature; Extracting the first latent state features through a preset residual network to obtain target speech features; Performing pooling processing on the target speech features through a preset pooling network to obtain a target speech vector; Decoding the target speech vector to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker; The pooling network includes a first convolutional layer, a pooling layer, and a second convolutional layer. The target speech feature is pooled by the preset pooling network to obtain a target speech vector, including: Performing dimensionality conversion processing on the target speech feature through the first convolution layer to obtain a second latent state feature, where the second latent state feature is a latent state representing the target Mel-cepstral coefficient; Calculating the hidden state mean and hidden state standard deviation of the second hidden state through the pooling layer and the preset time series dimension, and performing pooling processing on the second hidden state features according to the hidden state mean and the hidden state standard deviation to obtain an intermediate vector, wherein the intermediate vector is a sentence-level vector that is independent of time series; The intermediate vector is subjected to dimension change processing by the second convolutional layer to obtain the target speech vector.
2. The speech recognition method according to claim 1, wherein: The step of extracting the first latent state features through a preset residual network to obtain target speech features includes: Performing convolution processing on the first hidden state feature through the residual network to obtain a first convolution feature vector; Activate the first hidden state feature by using the activation function of the residual network to obtain a target activation feature vector; The first convolution feature and the target activation feature vector are summed according to a preset weight parameter to obtain the target speech feature.
3. The speech recognition method according to claim 2, wherein: The activation function includes a first function and a second function, and the step of activating the first hidden state feature by the activation function of the residual network to obtain a target activation feature vector includes: Activate the first latent state feature using the first function to obtain a first activated feature vector; Activate the first latent state feature using the second function to obtain a second activated feature vector; Performing a dot product process on the first activation feature vector and the second activation feature vector to obtain the target activation feature vector.
4. The speech recognition method according to claim 1, wherein: The step of decoding the target speech vector to obtain a target person label corresponding to the target speech data includes: Performing label probability calculation on the target speech vector using a preset function and a preset character label to obtain a label probability vector corresponding to each preset character label; The preset person label corresponding to the label probability vector with the largest value is selected as the target person label.
5. The speech recognition method according to any one of claims 1 to 4, characterized in that: The step of performing spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients includes: Performing spectrum calculation on the target speech data to obtain a target spectrogram; The target spectrogram is filtered to obtain target Mel-frequency cepstral coefficients.
6. A speech recognition device, characterized in that: The device comprises: A data acquisition module, used to acquire target speech data of a target speaker; An analysis module is used to perform spectrum analysis on the target speech data to obtain target Mel-frequency cepstral coefficients; A dimension conversion module, configured to perform dimension conversion processing on the target Mel-frequency cepstral coefficient to obtain a first latent state feature; A feature extraction module, configured to extract the first latent state features through a preset residual network to obtain target speech features; A pooling module, configured to perform pooling processing on the target speech features through a preset pooling network to obtain a target speech vector; A decoding module, configured to decode the target speech vector to obtain a target person label corresponding to the target speech data, wherein the target person label is used to represent the identity of the target speaker; The pooling network includes a first convolutional layer, a pooling layer, and a second convolutional layer. The target speech feature is pooled by the preset pooling network to obtain a target speech vector, including: Performing dimensionality conversion processing on the target speech feature through the first convolution layer to obtain a second latent state feature, where the second latent state feature is a latent state representing the target Mel-cepstral coefficient; Calculating the hidden state mean and hidden state standard deviation of the second hidden state through the pooling layer and the preset time series dimension, and performing pooling processing on the second hidden state features according to the hidden state mean and the hidden state standard deviation to obtain an intermediate vector, wherein the intermediate vector is a sentence-level vector that is independent of time series; The intermediate vector is subjected to dimension change processing by the second convolutional layer to obtain the target speech vector.
7. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the speech recognition method according to any one of claims 1 to 5 are realized.
8. A storage medium, which is a computer-readable storage medium and is used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the speech recognition method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speaker confirmation method and device based on residual time delay network, equipment and medium
CN110232932A
MFCC coefficient-based pig sound recognition system and method
CN114203187A