A speech recognition method and system

By using the depth residual shrink network in the preset depth residual shrink network model to remove irrelevant features in the speech signal, the problem of low recognition rate in a strong noise environment is solved, and a higher speech signal recognition rate and more accurate text information acquisition is achieved.

CN113889099BActive Publication Date: 2025-05-09SHANGHAI FUDAN KINGSTAR COMPUTER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111170413.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-08
Publication Date
2025-05-09
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

The existing deep learning-based speech recognition technology has a low recognition rate of speech signals in a highly noisy environment.

Method used

The depth residual shrinking network in the preset depth residual shrinking network model is used to filter the original voice signal, remove irrelevant features, and obtain text information of noise-free features.

Benefits of technology

The recognition rate of speech signals is improved in a strong noise environment, and more accurate text information is obtained by removing noise characteristics and environmental characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889099B_ABST
    Figure CN113889099B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method and system, which obtains an original speech signal, uses a deep residual shrinkage network in a preset deep residual shrinkage network model to filter the original speech signal to be recognized, obtains a target speech spectrum, extracts speech timing features from the target speech spectrum, classifies the speech timing features through a preset classification layer of the deep residual shrinkage network, obtains the character probability corresponding to the target speech spectrum, and predicts the character probability through a preset prediction model to obtain text information. Through the above, since the preset deep residual shrinkage network model incorporates a residual module and a soft threshold function, it has the characteristics of strong feature extraction capability and noise removal. The deep residual shrinkage network in the preset deep residual shrinkage network model is used to remove irrelevant features contained in the original speech spectrum, so that text information with features such as noise-free can be obtained in a strong noise environment, thereby improving the recognition rate of speech signals in a strong noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and more specifically, to a speech recognition method and system. Background Art

[0002] Speech recognition is an interdisciplinary subject that converts human speech sound waves into corresponding text output. Its more natural, convenient and efficient form of communication makes it one of the most important interfaces for human-computer interaction now and for a long time to come.

[0003] Speech is a time series signal. Speech recognition is based on the speech spectrum after time domain analysis. In the process of speech recognition, it is necessary to overcome the diversity faced by speech signals, including the diversity of the speaker's voice and the diversity of the environment. Existing end-to-end speech recognition technology mainly relies on deep learning technologies such as convolutional neural networks (CNN) and time series neural networks (RNN, GRU, LSTM, etc.) to solve the above problems and achieve a higher recognition rate.

[0004] However, the applicant has found that deep learning-based speech recognition technology has currently surpassed the human level in a clean environment and has a good recognition effect, but the recognition rate of speech signals in a strong noise environment is low. Summary of the invention

[0005] In view of this, the present application discloses a speech recognition method and system, which utilizes a deep residual shrinkage network in a preset deep residual shrinkage network model to remove irrelevant features contained in the original speech spectrum, so as to obtain text information with noise-free features in a noisy environment, aiming to improve the recognition rate of speech signals in a noisy environment.

[0006] In order to achieve the above purpose, the disclosed technical solution is as follows:

[0007] The first aspect of the present application discloses a speech recognition method, the method comprising:

[0008] Obtaining an original speech signal to be recognized;

[0009] The original speech signal to be recognized is filtered out by using a deep residual shrinkage network in a preset deep residual shrinkage network model to obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating the deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features include at least noise features and environmental features;

[0010] Extracting speech timing features from the target speech spectrum;

[0011] The speech time sequence features are classified by a preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum;

[0012] The character probability is predicted by a preset prediction model to obtain text information.

[0013] Preferably, the filtering process of the original speech signal to be recognized by using a deep residual shrinkage network in a preset deep residual shrinkage network model to obtain a target speech spectrum includes:

[0014] Preprocessing the original speech signal to be recognized by using the spectrum function of the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the original speech spectrum;

[0015] The irrelevant features contained in the original speech spectrum are removed by a preset soft threshold function of the deep residual shrinkage network to obtain a target speech spectrum.

[0016] Preferably, the extracting speech timing features from the target speech spectrum includes:

[0017] Extracting speech timing features from the target speech spectrum through the recurrent neural network layer of the preset deep residual shrinkage network model; the recurrent neural network layer includes a unidirectional recurrent neural network layer or a bidirectional recurrent neural network layer;

[0018] If the recurrent neural network layer is a unidirectional recurrent neural network layer, extracting speech timing features from the target speech spectrum through the unidirectional recurrent neural network layer;

[0019] If the recurrent neural network layer is a bidirectional recurrent neural network layer, the speech timing features are extracted from the target speech spectrum through the bidirectional recurrent neural network layer.

[0020] Preferably, the preset classification layer includes a fully connected layer and a logistic regression layer, the fully connected layer includes a first fully connected layer and a second fully connected layer, and the preset classification layer of the deep residual shrinkage network is used to classify the speech time series features to obtain the character probability corresponding to the target speech spectrum, including:

[0021] Inputting the speech time series features into the first fully connected layer and the second fully connected layer to obtain a speech output vector;

[0022] The speech output vector is input into the logistic regression layer for classification to obtain the first character probability and the second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in the preset speech model.

[0023] Preferably, the predicting the character probability by a preset prediction model to obtain text information includes:

[0024] Obtain a first character corresponding to the first character probability and a second character corresponding to the second character probability;

[0025] Combine the first character and the second character to obtain a character string;

[0026] The character string is calculated by using a preset function and a preset algorithm to obtain text information.

[0027] A second aspect of the present application discloses a speech recognition system, the system comprising:

[0028] An acquisition unit, used for acquiring an original speech signal to be recognized;

[0029] A filtering unit is used to filter the original speech signal to be recognized by using a deep residual shrinkage network in a preset deep residual shrinkage network model to obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating the deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features include at least noise features and environmental features;

[0030] An extraction unit, used for extracting speech time sequence features from the target speech spectrum;

[0031] A classification unit, used to classify the speech time series features through a preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum;

[0032] The prediction unit is used to predict the character probability through a preset prediction model to obtain text information.

[0033] Preferably, the filtering unit comprises:

[0034] A preprocessing module, used to preprocess the original speech signal to be recognized by using the spectrum function of the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the original speech spectrum;

[0035] The removal module is used to remove the noise features contained in the original speech spectrum through a preset soft threshold function of the deep residual shrinkage network to obtain a target speech spectrum.

[0036] Preferably, the extraction unit comprises:

[0037] A first extraction module is used to extract speech timing features from the target speech spectrum through a recurrent neural network layer of the preset deep residual shrinkage network model; the recurrent neural network layer includes a unidirectional recurrent neural network layer or a bidirectional recurrent neural network layer;

[0038] A second extraction module is used for extracting speech timing features from the target speech spectrum through the unidirectional recurrent neural network layer if the recurrent neural network layer is a unidirectional recurrent neural network layer;

[0039] The third extraction module is used to extract speech timing features from the target speech spectrum through the bidirectional recurrent neural network layer if the recurrent neural network layer is a bidirectional recurrent neural network layer.

[0040] Preferably, the classification unit includes:

[0041] An input module, used for inputting the speech time series features into the first fully connected layer and the second fully connected layer to obtain a speech output vector;

[0042] A classification module is used to input the speech output vector into the classification layer for classification to obtain a first character probability and a second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in a preset speech model.

[0043] Preferably, the prediction unit comprises:

[0044] An acquisition module, used for acquiring a first character corresponding to the first character probability and a second character corresponding to the second character probability;

[0045] A combining module, used for combining the first character and the second character to obtain a character string;

[0046] The calculation module is used to calculate the character string through a preset function and a preset algorithm to obtain text information.

[0047] It can be known from the above technical scheme that the present application discloses a speech recognition method and system, which obtains an original speech signal to be recognized, and uses a deep residual shrinkage network in a preset deep residual shrinkage network model to filter out the original speech signal to be recognized to obtain a target speech spectrum. The preset deep residual shrinkage network model is a model constructed by integrating a deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features include at least noise features and environmental features, and speech timing features are extracted from the target speech spectrum. The speech timing features are classified through a preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum, and the character probability is predicted through a preset prediction model to obtain text information. Through the above scheme, since the residual module and soft threshold function are integrated into the preset deep residual shrinkage network model, it has the characteristics of strong feature extraction and noise removal. The deep residual shrinkage network in the preset deep residual shrinkage network model is used to remove irrelevant features contained in the original speech spectrum, so that text information with noise-free features can be obtained in a strong noise environment, thereby improving the recognition rate of speech signals in a strong noise environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0049] Figure 1 A flowchart of a speech recognition method disclosed in an embodiment of the present application;

[0050] Figure 2 A schematic diagram of the structure of a preset deep residual shrinkage network model disclosed in an embodiment of the present application;

[0051] Figure 3 A schematic diagram of a soft threshold function disclosed in an embodiment of the present application;

[0052] Figure 4 This is a schematic diagram of the structure of SENet in the residual mode disclosed in the embodiment of the present application;

[0053] Figure 5 A schematic diagram of the structure of a speech recognition system disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0055] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.

[0056] As can be seen from the background technology, the applicant has found that the speech recognition technology based on deep learning has surpassed the human level in a clean environment and has a good recognition effect, but the recognition rate of speech signals in a strong noise environment is low.

[0057] Due to the limitations of the number of convolutional neural network layers and the CNN's own ability to extract features, the applicant's research has concluded that the recognition rate of speech signals in a strong noise environment is low.

[0058] In order to solve the above problems, the embodiment of the present application discloses a speech recognition method and system. Since the residual module and soft threshold function are integrated into the preset deep residual shrinkage network model, it has the characteristics of strong feature extraction and noise removal. The original speech signal is input into the preset deep residual shrinkage network model to remove irrelevant features contained in the original speech spectrum, so that text information with features such as noise-free can be obtained in a strong noise environment, thereby improving the recognition rate of speech signals in a strong noise environment. The specific implementation method is described by the following embodiment.

[0059] refer to Figure 1 FIG. 1 is a flow chart of a speech recognition method disclosed in an embodiment of the present application. The speech recognition method is applied to a server and mainly includes the following steps:

[0060] S101: Acquire an original speech signal to be recognized.

[0061] In S101, the original voice signal to be recognized is obtained by collecting the user's voice. Specifically, through the request interface provided to the outside, the sound file sent by the user based on the network communication module socket is obtained, and the original voice signal to be recognized is obtained through the sound file.

[0062] Among them, the original speech signal is a common time series, which is encoded in the form of a discrete signal and then stored using a certain sound file format.

[0063] The format type of the sound file can be a WAV format file, or a Microsoft audio format (Windows Media Audio, WMA) file, etc. The specific type of the sound file is not specifically limited in this application.

[0064] S102: Utilize a deep residual shrinkage network in a preset deep residual shrinkage network model to filter out the original speech signal to be recognized, and obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating a deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; irrelevant features include at least noise features and environmental features.

[0065] In S102, the process of filtering the original speech signal to be recognized by using the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the target speech spectrum is as follows:

[0066] Firstly, the spectral function Spectrogram of the deep residual shrinkage network in the preset deep residual shrinkage network model is used to preprocess the original speech signal to obtain the original speech spectrum. Then, the irrelevant features contained in the original speech spectrum are removed by the preset soft threshold function of the deep residual shrinkage network to obtain the target speech spectrum.

[0067] Among them, the spectrum function Spectrogram of the preset deep residual shrinkage network model (DRSN-DeepSpeech) is used to perform short-time Fourier transform on the original speech signal to be recognized to obtain the original speech spectrum.

[0068] Preprocessing involves processing the original speech signal to be recognized by framing, windowing, feature extraction, etc.

[0069] The irrelevant features include noise features, environmental features, etc., wherein the environmental features include far-field, near-field, and other features.

[0070] For the specific preset structure of the deep residual shrinkage network model, please refer to Figure 2 shown.

[0071] Figure 2In the deep residual shrinkage network model, the preset deep residual shrinkage network model is based on the open source model DeepSpeech2, which replaces the original convolutional neural network (CNN) input layer with a deep residual shrinkage network with Channel-Wise (DRSN-CW) for strong noise signals. The deep residual shrinkage network model adds cross-layer identity paths.

[0072] Figure 2 In the language recognition scenario, the first two convolutional layers convert irrelevant features such as environmental features or noise features into values ​​close to zero. After removing irrelevant features, the obtained useful features (target speech spectrum) are converted into values ​​far from zero. Then a set of thresholds are automatically learned, and irrelevant features are deleted using soft thresholding (soft threshold function), so that the target speech spectrum is retained.

[0073] The original speech spectrum is passed into the DRSN-CW of the deep residual contraction network to remove irrelevant features such as diversity and noise.

[0074] Figure 2 In the figure, Spectrogram is the spectrum function; BatchNormalization and BN are both batch normalization. Each batch of data is normalized during forward propagation, which is conducive to training convergence; RNN is a recurrent neural network, which extracts audio time series features through RNN; Lookahead Convolution is a forward-looking convolutional network, which integrates the information and features of future sequences through Lookahead Convolution; FC is a fully connected network; Connectionist Temporal Classification (CTC), through CTC, the input and output of the sequence do not need to be aligned one by one. The CTC function is a loss function used to measure the difference between the input sequence data and the actual output after passing through the neural network; ReLU and Sigmoid are both activation functions; C is the number of channels of the feature map; W is the width of the feature map; 1 is the height of the feature map, FullyConnected is a fully connected network, Vanilla or GRU, Uni or Bi and directional are various types of convolutional neural networks RNN, Conv is convolution, K is the number of convolution kernels, M is the number of network units in the fully connected layer, Absolute is the absolute value of the feature, and GAP is the global average pooling is global average pooling, x is the signal feature, and Identityshortcut is the identity path representing the residual module.

[0075] The entire network model is based on three parts: deep residual network, soft threshold function and attention mechanism. The deep residual network uses cross-layer identity path to greatly reduce the difficulty of training deep networks and solve the problem of network degradation, so that deeper networks have better feature extraction capabilities.

[0076] The expression of the soft threshold function is shown in formula (1).

[0077]

[0078] Among them, y is the optimization variable, x is the coefficient of wavelet transform, and τ is the threshold.

[0079] First, a positive threshold needs to be set. This threshold cannot be too large, that is, it cannot be greater than the maximum absolute value of the input data. Then, the soft threshold function will set the input data with an absolute value lower than this threshold to zero, and shrink the input data with an absolute value greater than this threshold towards zero. The relationship between input and output is as follows: Figure 3 shown.

[0080] Figure 3 In the figure, the soft threshold function is λ=2; x is the coefficient of wavelet transform, and the x-axis includes -3τ, -2τ, -τ, 0, τ, 2τ and 3τ; y is the optimization variable, and the y-axis includes -2τ, -τ, 0, τ, 2τ; τ is the threshold.

[0081] The Squeeze-and-Excitation Networks (SENet) algorithm is used to automatically learn and determine the threshold τ without manual operation.

[0082] SENet is a very classic attention algorithm. It learns a set of weight coefficients through a small network. Specifically, it first evaluates the importance of each feature channel and then assigns appropriate weights to each feature channel according to its importance.

[0083] The structure of SENet in residual mode can be referenced Figure 4 shown.

[0084] Figure 4 In the figure, X is the feature map, H is the height of the feature map, W is the width of the feature map, C is the number of channels of the feature map, Residual is the residual, Globalpooling is the global pooling, FC is the fully connected network, ReLU and Sigmoid are both activation functions, Scale is the scaling, r is the dimensionality reduction coefficient, is the signal characteristic of the final output.

[0085] The feature map X is first compressed into a one-dimensional vector, then input into two fully connected layers (the first fully connected layer and the second fully connected layer), and finally passes through the sigmoid function to obtain a scaling factor α between (0, 1). The expression of α is shown in formula (2).

[0086]

[0087] Among them, e is a natural constant, -Z is a negative eigenvalue, and α is a scaling factor.

[0088] The threshold τ can be calculated by formula (2). The expression of the threshold τ is shown in formula (3).

[0089] τ=α·average|X i,j,c | (3)

[0090] Among them, X is the feature map, i is the width of the feature map, j is the height of the feature map, c is the channel of the feature map, α is the scaling factor, τ is the threshold, and average is the average of the feature map.

[0091] Formula (3) ensures that all thresholds τ are positive, but not too large, thus avoiding the situation where all outputs are zero.

[0092] The speech signal is processed by DRSN-CW and then passed into the RNN. For large-scale speech data, in order to fully and effectively utilize the data, it is usually necessary to increase the nodes and depth of the RNN, which will make it more difficult to descend the gradient during training. Therefore, batch normalization not only speeds up the convergence of RNN training, but also improves the generalization error most of the time. This algorithm applies sequence normalization and only performs batch normalization on vertical connections.

[0093]

[0094] in, is the output value of layer l at time step t, f is the activation function, B is the batch normalization operation, W l is the corresponding weight, is the output value of layer l-1 at time step t, U l is the corresponding weight, is the output value of layer l at time step t-1.

[0095] S103: Extracting speech timing features from the target speech spectrum.

[0096] In S103, the speech of different people has different speech timing features. The speech timing features may include amplitude and zero crossing rate. For example, the amplitude can be used to distinguish voiced and unvoiced sounds, initials and finals, etc.; the zero crossing rate can be used to distinguish voiced and unvoiced sounds.

[0097] The process of extracting speech timing features from the target speech spectrum is shown in A1-A3.

[0098] A1: Extract speech timing features from the target speech spectrum through a recurrent neural network layer RNN of a preset deep residual shrinkage network model; the recurrent neural network layer RNN includes a unidirectional recurrent neural network layer RNN or a bidirectional recurrent neural network layer RNN.

[0099] Among them, a recurrent neural network is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain.

[0100] The unidirectional recurrent neural network layer can see different outputs at different times (such as time t, time t-1, time t+1), and the hidden layer at the previous time will affect the output at the current time. This structure is the unidirectional recurrent neural network structure.

[0101] In some tasks, the output at the current moment is not only related to the past information, but also to the information at the subsequent moments. For example, given a sentence, that is, a sequence of words, the part of speech of each word is related to the context, so a network layer can be added to transmit information in reverse time order to enhance the network's capabilities. So there is a bidirectional recurrent neural network (Bi-RNN), which consists of two layers of recurrent neural networks. Both layers of the network input sequences, but the information transmission direction is opposite.

[0102] A2: If the recurrent neural network layer is a unidirectional recurrent neural network layer, the speech timing features are extracted from the target speech spectrum through the unidirectional RNN layer.

[0103] A3: If the recurrent neural network layer is a bidirectional recurrent neural network layer, the speech timing features are extracted from the target speech spectrum through the bidirectional RNN layer.

[0104] S104: Classify the speech time series features through the preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum.

[0105] In S104, the preset classification layer of the deep residual shrinkage network plays the role of a classifier. The preset classification layer includes a fully connected layer and a logistic regression layer (Softmax logical regression, softmax). The number of fully connected layers can be multiple. The fully connected layer of this scheme includes a first weighted connection layer and a second fully connected layer.

[0106] Each node in the fully connected layer is connected to all nodes in the previous layer to combine the features extracted previously. Due to its fully connected nature, the fully connected layer generally has the most parameters.

[0107] Specifically, the process of classifying the speech time series features through the preset classification layer and obtaining the character probability corresponding to the target speech spectrum is as follows:

[0108] First, the speech time series features are input into the first fully connected layer and the second fully connected layer to obtain the speech output vector. Then, the speech output vector is input into the logistic regression layer for classification to obtain the first character probability and the second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in the preset speech model.

[0109] S105: Predicting the character probability by using a preset prediction model to obtain text information.

[0110] In S105, the preset prediction model is obtained by combining the CTC function and the language model trained on large-scale text.

[0111] The process of predicting character probabilities through the preset prediction model and obtaining text information is as follows:

[0112] First, a first character corresponding to the first character probability and a second character corresponding to the second character probability are obtained, and then the first character and the second character are combined to obtain a character string, and the character string is calculated by a preset function and a preset algorithm to obtain text information.

[0113] Among them, the preset function can be a CTC function or other functions. The specific preset function is determined by the technical personnel according to the actual situation. This application does not make any specific limitations. The preset function of this application is preferably a CTC function.

[0114] The preset algorithm can be a beam search (Breadth First Search, beam search) algorithm, or a best priority algorithm, etc. The specific preset algorithm is determined by technical personnel according to actual conditions, and this application does not make specific limitations. The preset algorithm of this application is preferably a beam search algorithm.

[0115] The Beam Search algorithm is a heuristic graph search algorithm, which is usually used when the graph solution space is relatively large. In order to reduce the space and time occupied by the search, some nodes with poor quality are cut off at each step of depth expansion, and some nodes with higher quality are retained. This reduces space consumption and improves time efficiency.

[0116] The calculation formula for predicting character probability using the preset prediction model is shown in formula (5).

[0117] Q(y)=log(P(y|x))+αlog(P LM (y))+β·wc(y) (5)

[0118] Among them, Q(y) is a variable, and character y is determined by maximizing this variable; P(y|x) is the probability of character y appearing in audio x; α is the weight for adjusting the language model and CTC function; P LM (y) is the probability of character y appearing in the language model; β is the predicted text length, that is, the predicted string length; wc(y) is the number of characters or words in the transcribed text (string).

[0119] Based on the speech recognition model DeepSpeech2, this solution uses a deep residual shrinkage model to deal with speech signal diversity and noise problems, which significantly improves the recognition accuracy of noisy speech. On the one hand, the deep residual shrinkage model incorporates a residual module to solve the grid degradation problem, and can use deeper network layers to enhance the model's ability to extract features; on the other hand, the deep residual shrinkage model implants the soft threshold function in traditional linear signal processing to filter out noise components in the signal. The algorithm combines the nonlinearity of deep learning with the linearity of traditional signal processing, which is more in line with the composition of actual speech signals. Therefore, it can significantly improve the model's ability to parse noisy speech and has stronger robustness.

[0120] In the embodiment of the present application, since the residual module and the soft threshold function are integrated into the preset deep residual shrinkage network model, it has the characteristics of strong feature extraction and noise removal. The deep residual shrinkage network in the preset deep residual shrinkage network model is used to remove irrelevant features contained in the original speech spectrum, so that text information with noise-free features can be obtained in a strong noise environment, thereby improving the recognition rate of speech signals in a strong noise environment.

[0121] Based on the above embodiments Figure 1 A speech recognition method is disclosed, and the present application embodiment also discloses a speech recognition system, such as Figure 5 As shown, the speech recognition system includes an acquisition unit 501, a filtering unit 502, an extraction unit 503, a classification unit 504 and a prediction unit 505.

[0122] The acquisition unit 501 is used to acquire an original speech signal to be recognized.

[0123] The filtering unit 502 is used to filter the original speech signal to be recognized by using the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating the deep residual shrinkage network into the deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features at least include noise features and environmental features.

[0124] The extraction unit 503 is used to extract speech time sequence features from the target speech spectrum.

[0125] The classification unit 504 is used to classify the speech time series features through a preset number of classification layers of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum.

[0126] The prediction unit 505 is used to predict the character probability through a preset prediction model to obtain text information.

[0127] Furthermore, the filtering unit 502 includes a pre-processing module and a removal module.

[0128] The preprocessing module is used to preprocess the original speech signal to be recognized by using the spectrum function of the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the original speech spectrum.

[0129] The removal module is used to remove irrelevant features contained in the original speech spectrum through a preset soft threshold function of the deep residual shrinkage network to obtain a target speech spectrum.

[0130] Furthermore, the extraction unit 503 includes a first extraction module, a second extraction module and a third extraction module.

[0131] The first extraction module is used to extract speech timing features from the target speech spectrum through a recurrent neural network layer of a preset deep residual shrinkage network model; the recurrent neural network layer includes a unidirectional recurrent neural network layer or a bidirectional recurrent neural network layer.

[0132] The second extraction module is used to extract speech timing features from the target speech spectrum through the unidirectional recurrent neural network layer if the recurrent neural network layer is a unidirectional recurrent neural network layer.

[0133] The third extraction module is used to extract speech timing features from the target speech spectrum through the bidirectional recurrent neural network layer if the recurrent neural network layer is a bidirectional recurrent neural network layer.

[0134] Furthermore, the classification unit includes an input module and a classification module.

[0135] The input module is used to input the speech time series features into the first fully connected layer and the second fully connected layer to obtain the speech output vector.

[0136] The classification module is used to input the speech output vector into the logistic regression layer for classification to obtain the first character probability and the second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in the preset speech model.

[0137] Furthermore, the prediction unit 505 includes an acquisition module, a combination module and a calculation module.

[0138] The acquisition module is used to acquire a first character corresponding to the first character probability and a second character corresponding to the second character probability.

[0139] The combining module is used to combine the first character and the second character to obtain a character string.

[0140] The calculation module is used to calculate the character string through a preset function and a preset algorithm to obtain text information.

[0141] In the embodiment of the present application, since the preset deep residual shrinkage network model has the characteristics of strong feature extraction capability and noise removal, the deep residual shrinkage network in the preset deep residual shrinkage network model is used to remove irrelevant features contained in the original speech spectrum, so that text information with noise-free features can be obtained in a strong noise environment, thereby improving the recognition rate of speech signals in a strong noise environment.

[0142] The specific implementation processes and derivative methods of the above-mentioned embodiments are all within the protection scope of this application.

[0143] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0144] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0145] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0146] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A speech recognition method, characterized in that: The method comprises: Obtaining an original speech signal to be recognized; The original speech signal to be recognized is filtered out by using a deep residual shrinkage network in a preset deep residual shrinkage network model to obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating the deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features include at least noise features and environmental features; Extracting speech timing features from the target speech spectrum; The speech time sequence features are classified by a preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum; Predicting the character probability by a preset prediction model to obtain text information; Among them, the preset classification layer includes a fully connected layer and a logistic regression layer, the fully connected layer includes a first fully connected layer and a second fully connected layer, and the preset classification layer of the deep residual shrinkage network is used to classify the speech time series features to obtain the character probability corresponding to the target speech spectrum, including: Inputting the speech time series features into the first fully connected layer and the second fully connected layer to obtain a speech output vector; The speech output vector is input into the logistic regression layer for classification to obtain the first character probability and the second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in the preset speech model.

2. The method according to claim 1, characterized in that The method uses a deep residual shrinkage network in a preset deep residual shrinkage network model to filter the original speech signal to be recognized to obtain a target speech spectrum, including: Preprocessing the original speech signal to be recognized by using the spectrum function of the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the original speech spectrum; The irrelevant features contained in the original speech spectrum are removed by a preset soft threshold function of the deep residual shrinkage network to obtain a target speech spectrum.

3. The method according to claim 1, characterized in that The extracting speech timing features from the target speech spectrum comprises: Extracting speech timing features from the target speech spectrum through the recurrent neural network layer of the preset deep residual shrinkage network model; the recurrent neural network layer includes a unidirectional recurrent neural network layer or a bidirectional recurrent neural network layer; If the recurrent neural network layer is a unidirectional recurrent neural network layer, extracting speech timing features from the target speech spectrum through the unidirectional recurrent neural network layer; If the recurrent neural network layer is a bidirectional recurrent neural network layer, the speech timing features are extracted from the target speech spectrum through the bidirectional recurrent neural network layer.

4. The method according to claim 1, characterized in that: The method of predicting the character probability by using a preset prediction model to obtain text information includes: Obtain a first character corresponding to the first character probability and a second character corresponding to the second character probability; Combine the first character and the second character to obtain a character string; The character string is calculated by using a preset function and a preset algorithm to obtain text information.

5. A speech recognition system, characterized in that: The system comprises: An acquisition unit, used for acquiring an original speech signal to be recognized; A filtering unit is used to filter the original speech signal to be recognized by using a deep residual shrinkage network in a preset deep residual shrinkage network model to obtain a target speech spectrum; the preset deep residual shrinkage network model is a model constructed by integrating the deep residual shrinkage network into a deep neural network; the target speech spectrum is used to indicate a speech spectrum that does not contain irrelevant features; the irrelevant features include at least noise features and environmental features; An extraction unit, used for extracting speech time sequence features from the target speech spectrum; A classification unit, used to classify the speech time series features through a preset classification layer of the deep residual shrinkage network to obtain the character probability corresponding to the target speech spectrum; the character probability is used to indicate the probability of occurrence of each character corresponding to the target speech spectrum; A prediction unit, used to predict the character probability by using a preset prediction model to obtain text information; The preset classification layer includes a fully connected layer and a logistic regression layer, and the fully connected layer includes a first fully connected layer and a second fully connected layer; The classification unit includes: An input module, used for inputting the speech time series features into the first fully connected layer and the second fully connected layer to obtain a speech output vector; A classification module is used to input the speech output vector into the classification layer for classification to obtain a first character probability and a second character probability corresponding to the target speech spectrum; the first character probability is used to indicate the probability of text information appearing in the audio; the second character probability is used to indicate the probability of text information appearing in a preset speech model.

6. The system according to claim 5, characterized in that The filtering unit comprises: A preprocessing module, used to preprocess the original speech signal to be recognized by using the spectrum function of the deep residual shrinkage network in the preset deep residual shrinkage network model to obtain the original speech spectrum; The removal module is used to remove the noise features contained in the original speech spectrum through a preset soft threshold function of the deep residual shrinkage network to obtain a target speech spectrum.

7. The system according to claim 5, characterized in that The extraction unit comprises: A first extraction module is used to extract speech timing features from the target speech spectrum through a recurrent neural network layer of the preset deep residual shrinkage network model; the recurrent neural network layer includes a unidirectional recurrent neural network layer or a bidirectional recurrent neural network layer; A second extraction module is used for extracting speech timing features from the target speech spectrum through the unidirectional recurrent neural network layer if the recurrent neural network layer is a unidirectional recurrent neural network layer; The third extraction module is used to extract speech timing features from the target speech spectrum through the bidirectional recurrent neural network layer if the recurrent neural network layer is a bidirectional recurrent neural network layer.

8. The system according to claim 5, characterized in that The prediction unit comprises: An acquisition module, used for acquiring a first character corresponding to the first character probability and a second character corresponding to the second character probability; A combining module, used for combining the first character and the second character to obtain a character string; The calculation module is used to calculate the character string through a preset function and a preset algorithm to obtain text information.

Citation Information

Patent Citations

  • A multimodal speech emotion recognition method based on enhanced residual neural network

    CN109460737A

  • Voice detection method and device

    CN111863036A