Speech Recognition Method, Apparatus, Device and Computer Readable Storage Medium

By classifying the voice signals in the telephone channel and detecting the front-end point and back-end point, and combining the acoustic model for text transfer, the problem of poor speech recognition effect in noisy environments is solved, and the accuracy of speech recognition and user interaction experience are improved.

CN114512128BActive Publication Date: 2025-07-04CHINA MERCHANTS BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210118048.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-08
Publication Date
2025-07-04
Estimated Expiration
2042-02-08

AI Technical Summary

Technical Problem

In noisy environments, the voice recognition effect of the telephone channel is poor, and the noise suppression and acoustic model optimization methods of the prior art are limited, which affects the accuracy and experience of the interaction between intelligent voice customer service and users.

Method used

The original speech signal is classified and processed through a pre-trained classification model, detect the front and back end points of the effective speech segment, and input them into the acoustic model for text transliteration, optimizing the classification results of speech frames.

Benefits of technology

It improves the accuracy of real-time voice recognition in the telephone channel and the recognition effect in the noisy environment, and optimizes the accuracy of voice recognition and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114512128B_ABST
    Figure CN114512128B_ABST
Patent Text Reader

Abstract

The present application discloses a voice recognition method, which includes the following steps: obtaining an original voice signal to be recognized; classifying an initial voice frame in the original voice signal based on a pre-trained classification model to obtain a voice frame classification result; detecting a front endpoint and a rear endpoint of a valid voice segment according to the voice frame classification result and in combination with a preset determination rule; and inputting the valid voice segment between the front endpoint and the rear endpoint into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result. The present invention also discloses a device, an apparatus and a computer-readable storage medium. In this embodiment, the initial voice frame is input into a pre-trained classification model to recognize the valid voice segment between the front endpoint and the rear endpoint, which improves the correctness of obtaining the target voice recognition result in real-time voice recognition in a telephone channel and optimizes the recognition effect of real-time voice in a noisy environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a voice recognition method, apparatus, device, and computer-readable storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, intelligent voice customer service is constantly replacing human customer service. In the work of intelligent voice customer service, the scenario of using automatic speech recognition technology to recognize the voice data of users and conduct voice interaction with users is becoming more and more common. The technical effect of speech recognition directly affects the accuracy of recognition and the interaction experience in the process of intelligent voice customer service interacting with users. Most of the existing intelligent voice customer services use telephone channel communication, and the audio sampling rate of 8KHZ is generally lower than 16KHZ or 44.1KHZ of channels such as mobile phones and PC terminals. Especially in a noisy environment, the voice quality of telephone channel communication is poor, and more information is lost in the telephone channel, resulting in poor speech recognition effect, which directly affects the accuracy of the interaction between intelligent voice customer service and users.

[0003] In order to improve the effect of speech recognition in a noisy environment, on the one hand, the existing technology is to remove part of the noise through noise suppression and other technologies in the signal preprocessing step to improve the voice quality in a noisy environment; on the other hand, it is to continuously optimize the acoustic model in the speech recognition step to improve the recognition accuracy.

[0004] In the existing technical solutions, the noise suppression method can reduce the noise signal in the speech frame to a certain extent and improve the accuracy of the voice activity detection module. However, the denoising algorithm will convert the original speech signal to a certain extent, affecting the recognition effect of the acoustic model; the method of optimizing the acoustic model has a high dependence on data, is more complex in operation, and has limited optimization effect on speech recognition in a noisy environment. The traditional energy-based voice activity detection method has poor effect in a noisy environment. Summary of the Invention

[0005] The main purpose of the present invention is to provide a voice recognition method, apparatus, and computer-readable storage medium, aiming to improve the effect of real-time voice recognition on the telephone channel in a noisy environment.

[0006] To achieve the above object, the present invention provides a voice recognition method, and the voice recognition method includes:

[0007] Obtain the original speech signal to be recognized;

[0008] Based on a pre-trained classification model, classify the initial speech frames in the original speech signal to obtain a speech frame classification result;

[0009] Based on the speech frame classification result and in combination with a preset determination rule, the front end point and the back end point of the valid speech segment are detected;

[0010] The valid speech segment between the front end point and the back end point is input into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

[0011] Optionally, after the step of obtaining the original speech signal to be recognized, the following steps are further included:

[0012] The byte stream of the original speech signal is intercepted and normalized to obtain the initial speech frame information in the form of a one-dimensional floating-point matrix of the original speech signal.

[0013] Optionally, the pre-trained classification model is a binary classification model, and the binary classification model includes: a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer; the step of classifying the initial speech frames in the original speech signal based on the pre-trained classification model to obtain a speech frame classification result includes:

[0014] The initial speech frame information in the form of a one-dimensional floating-point matrix is input into the pre-trained binary classification model;

[0015] The initial speech frames are subjected to feature extraction through the shallow feature extraction layer in the binary classification model to obtain primary features, where the shallow feature extraction layer includes a batch normalization (Batch Normalization) layer, a max pooling (MaxPooling) layer, and a one-dimensional convolutional layer;

[0016] The primary features are subjected to convolutional calculations of different scales through the stacked multi-scale one-dimensional convolutional residual layers to obtain the calculated advanced features, where the multi-scale one-dimensional convolutional residual layer contains m paths of one-dimensional convolutional residual layers of different scales, and each path of one-dimensional convolutional residual layer contains n one-dimensional convolutional residual blocks with the same convolutional kernel size and an average pooling (AvgPooling) layer, where m and n are positive integers;

[0017] The advanced features are linked to the integration layer of the binary classification model for integration to obtain a corresponding integration result;

[0018] The integration result is input into the fully connected layer and the softmax layer of the output layer to obtain a probability matrix of the initial speech frames belonging to different categories;

[0019] The obtained probability matrix is matched with a preset threshold to obtain the speech frame classification result of the initial speech frames.

[0020] Optionally, the step of detecting the start point and end point of the valid speech segment according to the speech frame classification result and in combination with a preset determination rule includes:

[0021] Based on the classification result of the initial speech frames in the original speech signal, obtain a classification result sequence of the original speech signal, where the classification result includes active speech frames and inactive speech frames;

[0022] According to the classification result sequence, in combination with a preset start point and end point determination rule, obtain the start point and end point of the valid speech segment of the original speech signal.

[0023] Optionally, the step of obtaining the start point and end point of the valid speech segment of the original speech signal according to the classification result sequence and in combination with a preset start point and end point determination rule includes:

[0024] Obtain a first classification result subsequence from the classification result sequence, where the first classification result subsequence includes: the classification results of N consecutive speech frames;

[0025] If the number of active speech frames in the first classification result subsequence reaches a first threshold, determine that the speech frame corresponding to the very front end of the first classification result subsequence is the start point of the valid speech segment;

[0026] Based on the start point, obtain a second classification result subsequence from the classification result sequence, where the second classification result subsequence includes: the classification results of M consecutive speech frames after the start point;

[0027] If the number of inactive speech frames in the second classification result subsequence reaches a second threshold, determine that the speech frame corresponding to the very end of the second classification result subsequence is the end point of the valid speech segment.

[0028] Optionally, the step of inputting the valid speech segment between the start point and the end point into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result includes:

[0029] Input the valid speech segment between the start point and the end point into a pre-trained acoustic model;

[0030] Decode the valid speech segment through the acoustic model to implement text transcription and obtain the target transcription result corresponding to the valid speech segment.

[0031] Optionally, before the step of obtaining the original speech signal to be recognized, the speech recognition method further includes:

[0032] Obtain a sample audio of training data, and extract sample features of the sample audio, where the sample audio has a corresponding classification result;

[0033] Based on the sample features, establish a sample feature data set;

[0034] Construct a multi-scale one-dimensional residual neural network, and perform deep learning on the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model.

[0035] Optionally, after the step of constructing a multi-scale one-dimensional residual neural network and performing deep learning on the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model, the speech recognition method further includes:

[0036] Test the initial binary classification model with telephone channel speech data to verify the classification effect of the initial binary classification model;

[0037] If the classification effect does not meet the preset standard, it is necessary to return to strengthen the data of the sample audio to obtain a sample audio of optimized training data;

[0038] Fine-tune the initial binary classification model according to the sample audio of the optimized training data to obtain a trained binary classification model;

[0039] If the classification effect meets the preset standard, use the initial binary classification model as the trained binary classification model.

[0040] In addition, to achieve the above object, the present invention also provides a speech recognition device, where the speech recognition device includes:

[0041] An acquisition module, configured to acquire an original speech signal to be recognized;

[0042] A determination module, configured to perform classification processing on an initial speech frame in the original speech signal based on a pre-trained classification model to obtain a speech frame classification result;

[0043] A detection module, configured to detect a front endpoint and a back endpoint of a valid speech segment according to the speech frame classification result and in combination with a preset determination rule;

[0044] A determination module, configured to input the valid speech segment between the front endpoint and the back endpoint into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

[0045] In addition, to achieve the above object, the present invention further provides an intelligent device, which includes a memory, a processor, and a speech recognition program stored on the memory and executable on the processor. When the speech recognition program is executed by the processor, the steps of the speech recognition method described above are implemented.

[0046] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a speech recognition program is stored. When the speech recognition program is executed by a processor, the steps of the speech recognition method described above are implemented.

[0047] The speech recognition method, device, equipment, and computer-readable storage medium proposed in the embodiments of the present invention obtain an original speech signal to be recognized; classify the initial speech frames in the original speech signal based on a pre-trained classification model to obtain a speech frame classification result; detect the front endpoint and the back endpoint of the valid speech segment according to the speech frame classification result and in combination with a preset determination rule; and input the valid speech segment between the front endpoint and the back endpoint into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

[0048] After the present invention obtains the original speech signal to be recognized, it classifies the speech signal. The steps of the classification process include judging the activity of each frame of the speech in the original speech signal through a pre-trained classification model, that is, obtaining the classification result of each speech frame in the original speech signal. After obtaining the corresponding classification results, these classification results are combined to obtain a classification result sequence corresponding to the original speech signal. The front endpoint and the back endpoint of the valid speech segment are judged according to the classification result sequence, and the valid speech segment between the front endpoint and the back endpoint is obtained. The valid speech segment is input into the acoustic model for speech recognition, and after obtaining the recognition result, text transcription is performed to obtain the target transcription result. This solution obtains the classification result of each speech frame in the original speech signal through the classification model, determines the front endpoint and the back endpoint to obtain the valid speech segment, and performs speech recognition of the acoustic model on the valid speech segment, improving the correctness of obtaining the target speech recognition result in real-time speech recognition in a telephone channel and optimizing the recognition effect of real-time speech in a noisy environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 Schematic diagram of the hardware structure of an implementation manner provided by an embodiment of the present invention;

[0052] Figure 2 Schematic diagram of the process of the first embodiment of the voice recognition method of the present application;

[0053] Figure 3 Schematic diagram of the refined process steps of the first embodiment of the voice recognition method of the present application;

[0054] Figure 4 Schematic diagram of the process of the second embodiment of the voice recognition method of the present application;

[0055] Figure 5 Schematic diagram of the process of the third embodiment of the voice recognition method of the present application;

[0056] Figure 6 Schematic diagram of the structure of the binary classification model in the third embodiment of the voice recognition method of the present application;

[0057] Figure 7 Schematic diagram of the structure of the one-dimensional residual convolution block in the binary classification model in the third embodiment of the voice recognition method of the present application;

[0058] Figure 8 Schematic diagram of the process of the fourth embodiment of the voice recognition method of the present application;

[0059] Figure 9 Schematic diagram of the refined process of detecting the front end point and the back end point of the effective speech segment according to the voice frame classification result and in combination with a preset determination rule in the fourth embodiment of the voice recognition method of the present application;

[0060] Figure 10 Schematic diagram of the process of the fifth embodiment of the voice recognition method of the present invention;

[0061] Figure 11 Schematic diagram of the device module of the embodiment of the voice recognition method of the present invention. Specific implementation manner

[0062] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0063] As Figure 1 shown, Figure 1It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention.

[0064] The device in the embodiment of the present invention can be a mobile terminal or a server device.

[0065] As shown in Figure 1 , the device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0066] Those skilled in the art can understand that Figure 1 the device structure shown in

[0067] does not constitute a limitation on the device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 1 As shown in

[0068] , the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a speech recognition program.

[0069] Among them, the operating system is a program for managing and controlling the real-time speech recognition device and software resources, and supports the operation of the network communication module, the user interface module, the real-time speech recognition program, and other programs or software; the network communication module is used to manage and control the network interface 1002; the user interface module is used to manage and control the user interface 1003. Figure 1 In the real-time speech recognition device shown in

[0070] , the real-time speech recognition device calls the real-time speech recognition program stored in the memory 1005 through the processor 1001 and executes the operations in the respective embodiments of the following speech recognition method.

[0071] Based on the above hardware structure, an embodiment of the speech recognition method of the present invention is proposed. Figure 2 , Figure 2This is a schematic flowchart of the first embodiment of the speech recognition method of the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here. The steps of the speech recognition method in this embodiment include:

[0072] Step S10: Obtain the original speech signal to be recognized;

[0073] Step S20: Based on a pre-trained classification model, classify the initial speech frames in the original speech signal to obtain a speech frame classification result;

[0074] Step S30: According to the speech frame classification result and in combination with a preset determination rule, detect the front end point and the back end point of the valid speech segment;

[0075] Step S40: Input the valid speech segment between the front end point and the back end point into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

[0076] In this embodiment, to solve the problem of poor real-time speech recognition effect in a telephone channel under a noisy environment, the present embodiment proposes to intercept the original speech signal to obtain the initial speech frames in the original speech signal, input the initial speech frames into a pre-trained classification model to obtain the corresponding speech frame classification result, combine the speech frame classification result with a preset determination rule to obtain the front end point and the back end point of the valid speech segment, and input the valid speech segment between the front end point and the back end point into the acoustic model for text transcription to obtain the target transcription result, improving the correctness of obtaining the target speech recognition result in real-time speech recognition in a telephone channel and optimizing the recognition effect of real-time speech in a noisy environment.

[0077] The following is a detailed description of each step:

[0078] Step S10: Obtain the original speech signal to be recognized;

[0079] In this embodiment, the speech signal to be recognized refers to the real-time speech audio data obtained from the user side during a real-time voice call through a telephone channel. The real-time speech audio data is converted to obtain the original speech signal to be recognized, and the original speech signal to be recognized in the form of a speech signal byte stream is obtained.

[0080] The original speech signal can be processed to obtain the initial speech frames therein.

[0081] Step S20: Based on a pre-trained classification model, classify the initial speech frames in the original speech signal to obtain a speech frame classification result;

[0082] In this embodiment, the initial speech frames in the original speech signal are classified based on a pre-trained classification model, and the corresponding speech frame classification results are obtained after the processing.

[0083] Among them, the pre-trained classification model is a binary classification model. The initial speech frames corresponding to the original speech signal in the form of a speech signal byte stream are classified using this binary classification model to obtain the corresponding speech frame classification results. The above speech frame classification results include active speech frames and inactive speech frames.

[0084] Specifically, the ways to obtain the classification results may specifically include:

[0085] Through each different feature extraction layer of the classification model, the features of each speech frame of the original speech signal are obtained. These continuously extracted speech frame features are connected to the integration layer. After the features of the speech frames are integrated in the integration layer, the integrated features are output from the classification model. The output result is the speech frame classification result corresponding to each speech frame of the original speech signal. This speech frame classification result includes active speech frames and inactive speech frames.

[0086] Step S30, according to the speech frame classification results and in combination with a preset determination rule, detect the front endpoint and the back endpoint of the valid speech segment;

[0087] In this embodiment, to determine the valid speech segment in the original speech signal, on the one hand, it is necessary to determine the activity of each speech frame in the valid speech segment to obtain the speech frame classification result corresponding to each speech frame. On the other hand, it is necessary to detect the front endpoint and the back endpoint of the valid speech segment, obtain the valid speech segment between the front endpoint and the back endpoint, and perform speech recognition on the valid speech segment between the front endpoint and the back endpoint.

[0088] This embodiment takes into account that in the scenario where a user on a telephone channel has a conversation with an intelligent voice customer service, during the process of the user having a conversation with the intelligent voice customer service, situations such as the user pausing for a period of time, occasional loud noise and human voices in the background environment often occur. Therefore, it is one-sided to only use the classification results obtained from the speech frame classification model as the sole criterion for determining the valid speech segment. Using this criterion to determine the validity of the speech segment cannot optimize the real-time speech recognition effect. Therefore, to determine the validity of the speech segment, it is also necessary to pre-establish corresponding determination rules according to the usage habits of different users, detect the front endpoint and the back endpoint of the valid speech segment, and then obtain the valid speech segment between the front endpoint and the back endpoint.

[0089] Step S40, input the valid speech segment between the front endpoint and the back endpoint into a pre-trained acoustic model for text transcription to obtain the corresponding target transcription result;

[0090] Using the obtained start point and end point of the valid speech segment, obtain the valid speech segment between the start point and the end point, input the valid speech segment into a pre-trained acoustic model, decode the valid speech segment, and implement text transcription to obtain the corresponding target transcription result.

[0091] In the speech recognition method proposed in this solution, by combining the classification result of the activity determination of valid speech frames and the start point and end point of the valid speech segment, the classification result of real-time speech is optimized. By obtaining the valid speech segment for text transcription, the recognition result is made more accurate, the real-time speech recognition effect is also optimized, and at the same time, the correctness of obtaining the target speech recognition result in real-time speech recognition in a telephone channel is improved, and the recognition effect of real-time speech in a noisy environment is optimized.

[0092] In this embodiment, to solve the problem of poor real-time speech recognition effect in a telephone channel in a noisy environment, the speech byte stream of the initial speech frame corresponding to the original speech signal is input into a classification model for activity determination, so that each speech frame of the real-time input original speech signal has a corresponding classification result, and the classification result is combined for recognition during the speech recognition process to improve the accuracy of real-time speech recognition and optimize the real-time speech recognition effect.

[0093] Furthermore, according to the speech frame classification result and a preset determination rule, the start point and end point of the valid speech segment are obtained. The above-mentioned preset start point and end point determination rule can be formulated by combining different usage habits of users, making the endpoint determination rule of the valid speech segment for speech recognition more flexible. Based on the above classification result and the preset determination rule of the start point and end point, the valid speech segment between the start point and the end point is obtained, and the valid speech segment between the start point and the end point is input into the acoustic model for text transcription to obtain the target transcription result.

[0094] In the speech recognition method proposed in this solution, by combining the classification result of the activity determination of valid speech frames and the start point and end point of the valid speech segment, the classification result of real-time speech is optimized, the valid speech segment for text transcription is obtained, the recognition result is made more accurate, the real-time speech recognition effect is also optimized, and at the same time, the correctness of obtaining the target speech recognition result in real-time speech recognition in a telephone channel is improved, and the recognition effect of real-time speech in a noisy environment is optimized.

[0095] The specific detailed process steps of this embodiment can be referred to Figure 3 as shown.

[0096] Furthermore, based on the first embodiment, a second embodiment of the speech recognition method of the present invention is proposed. The process of this embodiment is referred to Figure 4The difference between the second embodiment of the speech recognition method and the first embodiment of the speech recognition method is that after step S10, the following steps are further included:

[0097] Step S101: Intercept and normalize the byte stream of the original voice signal to obtain the initial voice frame information in the form of a one-dimensional floating-point matrix of the original voice signal.

[0098] Obtain the audio data of real-time voice from the telephone channel. After processing this audio data, obtain the corresponding original voice signal, cache the voice signal byte stream corresponding to this original voice signal, intercept the voice signal byte stream, and the intercepted length is a fixed value. This fixed length can determine a threshold according to the actual situation during actual use. When the number of bytes in the cache reaches the threshold, process the original voice signal byte stream in the cache, process the byte stream according to a fixed size. The voice signal bytes of the preset threshold length are an initial voice frame. Normalize the byte stream of these initial voice frames to generate a one-dimensional floating-point matrix corresponding to the initial voice frame, and obtain an initial voice frame in the form of a one-dimensional floating-point matrix.

[0099] In this embodiment, by performing byte stream conversion on the original voice signal, the voice signal byte stream corresponding to the original voice signal is obtained, and the voice signal byte stream is intercepted with a fixed length to obtain a voice signal byte stream of a fixed length. The above fixed length is set according to a preset threshold, and the preset threshold can be dynamically set.

[0100] The speech recognition method proposed in this solution makes the form of the voice signal byte stream for speech recognition more standardized, reduces the recognition difficulty of real-time voice signals during the recognition process, and optimizes the recognition effect of real-time voice to a certain extent.

[0101] Further, based on the first embodiment and the second embodiment, a third embodiment of the speech recognition method of the present invention is proposed. The process of this embodiment refers to Figure 5 。

[0102] The difference between the third embodiment of this speech recognition method and other embodiments of the speech recognition method is that in this embodiment, step S20 above is refined, that is, based on a pre-trained classification model, the initial voice frames in the original voice signal are classified to obtain a voice frame classification result.

[0103] Specifically, in this embodiment, the pre-trained classification model is a binary classification model, and the binary classification model includes: a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer; the step of classifying the initial voice frames in the original voice signal based on the pre-trained classification model to obtain a voice frame classification result includes:

[0104] Step S201: Input the initial speech frame information in the form of a one-dimensional floating-point matrix into a pre-trained binary classification model.

[0105] In one embodiment, the initial speech frame information in the form of a one-dimensional floating-point matrix corresponding to the acquired original speech signal is input into a pre-trained classification model. The pre-trained classification model includes the above binary classification model. For example, the VAD classification model. The above VAD classification model includes a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer. The above initial speech frame information is input into the shallow feature extraction layer in the VAD classification model, so that the trained VAD classification model receives the initial speech frame information in the form of a one-dimensional floating-point matrix corresponding to the original speech signal.

[0106] Step S202: Extract features from the initial speech frame through the shallow feature extraction layer in the binary classification model to obtain primary features. The shallow extraction layer includes a Batch Normalization layer, a MaxPooling layer, and a one-dimensional convolutional layer.

[0107] Specifically, after making feature selection, the MaxPooling layer selects features with higher classification recognition, providing non-linearity. According to relevant theories, the error in feature extraction mainly comes from two aspects: on the one hand, the increase in the variance of the estimated value caused by the limited neighborhood size; on the other hand, the shift of the estimated mean caused by the parameter error of the convolutional layer. Generally speaking, the MaxPooling layer can reduce the second error and retain more texture information.

[0108] In one embodiment, the above initial speech frame information is input into the shallow extraction layer in the VAD classification model, and the primary features of the speech frame information are respectively extracted through the one-dimensional Batch Normalization layer and the one-dimensional MaxPooling layer of two one-dimensional convolutional layers in the shallow extraction layer.

[0109] Step S203: Perform convolutional calculations of different scales on the primary features through the stacked multi-scale one-dimensional convolutional residual layer. The calculated high-level features are obtained. Among them, the multi-scale one-dimensional convolutional residual layer includes m paths of one-dimensional convolutional residual layers of different scales. Each path of one-dimensional convolutional residual layer includes n one-dimensional convolutional residual blocks with the same convolutional kernel size and an AvgPooling layer, where m and n are positive integers.

[0110] Among them, the stacked multi-scale one-dimensional convolutional residual layers perform convolutional calculations on the primary features at different scales. Each one-dimensional convolutional residual block uses a skip connection to add the input and the network output, alleviating the problem of vanishing gradients caused by increasing the depth in deep neural networks.

[0111] In addition, the AvgPooling layer is applied to global average pooling operations, which is often placed behind the convolutional layer to sample a feature and speed up the operation of the classification model. If the object for feature extraction in a feature extraction layer tends to be the overall characteristic, the AvgPooling layer has the function of preventing too much high-dimensional information from being lost.

[0112] In one embodiment, the primary features obtained from the shallow extraction layer are input into a multi-path different-scale convolutional residual network block in the VAD classification model, that is, the multi-scale one-dimensional convolutional residual layer. The multi-scale one-dimensional convolutional residual layer contains m paths of different-scale one-dimensional convolutional residual layers, and each path of one-dimensional convolutional residual layer contains n one-dimensional convolutional residual blocks with the same convolutional kernel size and an AvgPooling layer, where m and n are positive integers.

[0113] In the VAD classification model, the multi-scale one-dimensional convolutional residual layer extracts features through m paths of different-scale one-dimensional convolutional residual layers containing n one-dimensional convolutional residual blocks with different convolutional kernel sizes and an average pooling AvgPooling layer. After the feature extraction is completed, the one-dimensional AvgPooling layers after the one-dimensional convolutional residual blocks with different scales and the same convolutional kernel size calculate these extracted features to obtain the high-level features calculated by the multi-scale one-dimensional convolutional residual layer.

[0114] Step S204: Link the high-level features to the integration layer of the binary classification model for integration to obtain the corresponding integration result;

[0115] Step S205: Input the integration result into the fully connected layer and the softmax layer of the output layer to obtain the probability matrix of the initial speech frame belonging to different classes;

[0116] Step S206: Match the obtained probability matrix with a pre-set threshold to obtain the speech frame classification result of the initial speech frame.

[0117] Among them, the integration layer organically concentrates data features of different types, formats, and characteristic properties logically or physically, enabling better communication, sharing, and fusion of data between systems. The feature data after the organic concentration is used as the integration result obtained after the integration layer.

[0118] In addition, the fully connected layer is connected to all nodes in the previous layer at each node and is used to synthesize the features extracted previously. Due to its fully connected nature, the parameters of the fully connected layer are generally the most; the softmax layer of the normalization exponential function is an algorithm for solving multi-class regression problems and is a classifier widely used in the supervised learning part of deep networks in current deep learning research. It can "compress" a K-dimensional vector z containing arbitrary real numbers into another K-dimensional real vector σ(z) such that the range of each element is between (0, 1) and the sum of all elements is 1. This function is more than a probability matrix in multi-classification problems. It is a matrix used to describe the transition of a Markov chain. Each of its terms is a non-negative real number representing a probability. The probability matrix can be used to represent probabilities, and the result of matrix multiplication can be used to predict the probability of future events occurring. The obtained probability matrix can represent the probabilities of different classes of the initial speech frame.

[0119] In one embodiment, the obtained high-level features are connected to the integration layer for integration, and a probability matrix of the initial speech frame corresponding to the original speech signal is obtained through a fully connected layer and a softmax layer. The probability matrix is matched with a set threshold to obtain the classification result corresponding to the initial speech frame information of the input one-dimensional floating-point matrix.

[0120] The above initial speech frame classification results include active speech frames and inactive speech frames. Among them, an active speech frame is the current speech frame containing valid human voice speech; an inactive speech frame includes speech without valid human voice such as a silent frame and a noise frame.

[0121] As Figure 6 shown, it is a schematic structural diagram of a VAD classification model, which includes a shallow feature extraction layer, multiple one-dimensional convolutional residual layers with different scales, an integration layer, and an output layer.

[0122] Among them, the multiple one-dimensional convolutional residual network blocks with different scales include three paths:

[0123] The first path includes three cascaded one-dimensional convolutional residual blocks, where the convolutional kernel size is 1*3, and one one-dimensional AvgPooling layer;

[0124] The second path includes three cascaded one-dimensional convolutional residual blocks, where the convolutional kernel size is 1*5, and one one-dimensional AvgPooling layer;

[0125] The third path includes three cascaded one-dimensional convolutional residual blocks, where the convolutional kernel size is 1*7, and one one-dimensional AvgPooling layer.

[0126] Specifically, the specific structural schematic diagram of the one-dimensional residual convolutional block refers to Figure 7, each one-dimensional residual convolution block is composed of a cascade of a one-dimensional convolutional layer, a one-dimensional Batch Normalization layer, a ReLU layer, a one-dimensional convolutional layer, and a one-dimensional Batch Normalization layer. Then, the output of the second Batch Normalization layer and the input of the one-dimensional residual convolution block are added together and used as the input of the last ReLU layer. This structure is called a residual network, which can alleviate the gradient problem of the entire neural network to a certain extent during training and can learn more local information.

[0127] The above-mentioned primary features are secondarily extracted from the above-mentioned multi-path different-scale convolutional residual network blocks, and the features output from the multi-path different-scale convolutional residual network blocks are obtained. Then, the output features are transmitted to the next integration layer of the VAD classification model for integration of different features. After calculation through the fully connected layer and the Softmax layer, the speech frame classification result corresponding to the original speech signal is obtained.

[0128] In one embodiment, it is proposed to input the speech byte stream of the initial speech frame corresponding to the original speech signal into the classification model for activity determination. The method for determining activity is to classify each speech frame of the original speech signal through the VAD classification model, so that each speech frame of the real-time input original speech signal has a corresponding classification result, making the result obtained in the further recognition step more accurate. Moreover, through the VAD classification model including the shallow feature extraction layer, the multi-scale one-dimensional convolutional residual layer, the integration layer, and the output layer, the correctness of the classification result of each speech frame of the original speech signal is improved. During the speech recognition process, the recognition is combined with the classification result to improve the hit rate of the real-time speech recognition for active speech and optimize the recognition effect of the real-time speech.

[0129] Furthermore, based on the above-mentioned embodiments, a fourth embodiment of the speech recognition method of the present invention is proposed. The process of this embodiment is referred to Figure 8 as shown. In this embodiment, according to the speech frame classification result and in combination with the preset determination rule, the detailed process of detecting the front end point and the back end point of the effective speech segment can be referred to Figure 9 as shown.

[0130] The difference between this embodiment and each of the above-mentioned embodiments is that this embodiment refines the above-mentioned step S30 of detecting the front end point and the back end point of the effective speech segment according to the speech frame classification result and in combination with the preset determination rule, specifically including:

[0131] Step S301, based on the classification result of the initial speech frame in the original speech signal, obtain the classification result sequence of the original speech signal, where the classification result includes active speech frames and inactive speech frames.

[0132] Through the above-mentioned binary classification model, the activity of the initial speech frame of the original speech signal is judged. In this embodiment, the VAD classification model is used to judge each speech frame in the original speech signal to obtain the classification result corresponding to each speech frame. After obtaining the classification result of each speech frame, the classification result of the speech frame segment containing one or more speech frames is obtained, and the classification result of the speech frame segment is used as the classification result sequence corresponding to the original speech signal. The speech frame classification sequence is a sequence containing 0 and 1.

[0133] Step S302, according to the classification result sequence, combined with preset front endpoint and back endpoint determination rules, obtain the front endpoint and back endpoint of the valid speech segment of the original speech signal.

[0134] In the scenario where a telephone channel user is conversing with an intelligent voice customer service, there are often pauses in the conversation between the user and the intelligent voice customer service, and occasional loud noises and human voices in the background environment. Therefore, it is necessary to pre-formulate corresponding judgment rules based on the user's usage habits to detect the front and back endpoints of the valid voice segment of the original voice signal.

[0135] Further, in one embodiment, in step S302, obtaining the front endpoint and the rear endpoint of the valid speech segment of the original speech signal according to the classification result sequence in combination with preset front endpoint and rear endpoint determination rules includes:

[0136] Step a1: acquiring a first classification result subsequence from the classification result sequence, wherein the first classification result subsequence includes: classification results of N consecutive speech frames.

[0137] During a real-time voice call on a telephone channel, the original voice signal of the real-time voice passes through the classification module and is output as a voice frame classification result sequence represented by 0 and 1. The leading and trailing endpoints of the valid voice segment are determined, and the classification results of the current voice frame of the real-time voice call are cached. When the number of current voice frames of the real-time voice call is N, the N classification results are combined into a result sequence, and a first classification result subsequence is obtained. The first classification result subsequence contains the classification results of N consecutive voice frames.

[0138] Step a2: if the number of active speech frames in the first classification result subsequence reaches a first threshold, the speech frame corresponding to the front end of the first classification result subsequence is determined to be the front end point of the valid speech segment.

[0139] When the number of active speech frames in the above first classification result subsequence reaches the first preset threshold, it is determined that among the valid speech segments from the previous N frames to the current frame of the current frame, the previous N frames of the current frame are the front endpoints of the valid speech segment. For example, when it is necessary to determine the classification result sequence of three frames, assuming the current frame is the nth frame, the corresponding first classification result subsequence is: [the (n-2)th frame, the (n-1)th frame, the nth frame]. When the number of active speech frames in the sequence is 2, that is, when the number of active speech frames is equal to 2, it is determined that the (n-2)th frame is the front endpoint of the valid speech segment.

[0140] Step a3, based on the front endpoint, obtain a second classification result subsequence from the classification result sequence, where the second classification result subsequence includes: the classification results of consecutive M speech frames after the front endpoint.

[0141] After the front endpoint of the valid speech segment has been recognized, cache the current call speech frames of the real-time speech with a quantity of M after the front endpoint, form a result sequence with the M classification results, and obtain the second classification result subsequence, which contains the classification results of consecutive M speech frames.

[0142] Step a4, if the number of non-active speech frames in the second classification result subsequence reaches the second threshold, it is determined that the speech frame corresponding to the last end of the second classification result subsequence is the back endpoint of the valid speech segment.

[0143] When the number of non-active speech frames in the above second classification result subsequence reaches the second preset threshold, it is determined that among the M frames after the front endpoint speech frame to the front endpoint speech frame, the M frames after the front endpoint speech frame of the current frame are the back endpoints of the valid speech segment. For example, when the front endpoint of the speech stream is recognized, starting from the 6th frame after the front endpoint, cache the previous 5 frames of the current frame, and judge the classification result sequence of 6 frames including the current frame. When the number of 0s in the sequence is 5, that is, when the number of non-active speech frames is equal to 5, the current frame is the back endpoint of the valid speech segment.

[0144] In this embodiment, to solve the problem of poor real-time speech recognition effect of the telephone channel in a noisy environment, according to the speech frame classification results and combined with the preset determination rules, the front and back endpoints of the valid speech segment are obtained. The above preset front and back endpoint determination rules are formulated by combining the user's usage habits to make the speech recognition process more flexible.

[0145] Further, the valid speech segment between the front endpoint and the rear endpoint is passed into an acoustic model for text transcription to obtain a target transcription result. In this solution, by obtaining the valid speech segment between the front endpoint and the rear endpoint, the valid speech segment for text transcription is obtained, making the recognition result more accurate and optimizing the real-time speech recognition effect. In this embodiment, the speech recognition method proposed in this solution optimizes the classification result of real-time speech, improves the correctness of obtaining the target speech recognition result in real-time speech recognition in a telephone channel, and optimizes the recognition effect of real-time speech in a noisy environment.

[0146] Further, based on the above embodiments, a fifth embodiment of the speech recognition method of the present invention is proposed. The flowchart is referred to Figure 10 The difference between the fifth embodiment of the speech recognition method and the above embodiments is that before step S10, a classification model is further established, which specifically includes:

[0147] Step S50, obtaining a sample audio of training data, and extracting sample features of the sample audio, where the sample audio has a corresponding classification result;

[0148] Step S60, based on the sample features, establishing a sample feature data set;

[0149] Step S70, constructing a multi-scale one-dimensional residual neural network, and performing deep learning on the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model;

[0150] Step S80, testing the initial binary classification model with telephone channel voice data to verify the classification effect of the initial binary classification model;

[0151] Step S90, if the classification effect does not meet the preset standard, then return to perform data enhancement on the sample audio of the training data to obtain a sample audio of optimized training data;

[0152] Step S100, performing fine-tuning training on the initial binary classification model according to the sample audio of the optimized training data to obtain a trained binary classification model;

[0153] Step S110, if the classification effect meets the preset standard, then use the initial binary classification model as the trained binary classification model.

[0154] Before using the binary classification model to perform classification prediction on the input speech, it is necessary to create a binary classification model capable of activity determination and train the binary classification model. The creation and training process of the binary classification model includes steps such as training data preparation, model training, model testing, training corpus adjustment, and model fine-tuning. After the above steps, a trained binary classification model is obtained.

[0155] In one embodiment, before classifying speech frames in an original speech signal using a binary classification model, it is necessary to collect training data to train the classification results of the binary classification model to obtain the binary classification model.

[0156] The following elaborates on each step in detail:

[0157] Step S50: Obtain a sample audio of the training data, and extract the sample features of the sample audio, where the sample audio has a corresponding classification result;

[0158] Step S60: Based on the sample features, establish a sample feature data set;

[0159] Process the training data sample audio that needs to be trained for the binary classification model into a one-dimensional matrix, and process the one-dimensional matrix corresponding to the training data sample audio to extract the sample features of the sample audio of the training data.

[0160] Specifically, the sample audio of the training data includes four parts: human voice without noise, human voice with noise, noise without human voice, and noise with human voice. The training data sample audio is divided into human voice and noise, and noise and human voice have their corresponding features. Among them, the human voice without noise and the human voice with noise form positive samples, and the noise without human voice and the noise with human voice form negative samples. For the part of the human voice without noise, different noises are additionally superimposed according to a specific ratio. To ensure the balance of the training data sample audio, during the training of the classification model, the number of positive and negative sample audios in a batch should be kept equivalent. Collect these human voices and noises with corresponding sample features as the training data sample audio, and form a sample feature data set with the above-mentioned training data audio with sample features.

[0161] Step S70: Construct a multi-scale one-dimensional residual neural network, and train the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model;

[0162] In this embodiment, the above binary classification model uses a deep learning method. By constructing a multi-scale one-dimensional residual neural network and based on the above-mentioned collected sample feature data set, the multi-scale one-dimensional residual neural network is trained.

[0163] Step S80: Test the initial binary classification model with telephone channel speech data to verify the classification effect of the initial binary classification model;

[0164] Place the trained binary classification model in a real telephone channel to receive real-time voice data, and conduct classification tests on the real-time voice signals corresponding to the real-time voice data. Verify the classification effect of the initial binary classification model according to the classification results obtained by classifying the original voice signals of the real-time voice. If the classification effect of the binary classification model is not good, it is necessary to add real human voice segments and noise segments in the telephone channel as enhanced training data to adjust the training corpus of the binary classification model. By performing fine-tuning training on the binary classification model, the trained binary classification model is obtained.

[0165] Step S90, if the classification effect does not meet the preset standard, return to perform data enhancement on the sample audio to obtain the sample audio of the optimized training data.

[0166] Step S100, perform fine-tuning training on the initial binary classification model according to the sample audio of the optimized training data to obtain the trained binary classification model;

[0167] Step S110, if the classification effect meets the preset standard, use the initial binary classification model as the trained binary classification model.

[0168] In this embodiment, a classification test is performed on the above binary classification model. The trained classification model is used to perform real-time voice data classification tests in a real telephone channel. The activity of the voice frames in the original voice signal of the real-time voice is classified, and the classification results are used to verify the effect of the initial binary classification model in distinguishing each frame of the voice frame in the original voice signal as an active voice frame or a non-active voice frame in the telephone channel.

[0169] If the classification effect of the binary classification model does not meet the preset standard, it is necessary to add real human voice segments and noise segments in the telephone channel as enhanced training data to adjust the training corpus of the binary classification model. The adjustment methods include: performing binary classification annotation on the obtained original audio signal, and performing fine-tuning training on the adjusted training corpus through the binary classification model to obtain the fine-tuned trained binary classification model.

[0170] In this embodiment, by constructing a binary classification model for classifying the activity of the voice frames of the original voice signal, the classification results of each frame of the voice frames in the original voice signal for speech recognition are obtained. Classifying the activity of each frame of the voice frames in the original voice signal can improve the accuracy of speech recognition and the recognition effect of real-time voice in a high-noise environment in the telephone channel.

[0171] Before classifying and predicting the input speech using a classification model, it is necessary to create a binary classification model capable of activity determination. In this embodiment, the above binary classification model is a binary classification model trained with positive and negative sample audio data. By performing deep learning on the sample feature data set through the above classification model at multiple levels, the recognition effect of real-time speech in a telephone channel and a high-noise environment is improved. When training this model, including the creation and training process of the classification model, where the training process includes data preparation, model training, model testing, training corpus adjustment, and model fine-tuning, the binary classification model after passing the classification result test has a higher recognition accuracy for the speech frames in the original speech recognition signal, and also improves the recognition effect of real-time speech in a high-noise environment.

[0172] The present invention also provides a speech recognition device, which includes:

[0173] An acquisition module for acquiring the original speech signal to be recognized;

[0174] A determination module for classifying the initial speech frames in the original speech signal based on a pre-trained classification model to obtain a speech frame classification result;

[0175] A detection module for detecting the start point and end point of the valid speech segment according to the speech frame classification result and in combination with a preset determination rule;

[0176] A determination module for inputting the valid speech segment between the start point and the end point into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

[0177] The specific implementation manner of this embodiment is basically the same as that of the above embodiments of the speech recognition method, and will not be elaborated here.

[0178] In addition, to achieve the above object, the present invention also provides a speech recognition device, which includes:

[0179] A speech stream truncation module for performing fixed-length truncation processing on the input telephone speech stream;

[0180] Cache the byte stream of the input original speech signal, and a threshold can be determined according to the actual situation in the actual use process. When the number of bytes in the cache reaches the threshold, process the byte stream in the cache, intercept the byte stream according to a fixed size, that is, intercept a speech frame, and perform normalization processing on each frame of the byte stream to generate a corresponding one-dimensional floating-point matrix as the input of the subsequent module.

[0181] A speech frame classification module for classifying the input fixed-length speech frames to determine the category of the speech frame;

[0182] The voice frame classification module adopts a deep learning method and implements a binary classification model by constructing a multi-scale one-dimensional residual neural network. The input is a fixed-size voice byte stream, that is, a voice frame. Here, the size of the model input is 2048, and the output is the classification result of the voice frame, that is, an active voice frame or an inactive voice frame. An active voice frame means that the current frame contains valid human voice, and an inactive voice frame includes voice frames without valid human voice such as silence frames and noise frames.

[0183] Specifically, the above-mentioned multi-scale one-dimensional residual neural network includes a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer.

[0184] The front and end point determination module is used to determine the front end point and the end point of the valid speech segment according to the classification result of each frame of the continuous speech stream.

[0185] According to the classification result of each frame of speech, the front end point and the end point of the valid speech in the continuous real-time speech stream in the current telephone channel are detected, the non-valid speech segments are filtered, and only the valid speech segments are passed to the subsequent acoustic model for recognition, so as to improve the speech recognition effect.

[0186] In the voice frame classification module, the size of each frame of voice byte stream is 2048, the voice sampling rate of the telephone channel is 8KHZ, and the number of bits is 16 bits. It is calculated that each frame of voice stream is equivalent to 128 ms of voice. In the scenario where a user and an intelligent voice customer service have a conversation on a telephone channel, the user often has short pauses during the speaking process, and there are occasional noises and human voices with relatively large volumes in the background environment. Therefore, the classification result of each frame of speech by the voice frame classification module cannot be used as the basis for judging whether the frame belongs to a valid speech segment. It is necessary to formulate corresponding rules according to the user's usage habits to determine the front and end points of the valid speech segment. The output of the voice frame classification module is a sequence containing 0 and 1. When performing voice activity detection, it is necessary to detect the front and end points in real time according to the received voice stream.

[0187] In addition, an embodiment of the present invention also proposes a computer-readable storage medium.

[0188] The computer-readable storage medium stores a real-time speech recognition program. When the real-time speech recognition program is executed by a processor, the steps of the speech recognition method in any of the above embodiments are implemented. The specific implementation manner of the computer-readable storage medium of the present invention is basically the same as that of each embodiment of the above speech recognition method, and will not be repeated here.

[0189] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or system comprising such element.

[0190] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0191] From the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and executes the methods described in the various embodiments of the present invention.

[0192] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A voice recognition method, characterized in that, The speech recognition method includes: Obtain the original speech signal to be recognized; Based on a pre-trained classification model, classify the initial speech frames in the original speech signal to obtain a speech frame classification result; the pre-trained classification model is a binary classification model, and the binary classification model includes: a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer; According to the speech frame classification result and in combination with a preset determination rule, detect the front endpoint and the back endpoint of the effective speech segment; Input the effective speech segment between the front endpoint and the back endpoint into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

2. The voice recognition method according to claim 1, wherein After the step of obtaining the original speech signal to be recognized, it further includes: Intercept and normalize the byte stream of the original speech signal to obtain the initial speech frame information in the form of a one-dimensional floating-point matrix of the original speech signal.

3. The voice recognition method according to claim 2, wherein The step of, based on a pre-trained classification model, classifying the initial speech frames in the original speech signal to obtain a speech frame classification result includes: Input the initial speech frame information in the form of a one-dimensional floating-point matrix into a pre-trained binary classification model; Extract features from the initial speech frames through the shallow feature extraction layer in the binary classification model to obtain primary features, where the shallow feature extraction layer includes: a batch normalization (Batch Normalization) layer, a max pooling (MaxPooling) layer, and a one-dimensional convolutional layer; Perform convolution calculations of different scales on the primary features through the stacked multi-scale one-dimensional convolutional residual layers to obtain the calculated high-level features, where the multi-scale one-dimensional convolutional residual layer includes m paths of one-dimensional convolutional residual layers of different scales, and each path of one-dimensional convolutional residual layer includes n one-dimensional convolutional residual blocks with the same convolutional kernel size and an average pooling (AvgPooling) layer, where m and n are positive integers; Link the high-level features to the integration layer of the binary classification model for integration to obtain a corresponding integration result; Input the integration result into the fully connected layer and the softmax layer of the output layer to obtain a probability matrix of the initial speech frames belonging to different categories; Match the obtained probability matrix with a preset threshold to obtain the speech frame classification result of the initial speech frames.

4. The voice recognition method according to claim 1, wherein The step of, according to the speech frame classification result and in combination with a preset determination rule, detecting the front endpoint and the back endpoint of the effective speech segment includes: Based on the classification result of the initial speech frames in the original speech signal, obtain a classification result sequence of the original speech signal, where the classification result includes active speech frames and inactive speech frames; According to the classification result sequence and in combination with a preset front endpoint and back endpoint determination rule, obtain the front endpoint and the back endpoint of the effective speech segment of the original speech signal.

5. The voice recognition method according to claim 4, wherein The step of, according to the classification result sequence and in combination with a preset front endpoint and back endpoint determination rule, obtaining the front endpoint and the back endpoint of the effective speech segment of the original speech signal includes: Obtain a first classification result subsequence from the classification result sequence, where the first classification result subsequence includes: classification results of N consecutive speech frames; If the number of active speech frames in the first classification result subsequence reaches a first threshold, determine that the speech frame corresponding to the forefront of the first classification result subsequence is the front endpoint of the valid speech segment; Based on the front endpoint, obtain a second classification result subsequence from the classification result sequence, where the second classification result subsequence includes: classification results of M consecutive speech frames after the front endpoint; If the number of non-active speech frames in the second classification result subsequence reaches a second threshold, determine that the speech frame corresponding to the rearmost end of the second classification result subsequence is the rear endpoint of the valid speech segment.

6. The voice recognition method according to claim 1, characterized in that The step of inputting the valid speech segment between the front endpoint and the rear endpoint into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result includes: Input the valid speech segment between the front endpoint and the rear endpoint into a pre-trained acoustic model; Decode the valid speech segment through the acoustic model to implement text transcription and obtain the target transcription result corresponding to the valid speech segment.

7. The voice recognition method according to claim 3, wherein Before the step of obtaining the original speech signal to be recognized, the speech recognition method further includes: Obtain a sample audio of training data, and extract sample features of the sample audio, where the sample audio has corresponding classification results; Based on the sample features, establish a sample feature data set; Construct a multi-scale one-dimensional residual neural network, and perform deep learning on the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model.

8. The speech recognition method according to claim 7, wherein After the step of constructing the multi-scale one-dimensional residual neural network and performing deep learning on the multi-scale one-dimensional residual neural network based on the sample feature data set to obtain an initial binary classification model, the speech recognition method further includes: Test the initial binary classification model with telephone channel speech data to verify the classification effect of the initial binary classification model; If the classification effect does not reach the preset standard, return to perform data enhancement on the sample audio of the training data to obtain an optimized sample audio of the training data; Fine-tune and train the initial binary classification model according to the optimized sample audio of the training data to obtain a trained binary classification model; If the classification effect reaches the preset standard, use the initial binary classification model as the trained binary classification model.

9. A voice recognition device, characterized in that, The speech recognition device includes: An acquisition module for acquiring an original speech signal to be recognized; A determination module for classifying initial speech frames in the original speech signal based on a pre-trained classification model to obtain speech frame classification results; the pre-trained classification model is a binary classification model, and the binary classification model includes: a shallow feature extraction layer, a multi-scale one-dimensional convolutional residual layer, an integration layer, and an output layer; A detection module for detecting the front endpoint and the rear endpoint of the valid speech segment according to the speech frame classification results and in combination with a preset determination rule; A determination module, configured to input a valid speech segment between the front-end point and the back-end point into a pre-trained acoustic model for text transcription to obtain a corresponding target transcription result.

10. An intelligent device, characterized in that, The intelligent device includes a memory, a processor, and a speech recognition program stored on the memory and executable on the processor. When the speech recognition program is executed by the processor, the steps of the speech recognition method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, characterized in that, A speech recognition program is stored on the computer-readable storage medium. When the speech recognition program is executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Chinese speech recognition method based on deep residual

    CN110148408A