Speech recognition method, device, electronic device and storage medium
By using the effective interval and weight of the auxiliary voice block in streaming speech recognition, the spectrum information of the voice segment is recognized, which solves the problem of low recognition accuracy in streaming speech recognition by ULSTM, and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202010607003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-06-29
AI Technical Summary
In streaming speech recognition, the speech recognition accuracy of the one-way long-term short memory model (ULSTM) used in the prior art is low and cannot meet scenarios with high requirements for recognition accuracy.
By obtaining the spectrum information of the voice segment and using the effective interval and weight of the auxiliary voice block, the target voice block is recognized. The method includes inputting the spectrum information of the speech segment into the neural network model, which is set with the effective interval and weight of the auxiliary speech block to improve the recognition accuracy.
By considering the auxiliary voice blocks in the effective interval of the target voice block, the recognition accuracy of streaming voice recognition is improved, and scenarios with high requirements for recognition accuracy are met.
Smart Images

Figure CN113920998B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a speech recognition method, device, electronic equipment and storage medium. Background Art
[0002] With the continuous development of deep learning, speech recognition has evolved from the early hidden Markov model-mixture of gaussia (HMM-GMM) to the latest end-to-end neural network recognition model, and the accuracy of speech recognition has made great progress.
[0003] The structure of the common end-to-end neural network recognition model is based on the connectionist temporal classification (CTC) method, which includes a spectrum calculation layer, a feature extraction (Encoder) layer, a fully connected (FC) layer, and a normalized index (Softmax) layer. Among them, the output CTC probability obtains the corresponding recognition text through the Encoder layer. The common Encoders at present are the unidirectional long short-term memory model (ULSTM) and the bidirectional long short-term memory model (BLSTM). Among them, ULSTM only uses forward loop calculation in the process of time series modeling; BLSTM not only uses forward loop calculation in the process of time series modeling, but also uses reverse loop calculation. BLSTM fully considers the context information in the calculation process, so the recognition accuracy is higher than ULSTM.
[0004] However, in streaming speech recognition, the speech recognition device continuously receives speech segments and recognizes the received speech segments in real time. Since BLSTM requires the input of a complete speech segment, it is not suitable for streaming recognition scenarios. The encoder for streaming recognition usually uses ULSTM, but the speech recognition accuracy of ULSTM is low and cannot meet the scenarios with high recognition accuracy requirements. Summary of the invention
[0005] The embodiments of the present application provide a speech recognition method, device, electronic device and storage medium to solve the technical problem of low recognition accuracy in streaming speech recognition in the prior art.
[0006] In a first aspect, an embodiment of the present application provides a speech recognition method, comprising:
[0007] Acquire frequency spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block;
[0008] The target speech block is identified according to the spectrum information of the first speech segment, and the valid interval and weight of the auxiliary speech block.
[0009] In a possible design, identifying the target speech block according to the spectrum information of the first speech segment and the valid interval and weight of the auxiliary speech block includes:
[0010] The frequency spectrum information of the first speech segment is input into a neural network model, and a recognition result of the target speech block output by the neural network model is obtained, wherein the neural network model is provided with a valid interval and a weight of the auxiliary speech block.
[0011] In a possible design, before inputting the spectrum information of the first speech segment into the neural network model and obtaining the recognition result of the target speech block output by the neural network model, the method further includes:
[0012] The neural network model is trained using a sample set.
[0013] In a possible design, training the neural network model using the sample set includes:
[0014] sorting the sample set according to the length of the samples in the sample set;
[0015] According to the order of the sorted sample sets, the neural network model is trained using the sample sets.
[0016] In a possible design, training the neural network model using the sample set includes:
[0017] The neural network model is trained using the sample set using a connectionist temporal classification (CTC) function as a loss function.
[0018] In one possible design, the neural network model is a self-attention mechanism neural network model.
[0019] In a possible design, the sample set includes speech segments of various lengths and annotated texts corresponding to the speech segments of various lengths.
[0020] In a second aspect, an embodiment of the present application provides a speech recognition device, comprising:
[0021] An acquisition module, configured to acquire spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block;
[0022] The recognition module is used to recognize the target speech block according to the spectrum information of the first speech segment and the valid interval and weight of the auxiliary speech block.
[0023] In one possible design, the recognition module is specifically used to input the spectral information of the first speech segment into a neural network model, and obtain the recognition result of the target speech block output by the neural network model, and the neural network model is provided with a valid interval and weight of the auxiliary speech block.
[0024] In one possible design, the device further includes:
[0025] The training module is used to train the neural network model through a sample set.
[0026] In a possible design, the training module is specifically used to sort the sample set according to the length of the samples in the sample set; and train the neural network model through the sample set according to the order of the sorted sample set.
[0027] In one possible design, the training module is specifically used to use a connectionist temporal classification CTC function as a loss function to train the neural network model using the sample set.
[0028] In one possible design, the neural network model is a self-attention mechanism neural network model.
[0029] In a possible design, the sample set includes speech segments of various lengths and annotated texts corresponding to the speech segments of various lengths.
[0030] In a third aspect, the present application further provides an electronic device, including:
[0031] Processor; and
[0032] a memory for storing a computer program for the processor;
[0033] The processor is configured to implement any possible method in the first aspect by executing the computer program.
[0034] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored, and when the computer program is executed by a processor, any possible method in the first aspect is implemented.
[0035] The embodiment of the present application provides a speech recognition method, device, electronic device and storage medium, which first obtains the spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, and the auxiliary speech block is a speech block adjacent to the target speech block. Subsequently, the target speech block is recognized according to the spectrum information of the first speech segment, as well as the valid interval and weight of the auxiliary speech block. Since the auxiliary speech block in the valid interval of the target speech block is taken into account when performing streaming speech recognition, the recognition accuracy of streaming speech recognition can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0037] Figure 1 A schematic diagram of the existing CTC speech recognition architecture provided for this application;
[0038] Figure 2 A schematic diagram of an existing ULSTM neural network provided for this application;
[0039] Figure 3 A schematic diagram of an existing BLSTM neural network provided for this application;
[0040] Figure 4 A schematic diagram of an application scenario of a speech recognition method provided in an embodiment of the present application;
[0041] Figure 5 A flowchart of a speech recognition method provided in an embodiment of the present application;
[0042] Figure 6 A schematic diagram of a method for dividing speech segments provided in an embodiment of the present application;
[0043] Figure 7 A schematic diagram of temporal modeling of a self-attention mechanism neural network model provided in an embodiment of the present application;
[0044] Figure 8 is a flow chart of another speech recognition method provided in an embodiment of the present application;
[0045] Fig. 9 A schematic diagram of the weight value of Value in the prior art provided by this application;
[0046] Fig.10A schematic diagram of a weight value of Value provided in an embodiment of the present application;
[0047] Fig.11 It is a flowchart of another speech recognition method provided in an embodiment of the present application;
[0048] Fig.12 A schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application;
[0049] Fig.13 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0052] The structure of the currently common end-to-end neural network recognition model is based on the connectionist temporal classification (CTC) approach. Figure 1 A schematic diagram of the existing CTC speech recognition architecture provided for this application, such as Figure 1 As shown in the figure, the CTC speech recognition architecture includes a spectrum calculation layer, a feature extraction (Encoder) layer, a fully connected (FC) layer, and a normalized index (Softmax) layer.
[0053] The output CTC probability is passed through the Encoder layer to obtain the corresponding recognition text. Currently, the common Encoders are the unidirectional long short-term memory model (ULSTM) and the bidirectional long short-term memory model (BLSTM). Figure 2 A schematic diagram of an existing ULSTM neural network provided for this application, Figure 3 A schematic diagram of an existing BLSTM neural network provided for this application. Figure 2 and Figure 3 As shown in the figure, in the process of time series modeling, the neurons (cells) of ULSTM only use forward loop calculations, while in the process of time series modeling, the cells of BLSTM not only use forward loop calculations, but also reverse loop calculations. Since BLSTM fully considers the context information during the calculation process, its recognition accuracy is higher than that of ULSTM.
[0054] However, in streaming speech recognition, the speech recognition device continuously receives speech segments and recognizes the received speech segments in real time. Since BLSTM requires the input of a complete speech segment, it is not suitable for streaming recognition scenarios. The encoder for streaming recognition usually uses ULSTM, but the speech recognition accuracy of ULSTM is low and cannot meet the scenarios with high recognition accuracy requirements.
[0055] In view of the above problems, the present application provides a speech recognition method, device, electronic device and storage medium to improve the technical problem of low recognition accuracy during streaming speech recognition. The inventive concept of the present application is that when recognizing a target speech block, the target speech block can be assisted in recognition by an auxiliary speech block in the effective interval of the target speech block, thereby improving the recognition accuracy during streaming speech recognition.
[0056] The following describes the application scenarios of the speech recognition method provided in the embodiments of the present application. Figure 4 A schematic diagram of an application scenario of a speech recognition method provided in an embodiment of the present application. Figure 4 As shown, the terminal device 101 receives the user's streaming voice input, and then the terminal device 101 can send the voice segment in the streaming voice input to the server 102. After the server 102 recognizes the target voice block in the voice segment, it sends the recognition result of the target voice block to the terminal device 101.
[0057] Among them, the terminal device 101 is provided with an audio acquisition component.
[0058] The terminal device 101 may be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in a smart home, etc. In the embodiment of the present application, the device for realizing the function of the terminal may be a terminal, or a device capable of supporting the terminal to realize the function, such as a chip system, which may be installed in the terminal. In the embodiment of the present application, the chip system may be composed of a chip, or may include a chip and other discrete devices.
[0059] The server 102 may be a server or a server in a cloud service platform. The embodiment of the present application does not limit the type of server and may be specifically configured according to actual conditions.
[0060] It should be noted that the application scenarios of the technical solution of this application can be Figure 1 The present invention can be applied to the application scenarios in the present invention, but is not limited to this, and can also be applied to other scenarios that require speech recognition.
[0061] It can be understood that the above-mentioned speech recognition method can be implemented by the speech recognition device provided in the embodiment of the present application. The speech recognition device can be part or all of a certain device, for example, it can be a server or a processor in the server.
[0062] The following takes a server integrated or installed with relevant execution codes as an example to describe the technical solution of the embodiment of the present application in detail with a specific embodiment. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0063] Figure 5 This is a flow chart of a speech recognition method provided in an embodiment of the present application. The execution subject of this embodiment is a server. This embodiment involves a specific process of how to recognize a target speech block. Figure 5 As shown, the method includes:
[0064] S201. Acquire frequency spectrum information of a first speech segment, where the first speech segment includes a target speech block and an auxiliary speech block, where the auxiliary speech block is a speech block adjacent to the target speech block.
[0065] The first voice segment is any voice segment in the user's streaming voice input. It should be noted that the embodiment of the present application does not limit how to divide the voice segments in the streaming voice input, and can be specifically set according to actual conditions. In some embodiments, for streaming voice input, voice segments can be obtained in a partially overlapping manner, that is, the second half of the previous voice segment is used as the first half of the next voice segment to obtain the voice segment.
[0066] In addition, the embodiment of the present application does not limit how to divide the voice blocks in the voice segment. For example, the voice blocks can be divided according to the time length, for example, one voice block is one per second, or one voice block is one per 2 seconds. The target voice block is the voice block to be recognized in the voice segment. A voice segment may include one target voice block or multiple target voice blocks. The number of target voice blocks can be specifically set according to the actual situation.
[0067] For example, Figure 6 A schematic diagram of a method for dividing voice segments provided in an embodiment of the present application, such as Figure 6 The figure shows three voice segments obtained in sequence from the streaming voice input. The shaded blocks are target voice blocks, and the blank blocks are auxiliary voice blocks. Among them, the first voice segment being recognized is located between the second voice segment that has been successfully recognized and the third voice segment that has not been recognized. The target voice blocks in each voice segment can be combined into a complete streaming voice input. For each voice segment, it contains at least one target voice block and multiple auxiliary voice blocks. The auxiliary voice blocks are voice blocks adjacent to the target voice block, distributed in front of and behind the target voice block. The embodiment of the present application does not limit the number of auxiliary voices. For example, Figure 6 As shown, the interval of the auxiliary speech block can be set to (left, right) = (5, 3), that is, there are 5 auxiliary speech blocks before the target speech block, and there are 3 auxiliary speech blocks after the target speech block. In addition, adjacent speech segments can be partially overlapped, and the amount of overlap is related to the number of auxiliary speech blocks and target speech blocks. For example, Figure 6 As shown, for the first speech segment, the number of target speech blocks is 3, and the number of auxiliary speech blocks before the target speech blocks is 5, then the above 8 speech blocks overlap with the last 8 speech blocks of the second speech segment.
[0068] It should be noted that the embodiment of the present application does not limit how to determine the spectrum information from the audio file. For example, a spectrogram can be used to calculate the spectrum of the audio file to determine the spectrum information of the first voice segment.
[0069] S202: Identify the target speech block according to the spectrum information of the first speech segment, and the valid interval and weight of the auxiliary speech block.
[0070] In this step, after acquiring the spectrum information of the first voice segment, the server can identify the target voice block according to the spectrum information of the first voice segment and the effective interval and weight of the auxiliary voice block.
[0071] In an optional implementation, the server may input the spectrum information of the first speech segment into a neural network model, and obtain a recognition result of a target speech block output by the neural network model, wherein the neural network model is provided with a valid interval and weight of an auxiliary speech block.
[0072] Among them, the neural network model can be a self-attention mechanism (Self-Attention) neural network model.
[0073] The following is an explanation of the self-attention mechanism neural network model. Figure 7 A schematic diagram of temporal modeling of a self-attention mechanism neural network model provided in an embodiment of the present application. Figure 7 As shown in the figure, the self-attention mechanism neural network model first transforms the input spectrum information into three matrices, Query (Q), Key (K) and Value (V), through three different linear transformations. Then, matrix multiplication (MatMul), scale operation and normalized exponent (Softmax) operation are performed on Q and K respectively, and the result of the operation is MatMuled with V to obtain the output of the self-attention mechanism neural network model.
[0074] The Self-Attention neural network model used in the embodiment of the present application has better time series modeling capabilities than LSTM, thereby effectively improving the recognition accuracy. In addition, the Masked Self-Attention model has good computational parallelism, so the model training and reasoning speed are faster.
[0075] In addition, the present application sets the effective interval and weight of the auxiliary voice block in the neural network model, so that when the neural network module recognizes the target voice block in a certain voice segment of the streaming voice input, the auxiliary voice blocks within a certain range around the target voice block are retained and assigned different weights, which can meet the "recognition while speaking" requirement of the streaming voice input. In addition, compared with ULSTM, which only considers the recognition content before the target voice block, the present application can also improve the recognition accuracy of the target voice block through the auxiliary voice blocks on both sides.
[0076] The embodiment of the present application provides a speech recognition method, which first obtains the spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block. Subsequently, the target speech block is recognized according to the spectrum information of the first speech segment, as well as the valid interval and weight of the auxiliary speech block. Since the auxiliary speech block in the valid interval of the target speech block is taken into account when performing streaming speech recognition, the recognition accuracy of streaming speech recognition can be improved.
[0077] Based on the above embodiment, how to recognize the target speech block is specifically described below. Figure 8 is a flow chart of another speech recognition method provided by an embodiment of the present application, such as Figure 8 As shown, the speech recognition method includes:
[0078] S301, obtaining frequency spectrum information of a first speech segment, where the first speech segment includes a target speech block and an auxiliary speech block, where the auxiliary speech block is a speech block adjacent to the target speech block;
[0079] The technical terms, technical effects, technical features, and optional implementation methods of step S301 can be found in Figure 5 It is understood that the step S201 shown is repeated and will not be described again here.
[0080] S302, inputting the spectrum information of the first speech segment into the neural network model, and obtaining the recognition result of the target speech block output by the neural network model, wherein the neural network model is provided with the effective interval and weight of the auxiliary speech block.
[0081] Among them, the neural network model is a Masked Self-Attention neural network model.
[0082] In the step, after obtaining the spectrum information of the first speech segment, the server can input the spectrum information of the first speech segment into the neural network model, and obtain the recognition result of the target speech block output by the neural network model.
[0083] For example, if the spectrum information of the first speech segment is X=(x1, x2, ..., x t , ..., x N ), the server can convert X=(x1, x2, ..., x t , ..., x N) is input into the Masked Self-Attention neural network model. The Masked Self-Attention neural network model first transforms X into three matrices: Query (Q), Key (K), and Value (V). Then, the Masked Self-Attention neural network model Figure 6 The time series modeling shown calculates the three features and obtains the output of the Masked Self-Attention neural network model. The output of the Masked Self-Attention neural network model is shown in formula (1):
[0084]
[0085] Among them, d k is the number of columns of the Q and K matrices, that is, the vector dimension of Q and K at each moment. Softmax is a normalized exponential function.
[0086] Fig. 9 This is a schematic diagram of the weight value of Value in the prior art provided by this application, such as Fig. 9 As shown in the figure, when recognizing the target speech block output in the first speech segment at a certain moment, Self-Attention needs to calculate the similarity between the Query at that moment and the Key at all moments as the weight value of the Value, which is not conducive to streaming speech recognition.
[0087] Compared with the prior art, the neural network model provided in the embodiment of the present application is provided with valid intervals and weights of auxiliary speech blocks. Fig.10 A weight value diagram of a Value provided in an embodiment of the present application. Fig.10 As shown, the present application adds a matrix Mask on the basis of the standard Self-Attention, and the Mask includes the effective interval and weight of the auxiliary speech block. Therefore, when calculating the output at a certain moment, only the weight within the effective interval of the auxiliary speech block corresponding to the moment can be retained.
[0088] For example, Mask is an N x N matrix, which is constructed in such a way that the values within a certain range around the diagonal (the valid interval is [left, right]) are 0, and the values outside the range are -∞. Then Mask can be shown as formula (2):
[0089]
[0090] Among them, left and right are tunable parameters.
[0091] Based on the above Mask optimization, the output of the Masked Self-Attention neural network model can be optimized as shown in formula (3):
[0092]
[0093] Based on the above embodiment, the training process of the above neural network is described in detail below. Fig.11 is a flow chart of another speech recognition method provided in an embodiment of the present application, such as Fig.11 As shown, the speech recognition method includes:
[0094] S401, training the neural network model through the sample set.
[0095] The sample set includes speech segments of various lengths and annotated texts corresponding to the speech segments of various lengths.
[0096] In some embodiments, the server may first sort the sample set according to the length of the samples in the sample set, and then train the neural network model through the sample set according to the order of the sorted sample set.
[0097] The embodiments of the present application do not limit the loss function used for training. In some embodiments, a connectionist temporal classification (CTC) function can be used as a loss function to train the neural network model using a sample set.
[0098] For example, the training data of the neural network in the embodiment of the present application can be 5,000 hours of audio files and corresponding annotated texts in the customer service scenario. The sample set can be 5 million annotated speech segments of varying lengths between 1 second and 10 seconds randomly selected from the training data.
[0099] During the training process, the CTC function can be used as the loss function. In order to ensure the rapid convergence of the neural network model, the sample set can be sorted according to the sample length in the sample set, and then the neural network model can be trained through the sample set according to the order of the sorted sample set. That is, the first round of training is carried out in the order of sample length from small to large. After 20 rounds of training, the model with the best performance on the validation set is selected as the final neural network model.
[0100] S402: Acquire frequency spectrum information of a first speech segment, where the first speech segment includes a target speech block and an auxiliary speech block, where the auxiliary speech block is a speech block adjacent to the target speech block.
[0101] S403, inputting the spectrum information of the first speech segment into the neural network model, and obtaining the recognition result of the target speech block output by the neural network model, wherein the neural network model is provided with a valid interval and weight of the auxiliary speech block.
[0102] For the technical terms, technical effects, technical features, and optional implementation methods of steps S402-S403, please refer to Figure 8 The steps S301-S302 shown are understood, and the repeated contents will not be repeated here.
[0103] The embodiment of the present application provides a speech recognition method, which first obtains the spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block. Subsequently, the target speech block is recognized according to the spectrum information of the first speech segment, as well as the valid interval and weight of the auxiliary speech block. Since the auxiliary speech block in the valid interval of the target speech block is taken into account when performing streaming speech recognition, the recognition accuracy of streaming speech recognition can be improved.
[0104] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0105] Fig.12 The following is a schematic diagram of the structure of a speech recognition device provided in an embodiment of the present application. The speech recognition device can be implemented by software, hardware, or a combination of both, such as the server in the above embodiment, to execute the speech recognition method in the above embodiment. Fig.12 As shown, the speech recognition device comprises:
[0106] An acquisition module 501 is used to acquire spectrum information of a first speech segment, where the first speech segment includes a target speech block and an auxiliary speech block, where the auxiliary speech block is a speech block adjacent to the target speech block;
[0107] The recognition module 502 is used to recognize the target speech block according to the spectrum information of the first speech segment, and the effective interval and weight of the auxiliary speech block.
[0108] In one possible design, the recognition module 502 is specifically used to input the spectrum information of the first speech segment into the neural network model, and obtain the recognition result of the target speech block output by the neural network model, and the neural network model is set with the effective interval and weight of the auxiliary speech block.
[0109] In one possible design, the device further includes:
[0110] The training module 503 is used to train the neural network model through the sample set.
[0111] In a possible design, the training module 503 is specifically used to sort the sample sets according to the sample lengths in the sample sets; and train the neural network model through the sample sets according to the order of the sorted sample sets.
[0112] In one possible design, the training module 503 is specifically used to use the connectionist temporal classification CTC function as the loss function to train the neural network model through the sample set.
[0113] In one possible design, the neural network model is a self-attention mechanism neural network model.
[0114] In a possible design, the sample set includes speech segments of various lengths and annotated texts corresponding to the speech segments of various lengths.
[0115] Need to explain, Fig.12 The speech recognition device provided in the illustrated embodiment can be used to execute the method provided in any of the above embodiments. The specific implementation method and technical effect are similar and will not be described in detail here.
[0116] Fig.13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.13 As shown, electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0117] like Fig.13As shown, the electronic device includes: one or more processors 601, a memory 602, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are interconnected using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.13 A processor 601 is taken as an example.
[0118] The memory 602 is a non-transient computer-readable storage medium provided in the present application. The memory stores instructions executable by at least one processor to enable at least one processor to perform the speech recognition method provided in the present application. The non-transient computer-readable storage medium of the present application stores computer instructions, which are used to enable a computer to perform the speech recognition method provided in the present application.
[0119] The memory 602 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the speech recognition method in the embodiment of the present application (for example, Fig.11 The processor 601 executes various functional applications and data processing of the server by running the non-transient software programs, instructions and modules stored in the memory 602, that is, the speech recognition method in the above method embodiment is implemented.
[0120] The memory 602 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of the electronic device provided in accordance with an embodiment of the present application, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 602 may optionally include a memory remotely arranged relative to the processor 601, and these remote memories may be connected to the electronic device provided in the embodiment of the present application via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0121] The electronic device provided in the embodiment of the present application may further include: an input device 603 and an output device 606. The processor 601, the memory 602, the input device 603 and the output device 606 may be connected via a bus or other means. Fig.13 The example of connecting through bus is taken in the following.
[0122] The input device 603 can receive input digital or character information, and generate key signal input related to user settings and function control of the electronic device provided in the embodiment of the present application, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick and other input devices. The output device 606 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display and a plasma display. In some embodiments, the display device may be a touch screen.
[0123] Various implementations of the systems and techniques described herein can be realized in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for programmable processors and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or means (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0126] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0127] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0128] The present application also provides a chip, including a processor and an interface. The interface is used to input and output data or instructions processed by the processor. The processor is used to execute the method provided in the above method embodiment. The chip can be applied to a speech recognition device.
[0129] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes. Specifically, the computer-readable storage medium stores program information, and the program information is used for the above-mentioned speech recognition method.
[0130] The embodiment of the present application also provides a program, which, when executed by a processor, is used to execute the speech recognition method provided by the above method embodiment.
[0131] The embodiment of the present application also provides a program product, such as a computer-readable storage medium, in which instructions are stored. When the program product is run on a computer, the computer executes the speech recognition method provided by the above method embodiment.
[0132] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.
[0133] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A speech recognition method, characterized in that: The method comprises: Acquire frequency spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block; The spectral information of the first speech segment is input into the neural network model, and the recognition result of the target speech block output by the neural network model at a certain moment is obtained. The matrix of the neural network model is provided with the valid interval and weight of the auxiliary speech block to retain the weight of the valid interval of the auxiliary speech block corresponding to the moment.
2. The method according to claim 1, characterized in that: Before inputting the spectrum information of the first speech segment into the neural network model and obtaining the recognition result of the target speech block output by the neural network model, the method further includes: The neural network model is trained using a sample set.
3. The method according to claim 2, characterized in that The training of the neural network model by using the sample set includes: sorting the sample set according to the length of the samples in the sample set; According to the order of the sorted sample sets, the neural network model is trained using the sample sets.
4. The method according to claim 3, characterized in that: The training of the neural network model by using the sample set includes: The neural network model is trained using the sample set using a connectionist temporal classification (CTC) function as a loss function.
5. The method according to any one of claims 2 to 4, characterized in that: The neural network model is a self-attention mechanism neural network model.
6. The method according to any one of claims 2 to 4, characterized in that: The sample set includes speech segments of various lengths and annotated texts corresponding to the speech segments of various lengths.
7. A speech recognition device, characterized in that: The device comprises: An acquisition module, configured to acquire spectrum information of a first speech segment, wherein the first speech segment includes a target speech block and an auxiliary speech block, wherein the auxiliary speech block is a speech block adjacent to the target speech block; A recognition module, configured to recognize the target speech block according to the spectrum information of the first speech segment, and the valid interval and weight of the auxiliary speech block; The recognition module is specifically used to input the spectral information of the first speech segment into the neural network model, and obtain the recognition result of the target speech block output by the neural network model at a certain moment. The matrix of the neural network model is provided with the valid interval and weight of the auxiliary speech block to retain the weight of the valid interval of the auxiliary speech block corresponding to the moment.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Streaming phonetic transcription system based on self-attention mechanism
CN110473529A
End-to-end voice identification method, system and device and storage medium
CN110767218A