Speech Recognition Method, Apparatus, Electronic Device, Storage Medium, and Program Product

The modified Transformer structure with a neural network weight matrix and sparsified self-attention mechanism addresses the computational resource constraints of embedded devices, enhancing voice recognition speed and efficiency.

CN114678011BActive Publication Date: 2025-07-15KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210323478.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-07-15
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

Embedded voice recognition equipment has limited computing resources and requires rapid response. The self-attention module of the existing Transformer structure has a large amount of calculation, resulting in high system performance requirements.

Method used

The sparse processing based on the neural network weight matrix and the improved self-attention module are adopted. Through the sparse processing of neural network weight matrix and the simplified calculation process of self-attention value, the calculation complexity is reduced and the calculation speed is improved.

Benefits of technology

It significantly improves the speech recognition processing speed of embedded speech recognition devices, reduces the computational complexity, and meets the needs of fast response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114678011B_ABST
    Figure CN114678011B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a speech recognition method, apparatus, electronic device, storage medium and program product. The method includes: obtaining a speech signal to be subjected to speech recognition; performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network. The embodiment of the present invention can improve the calculation speed of the self-attention value, thereby improving the speed of the output result of the self-attention module of the improved Transformer structure, improving the processing speed of the acoustic model and the language model constructed by using the improved Transformer structure, and thus improving the processing speed of speech recognition as a whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a speech recognition method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] Currently, mainstream speech recognition systems mostly build acoustic models and language models based on the Transformer architecture. The Transformer architecture has a self-attention module for calculating self-attention. Figure 2 FIG. 1 is a schematic diagram of the calculation process of the self-attention module in the traditional Transformer architecture. When the self-attention module calculates self-attention, the following calculation process is adopted:

[0003] Q = XW q

[0004] K = XW k

[0005] V = XW v

[0006]

[0007] SelfAttention(Q, K, V) = AV

[0008] where Input X ∈ R n×m represents the input speech signal, n and m respectively represent the number of rows and columns of the speech signal, n can be the number of frames of the speech signal, and m can be the number of frequencies of the speech signal. X passes through different transformation matrices W q , W k , W v , to obtain d q 、d k 、d v respectively represent the number of columns of the transformation matrices W q , W k , W v . Softmax is used for normalization processing on the dimension of each row, A is the self-attention value, and SelfAttention(·) is the final output of the self-attention module. In the above formula, Figure 2 the Query in is represented as Q, the Key is represented as K, and the Value is represented as V.

[0009] The above steps require a large amount of computational processes and have high requirements for the performance of the system server. Compared with the online speech recognition server, the computational resources (such as memory and computing power) of the embedded speech recognition device are extremely limited, and at the same time, the system is required to have a faster response speed. Therefore, how to reduce the computational amount and improve the operation speed is an urgent problem to be solved in the application of embedded speech recognition. Summary of the Invention

[0010] To solve the problems in the prior art, embodiments of the present invention provide a speech recognition method, device, electronic device, storage medium, and program product.

[0011] Embodiments of the present invention provide a speech recognition method, including: obtaining a speech signal to be subjected to speech recognition; performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network.

[0012] According to a speech recognition method provided by an embodiment of the present invention, the performing speech recognition on the speech signal based on the acoustic model and the language model includes: performing acoustic processing on the speech signal to obtain acoustic features; using the acoustic features and the acoustic model to perform acoustic encoding to obtain a syllable sequence; performing word list matching on the syllable sequence and a word list to obtain a word sequence; and inputting the word sequence into the language model to output a language decoding result.

[0013] According to a speech recognition method provided by an embodiment of the present invention, the self-attention value is expressed as: A = softmax(f(g(W))), and the computational amount of f(g(W)) is less than Q = XW q ,K = XW k wherein, A represents the self-attention value, X represents the speech signal, W q 、W k represent transformation matrices, d k represents the number of columns of the transformation matrix W k ; W represents the weight matrix of the neural network, g(W) represents a function of W, and f(g(W)) represents a function of g(W).

[0014] According to a speech recognition method provided by an embodiment of the present invention, the self-attention value is expressed as: g(W) = WX or g(W) = W; wherein, X represents the speech signal.

[0015] According to a speech recognition method provided by an embodiment of the present invention, the weight matrix of the neural network is subjected to sparsification processing.

[0016] According to a speech recognition method provided by an embodiment of the present invention, the effective values of the weight matrix of the neural network exist on the diagonal and positions extending from the diagonal to both sides; and for each row of the weight matrix corresponding to the neural network, the sum of the number of effective values is less than the number of rows of the speech signal.

[0017] An embodiment of the present invention further provides a speech recognition device, including: a speech signal acquisition module, configured to: acquire a speech signal to be subjected to speech recognition; a speech recognition module, configured to: perform speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of the neural network.

[0018] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the speech recognition method as described in any one of the above are implemented.

[0019] An embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the speech recognition method as described in any one of the above are implemented.

[0020] An embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the speech recognition method as described in any one of the above are implemented.

[0021] The speech recognition method, device, electronic device, storage medium, and program product provided by the embodiments of the present invention acquire a speech signal to be subjected to speech recognition, and perform speech recognition on the speech signal based on an acoustic model and a language model. Among them, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of the neural network, which can improve the calculation speed of the self-attention value, thereby improving the output speed of the self-attention module of the improved Transformer structure, and improving the processing speed of the acoustic model and the language model constructed by using the improved Transformer structure, so as to improve the overall processing speed of speech recognition. Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 is one of the schematic flowcharts of the speech recognition method provided by an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of the calculation process of the self-attention module in the traditional Transformer structure;

[0025] Figure 3 is one of the schematic diagrams of the calculation process of the self-attention module in the speech recognition method provided by an embodiment of the present invention;

[0026] Figure 4 is the second schematic flowchart of the speech recognition method provided by the present invention;

[0027] Figure 5 is the second schematic diagram of the calculation process of the self-attention module in the speech recognition method provided by an embodiment of the present invention;

[0028] Figure 6 is a schematic diagram of the computational complexity of the speech recognition method provided by an embodiment of the present invention without sparsifying the weight matrix of the neural network;

[0029] Figure 7 is a schematic diagram of the sparsification process of the weight matrix of the neural network in the speech recognition method provided by an embodiment of the present invention;

[0030] Figure 8 is a schematic diagram of the structure of the speech recognition device provided by an embodiment of the present invention;

[0031] Figure 9 is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0033] The following will be combined with Figures 1-9Describe a speech recognition method, apparatus, electronic device, storage medium, and program product according to embodiments of the present invention.

[0034] Figure 1 It is one of the flow diagrams of the speech recognition method provided by the embodiments of the present invention. As Figure 1 shown, the method includes:

[0035] Step 101, obtain a speech signal to be subjected to speech recognition.

[0036] Obtain a speech signal to be subjected to speech recognition. The speech signal to be recognized may be an original speech signal or a speech signal after preliminary signal processing (such as denoising).

[0037] Step 102, perform speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the modified self-attention module of the Transformer structure is obtained based on the weight matrix of a neural network.

[0038] The present invention lies in the improvement of the self-attention module in the Transformer structure (not an entity structure, reflected in the calculation process of each functional module). Among them, the self-attention module is a functional module. The improvement points of the present invention include simplifying the calculation process of the self-attention value A in the self-attention module and reducing the calculation amount of the self-attention value A.

[0039] The improvement of the present invention lies in the acceleration of the calculation of the self-attention module in the Transformer structure. The self-attention module in the Transformer structure includes a self-Attention module (self-attention mechanism module) and an Encoder-DecoderAttention module (encoding-decoding attention module). Since the input parameters of the Encoder-Decoder Attention module are different from those of the self-Attention module, they are represented in different forms, and it is also used to calculate self-attention.

[0040] Perform speech recognition on the speech signal to be recognized based on an acoustic model and a language model. Among them, the acoustic model and the language model are constructed based on an improved Transformer structure. The improved Transformer structure is obtained by improving the self-attention module in the existing Transformer structure. The self-attention value A of the self-attention module in the improved Transformer structure is obtained based on the weight matrix of a neural network.

[0041] The self-attention value A is obtained through a neural network-based weight matrix. On the basis of effectively obtaining the self-attention value A, the computational complexity of the self-attention value A can be reduced, such that the computational complexity of the self-attention value A is less than that of the existing self-attention value A.

[0042] Figure 3 It is one of the schematic diagrams of the calculation process of the self-attention module in the speech recognition method provided by the embodiments of the present invention. As Figure 3 shown, the improvement of the present invention lies in the calculation of the self-attention value A, that is, Figure 3 the part within the dashed box in

[0043] The speech recognition method provided by the embodiments of the present invention includes obtaining a speech signal to be subjected to speech recognition, and performing speech recognition on the speech signal based on an acoustic model and a language model. Among them, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module. The self-attention value of the self-attention module in the modified Transformer structure is obtained based on a neural network-based weight matrix, which can improve the calculation speed of the self-attention value, and further improve the speed of the output result of the self-attention module of the modified Transformer structure, and improve the processing speed of the acoustic model and the language model constructed by using the modified Transformer structure, thereby improving the overall processing speed of speech recognition.

[0044] According to a speech recognition method provided by an embodiment of the present invention, the performing speech recognition on the speech signal based on the acoustic model and the language model includes: performing acoustic processing on the speech signal to obtain acoustic features; using the acoustic features and the acoustic model to perform acoustic encoding to obtain a syllable sequence; performing word list matching on the syllable sequence and a word list to obtain a word sequence; and inputting the word sequence into the language model to output a language decoding result.

[0045] Figure 4 It is the second schematic diagram of the flow of the speech recognition method provided by the present invention. The calculation of the self-attention module in the Transformer structure is improved based on the above-mentioned calculation method of the self-attention value, and an acoustic model and a language model are constructed by using the improved Transformer structure, and speech recognition is performed based on the constructed acoustic model and language model. As Figure 4 shown, the process of performing speech recognition based on the constructed acoustic model and language model includes:

[0046] After performing acoustic processing on the input speech signal, acoustic features are obtained; using the acoustic features and an acoustic model for acoustic encoding to obtain a syllable sequence; performing a vocabulary matching between the syllable sequence and a vocabulary to obtain a word sequence; inputting the word sequence into a language model to output a language decoding result, thereby completing the speech recognition process.

[0047] The computational complexity of the improved Transformer structure is significantly reduced, and it can be used for signal processing in acoustic models and language models to improve the response speed of the application device to speech signals, thereby achieving faster speech recognition tasks.

[0048] The speech recognition method provided by the embodiments of the present invention constructs an acoustic model and a language model based on the improved Transformer structure, and performs speech recognition based on the acoustic model and the language model, improving the processing speed of the acoustic model and the language model constructed using the Transformer structure, thereby improving the overall processing speed of speech recognition.

[0049] According to the speech recognition method provided by the embodiments of the present invention, the self-attention value is expressed as: A = softmax(f(g(W))), and the computational amount of f(g(W)) is less than Q = XW q ,K = XW k wherein, A represents the self-attention value, X represents the speech signal, W q 、W k represents a transformation matrix, d k represents the number of columns of the transformation matrix W k ; W represents the weight matrix of the neural network, g(W) represents a function of W, and f(g(W)) represents a function of g(W).

[0050] In the calculation of the self-attention value A in the embodiments of the present invention, the self-attention value A is also obtained through normalization processing. The improvement lies in that the object of the normalization processing is expressed as f(g(W)). In the calculation of the self-attention value A in the prior art, the object of the normalization processing is expressed as In the embodiments of the present invention, the computational amount of f(g(W)) is less than Thereby improving the calculation speed of the self-attention value A.

[0051] The speech recognition method provided by the embodiments of the present invention reduces the computational amount of the self-attention value by improving the object of the normalization processing, improves the calculation speed, and realizes the simplicity of improving the speech recognition speed.

[0052] According to a speech recognition method provided by the present invention, g(W) = WX or g(W) = W; wherein, X represents the speech signal.

[0053] g(W) can be expressed as WX, then the self-attention value A is expressed as: A = softmax(f(WX)), that is, the self-attention value A is obtained by normalizing the function of WX. f(WX) can be a non-linear function of WX. For example, f(WX) is expressed as sigmoid(WX), f(WX) = sigmoid(WX + b), etc. where b can take the value of the bias of the neural network, and sigmoid represents the activation function of the neural network. g(W) is expressed as WX, which reduces the computational amount of the self-attention value on the basis of retaining the input speech signal and improves the calculation speed of the self-attention value.

[0054] To further reduce the computational amount of the self-attention value A, f(WX) can also be expressed as a linear function of WX. For example, f(WX) is expressed as f(WX) = WX, f(WX) = WX + b, etc. where b can take the value of the bias of the neural network. By making f(WX) be a linear transformation function of WX, the computational amount of the self-attention value is further reduced on the basis of retaining the input speech signal, and the calculation speed of the self-attention value is improved.

[0055] To further reduce the computational amount of the self-attention value A, g(W) can also be independent of the input speech signal X, that is, the self-attention value A is generated by completely random initialization. The neural network updates the self-attention value A during training. g(W) can be expressed as W, that is, g(W) = W, then the self-attention value A is expressed as: A = softmax(f(W)), that is, the self-attention value A is obtained by normalizing the function of W.

[0056] Figure 5 It is the second schematic diagram of the calculation process of the self-attention module in the speech recognition method provided by the embodiments of the present invention. As Figure 5 shown, when g(W) = W, the calculation of the self-attention value A will not depend on the input speech signal X. By making g(W) = W and not performing the calculation of the input speech signal, the computational amount of the self-attention value is further reduced, and the calculation speed of the self-attention value is improved.

[0057] The speech recognition method provided by the embodiments of the present invention further reduces the computational amount of the self-attention value and improves the calculation speed of the self-attention value by making g(W) = WX or g(W) = W, thereby further improving the speed of speech recognition.

[0058] According to the speech recognition method provided by the embodiments of the present invention, the weight matrix of the neural network is sparsified.

[0059] During the calculation process of the self-attention value A, the weight matrix W of the neural network can be sparsified to further reduce the computational complexity of the self-attention value A. Sparsifying the weight matrix W of the neural network can be achieved by fixing certain positions of W to 0 values so that the data at these positions does not participate in the calculation process.

[0060] The speech recognition method provided by the embodiments of the present invention reduces the computational complexity of the self-attention value and improves the calculation speed of the self-attention value by sparsifying the weight matrix of the neural network.

[0061] According to a speech recognition method provided by an embodiment of the present invention, the valid values of the weight matrix of the neural network exist on the diagonal and the positions extending from the diagonal to both sides; and for each row of the weight matrix of the neural network, the sum of the number of valid values is less than the number of rows of the speech signal.

[0062] Figure 6 It is a schematic diagram of the computational complexity when the weight matrix of the neural network is not sparsified in the speech recognition method provided by the embodiments of the present invention. As Figure 6 shown, for example, W ∈ R n×n , where n represents the number of rows of the input speech signal and can be the number of frames of the speech signal. Then each frame of the signal needs to calculate the correlation with all signals, and the computational complexity is O(n 2 ), which requires high performance of the system server. Computational complexity is not the same as computational amount, but is related to the computational amount.

[0063] Figure 7 It is a schematic diagram of the sparsification process of the weight matrix of the neural network in the speech recognition method provided by the embodiments of the present invention. As Figure 7 shown, the left figure represents W before sparsification, W ∈ R n×n , and the right figure represents W after sparsification, W ∈ R a×n . Among them, a represents the maximum number of columns occupied by the valid values in each row of the weight matrix W of the neural network, and a is less than the number of rows n of the input speech signal, thereby reducing the computational complexity.

[0064] Moreover, according to the local correlation of speech, only the local self-attention value is calculated, and the valid values of W exist on the diagonal and the positions extending from the diagonal to both sides to achieve the calculation of the local attention value. If the number of rows of the input speech signal is expressed as the number of frames, then as Figure 7As shown in the right figure, the first valid value in the first row is used to calculate the correlation between the first frame and the first frame, and the second valid value in the first row is used to calculate the correlation between the first frame and the second frame; the first valid value in the second row is used to calculate the correlation between the second frame and the first frame, the second valid value in the second row is used to calculate the correlation between the second frame and the second frame, and the third valid value in the second row is used to calculate the correlation between the second frame and the third frame, thus ensuring the calculation of the local self-attention value.

[0065] The calculation complexity of the matrix calculation is reduced by making the maximum value a (the maximum number of columns occupied by the valid values in each row of the weight matrix W of the neural network) of the sum of the elements on the diagonal of each row and the number of elements adjacent to both sides less than the number of rows n of the speech signal.

[0066] W ∈ R a×n ,X ∈ R n×m ,where n and m respectively represent the number of rows and columns of the speech signal. n can be the number of frames of the speech signal, and m can be the number of frequencies of the speech signal. Taking the calculation of WX as an example, the features of each row of W are calculated with the features of each column of X. Since there are at most a valid values in each row of W, the calculation complexity is reduced from O(n 2 ) to O(an). If f(X) = W is adopted and the self-attention value A is calculated using the sparsified W, the calculation complexity of the self-attention value A will also be reduced. Although the calculation complexity is reduced, since the adjacent valid values are retained in W, the calculation of local correlation is ensured.

[0067] a can be randomly generated by the random number method. The optimal value of a can be selected through multiple rounds of experiments according to the speech recognition effect, such as determining the optimal value of a according to the processing speed and accuracy of speech recognition.

[0068] The speech recognition method provided by the embodiments of the present invention balances the improvement of the calculation speed and the requirement of speech recognition accuracy by making the valid values of W exist at the diagonal and the positions extending from the diagonal to both sides, and the sum of the number of valid values in each row is less than the number of rows of the speech signal.

[0069] The speech recognition device provided by the embodiments of the present invention will be described below. The speech recognition device described below can be correspondingly referred to the speech recognition method described above.

[0070] Figure 8 is the structural schematic diagram of the speech recognition device provided by the embodiments of the present invention. As Figure 8As shown, the device includes a voice signal acquisition module 10 and a voice recognition module 20, where: The voice signal acquisition module 10 is used to: acquire a voice signal to be subjected to voice recognition; The voice recognition module 20 is used to: perform voice recognition on the voice signal based on an acoustic model and a language model; Wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network.

[0071] The voice recognition device provided by the embodiment of the present invention acquires a voice signal to be subjected to voice recognition, and performs voice recognition on the voice signal based on an acoustic model and a language model. Wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network, which can improve the calculation speed of the self-attention value, thereby improving the speed of the output result of the self-attention module of the improved Transformer structure, and improving the processing speed of the acoustic model and the language model constructed by using the improved Transformer structure, so as to improve the processing speed of voice recognition as a whole.

[0072] According to a voice recognition device provided by an embodiment of the present invention, when the voice recognition module 20 is used to perform voice recognition on the voice signal based on an acoustic model and a language model, it is specifically used to: perform acoustic processing on the voice signal to obtain acoustic features; use the acoustic features and the acoustic model to perform acoustic encoding to obtain a syllable sequence; perform word list matching on the syllable sequence and the word list to obtain a word sequence; input the word sequence into the language model to output a language decoding result.

[0073] The voice recognition device provided by the embodiment of the present invention constructs an acoustic model and a language model based on an improved Transformer structure, and performs voice recognition based on the acoustic model and the language model, which improves the processing speed of the acoustic model and the language model constructed by using the Transformer structure, thereby improving the processing speed of voice recognition as a whole.

[0074] According to a voice recognition device provided by an embodiment of the present invention, the self-attention value is expressed as: A = softmax(f(g(W))), and the calculation amount of f(g(W)) is less than Q = XW q ,K = XW k ,wherein, A represents the self-attention value, X represents the voice signal, W q 、W krepresents the transformation matrix, d k represents the transformation matrix W k The number of columns of; W represents the weight matrix of the neural network, g(W) represents a function of W, and f(g(W)) represents a function of g(W).

[0075] The speech recognition device provided by the embodiments of the present invention reduces the calculation amount of self-attention values by improving the object for normalization processing, improves the calculation speed, and realizes the simplicity of improving the speech recognition speed.

[0076] According to a speech recognition device provided by an embodiment of the present invention, g(W)=WX or g(W)=W; where X represents the speech signal.

[0077] The speech recognition device provided by the embodiments of the present invention further reduces the calculation amount of self-attention values and improves the calculation speed of self-attention values by making g(W)=WX or g(W)=W, thereby further improving the speech recognition speed.

[0078] According to a speech recognition device provided by an embodiment of the present invention, the weight matrix of the neural network is sparsified.

[0079] The speech recognition device provided by the embodiments of the present invention reduces the calculation complexity of self-attention values and improves the calculation speed of self-attention values by sparsifying the weight matrix of the neural network.

[0080] According to a speech recognition device provided by an embodiment of the present invention, the effective values of the weight matrix of the neural network exist at the diagonal and positions extending from the diagonal to both sides; and for each row of the weight matrix of the neural network, the sum of the number of effective values is less than the number of rows of the speech signal.

[0081] The speech recognition device provided by the embodiments of the present invention balances the improvement of calculation speed and the requirement of speech recognition accuracy by making the effective values of W exist at the diagonal and positions extending from the diagonal to both sides, and the sum of the number of effective values in each row is less than the number of rows of the speech signal.

[0082] Figure 9 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, such as Figure 9As shown in the figure, the electronic device may include: a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete communication with each other through the communication bus 1040. The processor 1010 may call logic instructions in the memory 1030 to execute a speech recognition method, which includes: obtaining a speech signal to be recognized; performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network.

[0083] In addition, when the logic instructions in the above-mentioned memory 1030 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0084] On the other hand, an embodiment of the present invention further provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech recognition method provided by the above-mentioned various methods. The method includes: obtaining a speech signal to be recognized; performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network.

[0085] In another aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a speech recognition method provided by the above-mentioned various methods is implemented. The method includes: obtaining a speech signal to be subjected to speech recognition; performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network.

[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech recognition method, characterized in that, Including: Obtain a speech signal to be subjected to speech recognition; Performing speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network; the self-attention value is expressed as: A = softmax(f(g(W))), and the computational complexity of f(g(W)) is less than Q = XW q , K = XW k , where A represents the self-attention value, X represents the speech signal, W q , W k represents a transformation matrix, d k represents the number of columns of the transformation matrix W k ; W represents the weight matrix of the neural network, g(W) represents a function of W, f(g(W)) represents a function of g(W); g(W) = WX or g(W) = W; where X represents the speech signal.

2. The voice recognition method according to claim 1, wherein The speech recognition of the speech signal based on an acoustic model and a language model includes: Perform acoustic processing on the speech signal to obtain acoustic features; use the acoustic features and the acoustic model to perform acoustic encoding to obtain a syllable sequence; perform vocabulary matching on the syllable sequence and a vocabulary to obtain a word sequence; input the word sequence into the language model to output a language decoding result.

3. The voice recognition method according to claim 1, characterized in that The weight matrix of the neural network has been sparsified.

4. The speech recognition method according to claim 1, characterized in that The valid values of the weight matrix of the neural network exist on the diagonal and the positions extending from the diagonal to both sides; and for each row of the weight matrix corresponding to the neural network, the sum of the number of valid values is less than the number of rows of the speech signal.

5. A voice recognition device, characterized in that, Including: A speech signal acquisition module, configured to: obtain a speech signal to be subjected to speech recognition; A speech recognition module, configured to: perform speech recognition on the speech signal based on an acoustic model and a language model; wherein, the acoustic model and the language model are constructed based on a Transformer structure with a modified self-attention module, and the self-attention value of the self-attention module in the Transformer structure with the modified self-attention module is obtained based on the weight matrix of a neural network; the self-attention value is expressed as: A = softmax(f(g(W))), and the computational complexity of f(g(W)) is less than Q = XW q ,K = XW k ,wherein, A represents the self-attention value, X represents the speech signal, W q 、W k represents a transformation matrix, d k represents the number of columns of the transformation matrix W k ; W represents the weight matrix of the neural network, g(W) represents a function of W, f(g(W)) represents a function of g(W); g(W) = WX or g(W) = W; wherein, X represents the speech signal.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the speech recognition method according to any one of claims 1 to 4 are implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 4 are implemented.

8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Improved Transformer + CRF (Content Recognition Function)-based method for identifying old rattle named entity

    CN111783459A