Speech recognition method, device, equipment, readable storage medium and product
By performing pronunciation analysis and text structure analysis of the target speech, combining cross attention processing technology, integrating pronunciation features and text features, the problem of low speech recognition accuracy in the existing technology is solved, and a higher speech recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111465598.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-03
AI Technical Summary
In the existing speech recognition technology, the accuracy of speech recognition is low because it only includes the mapping relationship between text and text.
By obtaining the target speech for pronunciation analysis and text structure analysis, the first vector sequence and the second vector sequence are obtained, and cross-attention processing is performed, and pronunciation characteristics and text character characteristics are fused to improve the accuracy of speech recognition.
Through the above method, the problem that the pre-trained speech recognition model cannot be analyzed at the semantic level can be avoided, and the concept of context information can be supplemented, and the accuracy of speech recognition can be significantly improved.
Smart Images

Figure CN114333772B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition, and in particular to a speech recognition method, apparatus, device, readable storage medium and product. Background Art
[0002] With the development of artificial intelligence technology, speech recognition technology has made great progress and has been applied to various fields.
[0003] In the related art, in the process of speech recognition, speech recognition data is usually annotated manually, the manually annotated data is applied to the speech recognition model, and the speech recognition result is obtained using the trained speech recognition model.
[0004] However, in the related art, since only the mapping relationship between words is included, the accuracy of speech recognition is reduced to a certain extent. Summary of the invention
[0005] The embodiments of the present application provide a speech recognition method, device, equipment, readable storage medium and product, which improve the accuracy of speech recognition to a certain extent. The technical solution is as follows:
[0006] In one aspect, a speech recognition method is provided, the method comprising:
[0007] Acquire a target speech, where the target speech is the speech to be subjected to speech-to-text recognition;
[0008] Performing pronunciation analysis on the target speech to obtain a first vector sequence, where the first vector sequence is used to indicate pronunciation features corresponding to the target speech;
[0009] Performing text structure analysis on a character sequence corresponding to the target speech to obtain a second vector sequence, wherein the second vector sequence is used to indicate character sequence features corresponding to text characters in the target speech, and the character sequence is a result obtained by recognition using a pre-trained speech recognition model;
[0010] The first vector sequence and the second vector sequence are subjected to cross-attention processing to obtain a speech-to-text recognition result corresponding to the target speech, wherein the cross-attention processing is used to fuse the pronunciation features and the character sequence features.
[0011] In another aspect, a speech recognition device is provided, the device comprising:
[0012] An acquisition module is used to acquire a target speech, where the target speech is the speech to be subjected to speech-to-text recognition;
[0013] An analysis module, configured to perform pronunciation analysis on the target speech to obtain a first vector sequence, wherein the first vector sequence is used to indicate pronunciation features corresponding to the target speech;
[0014] The analysis module is further used to perform text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence, where the second vector sequence is used to indicate the character sequence features corresponding to the text characters in the target speech;
[0015] A fusion module is used to perform cross-attention processing on the first vector sequence and the second vector sequence to obtain a speech-to-text recognition result corresponding to the target speech, wherein the cross-attention processing is used to fuse the pronunciation features and the character sequence features.
[0016] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a speech recognition method as described in any of the above-mentioned embodiments of the present application.
[0017] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a speech recognition method as described in any of the above-mentioned embodiments of the present application.
[0018] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the speech recognition method described in any of the above embodiments.
[0019] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:
[0020] During the speech recognition process, the cross-attention network is used to perform contextual semantic analysis of the target speech, and then combined with the speech features of the target speech, the pre-trained speech recognition model is assisted in speech recognition of the target speech, avoiding the problem that the pre-trained speech recognition model cannot perform semantic analysis of the target speech, supplementing the recognition concept of contextual information, and further improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 It is a structural diagram of a voice interaction system in a related art provided by an exemplary embodiment of the present application;
[0023] Figure 2 is a structural diagram of a language recognition model provided by an exemplary embodiment of the present application;
[0024] Figure 3 is a schematic diagram of an implementation environment involved in a speech recognition method provided by an exemplary embodiment of the present application;
[0025] Figure 4 This is a structural block diagram of a vehicle-mounted voice product provided by an embodiment of the present application;
[0026] Figure 5 is a flowchart of the steps of a speech recognition method provided by an exemplary embodiment of the present application;
[0027] Figure 6 is a flowchart of the steps of a speech recognition method provided by another exemplary embodiment of the present application;
[0028] Figure 7 is a flowchart of the steps of a speech recognition method provided by another exemplary embodiment of the present application;
[0029] Figure 8 is a structural block diagram of a speech recognition device provided by an exemplary embodiment of the present application;
[0030] Fig. 9 is a structural block diagram of a speech recognition device provided by another exemplary embodiment of the present application;
[0031] Fig.10 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the present application more clear, the present application is further described in detail below in conjunction with the accompanying drawings.
[0033] The working principle and implementation environment of a speech recognition method provided by this application are described as follows:
[0034] Specific combination and Figure 1A flow chart of the voice interaction technology in the related art is shown, wherein the voice interaction process includes a microphone array 101, an acoustic front-end algorithm 102, a cloud recognition algorithm 103, an offline recognition algorithm 104, a fusion algorithm 105, and offline / cloud semantic information 106, and the received target voice is subjected to voice-to-text recognition to obtain the final voice-to-text recognition result. The whole process mainly includes two parts: voice recognition and semantic understanding, wherein voice recognition is responsible for converting voice signals into text, and semantic understanding is responsible for understanding the intention corresponding to the target voice, as shown below. Figure 1 The main functions of each part are briefly introduced.
[0035] The speech recognition technology mainly includes an acoustic front-end algorithm 102 and a cloud recognition algorithm 103. The acoustic front-end algorithm 102 mainly includes noise reduction suppression, sound source localization, echo cancellation and other processing on the target speech signal received by the microphone array 101. The cloud recognition algorithm 103 acoustic model mainly models the mapping relationship between the target speech signal and the pronunciation unit, and mainly includes an acoustic model and a language model. The acoustic model and the language model are integrated with an encoder and a decoder. The target speech is recognized by speech text through the encoder and the decoder. It is mainly responsible for modeling the mapping relationship from the pronunciation unit to the Chinese character. The decoder algorithm is mainly combined with the cloud recognition algorithm 103 to perform the entire speech-to-text search process, that is, to complete the semantic understanding process between the characters in the target speech.
[0036] The offline recognition algorithm 104 is mainly used to perform speech-to-text recognition on the received target speech in an offline scenario, including a fixed wake-up word wake-up engine, a customizable wake-up word wake-up engine, and an offline speech recognition engine.
[0037] The recognition results of the target speech by the cloud recognition algorithm 103 and the offline recognition algorithm 104 are combined to calculate the fusion algorithm, and then the speech recognition result corresponding to the target speech is finally determined by the offline / cloud semantic information 106.
[0038] In the related art, the decoder and encoder in the speech recognition model obtained during the pre-training process only include a self-attention network, which results in no speech information being received during the pre-training process, and the pre-trained speech model being unable to parse the speech information. It is difficult for the decoder to be initialized using the speech recognition model obtained during the pre-training process, and the training samples are limited, making it impossible to expand more training samples, resulting in a low recognition accuracy rate for the final speech recognition model.
[0039] In an embodiment of the present application, the decoder algorithm is optimized, and the optimized decoder algorithm is applied to the speech recognition model to assist in semantic recognition, so as to facilitate the decoder algorithm to initialize the speech recognition model and improve the accuracy of the speech recognition model.
[0040] Optionally, the speech recognition method involved in the embodiment of the present application is applied to the encoder and decoder in the speech recognition model. Figure 1 The working principle of the speech recognition model shown is introduced in detail.
[0041] The speech recognition model mainly includes an encoder 201 and a decoder 202. Optionally, the speech recognition model can be implemented as an end-to-end speech recognition model.
[0042] The encoder 201 is used to receive the target speech, and perform pronunciation analysis on the target speech to extract the speech features corresponding to the target speech, and obtain a first vector sequence, which includes M sub-layers, M is a positive integer, and each sub-layer includes a self-attention network (Multi-head Self Attention) and a feedforward neural network (Feed Forward), wherein the self-attention network is used to calculate the weighted sum of all speech features in the target speech for each feature, wherein a normative integration unit module (Add&Norm) is provided after the self-attention network, and a normative integration unit module is also provided after the feedforward neural network layer, and the normative integration unit is used to integrate and add the output of the self-attention network; in the embodiment of the present application, the self-attention network performs key-value weight calculation on the target speech, that is, the speech features in a query (query) target speech signal are mapped to a series (key-value), the query and each key are similarly calculated to obtain the weight, and then the weight is normalized using the softmax function, and finally the weight and the corresponding key value are weighted and summed to obtain the first vector sequence corresponding to the target speech.
[0043] The decoder 202 is used to process the text character sequence corresponding to the target speech to obtain the speech-text recognition result corresponding to the target speech, which includes a vectorization layer (Embedding Layer) and N+1 sublayers, where N is a positive integer. The vectorization layer is used to perform vectorization processing on the text characters corresponding to the target speech. The first N sublayers also include a self-attention network and a feedforward neural network. The N+1 layer includes a cross-attention network (Multi-head Cross Attention) and a feedforward neural network, where the cross-attention network is mainly used to perform contextual semantic analysis on the character sequence corresponding to the target speech.
[0044] The first vector sequence output by the encoder 201 is used as the input of the N+1th layer of the decoder 202, and then combined with the second vector sequence processed by the first N layers of the decoder 202 to obtain the speech text recognition result corresponding to the target speech.
[0045] Optionally, the working process corresponding to the decoder 202 can be implemented as iterative calculation, that is, the result currently output by the speech recognition model is used as the input for the next recognition process, so as to better achieve the purpose of parsing the context relationship corresponding to each text character in the target speech. For example, if the current output is "今" and it is used as the input for the next recognition, then in the next round of recognition process, the probability of outputting "天" is higher than that of outputting "田", and the speech text recognition result corresponding to the target speech is "今天".
[0046] In the embodiments of the present application, this method can be applied to the training process of the language recognition model, or can be directly applied to the speech recognition scenario for speech recognition. The present application does not limit this.
[0047] Combined Figure 3 The implementation environment involved in a speech recognition method shown in the present application will be described. As Figure 3 shown, the implementation environment includes a terminal device 301, a communication network 303, and a server 302. The server 303 integrates a language recognition model, including a decoder and an encoder in the language recognition model. The communication network can be implemented as a wired communication network or a wireless communication network. The present application does not limit this.
[0048] The user selects a target speech in the terminal device 301 for speech recognition and triggers a recognition instruction for the target speech.
[0049] The terminal device 301 receives the target speech and the speech recognition instruction. The target speech includes at least one of a voice segment recorded on site, an audio segment corresponding to a film and television segment, a music work, a weather forecast, a navigation voice, a voice corresponding to an online video / phone call, and a local voice. The speech recognition instruction is used to indicate text recognition of the received target speech. The speech recognition instruction can be triggered by a control displayed in the interface or by a voice wake-up method. The present application does not limit this.
[0050] The terminal device 301 uploads the target voice and voice recognition instructions to the server 302 through the communication network 303. The language recognition model in the server 302 uses the encoder and the decoder to perform text recognition on the target voice according to the voice recognition instructions, obtains the voice-text recognition result corresponding to the target voice, and displays the voice-text result around the area where the target voice is located. Exemplarily, when the target voice is displayed in the form of a conversation bubble on the current interface, the user long presses the conversation bubble to superimpose a selection option interface on the display area around the conversation bubble. The selection option interface is used to process the conversation bubble, including but not limited to a send option, a voice-to-text option, and a delete option. The voice-to-text option in the selection option interface is used to click on the voice-to-text option in the selection option interface. The terminal executes a voice-to-text event for the currently selected target voice. After the voice-to-text conversion is completed, the conversion result is displayed below the conversation bubble.
[0051] It should be noted that the execution subject of the above method may be the terminal device 301 , or the server 302 , or an interactive system of the terminal device 301 and the server 302 .
[0052] Illustratively, after the user inputs a voice command to the terminal device 301, the terminal device 301 sends the voice command to the server 302 for recognition. The language recognition model is integrated into the voice recognition framework in the server, and the target voice is recognized through the voice recognition framework. This application does not limit the implementation environment and execution subject of this method.
[0053] The above-mentioned terminal can be a mobile phone, a tablet computer, a desktop computer, a portable laptop computer, a smart TV, a smart home device, a car terminal and other terminal devices in various forms, and the embodiments of the present application are not limited to this.
[0054] It is worth noting that the above-mentioned servers can be independent physical servers, or they can be server clusters or distributed systems composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.
[0055] Among them, cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to realize data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which is used on demand and flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0056] In some embodiments, the above-mentioned server can also be implemented as a node in a blockchain system. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0057] Secondly, the application scenarios involved in the embodiments of the present application are briefly introduced:
[0058] Apply the speech recognition method provided in the embodiment of this application to travel scenarios, combine it with the Internet of Vehicles application, and improve the speech interaction system in travel scenarios. Figure 4 To explain, Figure 4 A structural block diagram of a voice product in a vehicle networking application scenario provided by an embodiment of the present application is shown. The above method is used in a voice recognition acoustic model, or the above method is used in the training process of a voice recognition acoustic model, and the voice recognition process serves the scenario of vehicle networking voice interaction.
[0059] The vehicle-mounted voice product includes a vehicle-mounted noise reduction module 401, a vehicle-mounted voice engine module 402, and a vehicle-mounted skill ecosystem module 403, wherein the vehicle-mounted noise reduction module 401 mainly performs noise reduction and echo elimination on the received voice signal, and can be used for noise suppression of wind noise, tire noise, music noise, and air conditioning noise, so that users can chat in the car; the vehicle-mounted voice engine module 402 is mainly for voice recognition and semantic understanding of the received voice signal, including a voice wake-up engine, cloud voice recognition, cloud semantic understanding, and an offline voice recognition engine; the vehicle-mounted skill ecosystem module 403 includes the type of received voice signal, that is, the voice signal can be realized as music, radio, news, navigation, surrounding food, telephone, car control, weather, etc. The cutting-edge technologies involved in the entire voice product include full-duplex, multi-tone zones, voiceprint recognition, and virtual people, and the embodiments of this application do not elaborate on the cutting-edge technologies involved.
[0060] It can be understood that the speech recognition method provided in the embodiment of the present application is not only applicable to vehicle-connected application scenarios, but can also be applied to any speech recognition scenarios, and the present application does not limit the application scenarios.
[0061] See also Figure 5 , Figure 5 is a flow chart of a speech recognition method provided in an embodiment of the present application, wherein the speech recognition method is applied to Figure 3 The terminal device 301 in the illustrated implementation environment is used as an example for explanation. The computer device includes a speech recognition model for speech recognition. The speech recognition model includes an encoder and a decoder, and includes the following steps.
[0062] Step 501, obtaining target speech.
[0063] In the embodiment of the present application, the recognition triggering method for the target voice includes but is not limited to:
[0064] First, an option control for identifying audio clips is provided in the application or online web page itself. When a selection operation is received through the selection control, a recognition control (option) is generated for converting speech to text of the audio content. In response to receiving a trigger operation on the recognition control, the terminal collects the target voice through a microphone or other audio collection device. For example, when a user browses an application or an online web page, he long presses the voice control and speaks, thereby collecting and acquiring the target voice.
[0065] Second, the application or online web page itself is used for speech recognition, that is, when the user wants to perform speech recognition on a certain audio content, he opens the application or online web page and uploads the target speech; optionally, the application is also suitable for offline speech recognition scenarios.
[0066] Optionally, the target voice is used to indicate the voice to be recognized by speech-to-text, including but not limited to voice clips recorded on site, audio clips corresponding to film and television clips, musical works, weather reports, navigation voices, voices corresponding to online videos / phone calls, and local recordings.
[0067] Step 502: Perform pronunciation analysis on the target speech to obtain a first vector sequence.
[0068] The first vector sequence is used to indicate the pronunciation features corresponding to the target speech.
[0069] Optionally, the target speech is analyzed for pronunciation by an encoder. The encoder receives the target speech and performs encoding processing on the target speech, wherein the encoder mainly includes M sublayers, M is a positive integer, and each sublayer includes a self-attention network and a feedforward neural network, wherein a standard integration unit module is also provided after the self-attention network, and a standard integration unit module is also provided after the feedforward neural network. The encoding process is as follows: in the first step, the encoder extracts the speech features corresponding to the target speech, and the speech features include the pronunciation of the target speech; in the second step, the speech features corresponding to the target are encoded using the self-attention network and the feedforward neural network to obtain a first vector sequence corresponding to the target speech, and the first vector sequence is used to indicate the pronunciation features corresponding to the target speech; exemplarily, the decoder receives the target speech a, performs encoding processing on the target speech, and obtains a first vector sequence [c1, c2, c3, c4, c5], where c1 is used to indicate the target The pronunciation unit of the character in the speech, wherein c1 can be used to represent the speech segment corresponding to a complete character (c1 is used to represent the entire speech segment corresponding to "I"), or c1 is used to represent one of the speech segments that constitute a complete character pronunciation (c1 and c2 are combined to form the speech segment of "I"). It should be noted that in the process of the encoder encoding the target speech, due to the different lengths of the target speech segments, the number of vectors in the first vector sequence that may be obtained is greater than the number corresponding to the speech segments of each character in the target speech. For example, if the target speech is "I am fine now", the first vector sequence that may be obtained includes five vectors, and these five vectors correspond to the speech segments of each of the five characters in the target speech; or, the first vector sequence obtained includes ten vectors, and these ten vectors form the speech segments corresponding to each of the five characters.
[0070] Step 503: Perform text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence.
[0071] Optionally, after receiving the target speech, the speech recognition model obtained in the pre-training process is used to perform speech-to-text recognition on the target speech to obtain a character sequence corresponding to the target speech. Exemplarily, the speech recognition model in the pre-training process is used to perform text recognition on the target speech [t1, t2, t3, t4, t5] to obtain a character sequence [x1, x2, x3, x4, x5].
[0072] The decoder performs text structure analysis on the character sequence, where the text structure analysis is used to indicate that contextual relationships between characters in the character sequence are analyzed.
[0073] Optionally, the decoder includes a vectorization layer and N+1 sublayers, where N is a positive integer, the first N sublayers include a self-attention network and a feedforward neural network, wherein a canonical integration unit module is arranged after the self-attention network, and a canonical integration unit module is also arranged after the feedforward neural network, and the N+1th sublayer includes a cross-attention network, wherein a canonical integration unit module is also arranged after the cross-attention network.
[0074] Optionally, the character sequence corresponding to the target speech is input into the decoder and converted into a vector sequence to obtain a mid-point vector sequence corresponding to the character features. For example, the character sequence [x1, x2, x3, x4, x5] of the target speech [t1, t2, t3, t4, t5] is vectorized to obtain a mid-point vector sequence [e1, e2, e3, e4, e5].
[0075] In an embodiment of the present application, the self-attention network in the first N sub-layers in the decoder performs self-attention processing on the aforementioned vector sequence [e1, e2, e3, e4, e5] to obtain a second vector sequence [o1, o2, o3, o3, o4, o5], wherein the self-attention processing includes self-attention calculation and feedforward neural network encoding.
[0076] Optionally, the first N sublayers in the decoder are used to perform text structure analysis on the character sequence of the target speech, including: inputting the character sequence into the first N sublayers, performing text analysis on the character sequence through the self-attention network and feedforward neural network of each sublayer, and obtaining a second vector sequence, which is used to indicate character sequence features corresponding to the text characters in the target speech.
[0077] Optionally, the N+1th layer in the decoder is used to perform a comprehensive analysis of the pronunciation feature structure in the target speech to obtain the contextual semantic relationship between each character in the target speech. For the specific workflow of the N+1th sublayer, please refer to the following step 504.
[0078] Optionally, after the first N sub-layers output the candidate vector sequence, the encoder calculates the similarity between each vector in the candidate vector sequence and the candidate vector sequence, normalizes and integrates the similarity, and determines the vector after the normalization and integration as the second vector feature. It should be noted that the self-attention network in the first N sub-layers mainly analyzes the contextual features between the character features obtained by the speech recognition model, and the method provided in the present application is to perform a secondary combination analysis of the text characters and the target speech whose contextual features have been determined, to obtain the contextual semantic features between the target speech and each character, thereby effectively improving the speech recognition accuracy.
[0079] Optionally, the execution order of step 502 and step 503 can be parallel or sequential, and the sequential execution includes executing step 502 first and then step 503, or executing step 503 first and then step 502, which is not limited in this application.
[0080] Step 504: perform cross-attention processing on the first vector sequence and the second vector sequence to obtain a speech-to-text recognition result corresponding to the target speech.
[0081] Optionally, the first vector sequence corresponding to the target speech obtained by the encoder is input into the N+1th sublayer of the decoder, and combined with the second vector sequence corresponding to the character features obtained by the first N layers to obtain a vector sequence corresponding to the recognition result of the target speech.
[0082] Optionally, the execution process of inputting the first vector sequence and the second vector sequence into the N+1th sublayer includes the following steps:
[0083] Determine the similarity relationship between the first vector sequence and the second vector sequence; normalize the similarity relationship to obtain a cross vector sequence of text characters corresponding to the second vector sequence in the first vector sequence, the cross vector sequence is used to indicate the contextual relationship between the i-th text character and the i+1-th text character in the target speech, i being a positive integer; feed-forward encode the cross vector sequence to obtain a semantic joint vector corresponding to the target speech; perform probability prediction on the semantic joint vector to obtain a probability prediction result corresponding to the character sequence; and obtain a speech recognition result corresponding to the target speech based on the probability prediction result. Exemplarily, the first vector sequence [c1, c2, c3, c4, c5] and the second vector sequence [o1, o2, o3, o3, o4, o5] are normalized for similarity to obtain the cross-vector sequence u1, u2, u3, u4, u5 of each attention o1, o2, o3, o3, o4, o5 in the first vector sequence [c1, c2, c3, c4, c5], and the cross-vector sequence u1, u2, u3, u4, u5 are feed-forward encoded to obtain the semantic joint vector [r1, r2, r3, r4, r5], where the vector corresponding to the character feature x2 is o2, then the cross-vector u2 corresponding to the forward character vector of x2 is used to represent the importance of character feature x2 for predicting character feature x3, that is, to represent the contextual relationship between character feature x2 and character feature x3.
[0084] Optionally, the decoder also includes a linear layer, which includes a classifier softmax. The linear layer is located after the N+1th sublayer, and uses the aforementioned semantic joint vector of the classifier to perform classification (probability) prediction to obtain the probability prediction results between each character in the target speech, and obtain the speech recognition result of the target speech based on the probability prediction results.
[0085] To summarize, the speech recognition method provided in the embodiment of the present application, during the speech recognition process, utilizes a cross-attention network to perform contextual semantic analysis of the target speech, and then combines the speech features of the target speech to assist the pre-trained speech recognition model in performing speech recognition on the target speech, thereby solving the problem that the pre-trained speech recognition model is unable to perform semantic-level analysis on the target speech, supplements the recognition concept of contextual information, and further improves the accuracy of speech recognition.
[0086] The speech recognition method provided in this embodiment retains the self-attention network in the first N sublayers of the decoder and the cross-attention network in the N+1th sublayer of the decoder. When the decoder is used for speech recognition, firstly, the decoder algorithm (structure) can support the function of initializing the language recognition model obtained during the pre-training process; secondly, compared with the decoder structure that includes both the self-attention network and the cross-attention network, the amount of parameters used by the decoder to calculate the character sequence is reduced, thereby effectively improving the recognition accuracy of the target speech.
[0087] See also Figure 6 , Figure 6 is a flowchart of another speech recognition method provided in an embodiment of the present application, the speech recognition method is applied to Figure 3 In the terminal device 301 in the illustrated implementation environment, the computer device includes a speech recognition model for speech recognition, and the speech recognition model includes an encoder and a decoder, including the following steps.
[0088] Step 601, obtaining target speech.
[0089] In an embodiment of the present application, the target voice is used to indicate the voice to be recognized by voice-to-text, including but not limited to voice clips recorded on site, audio clips corresponding to film and television clips, musical works, weather reports, navigation voices, voices corresponding to online videos / phone calls, and local recordings.
[0090] The execution process of this step is the same as step 501 and will not be repeated here.
[0091] Step 602: Perform pronunciation analysis on the target speech to obtain a first vector sequence.
[0092] Optionally, the encoder receives the target speech and performs encoding processing on the target speech, wherein the encoder mainly includes M sublayers, M is a positive integer, and each sublayer includes a self-attention network and a feedforward neural network, wherein the self-attention network and the feedforward neural network both include a standard integration unit module, and the encoding processing process is as follows: in the first step, the encoder extracts speech features corresponding to the target speech, and the speech features include the pronunciation of the target speech; in the second step, the speech features corresponding to the target are encoded using the self-attention network and the feedforward neural network to obtain a first vector sequence corresponding to the target speech, and the first vector sequence is used to indicate the pronunciation features corresponding to the target speech.
[0093] The execution process of this step is the same as step 502 and will not be repeated here.
[0094] Step 603: Perform text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence.
[0095] Optionally, after receiving the target voice, use the speech recognition model obtained during the pre-training process to perform speech text recognition on the target voice to obtain the character sequence corresponding to the target voice.
[0096] The execution process of this step is the same as that of step 503, and will not be elaborated here.
[0097] Step 604: Perform cross-attention processing on the first vector sequence and the second vector sequence to obtain the speech text recognition result corresponding to the target voice.
[0098] Optionally, input the first vector sequence corresponding to the target voice processed by the encoder into the (N + 1)-th sub-layer in the decoder, and combine it with the second vector sequence corresponding to the character features obtained from the previous N layers to obtain the semantic joint vector corresponding to the recognition result of the target voice.
[0099] In the embodiment of the present application, the semantic joint vector is input into the linear layer including the classifier softmax in an iterative calculation manner for probability prediction. The b-th output in the semantic joint vector is used as the input for the (b + 1)-th prediction. b is a positive integer. After all the vector probability predictions in the semantic joint vector are completed, the first probability prediction result corresponding to the target voice is determined.
[0100] Use the speech recognition model to perform probability prediction on the speech features (the first vector sequence) corresponding to the target voice to obtain the second probability prediction result corresponding to all the characters of the target voice.
[0101] Based on the first probability prediction result and the second probability prediction result, integrate and determine the target recognition result corresponding to the target voice, and feedback the target recognition result to the terminal display interface.
[0102] Optionally, integrating and determining the target recognition result corresponding to the target voice includes: performing weighted averaging on the first probability prediction result and the second probability prediction result, and using the result corresponding to the probability prediction value greater than or equal to the weighted average probability prediction value as the final recognition result value of the target voice; or, both the first probability prediction result and the second probability prediction result are used to represent at least two texts corresponding to each character in the target voice. For example, for the first character of the target voice, there are three recognized texts (jie, jie, jie), and the three texts "jie", "jie", and "jie" correspond to probability prediction values, which are used to represent the probability that the current character is the target text. Determine the character corresponding to the current target voice by using the text with a probability greater than the preset threshold, and determine the first character of the target voice as "jie".
[0103] To summarize, the speech recognition method provided in the embodiment of the present application, during the speech recognition process, utilizes a cross-attention network to perform contextual semantic analysis of the target speech, and then combines the speech features of the target speech to assist the pre-trained speech recognition model in performing speech recognition on the target speech, thereby solving the problem that the pre-trained speech recognition model is unable to perform semantic-level analysis on the target speech, supplements the recognition concept of contextual information, and further improves the accuracy of speech recognition.
[0104] See also Figure 7 , Figure 7 is a flowchart of another speech recognition method provided in an embodiment of the present application, the speech recognition method is applied to Figure 3 In the terminal device 301 in the illustrated implementation environment, the computer device includes a speech recognition model for speech recognition, and the speech recognition model includes an encoder and a decoder, including the following steps.
[0105] Step 701, obtaining a first speech recognition result corresponding to a target speech.
[0106] In an embodiment of the present application, the target voice is used to indicate the voice to be recognized by voice-to-text, including but not limited to voice clips recorded on site, audio clips corresponding to film and television clips, musical works, weather reports, navigation voices, voices corresponding to online videos / phone calls, and local recordings.
[0107] Optionally, a pre-trained speech recognition model is used to perform speech recognition on the target speech to obtain a first speech recognition result corresponding to the target speech, and the first speech recognition result includes pronunciation features and character features corresponding to the target speech.
[0108] Step 702, obtaining a second speech recognition result corresponding to the target speech.
[0109] In an embodiment of the present application, the speech recognition method provided in the embodiment of the present application is used to optimize the decoder structure; first, the encoder is used to perform pronunciation analysis on the target speech to obtain a first vector sequence. Please refer to step 502 for details of this step, which will not be repeated here; secondly, the first N sub-layers in the decoder are used to perform text structure analysis on the character features of the target speech to obtain a second vector sequence. Please refer to step 503 for details of this step, which will not be repeated here.
[0110] The first vector sequence is input into the N+1th sublayer of the decoder, and then combined with the second vector sequence, the probability prediction result corresponding to the target speech is obtained through the classifier, and the second speech recognition result corresponding to the target speech is determined based on the probability prediction result. For details of this step, please refer to step 504, which will not be repeated here.
[0111] Step 703: Determine a target recognition result corresponding to the target speech based on the first speech recognition result and the second speech recognition result.
[0112] Optionally, the first speech recognition result corresponds to a first weight value, and the second speech recognition result corresponds to a second weight value. Based on the first weight value, a first intermediate speech recognition result corresponding to the first speech recognition result is determined. Based on the second weight value, a second intermediate speech recognition result corresponding to the second speech recognition result is determined. Finally, based on the first intermediate speech recognition result and the second intermediate speech recognition result, a target recognition result corresponding to the target speech result is determined. In the embodiment of the present application, the speech recognition model obtained during the pre-training process is CTC, and the speech recognition result of the target speech is jointly determined in combination with the encoder and decoder structures provided in the embodiment of the present application. The first weight value occupied by CTC is set to 0.6, and the second weight value occupied by the encoder and decoder after the section is set to 0.4, where the value of M in the encoder is 12, and the value of N in the decoder is 6.
[0113] Optionally, the above-mentioned process of jointly determining the speech recognition result of the target speech can be applied to the training process of the speech recognition model. During the training process, the first weight value occupied by CTC is set to 0.4, and the first weight values occupied by the encoder and decoder are set to 0.6.
[0114] Optionally, the above-mentioned process of jointly determining the speech recognition result of the target speech is applied to the training process and the actual application process at the same time. During training, the first weight value corresponds to the weight value of CTC in the training process, which is 0.3. During decoding, the second weight value corresponds to the weight value of CTC in the decoding process. The speech recognition model is optimized and the parameters are adjusted based on the first weight value and the second weight value. This can not only improve the speech recognition model in the training stage to have more sufficient training samples without relying on manual annotation of data, but also improve the speech recognition accuracy of the speech recognition model. In addition, during the training process, the loss calculation of the speech recognition result can be performed based on the first weight value and the second weight value, and the parameters of the speech recognition model can be adjusted based on the loss calculation.
[0115] In the implementation of the present application, the decoder algorithm provided in the embodiment of the present application can also be used to initialize the speech recognition model, that is, the speech recognition model is initialized using the iterative calculation process of the decoder of N sub-layers, which includes determining the first parameter value corresponding to the target speech based on the joint semantic vector of the first vector sequence and the second vector sequence, and the first parameter value is used to indicate the parameter value corresponding to the speech recognition model after recognizing the target speech. The speech recognition model obtained in the pre-training process is initialized based on the first parameter value. First, a certain voice file support is provided for the speech recognition model, which can combine the contextual relationship between the character features and the target speech itself to improve the robustness of the speech recognition model; secondly, an initialization operation is provided to avoid the speech recognition model error becoming larger and larger, and to ensure that the speech recognition result will not have a large deviation. Optionally, the iterative calculation process includes: using the speech recognition model to recognize the target speech to obtain a character sequence, inputting the i-th text character in the character sequence into the N-layer sub-layer of the decoder for text structure analysis, and after the N-th sub-layer outputs the text character, the output text character is used as the input for outputting the i+1-th text character.
[0116] It should be noted that the decoder structure provided in the embodiment of the present application can perform separate initialization operations on other speech recognition models, or the decoder structure can be placed in a speech recognition model to perform initialization operations on itself. The present application does not limit this. In addition, after initialization using the decoder structure, in the process of recognizing the target speech, the amount of parameters required to calculate the speech recognition model is reduced to a certain extent, and the error rate of the speech recognition process is also reduced.
[0117] Optionally, the first N sublayers in the decoder also support initialization operations, which use the speech recognition model obtained during the pre-training process to initialize the decoder, re-perform semantic analysis on the character features of the target speech based on the initialized decoder, and then combine with the first vector sequence to obtain the speech text recognition result.
[0118] To summarize, the speech recognition method provided in the embodiment of the present application, during the speech recognition process, utilizes a cross-attention network to perform contextual semantic analysis of the target speech, and then combines the speech features of the target speech to assist the pre-trained speech recognition model in performing speech recognition on the target speech, thereby solving the problem that the pre-trained speech recognition model is unable to perform semantic-level analysis on the target speech, supplements the recognition concept of contextual information, and further improves the accuracy of speech recognition.
[0119] See also Figure 8 , Figure 8 This is a structural block diagram of an exploration potential evaluation device provided by an embodiment of the present application, and the device includes:
[0120] An acquisition module 801 is used to acquire a target speech, where the target speech is the speech to be subjected to speech-to-text recognition;
[0121] An analysis module 802 is used to perform pronunciation analysis on the target speech to obtain a first vector sequence, where the first vector sequence is used to indicate pronunciation features corresponding to the target speech;
[0122] The analysis module 802 is further used to perform text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence, where the second vector sequence is used to indicate the character sequence features corresponding to the text characters in the target speech;
[0123] The fusion module 803 is used to perform cross-attention processing on the first vector sequence and the second vector sequence to obtain a speech-to-text recognition result corresponding to the target speech, and the cross-attention processing is used to fuse the pronunciation features and the character sequence features.
[0124] In an alternative embodiment, see Fig. 9 , the device also includes a determination module 804.
[0125] The determining module 804 is used to determine a similarity relationship between the first vector sequence and the second vector sequence;
[0126] The analysis module 802 is used to normalize the similarity relationship to obtain a cross vector sequence of text characters corresponding to the second vector sequence in the first vector sequence, wherein the cross vector sequence is used to indicate a contextual relationship between the i-th text character and the i+1-th text character in the target speech, where i is a positive integer;
[0127] The recognition module 805 is used to recognize the cross vector sequence to obtain the speech text recognition result corresponding to the target speech.
[0128] In an alternative embodiment, see Fig. 9 , the determination module 804 is further used to perform feed-forward encoding on the cross vector sequence to obtain a semantic joint vector corresponding to the target speech;
[0129] A prediction module 806 is used to perform probability prediction on the semantic joint vector to obtain a probability prediction result corresponding to the character sequence;
[0130] The recognition module 805 is further configured to obtain a speech recognition result corresponding to the target speech based on the probability prediction result.
[0131] In an alternative embodiment, see Fig. 9The prediction module 806 is used to perform probability prediction on the semantic joint vector to obtain a first probability prediction result corresponding to the character sequence; perform probability prediction on the first vector sequence to obtain a second probability prediction result; and combine the first probability prediction result and the second probability prediction result to obtain the probability prediction result.
[0132] In an alternative embodiment, see Fig. 9 The acquisition module 801 is further used to acquire a first speech recognition result corresponding to the target speech, where the first speech recognition result is a result obtained by recognizing a speech recognition model obtained through pre-training;
[0133] The fusion module 803 is further used to determine a second speech recognition result corresponding to the target speech based on the first vector sequence and the second vector sequence;
[0134] The recognition module 805 is further configured to determine a target recognition result corresponding to the target speech based on the first speech recognition result and the second speech recognition result.
[0135] In an optional embodiment, the first speech recognition result corresponds to a first weight value, and the second speech recognition result corresponds to a second weight value;
[0136] The recognition module 805 is also used to determine a first intermediate speech recognition result corresponding to the first speech recognition result based on the first weight value; determine a second intermediate speech recognition result corresponding to the second speech recognition result based on the second weight value; and determine the target recognition result corresponding to the target speech based on the first intermediate speech recognition result and the second intermediate speech recognition result.
[0137] In an optional embodiment, the determination module 804 is further used to extract speech features in the target speech; perform self-attention processing and feedforward encoding on the speech features to obtain the first vector sequence.
[0138] In an optional embodiment, the determination module 804 is further used to perform vectorization processing on the character sequence to obtain a mid-point vector feature corresponding to the character sequence; and perform self-attention processing on the mid-point vector feature to obtain the second vector sequence.
[0139] To summarize, the speech recognition device provided in the embodiment of the present application, during the speech recognition process, utilizes a cross-attention network to perform contextual semantic analysis on the target speech, and then combines the speech features of the target speech to assist the pre-trained speech recognition model in performing speech recognition on the target speech, thereby solving the problem that the pre-trained speech recognition model is unable to perform semantic-level analysis on the target speech, supplements the recognition concept of contextual information, and further improves the accuracy of speech recognition.
[0140] It should be noted that the speech recognition device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the speech recognition device provided in the above embodiment and the speech recognition method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0141] Fig.10 The structure block diagram of a terminal device 301 provided by an exemplary embodiment of the present application is shown. The terminal device 301 may be a portable mobile terminal, such as a smart phone, a tablet computer, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal device 301 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0142] Typically, the terminal device 301 includes: a processor 1001 and a memory 1002 .
[0143] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0144] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one instruction, which is used to be executed by the processor 1001 to implement the web page embedding method provided in the method embodiment of the present application.
[0145] In some embodiments, the terminal device 301 may also optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002 and the peripheral device interface 1003 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 1003 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, a positioning assembly 1008 and a power supply 1009.
[0146] Those skilled in the art will understand that Fig.10 The structure shown in the figure does not constitute a limitation on the terminal device 301, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0147] The above-mentioned memory also includes one or more programs, and the one or more programs are stored in the memory and configured to be executed by the CPU.
[0148] An embodiment of the present application also provides a computer device, which includes a processor and a memory, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to implement the speech recognition method provided by the above-mentioned method embodiments.
[0149] An embodiment of the present application also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the speech recognition method provided by the above-mentioned method embodiments.
[0150] Optionally, the computer readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a solid state drive (SSD), or an optical disk. Among them, the random access memory may include a resistance random access memory (ReRAM) and a dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0151] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0152] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A speech recognition method, It is characterized in that The method comprises: Acquire a target speech, where the target speech is the speech to be subjected to speech-to-text recognition; Performing pronunciation analysis on the target speech to obtain a first vector sequence, where the first vector sequence is used to indicate pronunciation features corresponding to the target speech; Performing text structure analysis on a character sequence corresponding to the target speech to obtain a second vector sequence, wherein the second vector sequence is used to indicate character sequence features corresponding to text characters in the target speech, and the character sequence is a result obtained by recognition using a pre-trained speech recognition model; Determining a similarity relationship between the first vector sequence and the second vector sequence; Normalizing the similarity relationship to obtain a cross vector sequence of text characters corresponding to the second vector sequence in the first vector sequence, wherein the cross vector sequence is used to indicate a contextual relationship between an i-th text character and an i+1-th text character in the target speech, where i is a positive integer; The cross vector sequence is identified to obtain a speech text recognition result corresponding to the target speech.
2. The method according to claim 1, It is characterized in that The step of recognizing the cross vector sequence to obtain a speech text recognition result corresponding to the target speech includes: Performing feed-forward encoding on the cross vector sequence to obtain a semantic joint vector corresponding to the target speech; Probabilistically predicting the semantic joint vector to obtain a probability prediction result corresponding to the character sequence; Based on the probability prediction result, a speech recognition result corresponding to the target speech is obtained.
3. The method according to claim 2, It is characterized in that The performing probability prediction on the semantic joint vector to obtain a probability prediction result corresponding to the character sequence includes: Probabilistically predicting the semantic joint vector to obtain a first probability prediction result corresponding to the target speech; Performing probability prediction on the first vector sequence to obtain a second probability prediction result; The first probability prediction result and the second probability prediction result are combined to obtain the probability prediction result.
4. The method according to any one of claims 1 to 3, It is characterized in that The step of identifying the cross vector sequence to obtain a speech text recognition result corresponding to the target speech includes: Obtaining a first speech recognition result corresponding to the target speech, where the first speech recognition result is a result obtained by recognizing a speech recognition model obtained through pre-training; Identify the cross vector sequence to obtain a second speech recognition result corresponding to the target speech; The first speech recognition result and the second speech recognition result are combined to obtain a speech-to-text recognition result corresponding to the target speech.
5. The method according to claim 4, It is characterized in that The first speech recognition result corresponds to a first weight value, and the second speech recognition result corresponds to a second weight value; The combining the first speech recognition result and the second speech recognition result to obtain a speech-to-text recognition result corresponding to the target speech includes: Determining, based on the first weight value, a first intermediate speech recognition result corresponding to the first speech recognition result; Determining, based on the second weight value, a second intermediate speech recognition result corresponding to the second speech recognition result; The first intermediate speech recognition result and the second intermediate speech recognition result are combined to obtain the speech-to-text recognition result corresponding to the target speech.
6. The method according to any one of claims 1 to 3, It is characterized in that The performing pronunciation analysis on the target speech to obtain a first vector sequence includes: Extracting speech features in the target speech; The speech features are subjected to self-attention processing and feed-forward encoding to obtain the first vector sequence.
7. The method according to any one of claims 1 to 3, It is characterized in that The performing text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence includes: Performing vectorization processing on the character sequence to obtain a vector feature corresponding to the character sequence; Perform self-attention processing on the mid-turn vector features to obtain the second vector sequence.
8. A speech recognition device, It is characterized in that The device comprises: An acquisition module is used to acquire a target speech, where the target speech is the speech to be subjected to speech-to-text recognition; An analysis module, configured to perform pronunciation analysis on the target speech to obtain a first vector sequence, wherein the first vector sequence is used to indicate pronunciation features corresponding to the target speech; The analysis module is further used to perform text structure analysis on the character sequence corresponding to the target speech to obtain a second vector sequence, where the second vector sequence is used to indicate the character sequence features corresponding to the text characters in the target speech; A determination module, configured to determine a similarity relationship between the first vector sequence and the second vector sequence; The analysis module is further used to normalize the similarity relationship to obtain a cross vector sequence of text characters corresponding to the second vector sequence in the first vector sequence, wherein the cross vector sequence is used to indicate a contextual relationship between an i-th text character and an i+1-th text character in the target speech, where i is a positive integer; The recognition module is used to recognize the cross vector sequence and obtain the speech text recognition result corresponding to the target speech.
9. A computer device, It is characterized in that The computer device includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the speech recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, It is characterized in that The storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the speech recognition method according to any one of claims 1 to 7.
11. A computer program product, It is characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the speech recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition method and device based on self-attention mechanism and memory network
CN112599122A
Blind person navigation method and device based on auxiliary information
CN113091747A