Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3results about How to "Improve Speech Recognition Efficiency" patented technology

Speech recognition method, apparatus, device, and storage medium

The present application relates to the field of artificial intelligence and the field of financial technology, and discloses a speech recognition method, device and equipment and a storage medium, the method comprising: splicing a first embedding vector and a second embedding vector to generate a third embedding vector, reconstructing the third embedding vector to generate a fourth embedding vector; obtaining a predicted speech recognition text output by a speech recognition model based on the fourth embedding vector, obtaining a first loss value and a second loss value between the predicted speech recognition text and a preset speech recognition text; generating a total loss value of the predicted speech recognition text according to the first loss value and the second loss value, training the speech recognition model based on the total loss value; obtaining a current speech and a current field label sent by a customer service system, inputting the current speech and the current field label into the trained speech recognition model, and obtaining a current speech recognition text output by the trained speech recognition model. The present application is beneficial to improving the efficiency of speech recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition method and device, electronic equipment and storage medium

ActiveCN117219063Breduce the number of elementsReduce the amount of decoding calculationsPrediction probabilitySpeech sound
Embodiments of the present application provide a speech recognition method and device, electronic equipment and storage medium, at least applied to the field of artificial intelligence and speech recognition, wherein the method comprises: performing vector coding processing on the audio feature vector of the speech to be recognized to obtain an audio coding vector; performing classification processing on the audio coding vector to obtain a prediction probability distribution of each predicted character in a preset vocabulary corresponding to each speech frame in the speech to be recognized; performing pruning processing on the audio coding vector based on the prediction probability distribution to obtain a pruned audio coding vector; and performing speech recognition on the speech to be recognized based on the pruned audio coding vector to obtain a speech recognition result. Through the present application, the decoding calculation amount in the speech recognition process can be reduced, the decoding efficiency can be improved, and thus the speech recognition efficiency can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1

An online conference translation method and system

ActiveCN121480532BImprove Speech Recognition EfficiencyReduce identification uncertaintyNatural language translationSpeech recognitionLanguage speechSpeech sound
The application discloses an online conference translation method and system, and belongs to the technical field of speech recognition. The language space is constrained by the interactive side signal, and the recognition uncertainty is reduced. In the recording stage, a target language label generated by a sliding gesture is introduced. The label reflects the language of the speaker to whom the speech is directed. The language label of the speaker's mother tongue is combined, the candidate language set is optimally converged into a two-language set of "mother tongue + target language", and the search space of speech recognition in the language dimension is significantly reduced from the "complete language set of the conference" to the "two-language set". Based on this, the misrecognition probability caused by multi-language competition is reduced, especially the probability of misrecognizing foreign language terms as mother tongue homophonic words is reduced, and it can be quickly determined which languages are involved in each sentence of speech, so as to quickly determine which mixed language model is used for speech recognition, and the speech recognition efficiency of mixed language speech is improved.
Owner:QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD

Popular searches