Speech recognition method, apparatus, device, and storage medium

By obtaining embedding vectors from the speech recognition model trained for the customer service system and utilizing an attention mechanism, the problem of low speech recognition efficiency in the customer service system is solved, achieving efficient and reliable speech recognition text generation.

CN119107939BActive Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
1 Cites 0 Cited by

Patent Information

Application Number
CN202411178826.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-05-01
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

Existing speech recognition methods cannot meet the speech recognition needs of customer service systems, resulting in low efficiency. This is because general models lack specificity.

Method used

By acquiring the preset voice and domain labels of the customer service system, an embedding vector is generated and reconstructed using an attention mechanism. The speech recognition model is then trained by combining weight coefficients and loss values ​​to generate the current speech recognition text.

Benefits of technology

It improved the efficiency of voice recognition in the customer service system, reduced recognition time, and enhanced the reliability of voice-recognized text, avoiding the impact of human intervention.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence and the field of financial technology, and discloses a speech recognition method, device and equipment and a storage medium, the method comprising: splicing a first embedding vector and a second embedding vector to generate a third embedding vector, reconstructing the third embedding vector to generate a fourth embedding vector; obtaining a predicted speech recognition text output by a speech recognition model based on the fourth embedding vector, obtaining a first loss value and a second loss value between the predicted speech recognition text and a preset speech recognition text; generating a total loss value of the predicted speech recognition text according to the first loss value and the second loss value, training the speech recognition model based on the total loss value; obtaining a current speech and a current field label sent by a customer service system, inputting the current speech and the current field label into the trained speech recognition model, and obtaining a current speech recognition text output by the trained speech recognition model. The present application is beneficial to improving the efficiency of speech recognition.
Need to check novelty before this filing date? Find Prior Art