Audio recognition model training method and device, electronic equipment, computer readable storage medium and computer program product
CN121963703APending Publication Date: 2026-05-01MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MASHANG CONSUMER FINANCE CO LTD
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-01
AI Technical Summary
Technical Problem
Existing end-to-end audio recognition models are inadequate in recognizing rare or scene-specific words, especially in recognizing words such as personal names and place names.
Method used
By fusing audio features and phrase features of audio samples, phrase prediction and text prediction are performed using the first fused feature. The audio recognition model is trained by combining the loss values of the three prediction methods, thus fusing deep-level phrase information.
Benefits of technology
This improved the accuracy of the audio recognition model in recognizing rare words and words from specific scenarios, thereby enhancing the overall accuracy of audio recognition.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN121963703A_ABST
Abstract
The invention provides an audio recognition model training method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps of performing feature fusion on an audio feature and a related phrase feature to obtain a first fusion feature; performing phrase prediction based on the first fusion feature and the audio feature to obtain a first predicted phrase, and determining a first loss; performing text prediction based on the audio feature, the phrase feature and the first fusion feature to obtain a first prediction text, and determining a second loss; performing text prediction on the audio sample based on the audio feature and the first fusion feature to obtain a second prediction text, and determining a third loss based on the second prediction text and the text label; and training the audio recognition model based on the first loss, the second loss and the third loss. According to the invention, the accuracy of audio recognition can be improved.
Need to check novelty before this filing date? Find Prior Art