Audio recognition model training method and device, electronic equipment, computer readable storage medium and computer program product

CN121963703APending Publication Date: 2026-05-01MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing end-to-end audio recognition models are inadequate in recognizing rare or scene-specific words, especially in recognizing words such as personal names and place names.

Method used

By fusing audio features and phrase features of audio samples, phrase prediction and text prediction are performed using the first fused feature. The audio recognition model is trained by combining the loss values ​​of the three prediction methods, thus fusing deep-level phrase information.

Benefits of technology

This improved the accuracy of the audio recognition model in recognizing rare words and words from specific scenarios, thereby enhancing the overall accuracy of audio recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963703A_ABST
    Figure CN121963703A_ABST
Patent Text Reader

Abstract

The invention provides an audio recognition model training method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps of performing feature fusion on an audio feature and a related phrase feature to obtain a first fusion feature; performing phrase prediction based on the first fusion feature and the audio feature to obtain a first predicted phrase, and determining a first loss; performing text prediction based on the audio feature, the phrase feature and the first fusion feature to obtain a first prediction text, and determining a second loss; performing text prediction on the audio sample based on the audio feature and the first fusion feature to obtain a second prediction text, and determining a third loss based on the second prediction text and the text label; and training the audio recognition model based on the first loss, the second loss and the third loss. According to the invention, the accuracy of audio recognition can be improved.
Need to check novelty before this filing date? Find Prior Art