Speech recognition apparatus and method having named entity recognition mechanism
Patent Information
- Application Number
- TW114119149
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-05-20
Smart Images

Figure TWG2TB001910589_001 
Figure TWG2TB001910589_002 
Figure TWG2TB001910589_003
Abstract
Claims
1. A speech recognition device with a named entity recognition mechanism, comprising: an image processing circuit, including: a text recognition circuit configured to perform text recognition on each of a plurality of image frames to be recognized captured from a media stream, to generate a plurality of image-recognized text; and a first named entity recognition circuit configured to perform a first named entity recognition on the image-recognized text, to generate image recognition history information including at least one image recognition named entity; and a speech processing circuit, including: a speech recognition circuit configured to perform speech recognition on each of a plurality of audio input segments in the media stream based on the image recognition history information and the speech recognition history information including at least one speech recognition named entity, to generate a plurality of speech-recognized text; and a second named entity recognition circuit configured to perform a second named entity recognition on the speech-recognized text, to generate the speech recognition history information.
2. The voice recognition device as claimed in claim 1, wherein the image processing circuit further includes a first storage circuit configured to store the image recognition history information and access it by the voice recognition circuit; and the voice processing circuit further includes a second storage circuit configured to store the voice-recognized text for access by the second named entity recognition circuit, and to store the voice recognition history information for access by the voice recognition circuit.
3. The speech recognition device as claimed in claim 1, wherein the first named entity recognition circuit further comprises: a recognition circuit configured to perform a recognition process of the first named entity recognition on the image recognized text in a current iteration to generate at least one new named entity; and a post-processing circuit configured to perform a post-processing process of the first named entity recognition on the new named entity in the current iteration to: add the new named entity to the image recognition history information generated in the previous iteration to generate the image recognition history information of the current iteration when the new named entity is different from the image recognized named entity contained in the image recognition history information generated in the previous iteration; and not add the new named entity to the image recognition history information of the current iteration when the new named entity is the same as the image recognized named entity contained in the image recognition history information generated in the previous iteration.
4. The speech recognition device as described in claim 3, wherein when the number of image recognition named entities and newly added named entities included in the image recognition history information generated in the previous iteration exceeds a certain threshold, the post-processing circuit causes the newly added named entities to replace the image recognition named entity with the longest addition time, so as to generate the image recognition history information of the current iteration.
5. The speech recognition apparatus as claimed in claim 1, wherein the second named entity recognition circuit further comprises: a recognition circuit configured to perform a recognition process of the second named entity recognition on the speech-recognized text in a current iteration to generate at least one new named entity; and a post-processing circuit configured to perform a post-processing process of the second named entity recognition on the new named entity in the current iteration to: determine a similarity between the new named entity and the speech-recognition named entity when the new named entity is different from the speech-recognition named entity contained in the speech-recognition historical information generated in a previous iteration; and add the new named entity to the speech-recognition historical information generated in the previous iteration when the similarity is not greater than a similarity threshold to generate the speech-recognition historical information of the current iteration; When the similarity is greater than the similarity threshold, the new recognition named entity in the speech-recognized text is replaced by the corresponding speech-recognition named entity, and the new recognition named entity is not added to the speech-recognition historical information in the current iteration; and when the new recognition named entity is the same as the speech-recognition named entity contained in the speech-recognition historical information generated in the previous iteration, the new recognition named entity is not added to the speech-recognition historical information in the current iteration.
6. The speech recognition device as described in claim 5, wherein when the number of speech recognition named entities and newly added named entities included in the speech recognition history information generated in the previous iteration exceeds a quantity threshold, the post-processing circuit causes the newly added named entities to replace the speech recognition named entity with the longest addition time, so as to generate the speech recognition history information of the current iteration.
7. The speech recognition device as claimed in claim 1 further comprises: an image preprocessing circuit configured to preprocess the image to be recognized in the media stream by downsampling or upsampling, and then process it by the image processing circuit; a speech preprocessing circuit configured to resample an audio signal in the media stream to generate multiple resampled audio data; and a speech input buffer circuit configured to buffer the resampled audio data as the audio input segments, and then process them by the speech processing circuit.
8. The speech recognition device as claimed in claim 7, wherein each of the audio input segments corresponds to the resampled audio data accumulated in the speech input buffer circuit to a preset length.
9. The speech recognition device as claimed in claim 1, wherein the text recognition circuit operates according to an Optical Character Recognition (OCR) mechanism, and the speech recognition circuit operates according to an Automatic Speech Recognition (ASR) mechanism.
10. A speech recognition method with a named entity recognition mechanism, comprising: a text recognition circuit included in an image processing circuit performing text recognition on each of a plurality of image frames to be recognized captured from a media stream to generate a plurality of image-recognized text; a first named entity recognition circuit included in the image processing circuit performing first named entity recognition on the image-recognized text to generate image recognition history information including at least one image recognition named entity; a speech recognition circuit included in a speech processing circuit performing speech recognition on each of a plurality of audio input segments in the media stream based on the image recognition history information and the speech recognition history information including at least one speech recognition named entity to generate a plurality of speech-recognized text; and a second named entity recognition circuit included in the speech processing circuit performing second named entity recognition on the speech-recognized text to generate the speech recognition history information.
Citation Information
Patent Citations
Supplementing a media stream with additional information
US20190042852A1
Machine-learning based systems and methods for analyzing and distributing multimedia content
US20200007934A1
Audiovisual Content Screening for Locked Application Programming Interfaces
US20200077144A1
Game event recognition
US20220040570A1
Content filtering in media playing devices
US20240430526A1