Speech recognition apparatus and method having named entity recognition mechanism

TWI939021BActive Publication Date: 2026-09-11REALTEK SEMICON CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
TW114119149
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2026-09-11
Estimated Expiration
2045-05-20

Smart Images

  • Figure TWG2TB001910589_001
    Figure TWG2TB001910589_001
  • Figure TWG2TB001910589_002
    Figure TWG2TB001910589_002
  • Figure TWG2TB001910589_003
    Figure TWG2TB001910589_003
Patent Text Reader

Abstract

A speech recognition device with a named entity recognition mechanism. The image processing circuit includes a text recognition circuit and a first named entity recognition circuit. The text recognition circuit performs text recognition on the image to be recognized captured from the media stream to generate multiple image-recognized texts. The first named entity recognition circuit then performs named entity recognition to generate image recognition history information containing image-recognized named entities. The speech processing circuit includes a speech recognition circuit and a second named entity recognition circuit. The speech recognition circuit performs speech recognition on the audio input segment in the media stream based on the image recognition history information and the speech recognition history information containing speech-recognized named entities to generate multiple speech-recognized texts. The second named entity recognition circuit then performs named entity recognition to generate speech recognition history information.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A speech recognition device with a named entity recognition mechanism, comprising: an image processing circuit, including: a text recognition circuit configured to perform text recognition on each of a plurality of image frames to be recognized captured from a media stream, to generate a plurality of image-recognized text; and a first named entity recognition circuit configured to perform a first named entity recognition on the image-recognized text, to generate image recognition history information including at least one image recognition named entity; and a speech processing circuit, including: a speech recognition circuit configured to perform speech recognition on each of a plurality of audio input segments in the media stream based on the image recognition history information and the speech recognition history information including at least one speech recognition named entity, to generate a plurality of speech-recognized text; and a second named entity recognition circuit configured to perform a second named entity recognition on the speech-recognized text, to generate the speech recognition history information.

2. The voice recognition device as claimed in claim 1, wherein the image processing circuit further includes a first storage circuit configured to store the image recognition history information and access it by the voice recognition circuit; and the voice processing circuit further includes a second storage circuit configured to store the voice-recognized text for access by the second named entity recognition circuit, and to store the voice recognition history information for access by the voice recognition circuit.

3. The speech recognition device as claimed in claim 1, wherein the first named entity recognition circuit further comprises: a recognition circuit configured to perform a recognition process of the first named entity recognition on the image recognized text in a current iteration to generate at least one new named entity; and a post-processing circuit configured to perform a post-processing process of the first named entity recognition on the new named entity in the current iteration to: add the new named entity to the image recognition history information generated in the previous iteration to generate the image recognition history information of the current iteration when the new named entity is different from the image recognized named entity contained in the image recognition history information generated in the previous iteration; and not add the new named entity to the image recognition history information of the current iteration when the new named entity is the same as the image recognized named entity contained in the image recognition history information generated in the previous iteration.

4. The speech recognition device as described in claim 3, wherein when the number of image recognition named entities and newly added named entities included in the image recognition history information generated in the previous iteration exceeds a certain threshold, the post-processing circuit causes the newly added named entities to replace the image recognition named entity with the longest addition time, so as to generate the image recognition history information of the current iteration.

5. The speech recognition apparatus as claimed in claim 1, wherein the second named entity recognition circuit further comprises: a recognition circuit configured to perform a recognition process of the second named entity recognition on the speech-recognized text in a current iteration to generate at least one new named entity; and a post-processing circuit configured to perform a post-processing process of the second named entity recognition on the new named entity in the current iteration to: determine a similarity between the new named entity and the speech-recognition named entity when the new named entity is different from the speech-recognition named entity contained in the speech-recognition historical information generated in a previous iteration; and add the new named entity to the speech-recognition historical information generated in the previous iteration when the similarity is not greater than a similarity threshold to generate the speech-recognition historical information of the current iteration; When the similarity is greater than the similarity threshold, the new recognition named entity in the speech-recognized text is replaced by the corresponding speech-recognition named entity, and the new recognition named entity is not added to the speech-recognition historical information in the current iteration; and when the new recognition named entity is the same as the speech-recognition named entity contained in the speech-recognition historical information generated in the previous iteration, the new recognition named entity is not added to the speech-recognition historical information in the current iteration.

6. The speech recognition device as described in claim 5, wherein when the number of speech recognition named entities and newly added named entities included in the speech recognition history information generated in the previous iteration exceeds a quantity threshold, the post-processing circuit causes the newly added named entities to replace the speech recognition named entity with the longest addition time, so as to generate the speech recognition history information of the current iteration.

7. The speech recognition device as claimed in claim 1 further comprises: an image preprocessing circuit configured to preprocess the image to be recognized in the media stream by downsampling or upsampling, and then process it by the image processing circuit; a speech preprocessing circuit configured to resample an audio signal in the media stream to generate multiple resampled audio data; and a speech input buffer circuit configured to buffer the resampled audio data as the audio input segments, and then process them by the speech processing circuit.

8. The speech recognition device as claimed in claim 7, wherein each of the audio input segments corresponds to the resampled audio data accumulated in the speech input buffer circuit to a preset length.

9. The speech recognition device as claimed in claim 1, wherein the text recognition circuit operates according to an Optical Character Recognition (OCR) mechanism, and the speech recognition circuit operates according to an Automatic Speech Recognition (ASR) mechanism.

10. A speech recognition method with a named entity recognition mechanism, comprising: a text recognition circuit included in an image processing circuit performing text recognition on each of a plurality of image frames to be recognized captured from a media stream to generate a plurality of image-recognized text; a first named entity recognition circuit included in the image processing circuit performing first named entity recognition on the image-recognized text to generate image recognition history information including at least one image recognition named entity; a speech recognition circuit included in a speech processing circuit performing speech recognition on each of a plurality of audio input segments in the media stream based on the image recognition history information and the speech recognition history information including at least one speech recognition named entity to generate a plurality of speech-recognized text; and a second named entity recognition circuit included in the speech processing circuit performing second named entity recognition on the speech-recognized text to generate the speech recognition history information.

Citation Information

Patent Citations

  • Supplementing a media stream with additional information

    US20190042852A1

  • Machine-learning based systems and methods for analyzing and distributing multimedia content

    US20200007934A1

  • Audiovisual Content Screening for Locked Application Programming Interfaces

    US20200077144A1

  • Game event recognition

    US20220040570A1

  • Content filtering in media playing devices

    US20240430526A1