Speech recognition and highlighting system for presentation support

JP2026125550APending Publication Date: 2026-08-03合同会社CACAO
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
合同会社CACAO
Filing Date
2025-01-22
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0013】 本発明により、以下の効果を得ることができる 1.話者がリアルタイムで進行位置を把握でき、発表の質を向上させる。 2.視覚障害を持つ話者や多言語プレゼンテーションにおいてもスムーズな進行を可能とする。 3.シンプルな構成により、スマートフォンやタブレットで簡便に利用可能である。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026125550000001_ABST
    Figure 2026125550000001_ABST
Patent Text Reader

Abstract

This invention solves the problem of speakers losing track of their progress when reading text displayed on a screen during greetings, presentations, product descriptions, reports, and other situations, due to interruptions or looking away from the screen. It also resolves the problem of difficulty in smoothly narrating when presenting from a script for the first time or in an unfamiliar language. [Solution] This system collects the audio of a speaker reading text displayed on a screen using a microphone, converts it into text using speech recognition technology, and compares it with the displayed text in real time. It marks the part of the text the speaker has spoken and visually highlights the next phrase to be read. It also accurately guides the speaker by converting the text into audio and presenting it to them. As a result, the speaker can deliver a smooth presentation without losing track of their progress. It also prevents misreadings and forgotten words in greetings and presentations, improving the quality of the presentation through visual and auditory support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0003] , , , , , , , ,

[0005] ,

[0001] The present invention relates to a speech recognition technology that supports a speaker to proceed smoothly in presentations and greetings. This system has a function of analyzing the text displayed on a display by speech recognition technology and synchronizing it visually and auditorily.

Background Art

[0002] Conventionally, in greetings and presentations, the speaker proceeds while looking at the prepared manuscript, but there are the following problems 1. When the speaker looks away from the text displayed on the display, it is easy to lose the progress position. 2. In the case of a first-time manuscript or a presentation in an unfamiliar language, misreading or forgetting to say is likely to occur. 3. In conventional support systems, the real-time function of visually synchronizing the speech recognition result is insufficient.

[0003] In Patent Documents 1 to 3, support systems using speech recognition technology have been proposed, but all of them are different from the present invention in the following points: · Lack of a function of displaying the progress position in real time. · Marking of the read part and highlighting of the next clause are not performed.

[0004] A system for supporting a presentation using keywords registered in advance by a speaker is described. Using speech recognition technology, a function of confirming the match between the content spoken by the speaker and the registered keywords and displaying the result on a display is disclosed. However, in this document, technologies such as synchronizing the progress position of the entire text in real time and marking the read part and highlighting the next part to be read are not disclosed. [Patent Document 1]

[0005] While a system for assisting speakers in presentations and meetings has been described that uses speech recognition technology to analyze the speaker's utterances, extract keywords according to specified conditions, and visualize them, it does not describe technologies for real-time display of the speaker's position in the text or for auditory guidance of the next section to read. [Patent Document 2]

[0006] This document concerns a presentation support system using speech recognition technology, disclosing a technique that transcribes the speaker's voice into text in real time and displays it on a screen. However, it does not mention functions that visually or audibly guide the user's position within the displayed text, or techniques that mark completed sections and highlight the next section to be read. [Patent Document 3]

[0007] This invention provides a system that analyzes text displayed on a screen using speech recognition technology, marks the phrases in the displayed text that closely match the spoken phrases, and visually and audibly guides the user to the next section to read.

[0008] This allows speakers to deliver presentations and greetings smoothly without losing track of their progress even when looking away from the screen, providing a novel technology that effectively supports speakers by managing the text's progress in real time. [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2021-117405 [Patent Document 2] Japanese Patent Publication No. 2021-81452 [Patent Document 3] Japanese Patent Publication No. 2012-255866 [Overview of the Initiative] [Problems that the invention aims to solve]

[0010] The present invention solves the following problems by providing a system for enabling a speaker to smoothly read through the text displayed on a display. 1. Assistance to prevent the speaker from losing the progress position even when looking away from the screen. 2. Provision of a reading assistance function during greetings or presentations.

Means for Solving the Problems

[0011] The present invention provides an audio recognition and highlighting system including the following means. 1. Marking of the text progress position by audio recognition 2. Highlighting of the next clause to be read 3. Guidance of the progress position by voice

[0012] This system can operate on general-purpose devices such as smartphones, tablets, and personal computers.

Effects of the Invention

[0013] The following effects can be obtained by the present invention 1. The speaker can grasp the progress position in real time, improving the quality of the presentation. 2. Smooth progress is enabled even for speakers with visual impairments or in multilingual presentations. 3. With a simple configuration, it can be easily used on smartphones and tablets.

Brief Description of the Drawings

[0014] [Figure 1] Configuration diagram of the entire system [Figure 2] Structural diagram of the audio recognition character conversion unit [Figure 3] Structural diagram of the clause comparison unit [Figure 4] Structural diagram of the marking application determination unit [Figure 5] Structural diagram of the highlighting application determination unit [Figure 6] Structural diagram of the character audio conversion unit [Figure 7] Display example

Best Mode for Carrying Out the Invention

[0015] The present invention provides a system that supports a speaker to proceed smoothly during a presentation or a greeting. The following shows the embodiments for carrying out the present invention.

[0016] Configuration Diagram of the Whole System [Figure 1] The overall configuration of this system is as shown in [Figure 1], and it is a configuration that can be completed with a handy general-purpose device such as a smartphone, a tablet, or a personal computer.

[0017] This system is composed of a display, an audio recognition character conversion unit, a clause comparison unit, a marking application determination unit, an emphasized clause application determination unit, and a character audio conversion unit from

[0018] 1. Display 20: The text is displayed. It is a display on which a manuscript for a presentation or a greeting is displayed.

[0019] 2. Audio Recognition Character Conversion Unit: As shown in [Figure 2], the audio read by the speaker is picked up by a microphone and converted into characters using audio recognition technology.

[0020] 3. Clause Comparison Unit: As shown in [Figure 3], the character string obtained by the audio recognition character conversion unit is compared with the text displayed on the display, and clauses that almost match are identified.

[0021] 4. Marking Application Determination Unit: As shown in [Figure 4], marking is applied to the character strings that almost match identified by the clause comparison unit.

[0022] 5. Emphasis Application Determination Unit: As shown in [Figure 5], the next character string after the marked character string is enlarged, thickened, etc. to emphasize the clause to be read next. 701

[0023] 6. Character Audio Conversion Unit: As shown in [Figure 6], the emphasized clause is converted into audio and transmitted to the speaker through an earphone or the like.

Examples

[0024] The following are specific examples. 1. Example of display [Figure 7]

[0025] 2. Speech Recognition and Text Conversion - The speaker reads aloud, "Hello everyone. Thank you for gathering here today." - The microphone picks up the voice, and the speech recognition and text conversion unit converts it into text: "Hello everyone. Thank you for gathering here today."

[0026] 3. Sentence comparison and marking - The sentence comparison unit compares the converted string with the text displayed on the screen and applies a marking to the matching string "Hello everyone. Thank you for gathering here today."

[0027] 4. Emphasis - The phrase to be emphasized next, "Thank you," is highlighted by making it larger, bolder, or changing its color.

[0028] 5. Text-to-Speech Conversion - The text-to-speech conversion unit converts the highlighted phrase into speech and communicates the next phrase to be read, "Thank you," to the speaker through the earphones. [Explanation of symbols]

[0029] 401: Mike 402: Input speech waveform 403: Converted string (e.g., ABCD) 501: Text displayed on the screen (e.g., ABCD, EFG, HIJ) 601: Mark the part where the spoken text and the displayed text closely match. TIFF2026125550000002.tif974701: Examples of marked phrases (e.g., ABCD, EFG, HIJ) 801: Selected phrase for emphasis (e.g., EFG) 802: Waveform of the highlighted phrase converted to analog audio 803: Earphones 201: Screen displaying the text to be read aloud 202: An example where a phrase that closely matches the spoken phrase is marked. 203: An example where a mark is placed at the end of a phrase. 204: An example that highlights the phrase to be read aloud next.

Claims

1. This system compares text displayed on a screen with text generated by speech recognition in real time, and marks the speaker's position within the spoken phrase in real time.

2. The system according to claim 1, characterized in that it includes means for visually highlighting the next phrase following a marked phrase.

3. The system according to claim 2, characterized in that it comprises means for converting highlighted phrases into speech and presenting them through earphones or speakers worn by the speaker.

4. The system according to claim 3, characterized in that the marking and highlighting operate in real time so that the speaker does not lose track of their position even when they take their eyes off the display.

5. The system according to claim 4, characterized in that, in greetings, presentations, product descriptions, etc., a clear marking is added to the phrase that has been read.

6. A method for applying the system described in claim 5, characterized in that the current position of the displayed text is identified in real time, and the phrase to be read aloud next is highlighted and voice-guided.