Display system, display control device, and display control program
The display system addresses the challenge of recognizing and displaying multiple user voices by using a stereo conversion unit and display control device to separate and process stereo signals, enabling real-time conversion and visualization of conversations with language support and emphasis.
Patent Information
- Application Number
- JP2024030572
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-09-10
AI Technical Summary
Existing display systems struggle to individually recognize and display voices from multiple users due to limitations in device ports, such as USB terminals or audio jacks, which prevent connection of multiple microphones, making it difficult to distinguish and visualize two-way conversations.
A display system with a display unit, first and second microphones, a stereo conversion unit, and a display control device that separates and processes stereo signals into individual character information for display, allowing simultaneous recognition and display of multiple users' voices, with options for language conversion and emphasis.
Enables effective individual recognition and display of voices from multiple users, supporting real-time conversion and visualization of conversations across different languages, with enhanced clarity and emphasis on important information.
Smart Images

Figure 2025132783000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a display system, a display control device, and a display control program. [Background technology]
[0002] A system has been proposed that supports user communication in situations where multiple people are conversing across a partition (see, for example, Patent Document 1). In this system, a transparent display that functions as a partition is placed between two people. Microphones that pick up the voices of the people speaking are also placed on both sides of the transparent display. This allows the information processing device of the display system to convert the content of the conversation spoken into each microphone into text information and display it on the transparent display in real time. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-124615 Summary of the Invention [Problem to be solved by the invention]
[0004] To distinguish between multiple users and visualize and display a two-way conversation between them, multiple microphones are required. However, devices that receive voice information input may not be able to use multiple device ports for voice input and may only be able to use one device port. For example, some Android (registered trademark) devices are designed to recognize only one device port (a Universal Serial Bus (USB) terminal or an audio jack), making it difficult to connect multiple microphones to a single device port. Even if multiple microphones are connected to a single device port via a multi-tap (USB hub), some Android devices are designed to recognize only one device, making it difficult to connect multiple microphones. In such cases, it is difficult to individually recognize the voices of multiple users and display the two-way conversation on a display device.
[0005] Therefore, the purpose of the present disclosure, which has been made with these points in mind, is to provide a display system, a display control device, and a display control program that are capable of individually recognizing voices from multiple users, converting them into text information, and displaying them. [Means for solving the problem]
[0006] In one embodiment, (1) a display system includes a display unit that is arranged between multiple users when in use, a first microphone that is arranged on a first side of the display unit, a second microphone that is arranged on a second side opposite the first side of the display unit, a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, and a display control device that receives the stereo signals, separates the stereo signals into the first signal and the second signal, and displays first character information and second character information that are converted from audio signals included in the first signal and the second signal by voice recognition on the display unit.
[0007] (2) In the display system of (1) above, the display control device can display the first character information and the second character information while scrolling them sequentially on the display unit, and can display a boundary line between the first character information and the second character information.
[0008] (3) In the display system of (1) or (2) above, the display control device can convert the first character information in a first language into a second language different from the first language and display it on the display unit.
[0009] (4) In the display system of (3) above, the display control device can display the first character information converted from the latest utterances by the multiple users as a string of characters in the second language toward the second side and as a string of characters in the first language toward the first side, and can display the second character information converted from the latest utterances by the multiple users as a string of characters in the first language toward the first side and as a string of characters in the second language toward the second side.
[0010] (5) In any of the display systems (1) to (4) above, the display control device can display the latest first character information or the latest second character information converted from speech uttered by one or more of the plurality of users on the display unit in a predetermined manner different from the display manner of other character information.
[0011] (6) In any of the display systems (1) to (5) above, when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the audio levels of the audio signals contained in the first signal and the second signal, and based on the comparison result, can display only one of the first character information and the second character information on the display unit.
[0012] (7) In the display system of (6) above, the display control device can display on the display unit the first character information or the second character information obtained by converting the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has the higher audio level.
[0013] (8) In any of the display systems (1) to (5) above, when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device can compare the first character information with the second character information and, if they match by a predetermined percentage or more, display only one of the first character information and the second character information.
[0014] (9) In the display system of (8) above, the display control device can display on the display unit the first character information or the second character information obtained by voice recognition of the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has a higher audio level.
[0015] (10) In any of the display systems (1) to (5) above, when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device compares the first character information and the second character information corresponding to the speech within the speech section including the period in which the audio signals overlap, and can display only the first character information or the second character information that contains the greater amount of information on the display unit.
[0016] In one embodiment, (11) a display control device includes: a display unit that is placed among multiple users when in use; a communication unit that can communicate with a stereo conversion unit that converts a first signal input from a first microphone placed on a first side of the display unit and a second signal input from a second microphone placed on a second side opposite the first side of the display unit into a stereo signal; and a control unit that receives the input of the stereo signal, separates the stereo signal into the first signal and the second signal, and executes a process to display on the display unit first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal.
[0017] In one embodiment, (12) a display control program is a program of a display control device that controls a display system including a display unit that is placed between multiple users when in use, a first microphone that is placed on a first side of the display unit, a second microphone that is placed on a second side opposite the first side of the display unit, and a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, and causes a control unit of the display control device to execute a process of receiving input of the stereo signals, separating the stereo signals into the first signal and the second signal, and displaying first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal, respectively, on the display unit. [Effects of the Invention]
[0018] According to the embodiments of the present disclosure, it is possible to provide a display system, a display control device, and a display control program that individually recognize voices from multiple users, convert the voices into text information, and display the text information. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is a diagram illustrating an example of a usage scene of a display system according to an embodiment. [Figure 2] FIG. 2 is a top view of the display system of FIG. 1. [Figure 3]FIG. 2 is a block diagram showing a schematic configuration of the display system of FIG. [Figure 4] FIG. 2 is a diagram illustrating an outline of information processing by the display system of FIG. [Figure 5] FIG. 10 is a diagram showing an example of a display on a transparent screen as seen from the first user's side. [Figure 6] FIG. 10 is a diagram showing an example of a display on the transparent screen as seen from the side of a second user. [Figure 7] 4 is a flowchart showing a first example of a display control process executed by a control unit of the display control device. FIG. [Figure 8] FIG. 8 is a flowchart showing a first example of the first character information display process of FIG. 7. [Figure 9] FIG. 8 is a flowchart showing a first example of the second character information display process of FIG. 7. [Figure 10] FIG. 10 is a flowchart showing a second example of the display control process executed by the control unit of the display control device. [Figure 11] 11 is a flowchart showing an example of a signal level adjustment process of the first signal in FIG. 10. FIG. [Figure 12] 11 is a flowchart showing an example of a signal level adjustment process of the second signal in FIG. 10. FIG. [Figure 13] FIG. 10 is a flowchart showing a third example of the display control process executed by the control unit of the display control device. [Figure 14] FIG. 10 is a flowchart showing a third example of the display control process executed by the control unit of the display control device. [Figure 15] FIG. 10 is a flowchart showing a fourth example of a display control process executed by a control unit of the display control device. [Figure 16] FIG. 10 is a flowchart showing a fourth example of a display control process executed by a control unit of the display control device. [Figure 17] FIG. 10 is a diagram showing a second example of the display on the transparent screen as seen from the second user's side. [Figure 18] FIG. 18 is a flowchart showing a second example of the first character information display process corresponding to FIG. 17. [Figure 19] FIG. 18 is a flowchart showing a second example of the second character information display process corresponding to FIG. 17. [Figure 20]FIG. 10 is a diagram showing a third example of a display on the transparent screen as seen from the second user's side. [Figure 21] FIG. 21 is a flowchart showing a third example of the first character information display process corresponding to FIG. 20. [Figure 22] FIG. 21 is a flowchart showing a third example of the second character information display process corresponding to FIG. 20. [Figure 23] FIG. 10 is a diagram showing a fourth example of a display on the transparent screen as seen from the second user's side. [Figure 24] FIG. 10 is a diagram showing a fourth example of a display on the transparent screen as seen from the second user's side. [Figure 25] FIG. 25 is a flowchart showing a fourth example of the first character information display process corresponding to FIGS. 23 and 24. [Figure 26] FIG. 25 is a flowchart showing a fourth example of the second character information display process corresponding to FIGS. 23 and 24. [Figure 27] FIG. 10 is a flow diagram of a process for clearing an image. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The drawings used in the following description are schematic. The dimensions and ratios in the drawings do not necessarily correspond to the actual dimensions and ratios.
[0021] (Display system configuration) As shown in FIGS. 1 and 2, a display system 1 according to an embodiment is a system that supports a dialogue between a first user U1 and a second user U2. As an example, the first user U1 is a staff member at a tourist information center. The second user U2 is a traveler visiting the tourist information center. In the following embodiment, the first user U1 speaks a first language. The first language is, for example, Japanese. The second user U2 speaks a second language different from the first language. The second language is, for example, English. The display system 1 converts the speech of the first user U1 into the second language and the speech of the second user U2 into the first language, and displays the converted speech. However, the display system 1 is not limited to a system that includes such speech translation. The display system 1 may display the speech of the first user U1 and the second user U2 without translating them into another language. In this case, the first user U1 may be, for example, a staff member at a government, local government, or public institution. The second user U2 may be, for example, a general citizen visiting a government, local government, or public institution. The display system 1 may also be used at the counters of financial institutions, medical institutions, public transportation facilities, private business offices, and office conference rooms.
[0022] An acrylic plate, a vinyl curtain, or the like is installed between the first user U1 and the second user U2. In recent years, acrylic plates or vinyl curtains have been installed as a measure against droplets. In the display system 1, the acrylic plate, the vinyl curtain, or the like is used as a base material 10, and a transparent screen 5 is placed on this base material 10. The projector 30 projects the speech content of the first user U1 and the second user U2 onto this transparent screen 5. This allows the first user U1 and the second user U2, who are conversing across the transparent screen 5, to visually confirm the speech content even if they cannot hear what the other user is saying. The first user U1 side of the transparent screen 5 is the first side. The second user U2 side is the second side.
[0023] The transparent screen 5 can be placed anywhere on the base material 10. The transparent screen 5 may be placed in a position that does not prevent the first user U1 and the second user U2 from seeing each other's faces. For example, the transparent screen 5 may be placed in a position where the first user U1 and the second user U2 lower their lines of sight downward from the horizontal direction.
[0024] As shown in FIGS. 1 to 3 , the display system 1 includes a transparent screen 5, a display control device 20, a projector 30, a first microphone 40a, a second microphone 40b, and a stereo conversion unit 50. The transparent screen 5 and the projector 30 constitute a display unit. Hereinafter, the transparent screen 5 and the projector 30 may be collectively referred to as the display unit. The first microphone 40a and the second microphone 40b are each connected to the stereo conversion unit 50 so as to be able to communicate with each other via a wired or wireless connection. The stereo conversion unit 50 and the display control device 20 are connected so as to be able to communicate with each other via a wired or wireless connection. The display control device 20 and the projector 30 are connected so as to be able to communicate with each other via a wired or wireless connection. The display control device 20 is configured to be able to communicate with an external cloud server 60 that provides voice recognition processing via a communication line.
[0025] The transparent screen 5 can be a film-, sheet-, or plate-like member for projector projection that can be attached to the base material 10. The transparent screen 5 diffuses a portion of the light incident from the projector 30 to the incident side and the exit side. The first user U1 and the second user U2 can recognize the image projected from the projector 30 when the light diffused by the transparent screen 5 enters their field of vision. The shape of the transparent screen 5 can be, for example, rectangular, but is not limited to this. The transparent screen 5 can have various shapes.
[0026] The first microphone 40a and the second microphone 40b are monaural microphones. The first microphone 40a is disposed on the side of the transparent screen 5 facing the first user U1, i.e., the first side. For example, the first microphone 40a may be disposed on a desk in front of the first user U1. The first microphone 40a mainly converts the sound emitted by the first user U1 into a first signal and outputs it to the stereo conversion unit 50. The second microphone 40b is disposed on the side of the transparent screen 5 facing the second user U2, i.e., the second side. For example, the second microphone 40b may be disposed on a desk in front of the second user U2. The second microphone 40b mainly converts the sound emitted by the second user U2 into a second signal and outputs it to the stereo conversion unit 50. The first signal and the second signal are monaural signals.
[0027] The first signal and the second signal are sound signals detected by the first microphone 40a and the second microphone 40b, and include noise and sound information other than human voices. In the present application, the voice signal is a signal of a person's voice. When the first user U1 and the second user U2 speak to the first microphones 40a and 40b, the first signal or the second signal includes a voice signal.
[0028] The stereo converter 50 converts the first signal input from the first microphone 40a and the second signal input from the second microphone 40b into a stereo signal. The stereo signal includes two channel signals corresponding to left and right sounds, respectively. For example, the stereo converter 50 can output a stereo signal to the display control device 20, with the first signal acquired from the first microphone 40a accommodated in the left channel and the second signal acquired from the second microphone 40b accommodated in the right channel. The stereo signal is information obtained by converting information from two channels into one-dimensional array information based on a specific sampling rate and encoding information. By using the stereo converter 50, the first signal and the second signal can be combined and input to the display control device 20.
[0029] The display control device 20 separates the stereo signal received from the stereo conversion unit 50 into a first signal and a second signal. The display control device 20 separates the one-dimensionally arranged stereo signal into two signals, a left channel and a right channel, in accordance with the specifications of the stereo conversion unit 50. For example, the left channel signal is the first signal, and the right channel signal is the second signal.
[0030] The display control device 20 is configured to be able to transmit the first signal and the second signal to the cloud server 60, respectively. The display control device 20 acquires character information obtained by converting the audio signals included in the first signal and the second signal by speech recognition processing in the cloud server 60. The display control device 20 can further transmit character information in a first language to the cloud server 60 and acquire character information converted into a second language. The display control device 20 can also transmit character information in the second language to the cloud server 60 and acquire character information converted into the first language. Hereinafter, character information converted from the audio signal included in the first signal will be referred to as first character information. Character information converted from the audio signal included in the second signal will be referred to as second character information.
[0031] The display control device 20 can process the first character information and / or the second character information. The display control device 20 can generate or acquire image information based on the first character information and / or the second character information. The display control device 20 transmits an image signal based on the character information to the projector 30, which projects and displays the character information as a character string on the transparent screen 5. In this application, a character string means an image in which the character information is projected onto the transparent screen 5, which is a display unit. The character string is displayed at a set position on the display unit, with attributes such as color and size added to the character information.
[0032] For example, a general-purpose information device such as a mobile phone (smartphone) or a personal computer, or a dedicated information device, can be used as the display control device 20. The display control device 20 includes a communication unit 21, a control unit 22, a storage unit 23, and an input unit 24.
[0033] The communication unit 21 communicates with devices external to the display control device 20 via wireless or wired communication means. The external devices include the stereo conversion unit 50, the projector 30, and the cloud server 60. The communication unit 21 may support communication using various communication methods, such as a serial communication standard, a wired LAN (local area network) standard, a wireless LAN standard such as Wi-Fi, and mobile communication standards such as 4G (4th Generation) and 5G (5th Generation). The communication unit 21 can also transmit a first signal and a second signal to the cloud server 60 and receive text information converted from audio signals included in the first signal and the second signal.
[0034] The control unit 22 includes one or more processors. Processors include general-purpose processors that load specific programs to execute specific functions, and dedicated processors specialized for specific processing. Dedicated processors include application-specific integrated circuits (ASICs). Processors include programmable logic devices (PLDs). PLDs include field-programmable gate arrays (FPGAs). The control unit 22 may be either a system-on-a-chip (SoC) or a system in a package (SiP) in which one or more processors work together.
[0035] The control unit 22 controls the entire display control device 20 and executes various information processes. The control unit 22 separates a first signal and a second signal from the stereo signal acquired from the stereo conversion unit 50. The control unit 22 is configured to transmit the first signal and the second signal to the cloud server 60 via the communication unit 21. The control unit 22 is configured to acquire, via the communication unit 21, first character information and second character information converted from the audio signals included in the first signal and the second signal by the cloud server 60. The control unit 22 may further be configured to transmit the acquired first character information and second character information to the cloud server 60 and acquire the first character information and second character information translated into another language.
[0036] The control unit 22 executes a display process to display the first character information and the second character information acquired from the cloud server 60 on the projector 30. The control unit 22 can control the display position and display mode of the first character information and the second character information on the transparent screen 5.
[0037] The control unit 22 can cause the projector 30 to project character strings corresponding to the first character information and the second character information with left-right reversal. In the present application, the state in which the displayed characters appear as "normal characters" to the first user U1 is the normal projection state of the projector 30. In this case, the characters viewed by the second user U2 are "mirror characters." "Normal characters" are characters displayed in a normal manner. "Mirror characters" are characters in which the left and right sides of normal characters are reversed. Characters projected from the projector 30 with left-right reversal appear as "mirror characters" to the first user U1 and as "normal characters" to the second user U2.
[0038] When the first character information includes a predetermined word to be emphasized, the control unit 22 performs a process of emphasizing the word in the display process. Hereinafter, a word to be emphasized is referred to as an emphasized word. When the first character information includes an emphasized word, the control unit 22 can, for example, change the color of the character string corresponding to the word to a more conspicuous color. Keywords that are important in a conversation are stored in advance in the storage unit 23 as emphasized words.
[0039] When the first character information includes a word to be illustrated, the control unit 22 can cause a still image or a moving image related to the illustrated object to be displayed on the display unit by an illustration process. Hereinafter, the word to be illustrated will be referred to as an illustrated word. An illustrated word is a specific word for which a still image and / or a moving image for explanation can be displayed. Data on the still image and the moving image corresponding to the illustrated word is stored in the memory unit 23 as illustration display data.
[0040] The storage unit 23 may include at least one of a semiconductor storage device, a magnetic storage device, and an optical storage device. The semiconductor storage device may include a volatile memory such as a dynamic random access memory (DRAM) and a static random access memory (SRAM), and a nonvolatile memory such as a read only memory (ROM) and a flash memory. The semiconductor storage device includes a solid state drive (SSD) that uses a flash memory. The magnetic storage device includes a magnetic tape, a floppy disk (registered trademark), a hard disk, etc. The optical storage device includes, for example, a compact disc (CD), a digital versatile disc (DVD), and a Blu-ray (registered trademark).
[0041] The storage unit 23 is configured to store programs executed by the control unit 22, information necessary for the processing executed by the control unit 22, and information obtained as a result of the execution by the control unit 22. The storage unit 23 may store emphasis word information used when emphasizing first character information received from the cloud server 60, as well as graphic words and graphic display data used when displaying graphic information based on the first character information.
[0042] The input unit 24 accepts operations of the first user U1 on the display control device 20. The input unit 24 includes, for example, a touch panel, a keyboard, a mouse, and a pen input device. The input unit 24 may be built into the display control device 20. The input unit 24 may also be an external device attached to the display control device 20.
[0043] In addition to the above, the display control device 20 may include a display unit such as a liquid crystal display (LCD), an organic electroluminescence (EL) display, an inorganic EL display, etc. The display unit may display various information related to the operation of the display control device 20.
[0044] The projector 30 is placed beside a first user U1, who is, for example, a counter clerk, and projects an image onto the transparent screen 5 based on an image signal received from the display control device 20. The projector 30 may project still images and moving images for illustration onto predetermined positions on the transparent screen 5, in addition to images of character strings corresponding to the first character information and the second character information.
[0045] The cloud server 60 is a server that can be accessed via a network such as the Internet. The cloud server 60 of this embodiment is a voice recognition server that provides voice recognition services. When the cloud server 60 receives a voice signal, it performs voice recognition processing on the received voice signal. The cloud server 60 does not perform voice recognition processing if the received signal does not contain human voice or if the level of the voice signal is lower than a predetermined value. The cloud server 60 transmits text information obtained as a result of voice recognition to the sender of the voice signal.
[0046] Cloud server 60 may further provide a translation service for translating between different languages. Cloud server 60 may receive text information in a first language and transmit the text information translated into a second language to the sender. Cloud server 60 may also receive text information in a second language and transmit the information translated into the first language to the sender.
[0047] (Overview of display system processing) FIG. 4 shows a schematic diagram of how the speech of the first user U1 and the second user U2 is displayed on the transparent screen 5 by the display system 1. As shown in FIG.
[0048] When the first user U1 speaks into the first microphone 40a, a first signal including the audio signal is transmitted to the display control device 20 via the stereo conversion unit 50. The display control device 20 transmits the first signal to the cloud server 60. The display control device 20 may transmit the left channel signal included in the stereo signal transmitted from the stereo conversion unit 50 to the cloud server 60 by streaming. The display control device 20 may transmit the first signal to the cloud server 60 only if the first signal includes an audio signal having a signal level equal to or higher than a predetermined signal level.
[0049] The cloud server 60 converts the voice signal included in the received first signal into first character information, which is text data, using natural language processing such as morphological element analysis and machine learning. The cloud server 60 transmits the first character information converted from the voice signal to the display control device 20.
[0050] When translating the first character information, the display control device 20 transmits the received first character information in the first language to the cloud server 60. The cloud server 60 translates the first character information in the first language into the second language. The cloud server 60 transmits the first character information translated into the second language to the display control device 20.
[0051] The display control device 20 performs highlighting and illustration processing, etc., as necessary, on the first character information in the first language or the first character information translated into the second language acquired from the cloud server 60. The display control device 20 transmits image information of character strings to be displayed corresponding to the first character information and, if necessary, image information for illustration to the projector 30. The highlighting and illustration processing may be performed to support the first user U1 in presenting information to the second user U2. For example, when the first user U1 explains to the second user U2 how to get to Tokyo Station, the word "Tokyo" registered as an emphasis word in the first character information may be highlighted. Furthermore, when "Tokyo" is included as an illustration word, image information such as a map or a route map of the Tokyo area may be displayed.
[0052] The projector 30 projects the received text and image information onto the transparent screen 5. Here, the display control device 20 may perform display processing to project the first character information corresponding to the voice uttered by the first user U1 in an inverted manner so that it appears as normal characters to the second user U2. Furthermore, when the first user U1 checks the voice uttered by himself, the first character information may be projected so that it appears as normal characters to the first user.
[0053] The above describes the case where the first user U1 speaks into the first microphone 40a, but the same process flow is also executed when the second user U2 speaks into the second microphone 40b. However, if the second user U2 is a traveler visiting a tourist information center, the display control device 20 does not need to perform the highlighting process and the illustration process. Furthermore, the second character information corresponding to the voice uttered by the second user U2 may be displayed so that it appears as correct characters to the first user U1.
[0054] (First display example of the display system) Next, an example of an image displayed on the transparent screen 5 by the projector 30 based on an image signal from the display control device 20 will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a diagram showing an example of the transparent screen 5 as viewed from the first user U1 side. Fig. 6 is a diagram showing an example of the transparent screen 5 displaying the same content as Fig. 5 as viewed from the second user U2 side.
[0055] 5 and 6, the transparent screen 5 is divided into a first display area A1, a second display area A2, and a third display area A3. The first display area A1 is an area that displays first character information and second character information while scrolling sequentially upward in accordance with the speech of the first user U1 and the second user U2. In this example, the speech of the first user U1 is translated from Japanese, which is a first language, to English, which is a second language, and then displayed. The speech of the second user U2 is translated from English, which is the second language, to Japanese, which is the first language, and then displayed. In the first display area A1, the first character information corresponding to the speech of the first user U1 (e.g., "Hello. How can I help you?") is displayed in correct characters as seen by the second user U2. In addition, in the first display area A1, the second character information corresponding to the speech of the second user U2 (e.g., "Hello. Where can I go if I take the city sightseeing bus?") is displayed in correct characters as seen by the first user U1.
[0056] The second display area A2 is an area where still images and / or moving images are displayed when a graphical display is performed. For example, a map related to the conversation, transportation information, and a video guide to tourist spots are displayed in the second display area A2. The second display area A2 is also called a graphical display area. The second display area A2 is displayed so that it faces in the correct direction when viewed from the second user U2.
[0057] The third display area A3 is an area for displaying the most recent utterance of either the first user U1 or the second user U2. The third display area A3 is used by the speaker, the first user U1 or the second user U2, to confirm that their speech has been correctly recognized. Therefore, in the third display area A3, the first character information is displayed in correct characters as seen by the first user U1. The second character information is displayed in correct characters as seen by the first user U1. For example, in the examples of FIGS. 5 and 6, the most recent utterance by the first user U1 is "I'm going to Kiyomizu-dera Temple and Kinkaku-ji Temple." This utterance is displayed in the second language in the first display area A1 in correct characters as seen by the second user U2. At the same time, the third display area A3 is displayed in the first language in correct characters as seen by the first user U1.
[0058] 5 and 6 are merely examples of the arrangement of the first display area A1, the second display area A2, and the third display area A3. The positions, sizes, shapes, etc. of the first display area A1, the second display area A2, and the third display area A3 may be set as appropriate.
[0059] (First example of display control method) A display control method executed by the control unit 22 of the display control device 20 will be described below with reference to FIG. 7. The display control device 20 may be configured to read and implement a display control program recorded on a non-transitory computer-readable medium to perform the processes performed by the control unit 22 described below. Non-transitory computer-readable media include, but are not limited to, magnetic storage media, optical storage media, magneto-optical storage media, and semiconductor storage media. Magnetic storage media include magnetic disks, hard disks, and magnetic tapes. Optical storage media include optical discs such as CDs (Compact Discs), DVDs, and Blu-ray (registered trademark) Discs. Semiconductor storage media include ROMs (Read Only Memory), EEPROMs (Electrically Erasable Programmable Read-Only Memory), flash memories, etc.
[0060] First, when the first user U1 starts the display system 1, the control unit 22 starts acquiring a stereo signal from the stereo conversion unit 50 via the communication unit 21 (step S101). The display system 1 may be started by turning on / off a switch provided in the display control device 20, turning on / off a switch of the stereo conversion unit 50, turning on / off the first microphone 40a and the second microphone 40b, etc.
[0061] The control unit 22 separates the stereo signal acquired by the stereo conversion unit 50 into a first signal acquired from the first microphone 40a and a second signal acquired from the second microphone 40b (step S102).
[0062] The control unit 22 executes a first character information display process (step S103) for the audio signal included in the first signal. The control unit 22 also executes a second character information display process (step S104) for the audio signal included in the second signal. When the first user U1 and the second user U2 alternately converse with each other and their voices do not overlap, only one of steps S103 and S104 is executed at each time. On the other hand, when the first user U1 and the second user U2 speak simultaneously, steps S103 and S104 may be executed in parallel, independently of each other. The processes of steps S103 and S104 will be described below with reference to the flow charts of FIGS. 8 and 9.
[0063] In the flow diagram of FIG. 8, first, the control unit 22 transmits a first signal to the cloud server 60 (step S201). The voice signal included in the first signal transmitted to the cloud server 60 is subjected to voice recognition processing by the cloud server 60. The cloud server 60 extracts human voice from the first signal and converts it into text information indicating the content of the speech using natural language processing. The cloud server 60 transmits the converted text information to the display control device 20, for example, in units of sentences. If the first signal does not include human voice, the cloud server 60 does not need to transmit the text information to the display control device 20.
[0064] The control unit 22 sequentially acquires the first character information converted from the first signal from the cloud server 60 (step S202).
[0065] The control unit 22 translates the acquired first character information from the first language to the second language (step S203). The control unit 22 may transmit the first character information to the cloud server 60 and acquire the first character information translated into the second language. Note that in the display system 1, it is not essential to translate from the first language to the second language. When the first user U1 and the second user U2 use the same language, the first character information display process executed by the control unit 22 does not need to include step S203.
[0066] Next, the control unit 22 may perform highlighting processing on the first character information (step S204). Specifically, the control unit 22 compares the first character information with emphasis words stored in the storage unit 23 to confirm whether the first character information includes emphasis words. The first character information may be information in a first language or information translated into a second language. If the character information includes emphasis words, the control unit 22 generates display data for a character string in which the emphasis words are displayed in an emphasised manner. Highlighting includes, for example, but is not limited to, changing the color of emphasis words included in the displayed character string to a more conspicuous color. Highlighting may include changing the size and / or brightness of emphasis words included in the displayed character string.
[0067] The control unit 22 may execute a graphic processing on the first character information (step S205). Specifically, the control unit 22 compares the first character information with the graphic words stored in the storage unit 23, and checks whether the first character information includes the graphic words. The first character information may be information in a first language or information translated into a second language. If the character information includes the graphic words, the control unit 22 acquires graphic display data, which is data of still images and / or moving images corresponding to the graphic words, from the storage unit 23.
[0068] Either step S204 or step S205 may be executed first. Furthermore, steps S204 and S205 are not essential processes. Either or both of steps S204 and S205 may not be executed.
[0069] Next, the control unit 22 generates image information for display and transmits it to the projector 30, thereby updating the display on the transparent screen 5 (step S206). The control unit 22 adds the latest highlighted first character information to the character string scrolled in the first display area A1. When the first character information is translated into the second language in step S203, the control unit 22 may display the first character information translated into the second language in the first display area A1. When highlighting processing is performed in step S204, the control unit 22 may highlight an emphasized word included in the character string displayed in the first display area A1. Furthermore, for confirmation by the first user U1, the control unit 22 may display the first character information in the first language in the third display area A3 so that it appears as correct characters to the first user U1. Furthermore, when the illustration processing is performed in step S205 and illustration display data is available, the control unit 22 may display an image based on the illustration display data in the second display area A2.
[0070] Next, the display process of the second character information in step S104 will be described with reference to the flowchart of FIG.
[0071] The processing of steps S301 to S303 in Fig. 9 is the same as the processing of steps S201 to S203 in Fig. 8, and therefore details will be omitted. The control unit 22 transmits a second signal including a voice signal of the second user U2 to the cloud server 60 (step S301). The control unit 22 acquires second character information from the cloud server 60, which is obtained by performing voice recognition on the voice signal included in the second signal and converting it into character information (S302). If necessary, the control unit 22 executes processing to translate the acquired second character information (S303).
[0072] In the example of Fig. 9, the control unit 22 does not perform highlighting processing and illustration processing on the second character information, which correspond to steps S204 and S205 of the first character information display processing of Fig. 8. In this embodiment, it is assumed that the second user U2 is the one receiving information and explanations from the first user U1. Therefore, in this embodiment, it is assumed that no processing is performed on the content of the utterance of the second user U2.
[0073] The control unit 22 transmits the second character information acquired in step S302 or the second character information in the second language translated in step S303 to the projector 30 as data of a character string to be displayed, and updates the display on the transparent screen 5 (step S304). The control unit 22 adds the character string of the latest second character information to the character string scrolled and displayed in the first display area A1 so that it is displayed as correct characters to the first user U1.
[0074] 7, unless the first user U1 instructs the end of the processing (step S105: No), the control unit 22 repeats the processing of steps S101 to S104. The first user U1 can instruct the display control device 20 to end the processing by, for example, turning off the power switch of the display control device 20. When the first user U1 instructs the end of the processing (step S105: Yes), the display control device 20 ends the processing.
[0075] When the first user U1 and the second user U2 alternately converse, the voice of the speaking user is subjected to speech recognition processing in real time, and the image displayed on the transparent screen 5 is sequentially updated. In the first display area A1, the contents of the conversations between the first user U1 and the second user U2 are sequentially added to the bottom, and text corresponding to the contents of past conversations is scrolled upward. By doing so, the displays shown in Figures 5 and 6 are possible.
[0076] As described above, the display system 1 shown in FIG. 3 is configured to convert the first signal from the first microphone 40a and the second signal from the second microphone 40b into a stereo signal by providing the stereo conversion unit 50. Furthermore, the control unit 22 of the display control device 20 is configured to separate the stereo signal into the first signal and the second signal. This makes it possible to input the first signal and the second signal to the display control device 20 even if the display control device 20 does not support input of multiple audio signals. Furthermore, the control unit 22 performs voice recognition processing on the first signal and the second signal as separate signals using the cloud server 60. Therefore, the display system 1 of the present disclosure can individually recognize voices from multiple users, convert them into text information, and display it in real time.
[0077] Although it has been described that the process ends when an instruction to end the process is received from the user in step S105, the control unit 22 may end the process as soon as the process of step S103 and / or step S104 is completed. In this case, the display control device 20 may start this control flow at regular intervals. That is, the control unit 22 may periodically and repeatedly execute the processes of steps S101 to S104.
[0078] (Second example of display control method) During a conversation between a first user U1 and a second user U2, both users may speak simultaneously. This is called a speech collision. A speech collision can occur, for example, when the timing of speech coincides or when one user interrupts the other. When a speech collision occurs, the speech of the first user U1 or the second user U2 displayed on the display unit may be cut off midway, or the page may scroll too quickly, resulting in unintended display. Therefore, the display control device 20 may prioritize the first signal or the second signal with the higher audio level and display it on the display unit. Below, a control method of the display control device 20 when performing such processing will be described with reference to FIGS. 10 to 12.
[0079] In the flow diagram of Fig. 10, steps S401, S402, S404, S406, and S407 are the same as S101, S102, S103, S104, and S105 in Fig. 7, respectively, and therefore description thereof will be omitted. The flow diagram of Fig. 10 differs from the flow diagram of Fig. 7 in that a signal level adjustment process (step S403) for a first signal is executed before a first character information display process (step S404). The flow diagram of Fig. 10 also differs from the flow diagram of Fig. 7 in that a signal level adjustment process (step S405) for a second signal is executed before a second character information display process (step S406).
[0080] The signal level adjustment process for the first signal (step S403) will be described with reference to Fig. 11. In this process, the display control device 20 has a first counter, which is set to a predetermined value in advance.
[0081] The control unit 22 refers to the first counter, and if the value of the first counter is greater than 0 (S501: Yes), the process proceeds to step S502.
[0082] In step S502, the control unit 22 compares the signal level of the first signal with the signal level of the second signal. When the voice of the first user U1 is louder than the voice of the second user U2, the signal level of the first signal is higher than the signal level of the second signal.
[0083] If the signal level of the first signal is equal to or greater than the signal level of the second signal (step S502: Yes), the control unit 22 sets the value of the first counter to a predetermined set value (step S503). If the value of the first counter is the predetermined set value, the control unit 22 does not change the value of the first counter.
[0084] If the signal level of the first signal is lower than the signal level of the second signal (step S502: No), the control unit 22 decrements the value of the first counter by 1 (step S504).
[0085] On the other hand, when the value of the first counter becomes 0 in step S501 (S501: No), the control unit 22 proceeds to step S505.
[0086] In step S505, the control unit 22 compares the signal level of the first signal with the signal level of the second signal.
[0087] When the signal level of the first signal is equal to or higher than the signal level of the second signal (step S505: Yes), the control unit 22 sets the value of the first counter to the set value (step S506).
[0088] If the signal level of the first signal is lower than the signal level of the second signal (step S505: No), the control unit 22 reduces the signal level of the first signal (step S507).
[0089] After steps S503, S504, and S506, if the signal level of the first signal is smaller than the predetermined value, the control unit 22 returns the signal level to the original level (step S508).
[0090] The signal level adjustment may be performed by changing the sensitivity of the first microphone 40a and the second microphone 40b. The volume adjustment values of the first microphone 40a and the second microphone 40b for setting the first microphone 40a and the second microphone 40b to the original signal level and the reduced signal level may be stored in the storage unit 23 and may be changeable later by settings, etc. An upper limit may be set for the original signal level.
[0091] After the processes of steps S507 and S508, the control unit 22 proceeds to the first character information display process S404 in FIG.
[0092] 11, control unit 22 compares the signal level of the first signal with the signal level of the second signal each time it performs adjustment processing of the first signal. If the state in which the signal level of the first signal is lower than the signal level of the second signal continues beyond a predetermined set value, the signal level of the first signal is set lower than the predetermined value (S507). On the other hand, if the signal level of the first signal becomes higher than the signal level of the second signal, the signal level of the first signal is returned to the original set value (step S508).
[0093] When the voice signal level included in the first signal falls below a predetermined value, the cloud server 60 stops the voice recognition process of the first signal in step S404. Therefore, if the state in which the signal level of the first signal is lower than the signal level of the second signal continues beyond the set value of the first counter, the speech content of the first user U1 will no longer be displayed on the display unit. This makes it possible to prevent the speech content of the first user U1 from being unintentionally displayed on the display unit, mixing the speech content of the second user U2 with the speech content of the first user U1.
[0094] The flowchart in FIG. 12 illustrates the signal level adjustment process for the second signal (step S405) in the flowchart in FIG. 10. In the flowchart in FIG. 12, a second counter different from the first counter is used instead of the first counter. The flowchart in FIG. 12 is similar to the flowchart in FIG. 11, except that the first signal and the second signal are swapped. Thus, according to the process in the flowchart in FIG. 12, the control unit 22 compares the signal level of the first signal with the signal level of the second signal each time it performs the adjustment process for the second signal. If the state in which the signal level of the second signal is lower than the signal level of the first signal continues beyond the set value of the second counter, the signal level of the second signal is set lower than a predetermined value (S607). On the other hand, if the signal level of the second signal becomes higher than the signal level of the first signal, the signal level of the second signal is restored to its original value (step S608).
[0095] When the signal level of the second signal falls below a predetermined value, the cloud server 60 stops the speech recognition process of the speech signal included in the second signal in step S406. Therefore, if the signal level of the second signal remains lower than the signal level of the first signal for a long period of time, the speech content of the second user U2 will no longer be displayed on the display unit. This makes it possible to prevent the speech content of the second user U2 from being unintentionally displayed mixed with the speech content of the first user U1 on the display unit.
[0096] Furthermore, in this display control method, by appropriately setting a setting value for the first counter, speech recognition processing of the first signal is stopped when the speech level of the first user U1 remains lower than the speech level of the second user U2 for a predetermined number of consecutive times. Therefore, the predetermined setting value for the first counter corresponds to the time until speech recognition processing of the first signal is stopped. The same applies to speech recognition processing of the second signal. As a result, when the conversations of the first user U1 and the second user U2 temporarily overlap, even if the speech level of either the first user U1 or the second user U2 is relatively low, both the speech contents of the first user U1 and the second user U2 are displayed on the display unit for a certain period of time.
[0097] In the flowcharts of FIGS. 10 to 12, when a speech collision occurs, the control unit 22 stops the speech recognition process by the cloud server 60 by reducing the signal level of either the first signal or the second signal to a predetermined value or below. However, the method of stopping the display of the speech content of a user with a low speech level is not limited to this. For example, in step S507 of the flowchart of FIG. 11, the control unit 22 may stop the transmission of the first signal to the cloud server 60 instead of reducing the signal level of the first signal. Furthermore, when the signal level of the first signal remains low, the control unit 22 may not execute the display process of displaying the first text information in step S404. The same or a similar method may also be used to stop the speech of the second user U2 when the speech level of the second user U2 is low.
[0098] As described above, according to the display control method of the second example, even when a speech conflict occurs between the first user U1 and the second user U2, the display system 1 can display only the speech of the user with the higher voice level. In this way, for example, in a case where the second user U2 is nodding and speaking softly while the first user U1 is explaining, or in a case where the second microphone 40b unintentionally detects the voice of a nearby third party, the speech of the first user U1 can be preferentially displayed.
[0099] In the present embodiment, steps S403 and S404 and steps S405 and S406 have been described as being processed in parallel, but steps S405 and S406 may be processed after steps S403 and S404 have been processed. Furthermore, in step S407, the processing has been described as ending when an instruction to end the processing is received from the user, but the control unit 22 may end the processing as soon as the processing of step S404 and / or step S406 is completed. In this case, the display control device 20 may start this control flow at regular intervals.
[0100] (Third example of display control method) While the first user U1 is speaking, the second microphone 40b may detect the voice of the first user U1 in addition to the first microphone 40a. Also, while the second user U2 is speaking, the first microphone 40a may detect the voice of the first user U1 in addition to the second microphone 40b. In such cases, the voice of the same user may be included in the first signal and the second signal, resulting in duplicate display of identical voice-recognized text information and the incorrect speaker being displayed on the display unit. Therefore, the display control device 20 may compare the first text information recognized from the first signal with the second text information recognized from the second signal, and if there is a match of a predetermined percentage or more, display only one of the first text information and the second text information. For example, the control unit 22 may calculate a correlation value between the first text information and the second text information, and if the correlation value is higher than a predetermined value, prioritize and display the text information converted from the signal with the higher voice level on the display unit. Below, a control method of the display control device 20 when performing such processing is described with reference to the flowcharts of FIGS. 13 and 14.
[0101] In Fig. 13, steps S701 and S702 are the same as steps S101 and S102 in Fig. 7. Furthermore, steps S703 and S704 are the same as steps S201 and S202 in Fig. 8. Furthermore, steps S706 and S707 are the same as steps S301 and S302 in Fig. 9. Therefore, explanations of steps S701 to S704, S706, and S707 will be omitted.
[0102] In steps S705 and S708, control unit 22 waits to acquire the first character information or the second character information, whichever comes later. For example, control unit 22 waits until it acquires the first character information and the second character information for each sentence. The acquired character information may be written in hiragana, katakana, and / or kanji, or in English.
[0103] In step S709, the control unit 22 calculates a correlation value between the acquired first character information and second character information (step S709). The correlation value may be a numerical value evaluated based on the matching rate of characters and / or words, for example.
[0104] If the correlation value is equal to or less than the predetermined threshold (step S710: No), the control unit 22 executes the processes of steps S711 to S714 for the first character information and the processes of steps S715 and S716 for the second character information in parallel. The processes of steps S711 to S714 are the same as the processes of steps S203 to S206 in Fig. 8. The processes of steps S715 and S716 are the same as the processes of steps S303 and S304 in Fig. 9. Therefore, a description of the processes of steps S711 to S716 will be omitted.
[0105] On the other hand, if the correlation value is greater than the predetermined threshold (step S710: Yes), it is considered that the voice of either the first user U1 or the second user U2 is detected by both the first microphone 40a and the second microphone 40b. Therefore, the control unit 22 compares the signal level of the first signal with the signal level of the second signal (step S717). If the signal level of the first signal is higher than the signal level of the second signal at step S717 (S717: Yes), the control unit 22 executes the processes of steps S711 to S714 on the first character information. If the signal level of the first signal is equal to or lower than the signal level of the second signal at step S717 (S717: No), the control unit 22 executes the processes of steps S715 and S716 on the second character information.
[0106] After steps S714 and S716, unless the first user U1 instructs the end of the process (step S718: No), the control unit 22 returns to step S701. When the first user U1 instructs the end of the process (step S718: Yes), the control device 20 ends the process.
[0107] As described above, according to the third example display control method, the control unit 22 calculates a correlation value between the first character information and the second character information obtained by performing voice recognition processing on the voices of the first user U1 and the second user U2. If the correlation value is equal to or less than a predetermined threshold, the first character information and the second character information are generated by different users, and the control unit 22 processes the first character information and the second character information separately and displays them on the display unit. On the other hand, if the correlation value is greater than the predetermined threshold, the first character information and the second character information are considered to be generated by the same user. In this case, the control unit 22 displays only the character information corresponding to the voice signal with the higher signal level on the display unit. Generally, the voice signal detected by the microphone closer to the speaking user has a higher signal level than the voice signal detected by the microphone farther away. Therefore, even if the second microphone 40b detects the voice of the first user U1, only the first character information based on the first signal detected by the first microphone 40a is displayed on the display unit. Similarly, even if the first microphone 40a detects the voice of the second user U2, only the second character information based on the second signal detected by the second microphone 40b is displayed on the display unit. In this way, the control unit 22 can prevent the same character information from being displayed twice and the wrong speaker from being displayed on the display unit.
[0108] Although it has been described that the process ends when an instruction to end the process is received from the user in step S718, the control unit 22 may end the process upon completion of the process in step S714 and / or step S716. In this case, the display control device 20 may start this control flow at regular intervals.
[0109] In the third example of the control method described above, the first character information and the second character information resulting from speech recognition of the voice signals detected by the first microphone 40a and the second microphone 40b are compared to determine whether they are from the same user. However, the method for determining whether the voice signals detected by the first microphone 40a and the second microphone 40b are from the same user is not limited to this.
[0110] For example, the control unit 22 may compare frequency information between a first signal acquired from the first microphone 40a and a second signal acquired from the second microphone 40b, and if the correlation value is equal to or greater than a predetermined value, determine that the audio signals contained in the first signal and the second signal are from the same user. If the average signal levels do not match when calculating the correlation value, the correlation value will be small. Therefore, it is preferable to calculate the correlation value after matching the overall signal levels by calculating the average signal level ratio of the acquired first and second signals and normalizing the signal levels.
[0111] Furthermore, when comparing the first and second signals, the audio signal may be delayed depending on the distance from the speaker, depending on the microphone installation position. A time delay between the first and second signals may affect the accuracy of calculating the correlation value of the audio signals. Therefore, the display system 1 may be equipped with a calibration mode to measure the delay of the first microphone 40a and the second microphone 40b in advance. Measurement is performed by turning on the first microphone 40a and the second microphone 40b and having the first user U1 and the second user U2 speak at a higher signal level than normal so that the same voice is detected by both microphones. By checking the recorded data and comparing the start times of the audio signals generated by the speech, the extent of the delay due to the positions of the first microphone 40a and the second microphone 40b can be confirmed. Furthermore, by measuring the difference in signal level between the first microphone 40a and the second microphone 40b, a signal level threshold for the audio signal not to be displayed on the display unit may be determined.
[0112] (Fourth example of display control method) When a first user U1 and a second user U2 speak simultaneously during a conversation, it is rare for them to continue the conversation. Therefore, it is often preferable to prioritize and display the content of the speech of the user who has been speaking for a long time. Therefore, when the speech of the first user U1 and the second user U2 is compared and speech is input simultaneously, the control unit 22 may display only the speech that has the larger number of characters or words converted from the speech signal on the display unit. Below, a control method of the display control device 20 when performing such processing will be described with reference to the flow charts of FIGS. 15 and 16.
[0113] 15 and 16, steps S801 to S806 are the same as steps S701 to S704, S706 and S707, respectively, in Fig. 13. Therefore, a description of steps S801 to S806 will be omitted.
[0114] In step S807, the control unit 22 calculates the information amount n1 of the first character information and the information amount n2 of the second character information within a specific speech interval. The speech interval refers to a period during which the audio signal of the first signal or the second signal continues, including the period during which the audio signals overlap when the audio signals included in the first signal and the second signal overlap in time. The speech interval may correspond to a longer sentence of speech between the first user U1 and the second user U2. The information amount n1 of the first character information and the information amount n2 of the second character information may be the number of characters of the first character information and the number of characters of the second character information, respectively.
[0115] If the information amount n1 of the first character information is greater than the information amount n2 of the second character information (step S808: Yes), the control unit 22 executes the processes of steps S809 to S812 on the first character information. If the information amount n1 of the first character information is equal to or less than the information amount n2 of the second character information (step S808: No), the control unit 22 executes the processes of steps S813 and S814 on the second character information. The processes of steps S809 to S812 are the same as the processes of steps S203 to S206 in FIG. 8. Furthermore, the processes of steps S813 and S814 are the same as the processes of steps S303 and S304 in FIG. 9. Therefore, a description of the processes of steps S809 to S814 will be omitted.
[0116] After steps S812 and S814, unless the first user U1 instructs the end of the process (step S815: No), the control unit 22 returns to step S801 and repeats the above process. When the first user U1 instructs the end of the process (step S815: Yes), the control device 20 ends the process.
[0117] As described above, according to the display control method of the fourth example, the control unit 22 calculates the information amounts n1 and n2 of the first character information and the second character information obtained by performing voice recognition processing on the voices of the first user U1 and the second user U2. The control unit 22 compares the information amounts n1 and n2 and performs display processing only on the character information with the larger information amount n1 or n2. By doing so, even when the first user U1 and the second user U2 speak simultaneously, the control unit 22 can display on the display unit only the speech of the user who continued the conversation.
[0118] Regardless of the above, the control unit 22 may be configured to display the first character information or the second character information with the smaller amount of information on the display unit when the amount of information of the first character information or the second character information with the smaller amount of information is equal to or greater than a predetermined value. In this case, the control unit 22 may display the character information with the smaller amount of information before and after the character information with the larger amount of information, without inserting the character information with the smaller amount of information in the middle of the character information with the larger amount of information. That is, for example, if the second user U2 speaks during a period in which the first user U1 is speaking, the control unit 22 may display the content of the second user U2's utterance before or after the display of the entire content of the first user U1's utterance on the display unit. In this way, any utterance of a certain length or more can be displayed on the display unit.
[0119] In the fourth example of the display control method, the information amount n1 of the first character information and the information amount n2 of the second character information are the number of characters. However, the information amounts n1 and n2 are not limited to the number of characters. For example, the information amounts n1 and n2 may be the number of words.
[0120] Furthermore, instead of the amounts of information n1 and n2, the control unit 22 may compare the duration of the speech between the first user U1 and the second user U2. That is, when the control unit 22 simultaneously acquires a first signal and a second signal containing a voice signal, the control unit 22 detects the duration of the voice signal contained in the first signal and the second signal from the first signal and the second signal. The control unit 22 may display text information based on the signal containing the longer voice signal, out of the first signal and the second signal, on the display unit.
[0121] Although it has been described that the process ends when an instruction to end the process is received from the user in step S815, the control unit 22 may end the process upon completion of the process in step S812 and / or step S814. In this case, the display control device 20 may start this control flow at regular intervals.
[0122] Two or more of the methods described in the second to fourth examples of the display control method may be used in combination.
[0123] The following describes variations of the display that can be displayed on the transparent screen 5 of the display system 1.
[0124] (Example of the second display on the display system) A second display example of the display system 1 is shown in Fig. 17. Fig. 17 shows the display of the second display example on the transparent screen 5 as seen from the side of the second user U2. The second display example of the display system 1 is the first display example shown in Figs. 5 and 6, except that a horizontal line indicating a boundary is displayed between the character string based on the first character information and the character string based on the second character information displayed in the first display area A1. That is, according to this display method, the boundary is indicated by a line each time a different user speaks the character information displayed in the first display area A1.
[0125] Figures 18 and 19 are flow charts executed by the control unit 22 to display the second display example. The first character information display process in Figure 18 and the second character information display process in Figure 19 are executed as the processes of steps S103 and S104 in the flow chart of Figure 7, instead of the processes of the flow charts of Figures 8 and 9.
[0126] Steps S901 to S905 of the first character information display process in FIG. 18 are the same as steps S201 to S205 in the flowchart of FIG. 8, and therefore a description thereof will be omitted.
[0127] In step S906, the control unit 22 determines whether a character string corresponding to the second character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the first character information is to be displayed next (step S906).
[0128] If a character string corresponding to the second character information was previously displayed (step S906: Yes), the control unit 22 displays a line that serves as a boundary line immediately below the character string corresponding to the immediately preceding second character information (step S907).The control unit 22 displays a character string corresponding to the first character information below the line that serves as the boundary line, and displays an image for illustration in the second display area A2 if necessary (step S908).
[0129] In step S906, if a character string corresponding to the second character information has not been displayed immediately before (step S906: No), control unit 22 does not place a line that serves as a boundary line. Control unit 22 displays the character string corresponding to the first character information in first display area A1 as is, and displays an image in second display area A2 as necessary (step S908).
[0130] Steps S1001 to S1003 of the second character information display process in FIG. 19 are the same as steps S301 to S303 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0131] In step S1004, the control unit 22 determines whether a character string corresponding to the first character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the second character information is to be displayed next.
[0132] If a character string corresponding to the first character information has been displayed immediately before (step S1004: Yes), control unit 22 displays a line that serves as a boundary line immediately below the character string corresponding to the immediately preceding first character information (step S1005).Control unit 22 then displays a character string corresponding to the second character information below the line that serves as the boundary line (step S1006).
[0133] In step S1004, if a character string corresponding to the first character information has not been displayed immediately before (step S1004: No), the control unit 22 displays a character string corresponding to the second character information in the first display area A1 without placing a straight line as a boundary line (step S1006).
[0134] In this way, breaks in the conversation can be displayed in an easy-to-understand manner for the first user U1 and the second user U2 in the first display area A1 of the display unit.
[0135] (Third display example of the display system) A third display example of the display system 1 is shown in FIG. 20. FIG. 20 shows the third display example of the transparent screen 5 as seen from the second user U2 side. In the third display example of the display system 1, the most recent character string among the character strings corresponding to the first character information or the second character information displayed in the first display area A1 in the first display example shown in FIGS. 5 and 6 is highlighted in a specific manner. The specific manner includes, for example, using a priority color for the display color of the character string. In FIG. 20, the character string marked "highlighted" is the most recent character string highlighted relative to the character string marked "normal display."
[0136] As an example, if the first display area A1 is blank, when the first user U1 or the second user U2 starts a conversation, the text information is displayed in the priority color. As long as the same user's voice continues, the first display area A1 is displayed in the priority color. When the user who is speaking changes, the previous conversation is changed to the normal color, and only the most recent conversation is displayed in the priority color. The normal color can be, for example, white. The priority color can be, for example, yellow.
[0137] Specific ways to highlight a character string include changing the color, making the character bold, increasing the character size, changing the font style, changing the transparency of the character, etc.
[0138] Figures 21 and 22 are flowcharts executed by the control unit 22 to display the third display example. The processing of the flowcharts of the first character display processing in Figure 21 and the second character display processing in Figure 22 is executed as the processing of step S103 and step S104 in the flowchart of Figure 7, instead of the processing of the flowcharts of Figures 8 and 9.
[0139] Steps S1101 to S1105 in FIG. 21 are the same as steps S201 to S205 in the flowchart of FIG. 8, and therefore a description thereof will be omitted.
[0140] In step S1106, the control unit 22 determines whether a character string corresponding to the second character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the first character information is to be displayed next.
[0141] If a character string corresponding to the second character information was displayed immediately before (step S1106: Yes), the control unit 22 displays the character string of the previously displayed character information in the normal display mode (S1107).The control unit 22 displays a character string corresponding to the latest first character information in a specific mode below the previous character string displayed in the normal display mode, and displays an image in the second display area A2 if necessary (step S1108).
[0142] In step S1106, if a character string corresponding to the second character information has not been displayed immediately before (step S1106: No), a character string corresponding to the first character information is displayed in a specific manner in the first display area A1, and an image is displayed in the second display area A2 if necessary (step S1108). In this case, a character string corresponding to the first character information displayed in a specific manner corresponding to the most recent utterance is added below the character string corresponding to the first character information displayed in a specific manner corresponding to the past utterance.
[0143] Steps S1201 to S1203 in FIG. 22 are the same as steps S301 to S303 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0144] In step S1204, the control unit 22 determines whether a character string corresponding to the first character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the second character information is to be displayed next.
[0145] If a character string corresponding to the first character information was displayed immediately before (step S1204: Yes), control unit 22 displays the character string of the previously displayed character information in a normal display (S1205).Control unit 22 displays a character string corresponding to the latest second character information in a specific manner below the previous character string displayed in the normal display (step S1206).
[0146] In step S1204, when a character string corresponding to the first character information has not been displayed immediately before (step S1204: No), the previously displayed character information is not displayed normally, and a character string corresponding to the second character information is displayed in a specific manner in the first display area A1 (step S1206). In this case, a character string corresponding to the second character information displayed in a specific manner corresponding to the latest utterance is added below the character string corresponding to the second character information displayed in a specific manner corresponding to the past utterance.
[0147] In this way, the first display area A1 always displays a character string corresponding to the latest first character information or second character information in a specific manner. This emphasizes the latest conversation between the first user U1 and the second user U2, allowing the first user U1 and the second user U2 to find the latest conversation at a glance. As a result, the first user U1 and the second user U2 can concentrate on the latest conversation.
[0148] (Fourth display example of the display system) 23 and 24 show a fourth display example of the transparent screen 5 of the display system 1. FIGS. 23 and 24 show views of the sequential display of conversation content in the fourth display example as seen from the second user U2 side. In the fourth display example of the display system 1, the third display area A3 in the first display example shown in FIGS. 5 and 6 can be used by the second user U2 to check whether the content of his or her speech has been correctly recognized.
[0149] More specifically, as shown in FIG. 23 , the control unit 22 of the display control device 20 displays the first character information converted from the most recent utterance by the first user U1 in the first display area A1 in the second language so that it appears as correct characters when viewed from the second user U2. The control unit 22 also displays the first character information in the third display area A3 in the first language so that it appears as correct characters when viewed from the first user U1. Furthermore, as shown in FIG. 24 , the control unit 22 displays the second character information converted from the most recent utterance by the second user U2 in the first language so that it appears as correct characters when viewed from the first user U1, and in the second language so that it appears as correct characters when viewed from the second user U2. Also, in FIGS. 23 and 24 , the language from which the display in the first display area A1 has been converted is displayed. For example, the vertical display of “English” and “Japanese” in the third display area A3 indicates that the display in the first display area A1 has been translated from English to Japanese.
[0150] The display methods of Figures 23 and 24 are characterized in that they allow not only the first user U1, who is the counter clerk, but also the second user U2, who is the visitor, to check whether the content of the utterance has been correctly recognized by voice.
[0151] Figures 25 and 26 are flow diagrams of processing executed by the control unit 22 to display the fourth display example. The flow diagrams of the first character information display processing in Figure 25 and the second character information display processing in Figure 26 are executed as the processing of steps S103 and S104 in the flow diagram of Figure 7, instead of the processing of the flow diagrams of Figures 8 and 9.
[0152] The processing of steps S1301 to S1305 in Fig. 25 is the same as the processing of steps S201 to S205 in the flowchart of Fig. 8, and therefore description thereof will be omitted. Step S1306 in Fig. 25 is the same as step S206 in the flowchart of Fig. 8, except that processing for displaying first character information in the first language in the third display area A3 is not performed.
[0153] In step S1307, the control unit 22 clears the display in the third display area A3 and displays the first character information in the first language in the correct characters toward the first user U1. Furthermore, the control unit 22 displays in the third display area A3 a display indicating that the display translated from the first language to the second language is displayed in the first display area A1.
[0154] The processing from steps S1401 to S1404 in FIG. 26 is the same as the processing from steps S301 to S304 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0155] In step S1405, the control unit 22 clears the display in the third display area A3 and displays the second character information in the second language in the correct characters toward the second user U2. That is, the control unit 22 displays the second character information in the third display area A3 in inverted characters. Furthermore, the control unit 22 displays in the third display area A3 an indication that the display translated from the second language to the first language is displayed in the first display area A1.
[0156] In this way, the display system 1 displays text information of the speech recognition result to the current speaker of the first user U1 or the second user U2. This allows both the first user U1 and the second user U2 to confirm that speech recognition is being performed correctly for the content of their current speech when they each speak. Furthermore, in this display method, the third display area A3, which displays the speech recognition result of the current speech content, is shared by the first user U1 and the second user U2, allowing for efficient use of the display area of the display unit.
[0157] (Illustration removed) The display system 1 may have a function of erasing an image displayed in the second display area A2. For example, the control unit 22 of the display control device 20 may execute the process shown in the flowchart of Fig. 27 between steps S102 and S103 of the flowchart of Fig. 7. The flowchart of Fig. 27 will be described below.
[0158] First, the control unit 22 monitors whether or not a command to clear an image (image clear command) is included in the first signal acquired from the first microphone 40a via the stereo conversion unit 50 (S1501). For example, when the first user U1 wants to erase the image displayed in the second display area A2 of the transparent screen 5, the first user U1 verbally issues a voice command such as "image clear." The control unit 22 has a function of recognizing a specific voice command included in the first signal.
[0159] When the control unit 22 detects an image clear command from the first signal (step S1502: Yes), it erases the image displayed in the second display area A2 (step S1503). In this case, the control unit 22 may transmit the first signal, from which the audio of the image clear command has been removed, to the cloud server 60 in step S103 of Fig. 7. This prevents the image clear command uttered by the first user U1 from being displayed as text information in the first display area A1.
[0160] If the control unit 22 does not detect an image clear command from the first signal (step S1502: No), it does not perform any particular process.
[0161] Alternatively, the control unit 22 may transmit a first signal including an image clear command to the cloud server 60, and erase the image in the second display area A2 when the first character information generated by the voice recognition includes the image clear command. Such processing may be executed, for example, between steps S202 and S203 in the flowchart of FIG. 8.
[0162] As described above, the display system 1 allows the first user U1 to erase the image in the second display area A2 displaying the illustration by voice operation. Since the operation is performed by voice rather than by operating the input unit 24 of the display control device 20, the first user U1 can concentrate on the conversation. For example, if the input unit 24 is an external input device to the display control device 20 and has many buttons corresponding to the operations to be performed by the first user, the first user U1 may have difficulty operating the input unit 24. By employing voice operation, it is possible to prevent the conversation from being interrupted due to time consuming button operations. Furthermore, since the previous conversation content displayed in the first display area A1 is not erased, the first user U1 and the second user U2 can continue their conversation without any inconvenience while continuing to view the transparent screen 5.
[0163] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art would easily be able to make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included within the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies, and multiple components can be combined or divided into one.
[0164] In this disclosure, the terms "first" and "second" are identifiers for distinguishing the configuration. In this disclosure, the configurations distinguished by terms such as "first" and "second" can have their numbers swapped. For example, the first microphone 40a can swap the identifiers "first" and "second" with the second microphone 40b. The identifier swapping is performed simultaneously. The configurations remain distinguished even after the identifier swapping. Identifiers may be deleted. A configuration from which an identifier has been deleted is distinguished by a symbol. The identifiers "first" and "second" in this disclosure should not be used solely to interpret the order of the configurations or to justify the existence of an identifier with a smaller number.
[0165] The components and functional blocks included in each embodiment of the present disclosure may be arranged in the same hardware or different hardware as appropriate. The stereo conversion unit 50 in each embodiment of the present disclosure may be built into the display control device 20. The display control device 20 and the projector 30 may be integrally configured with the same hardware. Some or all of the communication unit 21, control unit 22, storage unit 23, and input unit 24 of the display control device 20 in each embodiment of the present disclosure may be included in the projector 30.
[0166] The display unit of the present disclosure is not limited to a configuration including the transparent screen 5 and the projector 30. The display unit may include, for example, a transparent organic light-emitting diode (OLED) display in which light-emitting elements are arranged within a transparent substrate.
[0167] In the above embodiment, the speech recognition process and the translation process are performed by the cloud server 60 located in a remote location, but the present invention is not limited to this. The speech recognition process and the translation process may be performed by separate servers rather than by the same server located in a remote location. At least one of the speech recognition process and the translation process may be performed by a server located near the display control device 20, or by the display control device 20 itself. [Explanation of symbols]
[0168] 1 Display System 5 Transparent screen (display) 10 Base material 20 Display control device 21 Communications Department 22 Control Unit 23 Memory section 24 Input section 30 Projector (display unit) 40a 1st microphone 40b Second microphone 50 Stereo conversion section 60 Cloud Servers U1 First User U2 Second user A1 1st display area A2 Second display area A3 3rd display area
Claims
1. a display unit disposed between a plurality of users when in use; a first microphone disposed on a first side of the display; a second microphone disposed on a second side of the display unit opposite to the first side; a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into a stereo signal; a display control device that receives the stereo signal, separates the stereo signal into the first signal and the second signal, and displays on the display unit first character information and second character information obtained by converting audio signals included in the first signal and the second signal by voice recognition; A display system comprising:
2. 2. The display system according to claim 1, wherein the display control device displays the first character information and the second character information on the display unit while scrolling them sequentially, and displays a boundary line between the first character information and the second character information.
3. The display system according to claim 1 , wherein the display control device converts the first character information in a first language into a second language different from the first language and displays the converted information on the display unit.
4. 4. The display system according to claim 3, wherein the display control device displays the first character information converted from the latest utterances by the plurality of users as a character string in the second language toward the second side and as a character string in the first language toward the first side, and displays the second character information converted from the latest utterances by the plurality of users as a character string in the first language toward the first side and as a character string in the second language toward the second side.
5. The display system according to claim 1, wherein the display control device displays the latest first character information or the latest second character information converted from speech uttered by one or more of the plurality of users on the display unit in a predetermined manner different from the display manner of other character information.
6. A display system described in any one of claims 1 to 5, wherein when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the audio levels of the audio signals contained in the first signal and the second signal, and displays only one of the first character information and the second character information on the display unit based on the comparison result.
7. The display system according to claim 6, wherein the display control device displays on the display unit the first character information or the second character information obtained by converting the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has the higher audio level.
8. 6. A display system according to claim 1, wherein when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device compares the first character information with the second character information, and if they match by a predetermined percentage or more, displays only one of the first character information and the second character information.
9. The display system according to claim 8 , wherein the display control device causes the display unit to display the first character information or the second character information obtained by voice recognition of the first signal or the second signal, whichever has a higher voice level.
10. 6. The display system according to claim 1, wherein when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device compares the first character information and the second character information corresponding to the speech within the speech section including the period in which the audio signals overlap, and causes the display unit to display only the first character information or the second character information that has the greater amount of information.
11. a display unit disposed between a plurality of users during use; and a communication unit capable of communicating with a stereo conversion unit that converts, into stereo signals, a first signal input from a first microphone disposed on a first side of the display unit and a second signal input from a second microphone disposed on a second side opposite to the first side of the display unit; a control unit that receives the input of the stereo signal, separates the stereo signal into the first signal and the second signal, and executes a process of displaying, on the display unit, first character information and second character information obtained by converting audio signals included in the first signal and the second signal by voice recognition; A display control device comprising:
12. A program for a display control device that controls a display system including a display unit that is placed between multiple users when in use, a first microphone that is placed on a first side of the display unit, a second microphone that is placed on a second side opposite the first side of the display unit, and a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, the program causing a control unit of the display control device to execute a process of receiving input of the stereo signals, separating the stereo signals into the first signal and the second signal, and displaying first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal, respectively, on the display unit.
Citation Information
Patent Citations
Conversation assistance apparatus, display system, information processing method, and display method
JP2023124615A