Display system, display control device, and display control program
The display system addresses the challenge of processing multiple user voices by using a stereo conversion unit to separate and recognize audio signals, facilitating real-time text display and managing overlapping speech, enhancing communication across language barriers.
Patent Information
- Application Number
- PCT/JP2024/035335
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2024-10-02
- Publication Date
- 2025-09-04
AI Technical Summary
Existing display systems struggle to individually recognize and display two-way dialogue between multiple users due to limitations in device ports, such as USB terminals or audio jacks, making it difficult to connect multiple microphones and process multiple audio signals simultaneously.
A display system with a stereo conversion unit that combines signals from multiple microphones into a stereo signal, which is then separated into individual signals for voice recognition and displayed on a transparent screen, allowing real-time conversion and display of text information in different languages or without translation, with control processes to manage overlapping speech and prioritize audio levels.
Enables real-time, individual recognition and display of multiple users' voices as text, supporting seamless communication across language barriers and handling simultaneous speech with clarity and accuracy.
Smart Images

Figure JP2024035335_04092025_PF_FP_ABST
Abstract
Description
Display system, display control device, and display control program CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to Japanese Patent Application No. 2024-030572, filed on February 29, 2024, the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to a display system, a display control device, and a display control program.
[0003] A system has been proposed that supports user communication in situations where multiple people are conversing across a partition (see, for example, Patent Document 1). In this system, a transparent display that functions as a partition is placed between two people. Furthermore, microphones that pick up the voices of the people speaking are placed on both sides of the transparent display. This allows an information processing device in the display system to convert the content of the conversation spoken into each microphone into text information and display it on the transparent display in real time.
[0004] Japanese Patent Application Laid-Open No. 2023-124615
[0005] In one embodiment, (1) a display system includes a display unit that is arranged between multiple users when in use, a first microphone that is arranged on a first side of the display unit, a second microphone that is arranged on a second side opposite the first side of the display unit, a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, and a display control device that receives the stereo signals, separates the stereo signals into the first signal and the second signal, and displays first character information and second character information that are converted from audio signals included in the first signal and the second signal by voice recognition on the display unit.
[0006] (2) In the display system of (1) above, the display control device can display the first character information and the second character information while scrolling them sequentially on the display unit, and can display a boundary line between the first character information and the second character information.
[0007] (3) In the display system of (1) or (2) above, the display control device can convert the first character information in a first language into a second language different from the first language and display it on the display unit.
[0008] (4) In the display system of (3) above, the display control device can display the first character information converted from the latest utterances by the multiple users as a string of characters in the second language toward the second side and as a string of characters in the first language toward the first side, and can display the second character information converted from the latest utterances by the multiple users as a string of characters in the first language toward the first side and as a string of characters in the second language toward the second side.
[0009] (5) In any of the display systems (1) to (4) above, the display control device can display the latest first character information or the latest second character information converted from speech uttered by one or more of the plurality of users on the display unit in a predetermined manner different from the display manner of other character information.
[0010] (6) In any of the display systems (1) to (5) above, when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the audio levels of the audio signals contained in the first signal and the second signal, and based on the comparison result, can display only one of the first character information and the second character information on the display unit.
[0011] (7) In the display system of (6) above, the display control device can display on the display unit the first character information or the second character information obtained by converting the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has the higher audio level.
[0012] (8) In any of the display systems (1) to (5) above, when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device can compare the first character information with the second character information and, if they match by a predetermined percentage or more, display only one of the first character information and the second character information.
[0013] (9) In the display system of (8) above, the display control device can display on the display unit the first character information or the second character information obtained by voice recognition of the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has a higher audio level.
[0014] (10) In any of the display systems (1) to (5) above, when the audio signal included in the first signal and the audio signal included in the second signal overlap in time, the display control device can compare the first character information and the second character information corresponding to the utterance within the speech section including the period in which the audio signals overlap, and display only the first character information or the second character information that contains the greater amount of information on the display unit.
[0015] In one embodiment, (11) a display control device includes: a display unit that is arranged between multiple users when in use; a communication unit that can communicate with a stereo conversion unit that converts a first signal input from a first microphone arranged on a first side of the display unit and a second signal input from a second microphone arranged on a second side opposite the first side of the display unit into a stereo signal; and a control unit that receives the input of the stereo signal, separates the stereo signal into the first signal and the second signal, and executes a process to display on the display unit first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal.
[0016] In one embodiment, (12) the display control program is a program of a display control device that controls a display system including a display unit that is placed between multiple users during use, a first microphone that is placed on a first side of the display unit, a second microphone that is placed on a second side opposite the first side of the display unit, and a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, and causes a control unit of the display control device to execute a process of receiving input of the stereo signals, separating the stereo signals into the first signal and the second signal, and displaying first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal, respectively, on the display unit.
[0017] 10 is a diagram illustrating an example of a usage scene of a display system according to an embodiment. FIG. 1 is a diagram illustrating a top view of the display system of FIG. 1. FIG. 1 is a block diagram illustrating a schematic configuration of the display system of FIG. 1. FIG. 1 is a diagram illustrating an overview of information processing by the display system of FIG. 1. FIG. 1 is a diagram illustrating an example of a display on a transparent screen as seen from a first user's side. FIG. 1 is a diagram illustrating an example of a display on a transparent screen as seen from a second user's side. FIG. 11 is a flow chart illustrating a first example of a display control process executed by a control unit of a display control device. FIG. 12 is a flow chart illustrating a first example of a first character information display process of FIG. 7. FIG. 13 is a flow chart illustrating a first example of a second character information display process of FIG. 7. FIG. 14 is a flow chart illustrating a second example of a display control process executed by a control unit of a display control device. FIG. 15 is a flow chart illustrating an example of a signal level adjustment process of a first signal of FIG. 10. FIG. 16 is a flow chart illustrating an example of a signal level adjustment process of a second signal of FIG. 15. FIG. 16 is a flow chart illustrating a third example of a display control process executed by a control unit of a display control device. FIG. 17 is a flow chart illustrating a third example of a display control process executed by a control unit of a display control device. FIG. 18 is a flow chart illustrating a fourth example of a display control process executed by a control unit of a display control device. FIG. 19 is a flow chart illustrating a fourth example of a display control process executed by a control unit of a display control device. FIG. 19 is a flow chart illustrating a second example of a display on a transparent screen as seen from a second user's side. FIG. 18 is a flow chart showing a second example of the first character information display process corresponding to FIG. 17. FIG. 18 is a flow chart showing a second example of the second character information display process corresponding to FIG. 17. FIG. 18 is a diagram showing a third example of the display of the transparent screen as seen from the side of the second user. FIG. 18 is a flow chart showing a third example of the first character information display process corresponding to FIG. 20. FIG. 18 is a flow chart showing a third example of the second character information display process corresponding to FIG. 20. FIG. 18 is a diagram showing a fourth example of the display of the transparent screen as seen from the side of the second user. FIG. 18 is a flow chart showing a fourth example of the display of the transparent screen as seen from the side of the second user. FIG. 18 is a flow chart showing a fourth example of the first character information display process corresponding to FIGS. 23 and 24. FIG. 18 is a flow chart showing a fourth example of the second ...
[0018] In a system supporting communication between multiple users, multiple microphones are required to distinguish between each user and visualize and display the two-way dialogue between the users. However, devices that receive voice information input may not be able to use multiple device ports for voice input and may only be able to use one device port. For example, some Android (registered trademark) devices are designed to recognize only one device port (a Universal Serial Bus (USB) terminal or an audio jack), making it difficult to connect multiple microphones to a single device port. Even if multiple microphones are connected to a single device port via a multi-tap (USB hub), some Android devices are designed to recognize only one device, making it difficult to connect multiple microphones. In such cases, it is difficult to individually recognize the voices of multiple users and display the two-way dialogue on a display device.
[0019] A display system, a display control device, and a display control program according to an embodiment of the present disclosure described below are capable of individually recognizing voices from multiple users, converting them into text information, and displaying the text information.
[0020] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The drawings used in the following description are schematic. The dimensions and ratios in the drawings do not necessarily correspond to the actual dimensions and ratios.
[0021] (Configuration of the Display System) As shown in FIGS. 1 and 2 , a display system 1 according to one embodiment is a system that supports a dialogue between a first user U1 and a second user U2. As an example, the first user U1 is a staff member at a tourist information center. The second user U2 is a traveler visiting the tourist information center. In the following embodiment, the first user U1 speaks a first language. The first language is, for example, Japanese. The second user U2 speaks a second language different from the first language. The second language is, for example, English. The display system 1 converts the speech of the first user U1 into the second language and the speech of the second user U2 into the first language, and displays the converted speech. However, the display system 1 is not limited to a system that includes such speech translation. The display system 1 may display the speech of the first user U1 and the second user U2 without translating them into another language. In this case, the first user U1 may be, for example, a staff member at a government, local government, or public institution. The second user U2 may be, for example, an ordinary citizen visiting a government, local government, or public institution. The display system 1 may also be used at the counters of financial institutions, medical institutions, public transportation facilities, private business offices, and office conference rooms.
[0022] An acrylic plate, vinyl curtain, or the like is installed between the first user U1 and the second user U2. In recent years, acrylic plates or vinyl curtains have been installed to prevent droplets from spreading. In the display system 1, the acrylic plate, vinyl curtain, or the like is used as a base material 10, and a transparent screen 5 is placed on the base material 10. The projector 30 projects the speech of the first user U1 and the second user U2 onto the transparent screen 5. This allows the first user U1 and the second user U2, who are conversing across the transparent screen 5, to visually confirm the speech even if they cannot hear what the other user is saying. The first user U1 side of the transparent screen 5 is the first side. The second user U2 side is the second side.
[0023] The transparent screen 5 can be placed anywhere on the base material 10. The transparent screen 5 may be placed in a position that does not prevent the first user U1 and the second user U2 from seeing each other's faces. For example, the transparent screen 5 may be placed in a position where the first user U1 and the second user U2 lower their lines of sight downward from the horizontal direction.
[0024] As shown in FIGS. 1 to 3 , the display system 1 includes a transparent screen 5, a display control device 20, a projector 30, a first microphone 40a, a second microphone 40b, and a stereo conversion unit 50. The transparent screen 5 and the projector 30 constitute a display unit. Hereinafter, the transparent screen 5 and the projector 30 may be collectively referred to as the display unit. The first microphone 40a and the second microphone 40b are each connected to the stereo conversion unit 50 so as to be able to communicate with each other via a wired or wireless connection. The stereo conversion unit 50 and the display control device 20 are connected so as to be able to communicate with each other via a wired or wireless connection. The display control device 20 and the projector 30 are connected so as to be able to communicate with each other via a wired or wireless connection. The display control device 20 is configured to be able to communicate with an external cloud server 60 that provides voice recognition processing via a communication line.
[0025] The transparent screen 5 can be a film-, sheet-, or plate-like member for projector projection that can be attached to the base material 10. The transparent screen 5 diffuses a portion of the light incident from the projector 30 to the incident side and the exit side. The first user U1 and the second user U2 can recognize the image projected from the projector 30 when the light diffused by the transparent screen 5 enters their field of vision. The shape of the transparent screen 5 can be, for example, rectangular, but is not limited to this. The transparent screen 5 can be in various shapes.
[0026] The first microphone 40a and the second microphone 40b are monaural microphones. The first microphone 40a is disposed on the side of the transparent screen 5 facing the first user U1, i.e., the first side. For example, the first microphone 40a may be disposed on a desk in front of the first user U1. The first microphone 40a mainly converts the sound emitted by the first user U1 into a first signal and outputs it to the stereo converter 50. The second microphone 40b is disposed on the side of the transparent screen 5 facing the second user U2, i.e., the second side. For example, the second microphone 40b may be disposed on a desk in front of the second user U2. The second microphone 40b mainly converts the sound emitted by the second user U2 into a second signal and outputs it to the stereo converter 50. The first signal and the second signal are monaural signals.
[0027] The first signal and the second signal are sound signals detected by the first microphone 40a and the second microphone 40b, and include noise and sound information other than human voices. In the present application, the voice signal is a signal of a person's voice. When the first user U1 and the second user U2 speak to the first microphones 40a and 40b, the first signal or the second signal includes a voice signal.
[0028] The stereo converter 50 converts the first signal input from the first microphone 40a and the second signal input from the second microphone 40b into a stereo signal. The stereo signal includes two channel signals corresponding to left and right sounds, respectively. For example, the stereo converter 50 can output a stereo signal to the display control device 20, with the first signal acquired from the first microphone 40a accommodated in the left channel and the second signal acquired from the second microphone 40b accommodated in the right channel. The stereo signal is information obtained by converting information from two channels into one-dimensional array information based on a specific sampling rate and encoding information. By using the stereo converter 50, the first signal and the second signal can be combined and input to the display control device 20.
[0029] The display control device 20 separates the stereo signal received from the stereo conversion unit 50 into a first signal and a second signal. The display control device 20 separates the one-dimensionally arranged stereo signal into two signals, a left channel and a right channel, in accordance with the specifications of the stereo conversion unit 50. For example, the left channel signal is the first signal, and the right channel signal is the second signal.
[0030] The display control device 20 is configured to be able to transmit the first signal and the second signal to the cloud server 60, respectively. The display control device 20 acquires character information obtained by converting the audio signals included in the first signal and the second signal by speech recognition processing in the cloud server 60. The display control device 20 can further transmit character information in a first language to the cloud server 60 and acquire character information converted into a second language. The display control device 20 can also transmit character information in the second language to the cloud server 60 and acquire character information converted into the first language. Hereinafter, character information converted from the audio signal included in the first signal will be referred to as first character information. Character information converted from the audio signal included in the second signal will be referred to as second character information.
[0031] The display control device 20 can process the first character information and / or the second character information. The display control device 20 can generate or acquire image information based on the first character information and / or the second character information. The display control device 20 transmits an image signal based on the character information to the projector 30, which projects and displays the character information as a character string on the transparent screen 5. In this application, a character string refers to an image in which the character information is projected onto the transparent screen 5, which is a display unit. The character string is displayed at a set position on the display unit as the character information with attributes such as color and size added.
[0032] The display control device 20 may be a general-purpose information device such as a mobile phone (smartphone) or a personal computer, or a dedicated information device. The display control device 20 includes a communication unit 21, a control unit 22, a storage unit 23, and an input unit 24.
[0033] The communication unit 21 communicates with devices external to the display control device 20 via wireless or wired communication means. The external devices include the stereo conversion unit 50, the projector 30, and the cloud server 60. The communication unit 21 may support communication using various communication methods, such as a serial communication standard, a wired local area network (LAN) standard, a wireless LAN standard such as Wi-Fi, and mobile communication standards such as 4G (4th Generation) and 5G (5th Generation). The communication unit 21 can also transmit a first signal and a second signal to the cloud server 60 and receive text information converted from audio signals included in the first signal and the second signal.
[0034] The control unit 22 includes one or more processors. Processors include general-purpose processors that load specific programs to execute specific functions, and dedicated processors specialized for specific processing. Dedicated processors include application-specific integrated circuits (ASICs). Processors include programmable logic devices (PLDs). PLDs include field-programmable gate arrays (FPGAs). The control unit 22 may be either a system-on-a-chip (SoC) or a system in a package (SiP) in which one or more processors work together.
[0035] The control unit 22 controls the entire display control device 20 and executes various information processes. The control unit 22 separates a first signal and a second signal from the stereo signal acquired from the stereo conversion unit 50. The control unit 22 is configured to transmit the first signal and the second signal to the cloud server 60 via the communication unit 21. The control unit 22 is configured to acquire, via the communication unit 21, first character information and second character information converted from the audio signals included in the first signal and the second signal by the cloud server 60. The control unit 22 may further be configured to transmit the acquired first character information and second character information to the cloud server 60 and acquire the first character information and second character information translated into another language.
[0036] The control unit 22 executes a display process to display the first character information and the second character information acquired from the cloud server 60 on the projector 30. The control unit 22 can control the display position and display mode of the first character information and the second character information on the transparent screen 5.
[0037] The control unit 22 can cause the projector 30 to project character strings corresponding to the first character information and the second character information with left-right reversal. In the present application, the state in which the displayed characters appear as "normal characters" to the first user U1 is considered to be the normal projection state of the projector 30. In this case, the characters viewed by the second user U2 are "mirror characters." "Normal characters" are characters displayed in a normal manner. "Mirror characters" are characters in which the left and right sides of normal characters are reversed. Characters projected from the projector 30 with left-right reversal appear as "mirror characters" to the first user U1 and as "normal characters" to the second user U2.
[0038] When the first character information includes a predetermined word to be emphasized, the control unit 22 performs a process of emphasizing the word in the display process. Hereinafter, a word to be emphasized is referred to as an emphasized word. When the first character information includes an emphasized word, the control unit 22 can, for example, change the color of the character string corresponding to the word to a more conspicuous color. Keywords that are important in a conversation are stored in advance in the storage unit 23 as emphasized words.
[0039] When the first character information includes a word to be illustrated, the control unit 22 can display a still image or a moving image related to the illustrated object on the display unit through an illustration process. Hereinafter, the word to be illustrated will be referred to as an illustrated word. An illustrated word is a specific word for which a still image and / or a moving image for explanation can be displayed. Data on the still image and the moving image corresponding to the illustrated word is stored in the memory unit 23 as illustration display data.
[0040] The storage unit 23 may include at least one of a semiconductor storage device, a magnetic storage device, and an optical storage device. The semiconductor storage device may include volatile memory such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory), and non-volatile memory such as ROM (Read Only Memory) and flash memory. The semiconductor storage device includes an SSD (Solid State Drive) that uses flash memory. The magnetic storage device includes a magnetic tape, a floppy disk (registered trademark), a hard disk, etc. The optical storage device includes, for example, a CD (Compact Disc), a DVD (Digital Versatile Disc), and a Blu-ray (registered trademark).
[0041] The storage unit 23 is configured to store programs executed by the control unit 22, information necessary for the processing executed by the control unit 22, and information obtained as a result of the execution by the control unit 22. The storage unit 23 may store emphasis word information used when emphasizing first character information received from the cloud server 60, as well as graphic words and graphic display data used when displaying graphic information based on the first character information.
[0042] The input unit 24 accepts operations by the first user U1 on the display control device 20. The input unit 24 includes, for example, a touch panel, a keyboard, a mouse, and a pen input device. The input unit 24 may be built into the display control device 20. The input unit 24 may also be an external device attached to the display control device 20.
[0043] In addition to the above, the display control device 20 may include a display unit such as a liquid crystal display (LCD), an organic electroluminescence (EL) display, an inorganic EL display, etc. The display unit may display various information related to the operation of the display control device 20.
[0044] The projector 30 is placed next to a first user U1, who is, for example, a counter clerk, and projects an image onto the transparent screen 5 based on an image signal received from the display control device 20. The projector 30 may project still images and videos for illustration at predetermined positions on the transparent screen 5, in addition to images of character strings corresponding to the first character information and the second character information.
[0045] The cloud server 60 is a server that can be accessed via a network such as the Internet. The cloud server 60 of this embodiment is a voice recognition server that provides voice recognition services. When the cloud server 60 receives a voice signal, it performs voice recognition processing on the received voice signal. The cloud server 60 does not perform voice recognition processing if the received signal does not contain human voice or if the level of the voice signal is lower than a predetermined value. The cloud server 60 transmits text information obtained as a result of voice recognition to the sender of the voice signal.
[0046] Cloud server 60 may further provide a translation service for translating between different languages. Cloud server 60 may receive text information in a first language and transmit the text information translated into a second language to the sender. Cloud server 60 may also receive text information in a second language and transmit the information translated into the first language to the sender.
[0047] (Outline of Processing of Display System) FIG. 4 shows a method in which the voices uttered by the first user U1 and the second user U2 are displayed on the transparent screen 5 by the display system 1. As shown in FIG.
[0048] When the first user U1 speaks into the first microphone 40a, a first signal including the audio signal is transmitted to the display control device 20 via the stereo conversion unit 50. The display control device 20 transmits the first signal to the cloud server 60. The display control device 20 may transmit the left channel signal included in the stereo signal transmitted from the stereo conversion unit 50 to the cloud server 60 by streaming. The display control device 20 may transmit the first signal to the cloud server 60 only if the first signal includes an audio signal having a predetermined signal level or higher.
[0049] The cloud server 60 converts the voice signal included in the received first signal into first character information, which is text data, using natural language processing such as morphological element analysis and machine learning. The cloud server 60 transmits the first character information converted from the voice signal to the display control device 20.
[0050] When translating the first character information, the display control device 20 transmits the received first character information in the first language to the cloud server 60. The cloud server 60 translates the first character information in the first language into the second language. The cloud server 60 transmits the first character information translated into the second language to the display control device 20.
[0051] The display control device 20 performs highlighting and illustration processing, etc., as necessary, on the first character information in the first language or the first character information translated into the second language acquired from the cloud server 60. The display control device 20 transmits image information of character strings to be displayed corresponding to the first character information and, if necessary, image information for illustration to the projector 30. The highlighting and illustration processing may be performed to support the first user U1 in presenting information to the second user U2. For example, when the first user U1 explains to the second user U2 how to get to Tokyo Station, "Tokyo," which is registered as an emphasis word in the first character information, may be highlighted. Furthermore, when "Tokyo" is included as an illustration word, image information such as a map or a route map of the Tokyo area may be displayed.
[0052] The projector 30 projects the received text and image information onto the transparent screen 5. Here, the display control device 20 may perform display processing to project the first character information corresponding to the voice uttered by the first user U1 in an inverted manner so that the first character information appears as normal characters to the second user U2. Furthermore, when the first user U1 checks the voice uttered by the first user U1, the first character information may be projected so that the first user U2 sees the normal characters.
[0053] The above describes the case where the first user U1 speaks into the first microphone 40a, but the same process flow is also executed when the second user U2 speaks into the second microphone 40b. However, if the second user U2 is a traveler visiting a tourist information center, the display control device 20 does not need to perform the highlighting process and the illustration process. Furthermore, the second character information corresponding to the voice uttered by the second user U2 may be displayed so that it appears as regular characters to the first user U1.
[0054] (First Display Example of Display System) Next, an example of an image displayed on the transparent screen 5 by the projector 30 based on an image signal from the display control device 20 will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a diagram showing an example of the transparent screen 5 as viewed from the first user U1 side. Fig. 6 is a diagram showing an example of the transparent screen 5 displaying the same content as Fig. 5 as viewed from the second user U2 side.
[0055] 5 and 6 , the transparent screen 5 is divided into a first display area A1, a second display area A2, and a third display area A3. The first display area A1 is an area in which first character information and second character information are displayed while scrolling sequentially upward in accordance with the speech of the first user U1 and the second user U2. In this example, the speech of the first user U1 is translated from Japanese, which is a first language, to English, which is a second language, and then displayed. The speech of the second user U2 is translated from English, which is the second language, to Japanese, which is the first language, and then displayed. In the first display area A1, the first character information corresponding to the speech of the first user U1 (e.g., "Hello. How can I help you?") is displayed in normal characters as seen by the second user U2. In addition, in the first display area A1, the second character information corresponding to the speech of the second user U2 (for example, "Hello. Where can I go if I take the city sightseeing bus?") is displayed in correct characters as seen by the first user U1.
[0056] The second display area A2 is an area where still images and / or moving images are displayed when a diagram is displayed. For example, a map related to the conversation, transportation information, and a video guide to a tourist spot are displayed in the second display area A2. The second display area A2 is also called a diagram display area. The second display area A2 is displayed so that it is oriented correctly when viewed from the second user U2.
[0057] The third display area A3 is an area for displaying the most recent utterance of either the first user U1 or the second user U2. The third display area A3 is used by the speaker, the first user U1 or the second user U2, to confirm that their speech is being correctly recognized. Therefore, in the third display area A3, the first character information is displayed in correct characters as seen by the first user U1. The second character information is displayed in correct characters as seen by the first user U1. For example, in the examples of FIGS. 5 and 6 , the most recent utterance by the first user U1 is "I'm going to Kiyomizu-dera Temple and Kinkaku-ji Temple." This utterance is displayed in the second language in the first display area A1 in correct characters as seen by the second user U2. At the same time, the third display area A3 displays the first language in correct characters as seen by the first user U1.
[0058] 5 and 6 are merely examples of the arrangement of the first display area A1, the second display area A2, and the third display area A3. The positions, sizes, shapes, etc. of the first display area A1, the second display area A2, and the third display area A3 may be set as appropriate.
[0059] (First Example of Display Control Method) A display control method executed by the control unit 22 of the display control device 20 will be described below with reference to FIG. 7 . The display control device 20 may be configured to read and implement a display control program recorded on a non-transitory computer-readable medium to perform the processing performed by the control unit 22 described below. Non-transitory computer-readable media include, but are not limited to, magnetic storage media, optical storage media, magneto-optical storage media, and semiconductor storage media. Magnetic storage media include magnetic disks, hard disks, and magnetic tapes. Optical storage media include optical disks such as CDs (Compact Discs), DVDs, and Blu-ray (registered trademark) Discs. Semiconductor storage media include read-only memories (ROMs), electrically erasable programmable read-only memories (EEPROMs), flash memories, and the like.
[0060] First, when the first user U1 starts the display system 1, the control unit 22 starts acquiring a stereo signal from the stereo conversion unit 50 via the communication unit 21 (step S101). The display system 1 may be started by turning on / off a switch provided in the display control device 20, turning on / off a switch of the stereo conversion unit 50, turning on / off the first microphone 40a and the second microphone 40b, etc.
[0061] The control unit 22 separates the stereo signal acquired by the stereo conversion unit 50 into a first signal acquired from the first microphone 40a and a second signal acquired from the second microphone 40b (step S102).
[0062] The control unit 22 executes a first character information display process (step S103) for the audio signal included in the first signal. The control unit 22 also executes a second character information display process (step S104) for the audio signal included in the second signal. When the first user U1 and the second user U2 alternately converse with each other and their voices do not overlap, only one of steps S103 and S104 is executed at each time. On the other hand, when the first user U1 and the second user U2 speak simultaneously, steps S103 and S104 may be executed in parallel and independently of each other. The processes of steps S103 and S104 will be described below with reference to the flowcharts of FIGS. 8 and 9.
[0063] In the flow diagram of FIG. 8 , first, the control unit 22 transmits a first signal to the cloud server 60 (step S201). The voice signal included in the first signal transmitted to the cloud server 60 is subjected to voice recognition processing by the cloud server 60. The cloud server 60 extracts human voice from the first signal and converts it into text information indicating the content of the speech using natural language processing. The cloud server 60 transmits the converted text information to the display control device 20, for example, in units of sentences. If the first signal does not include human voice, the cloud server 60 does not need to transmit the text information to the display control device 20.
[0064] The control unit 22 sequentially acquires the first character information converted from the first signal from the cloud server 60 (step S202).
[0065] The control unit 22 translates the acquired first character information from the first language to the second language (step S203). The control unit 22 may transmit the first character information to the cloud server 60 and acquire the first character information translated into the second language. Note that, in the display system 1, translation from the first language to the second language is not required. When the first user U1 and the second user U2 speak the same language, the first character information display process executed by the control unit 22 does not need to include step S203.
[0066] Next, the control unit 22 may perform a highlighting process on the first character information (step S204). Specifically, the control unit 22 compares the first character information with the emphasis words stored in the storage unit 23 to confirm whether the first character information includes the emphasis words. The first character information may be information in a first language or information translated into a second language. If the character information includes the emphasis words, the control unit 22 generates display data for a character string in which the emphasis words are highlighted. Highlighting includes, for example, but is not limited to, changing the color of the emphasis words included in the displayed character string to a more conspicuous color. Highlighting may include changing the size and / or brightness of the emphasis words included in the displayed character string.
[0067] The control unit 22 may execute a graphic processing on the first character information (step S205). Specifically, the control unit 22 compares the first character information with the graphic words stored in the storage unit 23 to confirm whether the first character information includes the graphic words. The first character information may be information in the first language or information translated into the second language. If the character information includes the graphic words, the control unit 22 acquires graphic display data, which is data of still images and / or moving images corresponding to the graphic words, from the storage unit 23.
[0068] Either step S204 or step S205 may be executed first. Furthermore, steps S204 and S205 are not essential processes. Either or both of steps S204 and S205 may not be executed.
[0069] Next, the control unit 22 generates image information for display and transmits it to the projector 30, thereby updating the display on the transparent screen 5 (step S206). The control unit 22 adds the latest highlighted first character information to the character string scrolled in the first display area A1. When the first character information is translated into the second language in step S203, the control unit 22 may display the translated first character information in the second language in the first display area A1. When a highlighting process is performed in step S204, the control unit 22 may highlight an emphasized word included in the character string displayed in the first display area A1. Furthermore, for confirmation by the first user U1, the control unit 22 may display the first character information in the first language in the third display area A3 so that it appears as normal characters to the first user U1. Furthermore, when a graphic process is performed in step S205 and graphic display data is available, the control unit 22 may display an image based on the graphic display data in the second display area A2.
[0070] Next, the display process of the second character information in step S104 will be described with reference to the flowchart of FIG.
[0071] The processes of steps S301 to S303 in Fig. 9 are the same as the processes of steps S201 to S203 in Fig. 8 , and therefore details thereof will be omitted. The control unit 22 transmits a second signal including a voice signal of the second user U2 to the cloud server 60 (step S301). The control unit 22 acquires second character information from the cloud server 60, which is obtained by performing voice recognition on the voice signal included in the second signal and converting it into character information (S302). If necessary, the control unit 22 executes a process of translating the acquired second character information (S303).
[0072] In the example of Fig. 9, the control unit 22 does not perform highlighting and illustration processing on the second character information, which corresponds to steps S204 and S205 of the first character information display processing of Fig. 8. In this embodiment, it is assumed that the second user U2 is the one receiving information and explanations from the first user U1. Therefore, in this embodiment, it is assumed that no processing is performed on the content of the utterance of the second user U2.
[0073] The control unit 22 transmits the second character information acquired in step S302 or the second character information in the second language translated in step S303 to the projector 30 as data of a character string to be displayed, and updates the display on the transparent screen 5 (step S304). The control unit 22 adds the character string of the latest second character information to the character string scrolled in the first display area A1 so that it is displayed as correct characters to the first user U1.
[0074] 7, unless the first user U1 instructs the end of the process (step S105: No), the control unit 22 repeats the processes of steps S101 to S104. The first user U1 can instruct the display control device 20 to end the process by, for example, turning off the power switch of the display control device 20. When the first user U1 instructs the end of the process (step S105: Yes), the display control device 20 ends the process.
[0075] When the first user U1 and the second user U2 alternately converse, the voice of the speaking user is subjected to real-time voice recognition processing, and the image displayed on the transparent screen 5 is sequentially updated. In the first display area A1, the contents of the conversations between the first user U1 and the second user U2 are sequentially added to the bottom, and text corresponding to the contents of past conversations is scrolled upward. By doing so, the displays shown in Figures 5 and 6 are possible.
[0076] As described above, the display system 1 shown in FIG. 3 is configured to convert the first signal from the first microphone 40a and the second signal from the second microphone 40b into a stereo signal by providing the stereo conversion unit 50. Furthermore, the control unit 22 of the display control device 20 is configured to separate the stereo signal into the first signal and the second signal. This makes it possible to input the first signal and the second signal to the display control device 20 even if the display control device 20 does not support input of multiple audio signals. Furthermore, the control unit 22 performs voice recognition processing on the first signal and the second signal as separate signals using the cloud server 60. Therefore, the display system 1 of the present disclosure can individually recognize voices from multiple users, convert them into text information, and display it in real time.
[0077] Although it has been described that the process ends when a user's instruction to end the process is received in step S105, the control unit 22 may end the process upon completion of the process in step S103 and / or step S104. In this case, the display control device 20 may start this control flow at regular intervals. In other words, the control unit 22 may periodically and repeatedly execute the process from step S101 to step S104.
[0078] (Second Example of Display Control Method) During a conversation between a first user U1 and a second user U2, both users may speak simultaneously. This is called a speech collision. A speech collision can occur, for example, when the timing of speech coincides or when one user interrupts the other. When a speech collision occurs, the speech of the first user U1 or the second user U2 displayed on the display unit may be cut off midway, or the scrolling may advance suddenly, resulting in unintended display. Therefore, the display control device 20 may prioritize the first signal or the second signal with the higher audio level and display it on the display unit. Below, with reference to Figures 10 to 12, a control method of the display control device 20 when performing such processing will be described.
[0079] In the flow diagram of Fig. 10, steps S401, S402, S404, S406, and S407 are the same as steps S101, S102, S103, S104, and S105 in Fig. 7, respectively, and therefore description thereof will be omitted. The flow diagram of Fig. 10 differs from the flow diagram of Fig. 7 in that a signal level adjustment process (step S403) for a first signal is executed before a first character information display process (step S404). The flow diagram of Fig. 10 also differs from the flow diagram of Fig. 7 in that a signal level adjustment process (step S405) for a second signal is executed before a second character information display process (step S406).
[0080] The signal level adjustment process for the first signal (step S403) will be described with reference to Fig. 11. In this process, the display control device 20 has a first counter, which is set to a predetermined value in advance.
[0081] The control unit 22 refers to the first counter, and if the value of the first counter is greater than 0 (S501: Yes), the process proceeds to step S502.
[0082] In step S502, the control unit 22 compares the signal level of the first signal with the signal level of the second signal. When the voice of the first user U1 is louder than the voice of the second user U2, the signal level of the first signal is higher than the signal level of the second signal.
[0083] If the signal level of the first signal is equal to or greater than the signal level of the second signal (step S502: Yes), the control unit 22 sets the value of the first counter to a predetermined set value (step S503). If the value of the first counter is the predetermined set value, the control unit 22 does not change the value of the first counter.
[0084] If the signal level of the first signal is lower than the signal level of the second signal (step S502: No), the control unit 22 decrements the value of the first counter by 1 (step S504).
[0085] On the other hand, when the value of the first counter becomes 0 in step S501 (S501: No), the control unit 22 proceeds to step S505.
[0086] In step S505, the control unit 22 compares the signal level of the first signal with the signal level of the second signal.
[0087] When the signal level of the first signal is equal to or higher than the signal level of the second signal (step S505: Yes), the control unit 22 sets the value of the first counter to a set value (step S506).
[0088] When the signal level of the first signal is lower than the signal level of the second signal (step S505: No), the control unit 22 reduces the signal level of the first signal (step S507).
[0089] After steps S503, S504, and S506, if the signal level of the first signal is smaller than the predetermined value, the control unit 22 returns the signal level to the original level (step S508).
[0090] The signal level adjustment may be performed by changing the sensitivity of the first microphone 40a and the second microphone 40b. The volume adjustment values of the first microphone 40a and the second microphone 40b for setting the first microphone 40a and the second microphone 40b to the original signal level and the reduced signal level may be stored in the storage unit 23 and may be changeable later by settings, etc. An upper limit may be set for the original signal level.
[0091] After the processes of steps S507 and S508, the control unit 22 proceeds to the first character information display process S404 in FIG.
[0092] 11 , the control unit 22 compares the signal level of the first signal with the signal level of the second signal each time the control unit 22 performs the adjustment process for the first signal. If the signal level of the first signal remains lower than the signal level of the second signal for a period exceeding a predetermined set value, the signal level of the first signal is set lower than the predetermined value (S507). On the other hand, if the signal level of the first signal becomes higher than the signal level of the second signal, the signal level of the first signal is returned to the original set value (step S508).
[0093] When the voice signal level included in the first signal falls below a predetermined value, the cloud server 60 stops the voice recognition process of the first signal in step S404. Therefore, when the state in which the signal level of the first signal is lower than the signal level of the second signal continues beyond the set value of the first counter, the speech content of the first user U1 is no longer displayed on the display unit. This makes it possible to prevent the speech content of the first user U1 from being unintentionally displayed on the display unit, mixing the speech content of the second user U2 with the speech content of the first user U1.
[0094] The flowchart of FIG. 12 illustrates the signal level adjustment process of the second signal (step S405) in the flowchart of FIG. 10 . In the flowchart of FIG. 12 , a second counter different from the first counter is used instead of the first counter. The flowchart of FIG. 12 also reverses the first and second signals in the flowchart of FIG. 11 . Thus, according to the process of the flowchart of FIG. 12 , the control unit 22 compares the signal level of the first signal with the signal level of the second signal each time it performs the adjustment process of the second signal. If the signal level of the second signal remains lower than the signal level of the first signal for a period exceeding the set value of the second counter, the signal level of the second signal is set lower than a predetermined value (S607). On the other hand, if the signal level of the second signal becomes higher than the signal level of the first signal, the signal level of the second signal is restored to its original value (step S608).
[0095] When the signal level of the second signal falls below a predetermined value, the cloud server 60 stops the speech recognition process of the speech signal included in the second signal in step S406. Therefore, if the signal level of the second signal remains lower than the signal level of the first signal for a long period of time, the speech content of the second user U2 will no longer be displayed on the display unit. This prevents the speech content of the first user U1 from being unintentionally displayed mixed with the speech content of the second user U2 on the display unit.
[0096] Furthermore, in this display control method, by appropriately setting a first counter value, speech recognition processing of the first signal is stopped when the speech level of the first user U1 is lower than the speech level of the second user U2 for a predetermined number of consecutive times. Therefore, the predetermined value of the first counter corresponds to the time until speech recognition processing of the first signal is stopped. The same applies to speech recognition processing of the second signal. As a result, when the conversations of the first user U1 and the second user U2 temporarily overlap, even if the speech level of either the first user U1 or the second user U2 is relatively low, both the speech contents of the first user U1 and the second user U2 are displayed on the display unit for a certain period of time.
[0097] In the flow charts of FIGS. 10 to 12 , when a speech collision occurs, the control unit 22 stops the speech recognition process by the cloud server 60 by reducing the signal level of either the first signal or the second signal to a predetermined value or below. However, the method of stopping the display of the speech content of a user with a low speech level is not limited to this. For example, in step S507 of the flow chart of FIG. 11 , the control unit 22 may stop the transmission of the first signal to the cloud server 60 instead of reducing the signal level of the first signal. Furthermore, when the signal level of the first signal remains low, the control unit 22 may not execute the display process of displaying the first text information in step S404. The same or a similar method may also be used to stop the speech of the second user U2 when the second user U2's speech level is low.
[0098] As described above, according to the display control method of the second example, even when a speech conflict occurs between the first user U1 and the second user U2, the display system 1 can display only the speech of the user with the higher voice level. In this way, for example, in a case where the second user U2 is nodding and speaking softly while the first user U1 is explaining, or in a case where the second microphone 40b unintentionally detects the voice of a nearby third party, the speech of the first user U1 can be displayed preferentially.
[0099] In the present embodiment, steps S403 and S404 and steps S405 and S406 are described as being processed in parallel, but steps S405 and S406 may be processed after steps S403 and S404 are processed. Also, in step S407, the processing is described as ending when a user's instruction to end the processing is received, but the control unit 22 may end the processing as soon as the processing of step S404 and / or step S406 is completed. In this case, the display control device 20 may start this control flow at regular intervals.
[0100] (Third Example of Display Control Method) While the first user U1 is speaking, the second microphone 40b may detect the voice of the first user U1 in addition to the first microphone 40a. Also, while the second user U2 is speaking, the first microphone 40a may detect the voice of the first user U1 in addition to the second microphone 40b. In such cases, the voice of the same user may be included in the first signal and the second signal, resulting in duplicate display of identical voice-recognized text information and the incorrect speaker being displayed on the display unit. Therefore, the display control device 20 may compare the first text information recognized from the first signal with the second text information recognized from the second signal, and if there is a match of a predetermined percentage or more, display only one of the first text information and the second text information. For example, the control unit 22 may calculate a correlation value between the first text information and the second text information, and if the correlation value is higher than a predetermined value, prioritize and display the text information converted from the signal with the higher voice level on the display unit. Hereinafter, a control method of the display control device 20 when performing such processing will be described with reference to the flowcharts of FIGS.
[0101] In Fig. 13, steps S701 and S702 are the same as steps S101 and S102 in Fig. 7. Furthermore, steps S703 and S704 are the same as steps S201 and S202 in Fig. 8. Furthermore, steps S706 and S707 are the same as steps S301 and S302 in Fig. 9. Therefore, explanations of steps S701 to S704, S706, and S707 will be omitted.
[0102] In steps S705 and S708, the control unit 22 waits to acquire the first character information or the second character information, whichever comes first. For example, the control unit 22 waits until it acquires the first character information and the second character information for each sentence. The acquired character information may be written in hiragana, katakana, and / or kanji, or in English.
[0103] In step S709, the control unit 22 calculates a correlation value between the acquired first character information and second character information (step S709). The correlation value may be a numerical value evaluated based on the degree of match between characters and / or words, for example.
[0104] If the correlation value is equal to or less than the predetermined threshold (step S710: No), the control unit 22 executes the processes of steps S711 to S714 for the first character information and the processes of steps S715 and S716 for the second character information in parallel. The processes of steps S711 to S714 are the same as the processes of steps S203 to S206 in Fig. 8. The processes of steps S715 and S716 are the same as the processes of steps S303 and S304 in Fig. 9. Therefore, a description of the processes of steps S711 to S716 will be omitted.
[0105] On the other hand, if the correlation value is greater than the predetermined threshold (step S710: Yes), it is considered that the voice of either the first user U1 or the second user U2 is detected by both the first microphone 40a and the second microphone 40b. Therefore, the control unit 22 compares the signal level of the first signal with the signal level of the second signal (step S717). If the signal level of the first signal is higher than the signal level of the second signal at step S717 (S717: Yes), the control unit 22 executes the processes of steps S711 to S714 on the first character information. If the signal level of the first signal is equal to or lower than the signal level of the second signal at step S717 (S717: No), the control unit 22 executes the processes of steps S715 and S716 on the second character information.
[0106] After steps S714 and S716, the control unit 22 returns to step S701 unless the first user U1 instructs the end of the process (step S718: No). When the first user U1 instructs the end of the process (step S718: Yes), the control device 20 ends the process.
[0107] As described above, according to the third example display control method, the control unit 22 calculates a correlation value between the first character information and the second character information obtained by performing voice recognition processing on the voices of the first user U1 and the second user U2. If the correlation value is equal to or less than a predetermined threshold, the first character information and the second character information are derived from voices uttered by different users, and the control unit 22 processes the first character information and the second character information separately and displays them on the display unit. On the other hand, if the correlation value is greater than the predetermined threshold, the first character information and the second character information are deemed to be derived from voices uttered by the same user. In this case, the control unit 22 displays only the character information corresponding to the voice signal with the higher signal level on the display unit. Generally, the voice signal detected by the microphone closer to the speaking user has a higher signal level than the voice signal detected by the microphone farther away. Therefore, even if the second microphone 40b detects the voice of the first user U1, only the first character information based on the first signal detected by the first microphone 40a is displayed on the display unit. Similarly, even if the first microphone 40 a detects the voice of the second user U2, only the second character information based on the second signal detected by the second microphone 40 b is displayed on the display unit. In this way, the control unit 22 can prevent the same character information from being displayed twice and the wrong speaker from being displayed on the display unit.
[0108] Although it has been described that the process ends when a user's instruction to end the process is received in step S718, the control unit 22 may end the process upon completion of the process in step S714 and / or step S716. In this case, the display control device 20 may start this control flow at regular intervals.
[0109] In the third example of the control method described above, the first character information and the second character information resulting from speech recognition of the voice signals detected by the first microphone 40 a and the second microphone 40 b, respectively, are compared to determine whether the voice signals are from the same user. However, the method for determining whether the voice signals detected by the first microphone 40 a and the second microphone 40 b are from the same user is not limited to this.
[0110] For example, the control unit 22 may compare frequency information of a first signal acquired from the first microphone 40a and a second signal acquired from the second microphone 40b, and if the correlation value is equal to or greater than a predetermined value, determine that the audio signals contained in the first signal and the second signal are from the same user. If the average signal levels do not match when calculating the correlation value, the correlation value will be small. Therefore, it is preferable to calculate the correlation value after matching the overall signal levels by calculating the average signal level ratio of the acquired first signal and second signal and normalizing the signal levels.
[0111] Furthermore, when comparing the first and second signals, it is possible that the audio signal may be delayed depending on the distance from the speaker depending on the microphone installation position. A time delay between the first and second signals may affect the accuracy of calculating the correlation value of the audio signals. Therefore, the display system 1 may be equipped with a calibration mode to measure the delay of the first microphone 40a and the second microphone 40b in advance. Measurement is performed by turning on the first microphone 40a and the second microphone 40b and having the first user U1 and the second user U2 speak at a higher signal level than normal so that the same voice is detected by both microphones. By checking the recorded data and comparing the start times of the audio signals generated by the speech, the extent of the delay due to the positions of the first microphone 40a and the second microphone 40b can be confirmed. Furthermore, by measuring the difference in signal level between the first microphone 40a and the second microphone 40b, a signal level threshold for the audio signal not to be displayed on the display unit may be determined.
[0112] (Fourth Example of Display Control Method) When a first user U1 and a second user U2 speak simultaneously during a conversation, it is rare for both users to continue the conversation. Therefore, it is often preferable to prioritize and display the content of the speech of the user who has been speaking longer. Therefore, when the speech of the first user U1 and the second user U2 is compared and speech is input simultaneously, the control unit 22 may display only the speech that has the greater number of characters or words converted from the speech signal on the display unit. Below, a control method of the display control device 20 when performing such processing will be described with reference to the flow charts of FIGS. 15 and 16 .
[0113] 15 and 16, steps S801 to S806 are the same as steps S701 to S704, S706 and S707, respectively, in Fig. 13. Therefore, a description of steps S801 to S806 will be omitted.
[0114] In step S807, the control unit 22 calculates the information amount n of the first character information in the specific utterance section. 1 and the information amount n of the second character information 2 The speech section refers to a period during which the speech signal of the first signal or the second signal continues, including the period during which the speech signals overlap when the speech signals included in the first signal and the second signal overlap in time. The speech section may correspond to a sentence that is longer than the one spoken by the first user U1 and the second user U2. The information amount n of the first character information 1 and the information amount n of the second character information 2 may be the number of characters of the first character information and the number of characters of the second character information, respectively.
[0115] Amount of information n of the first character information 1 is the amount of information in the second character information n 2 If the information amount n of the first character information is larger than the predetermined amount n (step S808: Yes), the control unit 22 executes the processes of steps S809 to S812 on the first character information. 1 is the amount of information in the second character information n 2In the following cases (step S808: No), the control unit 22 executes the processes of steps S813 and S814 for the second character information. The processes of steps S809 to S812 are the same as the processes of steps S203 to S206 in Fig. 8. The processes of steps S813 and S814 are the same as the processes of steps S303 and S304 in Fig. 9. Therefore, the description of the processes of steps S809 to S814 will be omitted.
[0116] After steps S812 and S814, unless the first user U1 instructs the end of the process (step S815: No), the control unit 22 returns to step S801 and repeats the above process. When the first user U1 instructs the end of the process (step S815: Yes), the control device 20 ends the process.
[0117] As described above, according to the display control method of the fourth example, the control unit 22 performs voice recognition processing on the voices of the first user U1 and the second user U2 to obtain the information amount n of the first character information and the second character information. 1 , n 2 The control unit 22 calculates the amount of information n 1 , n 2 Compare the amount of information n 1 , n 2 In this way, even when the first user U1 and the second user U2 speak at the same time, the control unit 22 can display only the speech of the user who continued the conversation on the display unit.
[0118] Regardless of the above, the control unit 22 may be configured to display the first character information or the second character information with the smaller amount of information on the display unit if the amount of information in the first character information or the second character information is equal to or greater than a predetermined value. In this case, the control unit 22 may display the character information with the smaller amount of information before and after the character information with the larger amount of information, rather than inserting the character information with the smaller amount of information in the middle of the character information with the larger amount of information. That is, for example, if the second user U2 speaks during a period in which the first user U1 is speaking, the control unit 22 may display the content of the second user U2's utterance before or after the display of the entire content of the first user U1's utterance on the display unit. In this way, any utterance of a certain length or more can be displayed on the display unit.
[0119] In the fourth example of the display control method, the information amount n of the first character information is 1 and the information amount n of the second character information 2 However, the amount of information n 1 , n 2 is not limited to the number of characters. For example, the amount of information n 1 , n 2 can be used as the number of words.
[0120] Furthermore, the control unit 22 1 , n 2 Alternatively, the control unit 22 may compare the duration of the speech between the first user U1 and the second user U2. That is, when the control unit 22 simultaneously acquires a first signal and a second signal including a voice signal, the control unit 22 detects the duration of the voice signal included in the first signal and the second signal from the first signal and the second signal. The control unit 22 may display text information based on the signal of the first signal or the second signal that has the longer voice signal duration on the display unit.
[0121] Although it has been described that the process ends when a user's instruction to end the process is received in step S815, the control unit 22 may end the process upon completion of the process in step S812 and / or step S814. In this case, the display control device 20 may start this control flow at regular intervals.
[0122] Two or more of the methods described in the second to fourth examples of the display control method may be used in combination.
[0123] The following describes variations of the display that can be displayed on the transparent screen 5 of the display system 1.
[0124] (Second Display Example of Display System) A second display example of the display system 1 is shown in Fig. 17. Fig. 17 shows the display of the second display example on the transparent screen 5 as seen from the side of the second user U2. The second display example of the display system 1 is the first display example shown in Figs. 5 and 6, except that a horizontal line indicating a boundary is displayed between the character string based on the first character information and the character string based on the second character information displayed in the first display area A1. In other words, according to this display method, the boundary is indicated by a line each time a different user speaks the character information displayed in the first display area A1.
[0125] 18 and 19 are flowcharts executed by the control unit 22 to display the second display example. The first character information display process in Fig. 18 and the second character information display process in Fig. 19 are executed as the processes of step S103 and step S104 in the flowchart of Fig. 7, instead of the processes of the flowcharts of Fig. 8 and 9.
[0126] Steps S901 to S905 of the first character information display process in FIG. 18 are the same as steps S201 to S205 in the flowchart of FIG. 8, and therefore a description thereof will be omitted.
[0127] In step S906, the control unit 22 determines whether a character string corresponding to the second character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the first character information is to be displayed next (step S906).
[0128] If a character string corresponding to the second character information has been displayed immediately before (step S906: Yes), the control unit 22 displays a line that serves as a boundary line immediately below the character string corresponding to the immediately preceding second character information (step S907).The control unit 22 displays a character string corresponding to the first character information below the line that serves as the boundary line, and displays an image for illustration in the second display area A2 if necessary (step S908).
[0129] In step S906, if a character string corresponding to the second character information has not been displayed immediately before (step S906: No), the control unit 22 does not place a line that serves as a boundary line. The control unit 22 displays the character string corresponding to the first character information in the first display area A1 as is, and displays an image in the second display area A2 as necessary (step S908).
[0130] Steps S1001 to S1003 of the second character information display process in FIG. 19 are the same as steps S301 to S303 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0131] In step S1004, the control unit 22 determines whether a character string corresponding to the first character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the second character information is to be displayed next.
[0132] If a character string corresponding to the first character information has been displayed immediately before (step S1004: Yes), the control unit 22 displays a line that serves as a boundary line immediately below the character string corresponding to the immediately preceding first character information (step S1005).The control unit 22 displays a character string corresponding to the second character information below the line that serves as the boundary line (step S1006).
[0133] In step S1004, if a character string corresponding to the first character information has not been displayed immediately before (step S1004: No), the control unit 22 displays a character string corresponding to the second character information in the first display area A1 without placing a straight line as a boundary line (step S1006).
[0134] In this way, the breaks in the conversation can be displayed in an easy-to-understand manner for the first user U1 and the second user U2 in the first display area A1 of the display unit.
[0135] (Third Display Example of Display System) A third display example of the display system 1 is shown in FIG. 20. FIG. 20 shows the third display example of the transparent screen 5 as seen from the second user U2 side. In the third display example of the display system 1, the most recent character string among the character strings corresponding to the first character information or the second character information displayed in the first display area A1 in the first display example shown in FIGS. 5 and 6 is highlighted in a specific manner. The specific manner includes, for example, using a priority color for the display color of the character string. In FIG. 20, the character string marked "highlighted" is the most recent character string displayed in a highlighted manner relative to the character string marked "normal display."
[0136] As an example, when the first display area A1 is blank, if the first user U1 or the second user U2 starts a conversation, the text information is displayed in the priority color. As long as the same user's voice continues, the first display area A1 is displayed in the priority color. When the speaking user switches, the previous conversation is changed to the normal color, and only the most recent conversation is displayed in the priority color. The normal color can be, for example, white. The priority color can be, for example, yellow.
[0137] Specific ways to highlight a character string include changing the color, making the character bold, increasing the character size, changing the font style, changing the transparency of the character, etc.
[0138] 21 and 22 are flowcharts executed by the control unit 22 to display the third display example. The first character display process in Fig. 21 and the second character display process in Fig. 22 are executed as steps S103 and S104 in the flowchart of Fig. 7, instead of the processes in the flowcharts of Fig. 8 and 9.
[0139] Steps S1101 to S1105 in FIG. 21 are the same as steps S201 to S205 in the flowchart of FIG. 8, and therefore a description thereof will be omitted.
[0140] In step S1106, the control unit 22 determines whether a character string corresponding to the second character information is displayed immediately before the area of the first display area A1 where the character string corresponding to the first character information is to be displayed next.
[0141] If a character string corresponding to the second character information was displayed immediately before (step S1106: Yes), the control unit 22 displays the character string of the previously displayed character information in the normal display mode (step S1107).The control unit 22 displays a character string corresponding to the latest first character information in a specific mode below the previous character string displayed in the normal display mode, and displays an image in the second display area A2 if necessary (step S1108).
[0142] In step S1106, if a character string corresponding to the second character information has not been displayed immediately before (step S1106: No), a character string corresponding to the first character information is displayed in a specific manner in the first display area A1, and an image is displayed in the second display area A2 if necessary (step S1108). In this case, a character string corresponding to the first character information displayed in a specific manner corresponding to the most recent utterance is added below the character string corresponding to the first character information displayed in a specific manner corresponding to the past utterance.
[0143] Steps S1201 to S1203 in FIG. 22 are the same as steps S301 to S303 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0144] In step S1204, the control unit 22 determines whether a character string corresponding to the first character information is displayed immediately before the area of the first display area A1 in which the character string corresponding to the second character information is to be displayed next.
[0145] If a character string corresponding to the first character information was displayed immediately before (step S1204: Yes), the control unit 22 displays the character string of the previously displayed character information in a normal display mode (step S1205).The control unit 22 displays a character string corresponding to the latest second character information in a specific mode below the previous character string displayed in the normal display mode (step S1206).
[0146] In step S1204, if a character string corresponding to the first character information has not been displayed immediately before (step S1204: No), the previously displayed character information is not displayed normally, and a character string corresponding to the second character information is displayed in a specific manner in the first display area A1 (step S1206). In this case, a character string corresponding to the second character information displayed in a specific manner corresponding to the most recent utterance is added below the character string corresponding to the second character information displayed in a specific manner corresponding to the past utterance.
[0147] In this way, the first display area A1 always displays a character string corresponding to the latest first character information or second character information in a specific manner. This highlights the latest conversation between the first user U1 and the second user U2, allowing the first user U1 and the second user U2 to find the latest conversation at a glance. As a result, the first user U1 and the second user U2 can concentrate on the latest conversation.
[0148] (Fourth display example of the display system) A fourth display example of the transparent screen 5 of the display system 1 is shown in Figures 23 and 24. Figures 23 and 24 show the sequential display of the conversation content in the fourth display example as seen from the second user U2 side. In the fourth display example of the display system 1, the third display area A3 in the first display example shown in Figures 5 and 6 can be used by the second user U2 to check whether the content of his or her speech has been correctly recognized.
[0149] More specifically, as shown in FIG. 23 , the control unit 22 of the display control device 20 displays the first character information converted from the most recent utterance by the first user U1 in the first display area A1 in the second language as correct characters as seen by the second user U2. The control unit 22 also displays the first character information in the third display area A3 in the first language as correct characters as seen by the first user U1. Furthermore, as shown in FIG. 24 , the control unit 22 displays the second character information converted from the most recent utterance by the second user U2 in the first language as correct characters as seen by the first user U1, and in the second language as correct characters as seen by the second user U2. Also, in FIGS. 23 and 24 , the language from which the display in the first display area A1 has been converted is displayed. For example, the display of "English" and "Japanese" arranged vertically in the third display area A3 indicates that the display in the first display area A1 is a translation from English to Japanese.
[0150] The display methods of Figures 23 and 24 are characterized in that they allow not only the first user U1, who is the counter clerk, but also the second user U2, who is the visitor, to confirm whether the content of the utterance has been correctly recognized as voice.
[0151] 25 and 26 are flowcharts of the process executed by the control unit 22 to display the fourth display example. The flowcharts of the first character information display process in Fig. 25 and the second character information display process in Fig. 26 are executed as the processes of steps S103 and S104 in the flowchart of Fig. 7, instead of the processes of the flowcharts of Fig. 8 and 9.
[0152] The processing of steps S1301 to S1305 in Fig. 25 is the same as the processing of steps S201 to S205 in the flowchart of Fig. 8, and therefore description thereof will be omitted. Step S1306 in Fig. 25 is the same as step S206 in the flowchart of Fig. 8, except that processing for displaying first character information in the first language in the third display area A3 is not performed.
[0153] In step S1307, the control unit 22 clears the display in the third display area A3 and displays the first character information in the first language in the correct characters toward the first user U1. Furthermore, the control unit 22 displays in the third display area A3 a display indicating that the display translated from the first language to the second language is displayed in the first display area A1.
[0154] The processing from steps S1401 to S1404 in FIG. 26 is the same as the processing from steps S301 to S304 in the flowchart of FIG. 9, and therefore a description thereof will be omitted.
[0155] In step S1405, the control unit 22 clears the display in the third display area A3 and displays the second character information in the second language in the correct characters toward the second user U2. That is, the control unit 22 displays the second character information in the third display area A3 in reverse video. Furthermore, the control unit 22 displays in the third display area A3 a display indicating that the display translated from the second language to the first language is displayed in the first display area A1.
[0156] In this manner, the display system 1 displays text information of the speech recognition result to the current speaker of the first user U1 or the second user U2. This allows both the first user U1 and the second user U2 to confirm that speech recognition is being performed correctly for the content of their current speech. Furthermore, in this display method, the third display area A3, which displays the speech recognition result of the current speech content, is shared by the first user U1 and the second user U2, allowing for efficient use of the display area of the display unit.
[0157] (Deletion of Graphical Display) The display system 1 may have a function of deleting an image displayed in the second display area A2. For example, the control unit 22 of the display control device 20 may execute the process shown in the flowchart of Fig. 27 between steps S102 and S103 of the flowchart of Fig. 7. The flowchart of Fig. 27 will be described below.
[0158] First, the control unit 22 monitors whether or not a command to clear an image (image clear command) is included in the first signal acquired from the first microphone 40a via the stereo conversion unit 50 (S1501). For example, when the first user U1 wants to erase the image displayed in the second display area A2 of the transparent screen 5, the first user U1 verbally issues a voice command such as "image clear." The control unit 22 has a function of recognizing a specific voice command included in the first signal.
[0159] When the control unit 22 detects the image clear command from the first signal (step S1502: Yes), it erases the image displayed in the second display area A2 (step S1503). In this case, the control unit 22 may transmit the first signal, from which the audio of the image clear command has been removed, to the cloud server 60 in step S103 of Fig. 7. This prevents the image clear command uttered by the first user U1 from being displayed as text information in the first display area A1.
[0160] If the control unit 22 does not detect an image clear command from the first signal (step S1502: No), it does not perform any particular process.
[0161] Alternatively, the control unit 22 may transmit a first signal including an image clear command to the cloud server 60, and erase the image in the second display area A2 when the first character information generated by the voice recognition includes the image clear command. Such processing may be executed, for example, between steps S202 and S203 in the flowchart of FIG. 8 .
[0162] As described above, the display system 1 allows the first user U1 to erase the image in the second display area A2 displaying the illustration by voice operation. Because the operation is performed by voice rather than by operating the input unit 24 of the display control device 20, the first user U1 can concentrate on the conversation. For example, if the input unit 24 is an external input device to the display control device 20 and has multiple buttons corresponding to the operations to be performed by the first user, the first user U1 may have difficulty operating the input unit 24. By employing voice operation, it is possible to prevent the conversation from being interrupted due to time consuming button operations. Furthermore, because the past conversation content displayed in the first display area A1 is not erased, the first user U1 and the second user U2 can continue their conversation without any inconvenience while continuing to view the transparent screen 5.
[0163] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art would easily be able to make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included within the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies, and multiple components can be combined or divided into one.
[0164] In this disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the configuration. In this disclosure, the configurations distinguished by descriptions such as "first" and "second" can have their numbers swapped. For example, the first microphone 40a can swap the identifiers "first" and "second" with the second microphone 40b. The identifier swapping is performed simultaneously. The configurations remain distinguished even after the identifier swapping. Identifiers may be deleted. A configuration from which an identifier has been deleted is distinguished by a symbol. The identifiers "first" and "second" used in this disclosure should not be used solely to interpret the order of the configurations or to justify the existence of an identifier with a smaller number.
[0165] The components and functional blocks included in each embodiment of the present disclosure may be arranged in the same hardware or different hardware as appropriate. The stereo conversion unit 50 in each embodiment of the present disclosure may be built into the display control device 20. The display control device 20 and the projector 30 may be integrally configured using the same hardware. Some or all of the communication unit 21, control unit 22, storage unit 23, and input unit 24 of the display control device 20 in each embodiment of the present disclosure may be included in the projector 30.
[0166] The display unit of the present disclosure is not limited to a configuration including the transparent screen 5 and the projector 30. The display unit may include, for example, a transparent organic light-emitting diode (OLED) display in which light-emitting elements are arranged within a transparent substrate.
[0167] In the above embodiment, the speech recognition process and the translation process are performed by the cloud server 60 located in a remote location, but the present invention is not limited to this. The speech recognition process and the translation process may be performed by separate servers rather than the same server located in a remote location. At least one of the speech recognition process and the translation process may be performed by a server located near the display control device 20 or by the display control device 20 itself.
[0168] REFERENCE SIGNS LIST 1 Display system 5 Transparent screen (display unit) 10 Base material 20 Display control device 21 Communication unit 22 Control unit 23 Memory unit 24 Input unit 30 Projector (display unit) 40a First microphone 40b Second microphone 50 Stereo conversion unit 60 Cloud server U1 First user U2 Second user A1 First display area A2 Second display area A3 Third display area
Claims
1. A display system comprising: a display unit that is placed between multiple users when in use; a first microphone that is placed on a first side of the display unit; a second microphone that is placed on a second side opposite the first side of the display unit; a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals; and a display control device that receives the stereo signals, separates the stereo signals into the first signal and the second signal, and displays first character information and second character information that are converted by voice recognition from audio signals included in the first signal and the second signal on the display unit.
2. The display system according to claim 1, wherein the display control device displays the first character information and the second character information while scrolling them sequentially on the display unit, and displays a boundary line between the first character information and the second character information.
3. A display system according to claim 1 or 2, wherein the display control device converts the first character information in a first language into a second language different from the first language and displays the converted information on the display unit.
4. The display system described in claim 3, wherein the display control device displays the first character information converted from the latest utterances by the multiple users as a character string in the second language toward the second side and as a character string in the first language toward the first side, and displays the second character information converted from the latest utterances by the multiple users as a character string in the first language toward the first side and as a character string in the second language toward the second side.
5. A display system as claimed in any one of claims 1 to 4, wherein the display control device displays the latest first character information or the latest second character information converted from speech uttered by one or more of the plurality of users on the display unit in a predetermined manner different from the display manner of other character information.
6. A display system described in any one of claims 1 to 5, wherein when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the audio levels of the audio signals contained in the first signal and the second signal, and displays only one of the first character information and the second character information on the display unit based on the comparison result.
7. The display system according to claim 6, wherein the display control device causes the display unit to display the first character information or the second character information obtained by converting the audio signal contained in the first signal or the audio signal contained in the second signal, whichever has the higher audio level.
8. A display system as claimed in any one of claims 1 to 5, wherein when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the first character information with the second character information, and if they match by a predetermined percentage or more, displays only one of the first character information and the second character information.
9. The display system according to claim 8, wherein the display control device causes the display unit to display the first character information or the second character information obtained by voice recognition of the first signal or the second signal, whichever has a higher voice level.
10. A display system as claimed in any one of claims 1 to 5, wherein when the audio signal contained in the first signal and the audio signal contained in the second signal overlap in time, the display control device compares the first character information and the second character information corresponding to the speech within the speech section including the period in which the audio signals overlap, and causes the display unit to display only the first character information or the second character information which has the greater amount of information.
11. A display control device comprising: a display unit arranged between multiple users when in use; a communication unit capable of communicating with a stereo conversion unit that converts a first signal input from a first microphone arranged on a first side of the display unit and a second signal input from a second microphone arranged on a second side opposite the first side of the display unit into a stereo signal; and a control unit that receives the input of the stereo signal, separates the stereo signal into the first signal and the second signal, and executes processing to cause the display unit to display first character information and second character information obtained by converting audio signals included in the first signal and the second signal, respectively, through voice recognition.
12. A program for a display control device that controls a display system including a display unit that is placed between multiple users when in use, a first microphone that is placed on a first side of the display unit, a second microphone that is placed on a second side opposite the first side of the display unit, and a stereo conversion unit that converts a first signal input from the first microphone and a second signal input from the second microphone into stereo signals, the program causing a control unit of the display control device to execute a process of receiving input of the stereo signals, separating the stereo signals into the first signal and the second signal, and displaying first character information and second character information that are converted by voice recognition from the audio signals included in the first signal and the second signal, respectively, on the display unit.
Citation Information
Patent Citations
Conference supporting system, method and computer program
JP2006229903A
Translation display device
JP2019070915A
Conversation assistance apparatus, display system, information processing method, and display method
JP2023124615A
Apparatus and method for separating source by using advanced orthogonality on mutlichannel spectrogram
KR1020110009391A
Audio processing device, audio processing method, and audio processing program
WO2013145578A1