Electronic device, method for controlling electronic device, and program
By combining the shading part and voice recognition technology on the transparent screen, the speech content of different users can be distinguished and displayed, solving the problem of poor communication between hearing-impaired people and others and achieving smooth communication.
Patent Information
- Application Number
- CN202480008635.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-24
- Filing Date
- 2024-02-07
- Publication Date
- 2025-09-16
AI Technical Summary
Existing display systems have difficulty effectively distinguishing and displaying the speech content of different users when assisting hearing-impaired people to communicate with others, resulting in poor communication.
Using a combination of a transparent screen and a shading part, through voice recognition technology and a control system, it can distinguish and display the speech content of different users, and translate and emphasize it as needed to ensure that each user only sees the information they need.
It enables smooth communication between the hearing-impaired and others. Through the design of transparent screen and light-shielding part, it ensures that each user only sees the information they need, reduces interference, and improves the clarity and efficiency of communication.
Smart Images

Figure CN120660068A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority from Japanese Patent Application No. 2023-27828 filed in Japan on February 24, 2023, and the entire contents of that prior application are incorporated herein for reference. Technical Field
[0003] The present disclosure relates to an electronic device, a control method of an electronic device, and a program. Background Art
[0004] A display system has been proposed that automatically converts a speaker's voice into text and displays subtitles in real time on a transparent display placed between the speaker and a hearing-impaired listener (see, for example, non-patent document 1). This display system displays the original text for the listener on the upper portion of the transparent display and the original text for the speaker on the lower portion of the transparent display. The text on the lower portion of the transparent display is displayed for the speaker to confirm any erroneous conversions made by automatic speech recognition. Furthermore, a translation display device has been proposed that can smoothly assist conversations between speakers speaking different languages (see, for example, patent document 1). In this device, a single display unit displays a text string based on language A and a text string based on language B in display areas A1 and B1, which are divided into opposite directions. Therefore, the text is displayed in an orientation that is easy for each user to read.
[0005] Prior art literature
[0006] Non-patent literature
[0007] Non-Patent Literature 1: Kenta Yamamoto, Ippei Suzuki, Akihisa Shitara, and Yoichi Ochiai, 2021. "Transparent Subtitles: Real-time Subtitles on Transparent Displays for Deaf and Hearing-Impaired Viewing." [Online], January 27, 2021, in arXiv e-prints. 5 pages, Internet<URL:https: / / arxiv.org / abs / 2101.11326>
[0008] Patent Literature
[0009] Patent Document 1: Japanese Patent Application Publication No. 2019-70915 Summary of the Invention
[0010] In one embodiment, (1) a program is executed by an electronic device to control a display system that displays first text based on a first user's speech and second text based on a second user's speech on a display unit.
[0011] an obtaining step of obtaining at least one of the first character and the second character;
[0012] a determining step of determining, based on a given condition, whether at least one of the first character and the second character is based on the speech of the first user or the second user; and
[0013] The display step includes changing a display mode of at least one of the first character and the second character and displaying the character on the display unit so that the first character and the second character are displayed in different modes.
[0014] (2) In the procedure of (1) above,
[0015] In the discrimination step,
[0016] As the given condition, based on conditions related to detection performed by at least any one of a switch, an acceleration sensor, a gyro sensor, a magnetic sensor, a camera, and a plurality of microphones, it is determined whether at least one of the first text and the second text is based on the speech of the first user or the second user.
[0017] (3) In the procedure of (1) above,
[0018] In the display step,
[0019] One of the first character and the second character is displayed on the display unit more emphasized than the other.
[0020] (4) In the procedure of (1) above,
[0021] In the display step,
[0022] One of the first character and the second character is displayed on the display unit in a different size, thickness, color, font, background color, decorative text, and display position from the other.
[0023] (5) In the procedure of (1) above,
[0024] When the language of the first character is different from the language of the second character,
[0025] causing the electronic device to perform a translation step of translating the first text into a language of the second text,
[0026] In the displaying step, a result of translating the first character into the language of the second character is displayed as the first character.
[0027] (6) In the procedure of (1) above,
[0028] When the language of the first character and the language of the second character are different languages,
[0029] In the obtaining step, a result of translating the first character into the language of the second character is obtained as the first character.
[0030] In the displaying step, a result of translating the first character into the language of the second character is displayed as the first character.
[0031] (7) In the procedure described in any of (1) to (6) above,
[0032] In the display step,
[0033] One of the first character and the second character is flipped left-right and displayed in a second display area of the display unit.
[0034] (8) In the procedure described in any of (1) to (6) above,
[0035] The electronic device controlling the display system is caused to execute the following steps, wherein the display unit includes: a first display area having a light-shielding portion that at least partially blocks light; and a second display area that at least partially transmits light.
[0036] That is, in the display step,
[0037] flipping the first character or the second character left to right,
[0038] At least one of the first character and the second character is displayed in the first display area while being shielded by the shielding portion.
[0039] At least one of the first character and the second character is displayed in the second display area.
[0040] (9) In the procedure of (1) above,
[0041] In the display step,
[0042] When a given keyword is included in the first text, a given picture or video corresponding to the given keyword is displayed on the display unit.
[0043] When the second text includes the given keyword, the given picture or video is not displayed on the display unit.
[0044] (10) In the procedure of (1) above,
[0045] In the obtaining step, the first character and the second character are obtained separately; or,
[0046] The electronic device is caused to execute a storage step of storing the first character and the second character separately.
[0047] In one embodiment, (11) an electronic device includes a control unit that controls to display first characters based on utterances of a first user and second characters based on utterances of a second user on a display unit.
[0048] The control unit performs the following processing:
[0049] obtaining at least one of the first character and the second character,
[0050] Based on a given condition, determining whether at least one of the first character and the second character is based on the speech of the first user or the second user,
[0051] The display format of at least one of the first character and the second character is changed and displayed on the display unit so that the first character and the second character have different display formats.
[0052] In one embodiment, (12) a control method is provided for controlling an electronic device that displays first text based on a first user's speech and second text based on a second user's speech on a display unit. The control method includes the following steps.
[0053] The acquiring step acquires at least one of the first character and the second character.
[0054] The step of determining whether at least one of the first character and the second character is based on the speech of the first user or the second user based on a given condition.
[0055] The display step includes changing a display mode of at least one of the first character and the second character and displaying the character on the display unit so that the first character and the second character are displayed in different modes. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a perspective view illustrating a usage scenario of the projector display system according to one embodiment.
[0057] Figure 2 yes Figure 1 A schematic diagram of the projector display system.
[0058] Figure 3 1 is a diagram showing an example of a display on a transparent screen viewed from a first user side.
[0059] Figure 4 2 is a diagram showing an example of a display on a transparent screen viewed from a second user.
[0060] Figure 5 This is a diagram showing the relationship between the light shielding unit and the arrangement of the second user in one embodiment.
[0061] Figure 6 It is a diagram showing an example of the arrangement of a screen portion and a light shielding portion.
[0062] Figure 7 It is a diagram showing an example of the arrangement of a screen portion and a light shielding portion.
[0063] Figure 8 It is a diagram showing an example of the arrangement of a screen portion and a light shielding portion.
[0064] Figure 9 It is a diagram showing an example of the arrangement of a screen portion and a light shielding portion.
[0065] Figure 10 This is a schematic diagram of the configuration of a projector display system according to an embodiment of the present invention that utilizes an external cloud server for speech recognition.
[0066] Figure 11 This is a plan view showing an example of the arrangement of components of a projector display system.
[0067] Figure 12 This is a schematic configuration diagram of a projector display system including a detection unit according to one embodiment.
[0068] Figure 13 Yes Figure 12 A diagram showing an example of a sensor included in a detection unit.
[0069] Figure 14 This is a diagram illustrating a problem solved by a projector display system according to one embodiment.
[0070] Figure 15 This is a flowchart illustrating the operation of the projector display system according to one embodiment.
[0071] Figure 16A This is a diagram showing an operation example of a projector display system according to one embodiment.
[0072] Figure 16B This is a diagram showing an operation example of a projector display system according to one embodiment.
[0073] Figure 17AThis is a diagram showing an operation example of a projector display system according to one embodiment.
[0074] Figure 17B This is a diagram showing an operation example of a projector display system according to one embodiment.
[0075] Figure 18A This is a diagram showing an operation example of a projector display system according to one embodiment.
[0076] Figure 18B This is a diagram showing an operation example of a projector display system according to one embodiment.
[0077] Figure 19 This is a flowchart illustrating the operation of the projector display system according to one embodiment. DETAILED DESCRIPTION
[0078] When using a device such as the one described above to display text between two interlocutors on the same screen, for example, it is preferable to display the text appropriately to facilitate smooth communication between the interlocutors. The present disclosure provides an electronic device, a method for controlling an electronic device, and a program that facilitate smooth communication between interlocutors. According to one embodiment, the electronic device, the method for controlling an electronic device, and the program facilitate smooth communication between interlocutors.
[0079] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. The drawings used in the following description are schematic drawings. Dimensions and ratios in the drawings may not necessarily correspond to actual dimensions and ratios.
[0080] (Structure of projector display system)
[0081] like Figure 1 As shown, the projector display system 1 involved in one embodiment is a system that assists the conversation between a first user U1 and a second user U2 who has a hearing impairment, for example. In addition, the second user U2 is not limited to a user with a hearing impairment, and can also be a user who does not have a specific impairment, etc. The second user U2 can also be, for example, a person who wants to communicate in a language different from the language of the first user U1. Furthermore, the second user U2 can also be, for example, a person who wants to communicate in the same language as the first user U1. As an example, the first user U1 is a staff member in charge of a counter of a government, local government or public institution. The second user U2 is, for example, a person who visits a government, local government or public institution. The scenarios for using the projector display system 1 are not limited to the above scenarios. The projector display system 1 can also be used at counters of financial institutions, medical institutions, public transportation facilities, business stores of individual operators, and conference rooms in offices.
[0082] In recent years, due to the prevalence of infectious diseases, plastic curtains or acrylic panels, etc., are sometimes placed between counter attendants and customers in situations where they come into contact with customers to prevent droplet infection. In a projector display system 1, these plastic curtains or acrylic panels serve as a substrate 10, on which a transparent screen 5 is placed. The speech of a first user U1 is projected as text from a projector 30 onto the transparent screen 5. This allows a second user U2, located opposite the first user U1 across the transparent screen 5, to visually confirm the first user U1's speech, even if the first user U1 cannot or has difficulty hearing the speech. Furthermore, the speech of the second user U2 can also be projected as text onto the transparent screen 5 from the projector 30. This allows the first user U1, located opposite the second user U2 across the transparent screen 5, to visually confirm the second user U2's speech, even if the second user U2 cannot or has difficulty hearing the speech.
[0083] The transparent screen 5 can be placed at any position on the substrate 10. The transparent screen 5 can be placed at a position that does not prevent the first user U1 and the second user U2 from visually viewing each other's faces. For example, the transparent screen 5 can be placed at a position where the first user U1 and the second user U2 can lower their sight lines from the horizontal direction to the bottom.
[0084] The projector display system 1 is, for example, Figure 2 As shown, in addition to the transparent screen 5 and the projector 30, the display system further includes an electronic device 20 and a microphone 40. Thus, the display system according to one embodiment may include a display unit (eg, the transparent screen 5).
[0085] The transparent screen 5 includes a screen portion 11 and a light shielding portion 12 disposed on a base material 10 .
[0086] The screen portion 11 is a film-like, sheet-like, or plate-like member that can be attached to the substrate 10 and used for projecting by the projector. Generally speaking, the screen portion 11 itself is sometimes referred to as a "transparent screen." The screen portion 11 diffuses a portion of the light incident from the projector 30 toward the incident side and the outgoing side. The first user U1 and the second user U2 enter their fields of view through the light diffused by the screen portion 11, allowing them to recognize the image projected from the projector 30. The shape of the screen portion 11 can be, for example, rectangular, but is not limited thereto. The screen portion 11 can have various shapes.
[0087] The light shielding portion 12 at least partially blocks (shields) the transmission of light emitted from the projector 30. The light shielding portion 12 is, for example, a dichroic film that suppresses light of a specific wavelength. For example, a dichroic film that does not transmit red light can be used as the light shielding portion 12. The light shielding portion 12 can be arranged to overlap a portion of the screen portion 11. For example, the light shielding portion 12 can be arranged to overlap the surface of the screen portion 11. The shape of the light shielding portion 12 can be, for example, rectangular, but is not limited to this. The light shielding portion 12 can have various shapes. The portion of the screen portion 11 where the light shielding portion 12 is not arranged can be a portion that at least partially transmits light.
[0088] The microphone 40 is arranged on the side of the first user U1. For example, the microphone 40 can be arranged on the table in front of the first user U1. In addition, the microphone 40 can be installed on the headset equipped with the first user U1. The microphone 40 converts the voice uttered by the first user U1 into an electrical signal and outputs it to the electronic device 20. In addition, the microphone 40 is not only arranged on the side of the first user U1, but also on the side of the second user U2. Furthermore, the microphone 40 can also be a microphone array (array microphone) in which multiple microphones are arranged at a given position. According to the microphone array, based on the voice uttered by the first user U1 and / or the second user U2, the orientation (direction) of the first user U1 and / or the second user U2 relative to the microphone array can be determined.
[0089] The electronic device 20 receives the output from the microphone 40, recognizes the speech content of at least one of the first user U1 and the second user U2, and generates a display based on the recognized speech content. The electronic device 20 converts the display into an image signal and displays it on the projector 30. The display is, for example, text (character string). In addition, the display may also represent the content (character string) in a first language recognized as the speech content of the first user U1 translated into content (character string) in a second language different from the first language. Here, the second language may also be, for example, the language used by the second user. In the following, in the embodiments of the present disclosure, "text" sometimes includes "character string". In addition, in the embodiments of the present disclosure, "character string" sometimes includes "text". The electronic device 20 can use general-purpose devices such as mobile phones, smartphones, tablet terminals, and personal computers, or dedicated devices. The electronic device 20 is configured to include a voice recognition unit 21, a control unit 22, and a storage unit 23.
[0090] The speech recognition unit 21 performs speech recognition processing based on the speech signal output from the microphone 40. The speech recognition result is output as text data. The speech recognition processing can adopt various well-known technologies. The speech recognition unit 21 has a speech recognition engine that performs speech recognition processing using the speech recognition dictionary stored in the storage unit 23. The speech recognition dictionary includes an acoustic model, a language model, and a phonetic dictionary. The speech recognition unit 21 can include a dedicated or general-purpose processor. The processing of the speech recognition unit 21 can be performed by the same processor as the processor that performs the functions of the control unit 22 described below.
[0091] The control unit 22 controls the entire electronic device 20 and causes the text of the speech content of the first user U1 and / or the second user U2, recognized by the speech recognition unit 21, to be displayed as a display object on the projector 30. The control unit 22 can control the display position and display method of the display object on the transparent screen 5. As described later, the control unit 22 can project the text of the speech content of the first user U1 and / or the second user U2 onto the projector 30 as a first display object DO1 and a second display object DO2 that is a left-right flip of the first display object DO1.
[0092] The control unit 22 includes one or more processors. Processors include general-purpose processors that can read specific programs and execute specific functions, as well as dedicated processors dedicated to specific processing. Dedicated processors include ICs (ASICs; Application Specific Integrated Circuits) for specific purposes. Processors include programmable logic devices (PLDs; Programmable Logic Devices: Field Programmable Gate Arrays). PLDs include FPGAs (Field-Programmable Gate Arrays). The control unit 22 can be either an SoC (System-on-a-Chip) or a SiP (System in a Package) in which one or more processors work together.
[0093] The storage unit 23 is configured to store programs executed by the control unit 22, information required for processing performed by the control unit 22, and information obtained as a result of the processing performed by the control unit 22. The storage unit 23 can store the aforementioned speech recognition dictionary. For example, the storage unit 23 can store a dictionary required for converting recognized text information into Chinese characters. The storage unit 23 can include at least one of a semiconductor memory device, a magnetic storage device, and an optical storage device. Semiconductor memory devices can include volatile memory such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory), as well as non-volatile memory such as ROM (Read Only Memory) and flash memory. Semiconductor memory devices include SSDs (Solid State Drives) using flash memory. Magnetic storage devices include magnetic tapes, floppy disks (registered trademark), and hard disks. Optical storage devices include, for example, Compact Discs (CDs), Digital Versatile Discs (DVDs), and Blu-ray Discs (registered trademark).
[0094] The projector 30 is disposed at the first user U1 side and projects text, still images, and moving images onto the transparent screen 5. Alternatively, the projector 30 may be disposed at the second user U2 side and project text, still images, and moving images onto the transparent screen 5.
[0095] (Example of display on a projector display system)
[0096] Next, the image displayed on the transparent screen 5 by the projector 30 based on the image signal from the electronic device 20 will be described. In the transparent screen 5, the area on the screen portion 11 is divided into a first display area A1 and a second display area A2. The first display area A1 is the area where the screen portion 11 overlaps with the light shielding portion 12. The second display area A2 is the area where the screen portion 11 and the light shielding portion 12 do not overlap. The first display area A1 and the second display area A2 are used to display the speech of the first user U1 and / or the second user U2 as text in the first display area A1 and / or for the second user U2. The display displayed by the projector 30 in the second display area A2 is referred to as the second display object DO2. The second display object DO2 is a left-right flip of the first display object DO1. The display displayed by the projector 30 in the first display area A1 is referred to as the first display object DO1. The first display area A1 is an area that is displayed to the first user U1 so that the first user U1 can confirm the content of the second display object DO2 displayed to the second user U2.
[0097] like Figure 3 As shown, the transparent screen 5 displays a first display object DO1 in the first display area A1 when observed by the first user U1. In addition, a second display object DO2 is displayed in the second display area A2. For the first user U1, the second display object DO2 is displayed as mirror text. The first display object DO1 is displayed as normal text. "Mirror text" refers to text that is flipped left and right. "Normal text" is text displayed in the usual way. Normal text can refer to upright text. In the present disclosure, "upright" refers to a state in which text, etc. are displayed upright from the user's perspective. "Upright" is used in contrast to a state in which text, etc. are flipped left and right.
[0098] On the other hand, when the image projected by the same projector 30 onto the transparent screen 5 is observed from the second user U2 side, as shown in FIG. Figure 4 As shown, the second display object DO2 is displayed as normal text. Without the light shielding portion 12, the first display object DO1 would be displayed as mirrored text. The first display object DO1 is unnecessary text that does not need to be displayed to the second user U2. Because it is mirrored text, it may confuse the second user U2. However, by arranging the light shielding portion 12 overlapping the first display area A1 of the screen portion 11, the first display object DO1 is not displayed to the second user U2.
[0099] For example, when the shading portion 12 is a dichroic film that blocks the transmission of light with a red wavelength, the projector 30 projects the first display object DO1 using red light. As a result, the image light of the first display object DO1 is reflected by the shading portion 12 and is visually recognized by the first user U1. On the other hand, the image light of the first display object DO1 is difficult to transmit through the shading portion 12, and therefore difficult to be viewed by the second user U2. As a result, it is difficult for the second user U2 to visually recognize the first display object DO1 as a mirror image text. Therefore, the second user U2 is not easily distracted by the first display DO1. In addition, Figure 3 as well as Figure 4 The image shown is an example of a display by a projector display system, and the positions of the first display area A1 and the second display area A2 can be changed as appropriate.
[0100] (Example of using a viewing angle control film for the light shielding portion)
[0101] In one embodiment, as the light shielding portion 12, a film or component that limits the propagation angle of the transmitted light can be used instead of a dichroic film. The film that controls the propagation direction of the transmitted light includes a viewing angle control film. For example, Figure 5As shown, the shading portion 12 can be configured to limit the propagation range of the transmitted light to a range from 0° to a given angle θ in the vertical direction relative to the normal of the transparent screen 5, and prevent the light from propagating in other directions. The given angle θ can be set to 10°-30°, for example, but is not limited thereto. In this way, the shading portion 12 limits the viewing angle in the vertical direction to a range of 2θ. The shading portion 12 can have a Venetian blind effect. Figure 5 As shown, when the eyes of the second user U2 are outside the range, it is difficult for the second user U2 to visually recognize the second display object DO2 projected onto the light shielding portion 12. Therefore, the second user U2 is not easily disturbed by the first display object DO1.
[0102] (Configuration of transparent screen and light shielding part)
[0103] Figures 6 to 8 This diagram shows an example arrangement of the substrate 10, screen portion 11, and light shielding portion 12 in the transparent screen 5. The light shielding portion 12 may be a film or member with other light-shielding functions, instead of the dichroic film and viewing angle control film. The screen portion 11 and light shielding portion 12 have projector projection surfaces 11a and 12a, respectively. The substrate 10, screen portion 11, and light shielding portion 12 can be arranged in any order, as long as these surfaces face the projector 30 (i.e., the first user U1). The projector projection surfaces 11a and 12a are predetermined surfaces on which light from the projector 30 is projected.
[0104] exist Figure 6 In FIG, the transparent screen 5 is configured such that the screen portion 11 is arranged on the projector 30 side of the base material 10 , and the light shielding portion 12 is arranged on the projector 30 side of the screen portion 11 .
[0105] exist Figure 7 In the figure, the transparent screen 5 is configured such that the light shielding portion 12 is arranged on the projector 30 side of the base material 10 , and the screen portion 11 is arranged on the projector 30 side of the base material 10 and the light shielding portion 12 .
[0106] exist Figure 8 In the transparent screen 5 , the screen portion 11 is arranged on the projector 30 side of the base material 10 , and the light shielding portion 12 is arranged on the second user U2 side of the base material 10 .
[0107] exist Figures 6 to 8 Alternatively, the order of the substrate 10, screen portion 11, and light shielding portion 12 can be reversed left to right. A transparent adhesive can be used to bond the substrate 10, screen portion 11, and light shielding portion 12 to each other. Alternatively, the bonding of the substrate 10, screen portion 11, and light shielding portion 12 can be performed by electrostatically charging any of the surfaces to be bonded. Using electrostatic bonding offers the advantage of easy removal and / or relocation.
[0108] In addition, if Figure 9 As shown, the transparent screen 5 may be configured without using the base material 10 such as an acrylic plate or a plastic sheet, and the light shielding portion 12 may be attached to the screen portion 11. In this case, the screen portion 11 also functions as the base material 10.
[0109] (Other structural examples of transparent screen and light shielding portion)
[0110] In addition to the aforementioned configuration examples, the transparent screen 5 and / or light shielding portion 12 according to one embodiment can also adopt various configurations. For example, the light shielding portion 12 can be configured to be repositionable relative to the screen portion 11 or removable. Furthermore, the light shielding portion 12 can be supported by another member and positioned opposite the screen portion 11 at any position, rather than being attached to the substrate 10 and the screen portion 11. Furthermore, the light shielding portion 12 can be attached to a member such as a roller blind. The roller blind can be one suspended from a ceiling, for example, so that a transparent sheet wound into a roll can be pulled down and rolled up.
[0111] (Display Mode of the First Display Object DO1 and / or the Second Display Object)
[0112] The projector display system 1 can display a first display object DO1 and / or a second display object DO2 based on text information obtained through voice recognition of the speech of the first user U1 and / or the second user U2 in various ways. The display method of the first display object DO1 and / or the second display object DO2 is controlled by the electronic device 20. For example, a specific word can be registered as a given word in the storage unit 23 of the electronic device 20. If the speech of the first user U1 obtained as a result of voice recognition contains the given word registered in the storage unit 23, the control unit 22 of the electronic device 20 can change the word to be highlighted in the second display object DO2 displayed on the projector 30. This emphasis includes changing the color of the text to red, for example, or changing the font size of the text. Alternatively, if the speech of the second user U2 obtained as a result of voice recognition contains the given word registered in the storage unit 23, the control unit 22 of the electronic device 20 may not change the word to be highlighted in the second display object DO2 displayed on the projector 30.
[0113] Furthermore, the storage unit 23 of the electronic device 20 may store a given word that has been converted into an image such as a specific illustration or a given dynamic image. When the speech content of the first user U1 obtained as a result of voice recognition includes a given word stored in the storage unit 23, the control unit 22 of the electronic device 20 may convert the word in the second display object DO2 displayed on the projector 30 into an image such as an illustration or a given dynamic image. By converting the given word into an image such as an illustration or a given dynamic image, the second user U2 can easily and intuitively understand the second display object DO2. On the other hand, when the speech content of the second user U2 obtained as a result of voice recognition includes a given word registered in the storage unit 23, the control unit 22 of the electronic device 20 may not convert the word into an image such as an illustration or a given dynamic image.
[0114] There can be multiple first users U1 and / or second users U2. The voice recognition unit 21 or the control unit 22 of the electronic device 20 can be configured to identify the first user U1 and / or the second user U2 who is speaking from multiple people based on the pitch of the speaker's voice, the characteristics of the voiceprint, and the characteristics of the voices of other speakers. The control unit 22 can change the display mode of the text in at least one of the first display DO1 and the second display DO2 according to the determined speaker. Changing the display mode includes changing at least any one of the displayed color and the displayed position. In this way, the first user U1 and / or the second user U2 can easily know who is speaking in the first display DO1 and / or the second display DO2.
[0115] Furthermore, the electronic device 20 may have a translation function for translating text into a foreign language. In this case, the control unit 22 may include a translation engine. The storage unit 23 stores dictionary information required for translation, etc. The control unit 22 translates the speech content of the first user U1 obtained as a result of voice recognition into, for example, the language used by the second user U2, and obtains text information as the second display object DO2. Thus, even if the second user U2 is a person who speaks a different language from the first user U1, the projector display system 1 can display the speech content of the first user U1 in a language that the second user U2 can understand. In addition, the control unit 22 can also translate the speech content of the second user U2 obtained as a result of voice recognition into, for example, the language used by the first user U1, and obtain text information as the first display object DO1. Thus, even if the first user U1 is a person who speaks a different language from the second user U2, the projector display system 1 can display the speech content of the second user U2 in a language that the first user U1 can understand.
[0116] (Use of external cloud servers for speech recognition processing, etc.)
[0117] Next, the use of the external cloud server 60 (refer to Figure 10 as well as Figure 12 ) is described below. A "cloud server" is a server that can be accessed via a network such as the Internet. The cloud server 60 of the following embodiment includes, for example, a server that performs a voice recognition service provided as a paid service by a third party. When using the cloud server 60, it is assumed that a metered charge corresponding to the connection time or the data volume of the voice signal transmitted is performed. The projector display system is used to transmit a voice signal to the cloud server 60 when a given condition is met.
[0118] Furthermore, in one embodiment, the external cloud server 60 may also provide a translation function for translating text into other languages in place of, or in addition to, the speech recognition process. In this case, the cloud server 60 may include a translation engine. By providing the translation function on the cloud server 60, the electronic device 20, for example, can receive the results of text translation into a foreign language from the cloud server 60 even if it does not have a translation engine or other functions.
[0119] (Example of a system configuration for communicating with a cloud server)
[0120] Reference Figure 10 A projector display system 1A according to an embodiment of the present disclosure will be described. Figure 10 In FIG. 1 , components identical or similar to those included in the projector display system 1 are denoted by the same reference numerals as those of the components included in the projector display system 1 .
[0121] The projector display system 1A includes a transparent screen 5, an electronic device 20, a projector 30, and a microphone 40. The transparent screen 5 is Figure 10 Simplified in the figure, but similar to Figure 2 The transparent screen 5 shown also has a screen portion 11 and a light shielding portion 12 .
[0122] In one embodiment, the transparent screen 5, the electronic device 20, and the projector 30 can be as follows: Figure 11 As shown in the top view of . Figure 11 As shown, the first user U1 faces the direction of arrow d1, that is, toward the second user U2 through the transparent screen 5. Furthermore, it is assumed that the second user U2 faces the direction of arrow d2, that is, toward the first user U1 through the transparent screen 5. In short, the first user U1 and the second user U2 face each other through the transparent screen 5.
[0123] The transparent screen 5 can be formed integrally with a transparent substrate 10 such as an acrylic plate. The transparent screen 5 can be provided on the entire surface or a portion of one surface of the substrate. The transparent screen 5 is supported by one or more brackets 61 and is arranged on a table between the first user U1 and the second user U2 in a state of standing approximately vertically. The projector 30 is fixed to a positioning member 62 extending from one of the brackets 61. The positioning member 62 fixes the distance and orientation of the projector 30 relative to the transparent screen 5. In the illustrated example, the electronic device 20 is, for example, a smartphone or a portable information terminal. The microphone 40 can be built into the electronic device 20. The microphone 40 can also be a device different from the electronic device 20 that is electrically connected to the electronic device 20. In Figure 11 In the example shown, the microphone 40 is configured on the first user U1 side by being built into or connected to the electronic device 20. Alternatively, the microphone 40 may be configured on the second user U2 side by being connected to the electronic device 20. Furthermore, the microphone 40 may include a microphone configured on the first user U1 side by being built into or connected to the electronic device 20, and a microphone configured on the second user U2 side by being connected to the electronic device 20.
[0124] The electronic device 20 can also be used with Figure 2 The electronic device 20 of the projector display system 1 similarly includes a control unit 22 and a storage unit 23, but does not include a voice recognition unit 21. The electronic device 20 also includes a communication unit 27. Furthermore, the electronic device 20 may further include an input unit 28 and / or a timer unit 29.
[0125] return Figure 10 Continuing with the description of the projector display system 1A, the control unit 22 includes one or more processors, similarly to the control unit 22 of the projector display system 1. The control unit 22 is configured to control the entire electronic device 20 and transmit the voice signal acquired by the microphone 40 to the cloud server 60 via the communication unit 27. The control unit 22 is configured to acquire text information (text information) converted from the voice signal by the cloud server 60 via the communication unit 27. The control unit 22 may also be configured to acquire text information (text information) translated into a foreign language by the cloud server 60 via the communication unit 27. As in the case of the projector display system 1, the control unit 22 causes the projector 30 to display the received text information as a display object.
[0126] The storage unit 23 stores information required for processing executed by the control unit 22 and information obtained as a result of the processing by the control unit 22, similarly to the projector display system 1. The control unit 22 includes a storage device such as a semiconductor memory, similarly to the projector display system 1.
[0127] Communication unit 27 communicates with devices external to electronic device 20 via wireless or wired communication means. Communication unit 27 can implement various communication methods, including wired LAN (local area network) standards, wireless LAN standards such as Wi-Fi, and mobile communication standards such as 4G (4th Generation) and 5G (5th Generation). Communication unit 27 transmits audio signals and receives text data to cloud server 60.
[0128] The input unit 28 receives input and operation from the first user U1. The input unit 28 may include a switch provided by the electronic device 20, or a switch electrically connected to the electronic device 20. The switch as part of the input unit 28 may also be installed on the microphone 40. The input unit 28 may be linked to the switch of the microphone 40. The input unit 28 may also be, for example, a touch panel provided by the electronic device 20. Figure 12 In the example shown, it is assumed that the input unit 28 is located on the first user U1 side by being built into or connected to the electronic device 20. Alternatively, the input unit 28 may be located on the second user U2 side by being connected to the electronic device 20. Furthermore, the input unit 28 may include an input unit located on the first user U1 side by being built into or connected to the electronic device 20, and an input unit located on the second user U2 side by being connected to the electronic device 20. The input unit 28 located on the second user U2 side may also be attached to the microphone 40 located on the second user U2 side.
[0129] The timer 29 is configured to measure time. The timer 29 can be implemented, for example, by a real-time clock (RTC) built into the electronic device 20. The timer 29 can be provided as a function of the control unit 22. The timer 29 can function as one or more timers capable of measuring elapsed time starting at one or more time points.
[0130] (Example with a detection unit)
[0131] Reference Figure 12 A projector display system 1B including a detection unit according to an embodiment of the present disclosure will be described. Figure 10 The projector display system 1A shown can also communicate with an external cloud server 60 that provides voice recognition processing. Figure 10 The projector display system 1A shown differs in that it includes a detection unit 42. The configuration of the projector display system 1B other than the detection unit 42 is the same as or similar to that of the projector display system 1A, so only the differences from the projector display system 1A will be described below, and other descriptions will be omitted.
[0132] The detection unit 42 may also be provided separately from the electronic device 20 and connected to the electronic device 20 via wired and / or wireless communication. In addition, the detection unit 42 may also be provided (for example, built-in) in the electronic device 20 or the microphone 40. The detection unit 42 may include Figure 13 The detection unit 42 may include at least one of a switch 42a, an acceleration sensor and / or a gyro sensor 42b, a magnetic sensor 42c, and a camera 41. The detection unit 42 may include a sensor for detecting the presence of at least one of the first user U1 and the second user U2.
[0133] The switch 42a can be operated by the first user U1 and / or the second user U2. For example, the switch 42a can detect an operation pressed by the first user U1 and / or the second user U2. By being provided on one side of each of the first user U1 and the second user U2, the switch 42a can detect whether an operation based on one of the first user U1 and the second user U2 has been performed. The switch 42a can also be designed to switch on / off, for example. In addition, the switch 42a can also be designed to maintain an on state only when it is pressed, and maintain an off state when it is not pressed. The switch 42a can also be a touchpad that detects operations such as touching, clicking, sliding, swiping and / or pinching by the first user U1 and / or the second user U2. The switch 42a can be provided on either the electronic device 20 or the microphone 40. In addition, for example, the input unit 28 can also realize the function of the switch 42a. When the microphone 40 detects voice input, the first user U1 and / or the second user U2 operates the switch 42a, thereby detecting whether the voice detected by the microphone 40 is uttered by the first user U1 or the second user U2.
[0134] The accelerometer / gyro sensor 42b may include at least one of an accelerometer and a gyro sensor. The accelerometer / gyro sensor 42b may have the function of detecting the orientation of the electronic device 20 or the microphone 40. For example, the accelerometer / gyro sensor 42b may be configured to detect whether the microphone 40 is facing the first user U1 or the second user U2. When the microphone 40 detects voice input, the accelerometer / gyro sensor 42b detects the orientation of the microphone 40, thereby determining whether the voice detected by the microphone 40 is uttered by the first user U1 or the second user U2.
[0135] The magnetic sensor 42c, like the acceleration sensor / gyroscope sensor 42b, can have a function of detecting the orientation of the electronic device 20 or the microphone 40. For example, the magnetic sensor 42c can be configured to detect whether the microphone 40 is facing the first user U1 or the second user U2 by comparing the orientation of the magnetic sensor 42c with the direction of the geomagnetic field. When the microphone 40 detects the input of sound, the magnetic sensor 42c detects the orientation of the microphone 40, thereby enabling detection of whether the voice detected by the microphone 40 is emitted by the first user U1 or the second user U2.
[0136] The camera 41 can be one or two cameras that respectively photograph the first user U1 and the second user U2. In addition, the camera 41 can also be a camera with a wide field of view that photographs both the first user U1 and the second user U2. When there is a human image in the image of the camera 41, the presence of either or both of the first user U1 and the second user U2 can be detected. In addition, the camera 41 can also photograph the mouths of the first user U1 and the second user U2 respectively. When the microphone 40 detects the input of sound, by detecting the movements of the mouths of the first user U1 and the second user U2 photographed by the camera 41, it is possible to detect whether the voice detected by the microphone 40 is emitted by the first user U1 or the second user U2.
[0137] (Embodiment of a specific example of the projector display system 1)
[0138] Next, the technical problem solved by the projector display system 1 according to the embodiment of the present disclosure will be described. The first user U1 and the second user U2 can either use the same language or different languages. Hereinafter, an example in which the first user U1 and the second user U2 use the same language will be described.
[0139] The projector display system 1, for example, acquires the voice data of the Japanese language of the first user U1 via the microphone 40 or the like, and displays the voice data as text in a given area (second display area A2) of the screen unit 11. As a result, for example, as Figure 14 shown, in the second display area A2, the speech content of the first user U1 is displayed as text in Japanese such as "おはようございます。本日はどのようなご用件でしょうか". Figure 14 The situation of the display in the screen unit 11 as viewed from the second user U2 side is shown. The second user U2 can confirm the result of the speech recognition of the speech content of the first user U1 by visually observing the display (text in Japanese) in the second display area A2.
[0140] Next, the projector display system 1, for example, obtains voice data in Japanese of the second user U2 via a microphone 40 or the like, and displays the voice data as text in a given area (second display area A2) of the screen unit 11. As a result, for example, as Figure 14 shown, in the second display area A2, the speech content of the second user U2 is displayed as text in Japanese such as "福祉パスの申請に来ました。". The second user U2 can confirm the result of the speech recognition of his or her own speech content by visually observing the display (text in Japanese) in the second display area A2.
[0141] Furthermore, assume that the projector display system 1 obtains voice data in Japanese of the first user U1 and displays the voice data as text in a given area (second display area A2) of the screen unit 11. As a result, for example, as Figure 14 shown, in the second display area A2, the speech content of the first user U1 is displayed as text in Japanese such as "わかりました少々お待ちください".
[0142] As a result, as Figure 14 shown in the second display area A2, if the conversations of the first user U1 and the second user U2 are concentrated and displayed on one screen, it is difficult for the first user U1 and / or the second user U2 to clearly distinguish which party's speech each string is. That is, in the Figure 14 shown display, it is difficult for both the first user U1 and the second user U2 to distinguish their own speech from the other party's speech, and it is also difficult to grasp the history of the communication between the two parties in the conversation. Therefore, in the communication through such a displayed screen, there is a possibility that the communication between the interlocutors cannot proceed smoothly.
[0143] According to the projector display system 1 according to an embodiment, the text based on the speech of the first user and the text based on the speech of the second user are displayed in mutually different display manners. Therefore, according to the projector display system 1 according to an embodiment, for example, even if the texts of two interlocutors are displayed on the same screen, it is easy for the first user U1 and the second user U2 to distinguish which party's speech each string is. That is, according to the projector display system 1 according to an embodiment, it is easy for both the first user U1 and the second user U2 to distinguish their own speech from the other party's speech, and it is easy to grasp the history of the communication between the two parties in the conversation. Therefore, according to the projector display system 1 according to an embodiment, the communication between the interlocutors becomes smooth.
[0144] Hereinafter, the above-described operation of the projector display system 1 will be described. Figure 15 is a flowchart for explaining the operation of the projector display system 1.
[0145] In the following description, the language used by the first user U1 (speaker) is also recorded as the first language. In addition, the language used by the second user U2 (listener) is also recorded as the second language. In addition, similar to the above example, the language used by the first user U1 (speaker) is Japanese. That is, the first language is Japanese. On the other hand, the language used by the second user U2 (listener) is also Japanese. That is, the second language is Japanese. The first language and the second language can be set by the first user U1 and / or the second user U2 at the start or before the start of the operation of the projector display system 1, for example, or can be set at an appropriate time. In addition, in one example, the first language can also be automatically determined by, for example, voice recognition of the first user's language. Similarly, the second language can also be automatically determined by, for example, voice recognition of the second user's language.
[0146] when Figure 15 When the action shown in FIG. 1 is started, the projector display system 1 obtains the voice data based on the speech of the first user U1 and / or the voice data based on the speech of the second user U2 (step S11). Figure 15 Controls related to the actions shown.
[0147] In step S11, the control unit 22 can, for example, obtain voice data based on the speech of the first user U1 and / or the second user U2 detected by the microphone 40. For example, when the microphone 40 is respectively provided on the first user U1 side and the second user U2 side, each microphone 40 can detect the speech of the first user U1 or the second user U2. In addition, when the microphone 40 is provided only on the first user U1 side, when the second user U2 speaks, the microphone 40 is directed toward the second user U2, thereby the microphone 40 can also detect the speech of the second user U2. In this case, it is also possible that when the first user U1 speaks next, the microphone 40 is directed toward the first user U1, thereby the microphone 40 detects the speech of the first user U1.
[0148] Next, the projector display system 1 determines whether the acquired speech data based on the speech of the first user U1 and / or the second user U2 is based on the speech of the first user U1 or the second user U2 (step S12). In step S12, the projector display system 1 can determine whether the acquired speech data is based on the speech of the first user U1 or the second user U2 based on a given condition. Various conditions can be used as the given conditions. Specific examples of the given conditions are described below.
[0149] In one embodiment, the control unit 22 can determine whether the speech data is based on the speech of the first user U1 or the second user U2 based on conditions related to detection by the switch 42a. For example, if the switch 42a is provided on the first user U1 side and the second user U2 side, the first user U1 and the second user U2 can be instructed to turn on the switch 42a when speaking. For example, if the microphone 40 detects speech and the switch 42a provided on the first user U1 side is turned on, the data obtained by the control unit 22 can be determined to be speech data based on the speech of the first user U1. On the other hand, if the microphone 40 detects speech and the switch 42a provided on the second user U2 side is turned on, the data obtained by the control unit 22 can be determined to be speech data based on the speech of the second user U2.
[0150] Furthermore, for example, if the switch 42a is provided only on one of the first user U1 and the second user U2, it is also possible to instruct either the first user U1 or the second user U2 to switch the switch 42a on or off depending on the user who is speaking. For example, the switch 42a may be turned on when the first user U1 is speaking, and turned off otherwise. In this case, for example, if the switch 42a is on when the microphone 40 detects speech, the control unit 22 can determine that the data acquired is based on the speech of the first user U1. On the other hand, if the switch 42a is off when the microphone 40 detects speech, the control unit 22 can determine that the data acquired is based on the speech of someone other than the first user U1 (e.g., the second user U2). Furthermore, for example, the switch 42a may be turned on when the second user U2 is speaking, and turned off otherwise. In this case, for example, if the switch 42a is in the on state when the microphone 40 detects voice, the data obtained by the control unit 22 can be determined to be voice data based on the speech of the second user U2. On the other hand, for example, if the switch 42a is in the off state when the microphone 40 detects voice, the data obtained by the control unit 22 can be determined to be voice data based on the speech of someone other than the second user U2 (for example, the first user U1).
[0151] In one embodiment, the control unit 22 can determine whether the speech data is based on the speech of the first user U1 or the second user U2 based on conditions related to detection by at least one of the accelerometer / gyro sensor 42b and the magnetic sensor 42c. The orientation of the electronic device 20 and / or the microphone 40 can be determined based on at least one of the accelerometer / gyro sensor 42b and the magnetic sensor 42c. Therefore, if the microphone 40 is located only on the first user U1 side, the first user U1 can be instructed to orient the microphone 40 toward the first user U1 when the first user U1 is speaking. Alternatively, in this case, the first user U1 can be instructed to orient the microphone 40 toward the second user U2 when the second user U2 is speaking. For example, if the microphone 40 detects speech and is oriented toward the first user U1, the data acquired by the control unit 22 can be determined to be speech data based on the speech of the first user U1. On the other hand, for example, when the microphone 40 detects voice, if the microphone 40 is directed toward the second user U2 , the data acquired by the control unit 22 can be determined to be voice data based on the speech of the second user U2 .
[0152] In one embodiment, the control unit 22 can determine whether the speech data originates from the first user U1 or the second user U2 based on conditions related to detection by the camera 41. The control unit 22 can determine whether the first user U1 and / or the second user U2 are speaking based on images or moving images of their faces, etc., captured by the camera 41. For example, the control unit 22 can determine whether the first user U1 and / or the second user U2 are speaking based on images or videos of their mouths, etc., captured by the camera 41. In this case, if the microphone 40 is provided on both the first user U1 and second user U2 sides, for example, the first user U1 and / or the second user U2 do not need to personally operate an operating unit such as the switch 42a. For example, if the microphone 40 detects speech and the camera 41 detects movement of the first user U1's mouth, etc., the data acquired by the control unit 22 can be determined to be speech data originating from the first user U1's speech. On the other hand, for example, when the microphone 40 detects voice and the camera 41 detects the movement of the second user U2's mouth, the data acquired by the control unit 22 can be determined to be voice data based on the speech of the second user U2.
[0153] In one embodiment, the control unit 22 can determine whether the speech data originates from the speech of the first user U1 or the second user U2, based on conditions related to detection by multiple microphones, such as a microphone array (array microphone). The microphone 40, which includes multiple microphones, can determine the direction of the sound source producing the sound detected by these multiple microphones. Therefore, if such a microphone 40 detects speech originating from, for example, the first user U1, the control unit 22 can determine that the data acquired is speech data originating from the speech of the first user U1. On the other hand, if the microphone 40 detects speech originating from, for example, the second user U2, the control unit 22 can determine that the speech data is originating from the speech of the second user U2.
[0154] After step S12, the projector display system 1 converts the speech data based on the speech of the first user U1 and / or the speech data based on the speech of the second user U2 into text data for each part of speech or each word in the language (step S13). As described above, the control unit 22 may also cause the cloud server 60 to perform the operations of step S13. In this case, the control unit 22 transmits the speech data based on the speech of the first user U1 and / or the speech data based on the speech of the second user U2 to the cloud server 60 via the communication unit 27. The cloud server 60 then transmits the result of converting the speech data based on the speech of the first user U1 and / or the speech of the second user U2 into text data for each part of speech or each word in the language to the electronic device 20. In this case, the electronic device 20 may receive the text data from the cloud server 60 via the communication unit 27. Similarly, the cloud server 60 may also take over at least some of the operations of the control unit 22. In one embodiment, the operations of step S12 may be performed by the cloud server 60 in place of or in conjunction with the operations of the control unit 22. In this case, the control unit 22 transmits the voice data based on the speech of the first user U1 and / or the voice data based on the speech of the second user U2 acquired in step S11 to the cloud server 60 via the communication unit 27 .
[0155] Hereinafter, the text based on the speech of the first user U1, which is converted into text data in step S13, may sometimes be referred to as "first text." Similarly, the text based on the speech of the second user U2, which is converted into text data in step S13, may sometimes be referred to as "second text." In summary, in step S12, the electronic device 20 may perform the following determination step: based on given conditions, the electronic device 20 may determine which of the first and second users' speech at least one of the first and second texts is based on. Based on this processing, the electronic device 20 may, for example, perform the step of obtaining at least one of the first and second texts at a point in time up to step S13.
[0156] As described above, the control unit 22 may also determine, based on a given condition, whether at least one of the first character and the second character is based on the speech of the first user or the second user. Here, the given condition may be, for example, a condition related to detection by at least one of the switch 42a, the acceleration sensor / gyro sensor 42b, the magnetic sensor 42c, the camera 41, and the plurality of microphones.
[0157] Next, the control unit 22 determines whether the first character and the second character are in the same language (step S14). As described above, the case where the first character and the second character are in the same language (Japanese) ("Yes" in step S14) is described here. If the first character and the second character are in the same language, the control unit 22 displays the first character and the second character on the display unit (screen unit 11) in a manner that allows the first character and the second character to be displayed in different formats (step S15). In step S15, the control unit 22 may change the display format of at least one of the first character and the second character and display it on the display unit (screen unit 11). In this way, the electronic device 20 includes a display unit (e.g., screen unit 11) that displays the first character, which is the character based on the speech of the first user, and the second character, which is the character based on the speech of the second user.
[0158] In step S15, the control unit 22 may display one of the first and second characters on the display unit (screen unit 11) with greater emphasis than the other. Furthermore, in step S15, the control unit 22 may display one of the first and second characters on the display unit (screen unit 11) in a manner different from the other in at least one of size, weight, color, font, background color, decorative text, and display position.
[0159] For example, the control unit 22 may display the second character, which is the text based on the second user U2's speech, on the display unit (screen unit 11) in a manner that emphasizes the first character, which is the text based on the first user U1's speech. In this case, the control unit 22 may, for example, make the size of the second character larger than the size of the first character. Furthermore, the control unit 22 may, for example, make the font of the second character more distinctive than the font of the first character. Furthermore, the control unit 22 may, for example, set the color of the first character to black or white, which is used for normal text display, and set the color of the second character to a conspicuous color such as red or yellow. Furthermore, the control unit 22 may, for example, display the second character in italic or bold font, and display the first character in a font that does not use italic or bold font. Furthermore, the control unit 22 may, for example, display the second character underlined and display the first character ununderlined. Furthermore, the control unit 22 may make various changes to the display format of at least one of the first and second characters, so that the first and second characters have different display formats. Furthermore, the control unit 22 may appropriately combine two or more of the above-mentioned display modes so as to display the first character and the second character in different display modes.
[0160] Figure 16A as well as Figure 16B : is a diagram showing an example of display on the display unit (screen unit 11) after the process of step S15 is performed. Figure 16A as well as Figure 16B In the example, based on the content of the text spoken by the first user U1 and the second user U2 Figure 14 The same is true for . That is, Figure 16A as well as Figure 16B The projector display system 1 shows Figure 14 An example of the result of such a text. Figure 14 In the second display area A2 of the first user U1, the first text based on the speech of the first user U1 and the second text based on the speech of the second user U2 are displayed in a state where it is difficult to distinguish. Figure 16A In the first display area A1 and the second display area A2 shown, the first text based on the speech of the first user U1 is displayed in a normal manner, whereas the second text based on the speech of the second user U2 is displayed in a marked manner. Figure 16B In the second display area A2 shown, the first text based on the speech of the first user U1 is displayed in a normal manner, whereas the second text based on the speech of the second user U2 is displayed in a marked manner. Figure 16A as well as Figure 16BAmong them, it is easy to grasp that the text "I have come to apply for the welfare pass." marked with a mark is the second text, that is, the text based on the speech of the second user U2. On the other hand, in Figure 16A and Figure 16B Among them, it is easy to grasp that the texts "Good morning. What can I do for you today?" and "Understood. Please wait a moment." without marks are the first texts, that is, the texts based on the speech of the first user U1.
[0161] Figure 17A and Figure 17B are diagrams showing other display examples of the display unit (screen unit 11) after the process of step S15 is performed. In the first display area A1 and the second display area A2 shown in Figure 17A , the first text based on the speech of the first user U1 and the second text based on the speech of the second user U2 are displayed surrounded by different dialogue bubbles. In addition, in the second display area A2 shown in Figure 16B , the first text based on the speech of the first user U1 and the second text based on the speech of the second user U2 are displayed surrounded by different dialogue bubbles. Therefore, in Figure 16A and Figure 16B Among them, it is easy to grasp that the text "I have come to apply for the welfare pass." marked with a mark is the second text, that is, the text based on the speech of the second user U2. On the other hand, in Figure 16A and Figure 16B Among them, it is easy to grasp that the texts "Good morning. What can I do for you today?" and "Understood. Please wait a moment." without marks are the first texts, that is, the texts based on the speech of the first user U1.
[0162] Figure 16A and Figure 17A show the situation of observing the transparent screen 5 (or screen unit 11) from the first user U1. As shown in Figure 16A and Figure 17A , the first display object DO1 is displayed in the first display area A1. In addition, the second display object DO2 is displayed in the second display area A2. Relative to the first user U1, the second display object DO2 is displayed as mirror text. The first display object DO1 is displayed as normal text.
[0163] Figure 16B and Figure 17B show the situation of observing the image projected by the same projector 30 onto the transparent screen 5 (or screen unit 11) from the side of the second user U2. As shown in Figure 16B and Figure 17BAs shown in FIG. 1 , when viewed from the side of the second user U2, the second display object DO2 is displayed as normal text. In the first display area A1, if there is no light shielding portion 12, the first display object DO1 is displayed as mirror text. The first display object DO1 is unnecessary text that does not need to be displayed to the second user U2. Since it is mirror text, it may confuse the second user U2. However, by overlapping the light shielding portion 12 in the first display area A1 of the screen portion 11 (or the screen portion 11), as shown in FIG. Figure 16B as well as Figure 17B As shown, the first display object DO1 is not displayed to the second user U2.
[0164] Thus, in one embodiment, the control unit 22 may flip one of the first and second characters and display it on the display unit (screen unit 11) included in the display system. Furthermore, the display unit (screen unit 11 or transparent screen 5) included in the display system controlled by the electronic device 20 may include a first display area A1 having a transmissive light-blocking portion 12 that at least partially blocks light, and a second display area A2 that at least partially transmits light. In this case, the control unit 22 may flip one of the first and second characters and display it. Furthermore, the control unit 22 may display at least one of the first and second characters in the first display area A1 while being shielded by the light-blocking portion 12. Furthermore, the control unit 22 may display at least one of the first and second characters in the second display area A2. A display system according to one embodiment may, for example, include a display unit (screen unit 11 or transparent screen 5). Furthermore, a display system according to one embodiment may, for example, include a projector 30. Furthermore, a display system according to one embodiment may, for example, include a microphone 40. Furthermore, the display system according to one embodiment may include, for example, the electronic device 20. Furthermore, the display system according to one embodiment may include, for example, elements other than the aforementioned functional units, or may not include some of the aforementioned functional units.
[0165] As described above, according to the projector display system 1 according to one embodiment, even when text from two interlocutors is displayed on the same screen, for example, the first user U1 and the second user U2 can easily distinguish which user's speech character string is attributed to each other. In other words, according to the projector display system 1 according to one embodiment, both the first user U1 and the second user U2 can easily distinguish their own speech from that of the other user, making it easier to understand the history of their interactions during the conversation. Therefore, according to the projector display system 1 according to one embodiment, communication between the interlocutors becomes smooth.
[0166] (Example of translated text)
[0167] Next, we'll describe a scenario where the first language used by the first user U1 (speaker) differs from the second language used by the second user U2 (listener). Here, the first user U1 (speaker) speaks Japanese. In other words, their first language is Japanese. Meanwhile, the second user U2 (listener) speaks English. In other words, their second language is English.
[0168] exist Figure 15 In step S14, when the first character and the second character are not in the same language (that is, in different languages), the control unit 22 performs Figure 15 After the processing of step S16 shown, the processing of step S15 described above is executed. In step S16, the control unit 22 converts, or translates, the text data of the first character (Japanese) into the language of the second character (English). Alternatively, in step S16, the control unit 22 may convert, or translate, the text data of the second character (English) into the language of the first character (Japanese). As described above, the cloud server 60 may replace the control unit 22 and convert the text data of the first character and / or the second character into text data in different languages.
[0169] Figure 18A as well as Figure 18B 1 is a diagram showing a display example of the display unit (screen unit 11) after the process of step S15 is performed through step S16. Figure 18B In the second display area A2 shown, the electronic device 20 translates the text (first text and second text) based on the speech of the first user U1 and the second user U2 into English and displays it as a display object DO2 for the second user U2. In this way, the electronic device 20 can translate the first text (Japanese) based on the speech of the first user U1 into English and display it in the second display area A2. In addition, the electronic device 20 can directly display the second text (English) based on the speech of the second user U2 in English in the second display area A2. On the other hand, Figure 18A In the first display area A1 shown, the electronic device 20 directly displays the text (first text and second text) based on the speech of the first user U1 and the second user U2 in Japanese as a display object DO1 for the first user U1. In this way, the electronic device 20 can translate the second text (English) based on the speech of the second user U2 into Japanese and display it in the first display area A1. In addition, the electronic device 20 can directly display the first text (Japanese) based on the speech of the first user U1 in Japanese in the first display area A1. Figure 18A In the first display area A1 and the second display area A2 shown, the first text based on the speech of the first user U1 is displayed in a normal manner, whereas the second text based on the speech of the second user U2 is displayed in a marked manner. Figure 16BIn the second display area A2 shown, the first text based on the speech of the first user U1 is displayed in a normal manner, while the second text based on the speech of the second user U2 is displayed in a marked manner. Therefore, in Figure 18A and Figure 18B , it is possible to easily grasp that the text with the mark attached is based on the second text, that is, the text of the speech of the second user U2. On the other hand, in Figure 18A and Figure 18B , it is possible to easily grasp that the text without the mark is based on the first text, that is, the text of the speech of the first user U1.
[0170] In this way, when the language of the first text and the language of the second text are different, the electronic device 20 can also perform a translation step of translating the first text into the language of the second text. In addition, the electronic device 20 can also display the result of translating the first text into the language of the second text as the first text on the display unit (screen unit 11). In addition, when the language of the first text and the language of the second text are different languages, the electronic device 20 can also obtain the result of translating the first text into the language of the second text as the first text, for example, from the cloud server 60. In this case, the electronic device 20 can also display the result of translating the first text into the language of the second text as the first text on the display unit (screen unit 11).
[0171] In addition, as shown in Figure 18A and Figure 18B , when the languages of the first text and the second text are respectively discriminated, it is also possible to perform a display that prompts the language. For example, as shown in the second display area A2 of Figure 18A and Figure 18B , when the language of the second text is English, it is also possible to display a display object DO22 that implies that meaning. For example, it is also possible to display the meaning that the language of the second text is English in Japanese ("English") and English ("EN"). In addition, as shown in the first display area A1 of Figure 18A , when the language of the first text is Japanese, it is also possible to display a display object DO11 that implies that meaning. For example, it is also possible to display the meaning that the language of the first text is Japanese in Japanese ("Japanese") and English ("JA" or "JP", etc.). The language of the first text and the language of the second text can be set by the first user U1 and / or the second user U2, for example, at the start or before the start of the operation of the projector display system 1, or can be set at an appropriate timing. In addition, in one example, the language of the first text can also be automatically discriminated by voice recognition of the language of the first user. Similarly, the language of the second text can also be automatically discriminated by voice recognition of the language of the second user. The display method of the display object DO11 and the display object DO22 is not limited to Figure 18Aas well as Figure 18B The display mode shown in the figure may be any of various display modes. In addition, the positions (locations) where the display objects DO11 and DO22 are displayed on the display unit (screen unit 11) are not limited to the positions (locations) where the display objects DO11 and DO22 are displayed on the display unit (screen unit 11). Figure 18A as well as Figure 18B The examples shown can be set to various positions (parts).
[0172] (Example of displaying graphics related to keywords)
[0173] When a given keyword is displayed on the display unit (screen unit 11) based on the text of the speech of the first user U1 and the second user U2, the electronic device 20 may also display a given graphic associated with the given keyword. To perform this process, for example, the given keyword and the given graphic may be stored in association with each other in the storage unit 23 of the electronic device 20. Furthermore, when performing this process, the control unit 22 of the electronic device 20 may obtain information storing the given keyword and the given graphic in association with each other from, for example, the cloud server 60 via the communication unit 27.
[0174] Figure 19 This is a flowchart illustrating the operation of the projector display system 1 for displaying a given graphic related to a given keyword. Figure 19 The actions shown can be performed, for example, Figure 15 The actions shown are performed after step S15.
[0175] when Figure 19 When the operation shown in the figure starts, the control unit 22 determines whether the first character displayed on the display unit (screen unit 11) contains a given keyword (step S21). If the first character does not contain the given keyword in step S21, the control unit 22 can skip the processing shown in step S22 and end the process. Figure 19 On the other hand, if the first character contains a given keyword in step S21, the control unit 22 may end the process after executing the process shown in step S22. Figure 19 The operation shown. In step S22, the control unit 22 displays the image or video corresponding to the given keyword on the display unit (screen unit 11). Here, the image or video corresponding to the given keyword may be an image or video associated with the given keyword. The image or video associated with the given keyword may be read from the storage unit 23 or obtained from the cloud server 60 via the communication unit 27, for example.
[0176] exist Figure 16A as well as Figure 16BIn the example shown, for instance, in the storage unit 23 or the cloud server 60, a given keyword "福祉パス" is associated with a picture or video representing the appearance of this "福祉パス". In this case, when the electronic device 20 displays the text of "福祉パス" in the second display area A2, for example, it can display a picture or video representing the appearance of "福祉パス" in the third display area A3 on the display unit (screen unit 11). Here, the electronic device 20 can display an image or moving image representing the appearance of "福祉パス" on the display unit (screen unit 11) simultaneously or approximately simultaneously while displaying the text of "福祉パス". In addition, the electronic device 20 can also display an image or moving image representing the appearance of "福祉パス" on the display unit (screen unit 11) slightly before or slightly after displaying the text of "福祉パス".
[0177] In addition, in Figure 18A and Figure 18B In the example shown, for instance, in the storage unit 23 or the cloud server 60, a given keyword "bullet train" is associated with an image or moving image representing the appearance of "新干线". In this case, when the electronic device 20 displays the text of "bullet train" in the second display area A2, for example, it can display an image or moving image representing the appearance of "新干线" in the third display area A3 on the display unit (screen unit 11).
[0178] In addition, in the case of displaying a given graphic associated with a given keyword as described above, for example, it is also possible to easily understand the text based on the speech of the first user U1 and assist the second user U2. That is, the electronic device 20 can display a picture or video associated with the given keyword, for example, in the third display area A3 only when the first text based on the speech of the first user U1 contains the given keyword. On the other hand, even when the second text based on the speech of the second user U2 contains the given keyword, the electronic device 20 may not display the picture or video associated with the given keyword in the third display area A3, for example. Thus, since the given picture or video is displayed only in accompaniment with the speech of the service provider side (the first user U1), it is possible to suppress the situation where the frequent display of the given picture or video during the conversation between the interlocutors causes confusion in the smooth conversation.
[0179] Thus, when the first text contains the given keyword, the electronic device 20 can also display the given picture or video corresponding to the given keyword on the display unit (screen unit 11). On the other hand, when the second text contains the given keyword, the electronic device 20 may not display the given picture or video on the display unit (screen unit 11).
[0180] Furthermore, to perform the aforementioned processing, the electronic device 20 may also distinguish the first character information from the second character information, for example, by obtaining the information from a cloud server 60. The information thus obtained may be stored, for example, in the storage unit 23 of the electronic device 20 or in a working area within any storage unit. Furthermore, to perform the aforementioned processing, the electronic device 20 may also perform a storage step to separately store the first character information from the second character information. The information thus obtained may be stored, for example, in the storage unit 23 of the electronic device 20 or in a storage area within any storage unit.
[0181] In the above-described embodiment, the sizes and / or positions of the first display area A1, the second display area A2, and / or the third display area A3 are not limited to those shown in the illustrations and can be modified to any desired size and / or position to improve user convenience. Furthermore, the third display area A3, which displays an image or video corresponding to a given keyword, can also be an area at least partially contained within at least one of the first display area A1 and the second display area A2.
[0182] At least part of the control and processing described above may be executed by the projector display system 1 instead of the electronic device 20 or together with the electronic device 20 , or may be executed by the cloud server 60 .
[0183] According to a program according to one embodiment, the display format of at least one of the first and second characters is changed and displayed on the display unit (screen unit 11) so that the first and second characters have different display formats. Therefore, according to the program according to one embodiment, even if the characters of two interlocutors are displayed on the same screen, the first user U1 and the second user U2 can easily distinguish which user's speech is the user's. In other words, according to the program according to one embodiment, the first user U1 and the second user U2 can easily distinguish their own speech from the other user's speech, and can easily grasp the history of their communication during the conversation. Therefore, according to the program according to one embodiment, communication between the interlocutors is smooth. Furthermore, according to the projector display system 1 according to one embodiment, the display format of at least one of the first and second characters is changed and displayed on the display unit (screen unit 11) so that the first and second characters have different display formats. Therefore, according to the projector display system 1 according to one embodiment, communication between the interlocutors is smooth.
[0184] Although the embodiments of the present disclosure have been described based on the accompanying drawings and embodiments, it should be noted that those skilled in the art can easily make various modifications or corrections based on the present disclosure. Therefore, it should be noted that these modifications or corrections are included in the scope of the present disclosure. For example, the functions included in each structural part can be reconfigured in a logically non-contradictory manner, and multiple structural parts can be combined into one or divided. Although the embodiments of the present disclosure have been described with the device as the center, the embodiments of the present disclosure can also be implemented as a method including steps performed by each structural part of the device. The embodiments of the present disclosure can also be implemented as a method executed by a processor possessed by the device, a program, or a storage medium or recording medium having a program recorded thereon. It should be understood that the scope of the present disclosure also includes these.
[0185] In the present disclosure, the descriptions of "first" and "second" are identifiers used to distinguish the structures. The structures distinguished by the descriptions of "first" and "second" in the present disclosure can exchange the numbers in the structures. For example, the first display can exchange the identifiers "first" and "second" with the second display. The identifiers can be exchanged at the same time. The structures can also be distinguished after the identifiers are exchanged. The identifiers can be deleted. The structures with the identifiers deleted are distinguished by symbols. The interpretation of the order of the structures and the existence of identifiers with smaller numbers should not be interpreted based on the descriptions of identifiers such as "first" and "second" in the present disclosure.
[0186] The constituent elements and functional blocks included in each embodiment of the present disclosure can be appropriately configured on the same hardware or different hardware. Part or all of the microphone, camera and detection unit of each embodiment of the present disclosure can be included in the same hardware as the electronic device or projector. Part or all of the control unit, storage unit, communication unit, input unit and timing unit involved in each embodiment of the present disclosure can also be included in the projector as hardware or software. The projector display system using the cloud server of the present disclosure can also be used in a system that does not include a second display area and only uses the first display area to display the speech content of the first user as text information (text information) to the second user.
[0187] -Description of Reference Numerals-
[0188] 1. 1A, 1B Projector Display System
[0189] 5. Transparent screen
[0190] 10. Substrate
[0191] 11 Screen
[0192] 12 shading part
[0193] 20 Electronic devices
[0194] 21 Speech Recognition Unit
[0195] 22 Control Unit
[0196] 23 Storage
[0197] 24 Display unit
[0198] 27 Ministry of Communications
[0199] 28 Input
[0200] 29 Timekeeping Department
[0201] 30 Projector
[0202] 40 microphones
[0203] 42 Testing Department
[0204] 42a switch
[0205] 42b accelerometer / gyroscope sensor
[0206] 42c magnetic sensor
[0207] 41 Camera
[0208] 60 cloud servers
[0209] 61 bracket
[0210] 62 Positioning components
[0211] U1 First User
[0212] U2 Second User
[0213] A1 first display area
[0214] A2 Second display area
[0215] DO1 First Display
[0216] DO2 Second Display
[0217] θ Half field of view angle.
Claims
1. A program for an electronic device to control a display system, wherein the display system displays first text based on a first user's speech and second text based on a second user's speech on a display unit, the program causing the electronic device to execute: an acquiring step of acquiring at least one of the first character and the second character; a determining step of determining, based on a given condition, whether at least one of the first character and the second character is based on the speech of the first user or the second user; as well as The display step includes changing a display mode of at least one of the first character and the second character and displaying the character on the display unit so that the first character and the second character are displayed in different modes.
2. The program according to claim 1, wherein In the discrimination step, As the given condition, based on conditions related to detection performed by at least any one of a switch, an acceleration sensor, a gyro sensor, a magnetic sensor, a camera, and a plurality of microphones, it is determined whether at least one of the first text and the second text is based on the speech of the first user or the second user.
3. The program according to claim 1, wherein In the display step, One of the first character and the second character is displayed on the display unit more emphasized than the other.
4. The program according to claim 1, wherein In the display step, One of the first character and the second character is displayed on the display unit in a different size, thickness, color, font, background color, decorative text, and display position from the other.
5. The program according to claim 1, wherein When the language of the first character is different from the language of the second character, causing the electronic device to perform a translation step of translating the first text into a language of the second text, In the displaying step, a result of translating the first character into the language of the second character is displayed as the first character.
6. The program according to claim 1, wherein When the language of the first character and the language of the second character are different languages, In the obtaining step, a result of translating the first character into the language of the second character is obtained as the first character. In the displaying step, a result of translating the first character into the language of the second character is displayed as the first character.
7. The program according to any one of claims 1 to 6, wherein In the display step, one of the first character and the second character is flipped left-right and displayed on the display unit.
8. The program according to any one of claims 1 to 6, wherein The display unit of the display system includes: a first display area having a light-shielding portion that at least partially blocks light transmission; and a second display area that at least partially transmits light. In the display step, flipping the first character or the second character left to right, At least one of the first character and the second character is displayed in the first display area while being shielded by the light shielding portion. At least one of the first character and the second character is displayed in the second display area.
9. The program according to claim 1, wherein In the display step, When a given keyword is included in the first text, a given picture or video corresponding to the given keyword is displayed on the display unit. When the second text includes the given keyword, the given picture or video is not displayed on the display unit.
10. The program according to claim 1, wherein In the obtaining step, the first character and the second character are obtained separately; or The electronic device is caused to execute a storage step of storing the first character and the second character separately.
11. An electronic device comprising a control unit configured to control a display unit to display first characters based on a first user's speech and second characters based on a second user's speech. The control unit performs the following processing: obtaining at least one of the first character and the second character, Based on a given condition, determining whether at least one of the first character and the second character is based on the speech of the first user or the second user, The display format of at least one of the first character and the second character is changed and displayed on the display unit so that the first character and the second character have different display formats.
12. A control method for an electronic device that displays, on a display unit, first text based on a first user's speech and second text based on a second user's speech, comprising: an obtaining step of obtaining at least one of the first character and the second character; a determining step of determining, based on a given condition, whether at least one of the first character and the second character is based on the speech of the first user or the second user; as well as The display step includes changing a display mode of at least one of the first character and the second character and displaying the character on the display unit so that the first character and the second character are displayed in different modes.
Citation Information
Patent Citations
Translation display device
JP2019070915A
Decomposition method of nitrous oxide and decomposition device of nitrous oxide
JP2023027828A