Information terminal and voice output method for information terminal
By generating stereophonic signals to create an acoustic space corresponding to the terminal screen, the information terminal addresses the challenge of inadequate layout recognition for visually impaired users, enhancing their interaction with icons, menus, and document editing.
Patent Information
- Application Number
- PCT/JP2024/014171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-05
- Publication Date
- 2025-10-09
AI Technical Summary
Existing information terminals, such as personal computers and smartphones, do not adequately assist visually impaired users in quickly identifying and navigating the layout of icons, menus, and document editing points, despite having text-to-speech functionality.
The information terminal generates and outputs stereophonic signals to create an acoustic space corresponding to the terminal screen, using boundary sounds to help users recognize the layout and positions of icons, menus, and document sections.
This method enhances user accessibility by allowing visually impaired individuals to efficiently navigate and operate the terminal screen, improving their ability to interact with icons, menus, and edit documents.
Smart Images

Figure JP2024014171_09102025_PF_FP_ABST
Abstract
Description
Information terminal and audio output method for information terminal
[0001] The present invention relates to an information terminal such as a personal computer, a smartphone, or a wearable display, and to a sound output method for an information terminal.
[0002] Patent Document 1 discloses a voice output device that modulates and outputs voices that read text data. When outputting voice data corresponding to the text data, the voice output device modulates the voice data in accordance with character attribute information (character type, character size, character decoration) of the character strings that make up the text data.
[0003] JP 2016-170218 A
[0004] In recent years, information terminals such as personal computers, smartphones, and wearable displays have been equipped with features to improve accessibility for visually impaired users. One such feature is the text-to-speech function. This function makes it possible to read aloud the names of icons and menus on the information terminal screen using synthesized speech, as well as the text data in open documents using synthesized speech.
[0005] As a technique for reading text data aloud, a technique is known that modulates audio data according to character attribute information of the text data to make it easier to understand the contents of the document, as shown in Patent Document 1. However, even if such a technique is used, sufficient accessibility is not necessarily achieved.
[0006] Specifically, users such as visually impaired people desire to be able to quickly reach the edited portion when editing a document, such as by replacing or inserting text. Similarly, users desire to be able to quickly reach the target portion when operating icons or menus on a terminal screen. To achieve this, it is desirable to have a system that allows users to efficiently recognize the layout on the terminal screen. However, the technology of Patent Document 1 does not allow users to recognize such layout on the terminal screen.
[0007] The present invention has been made in view of the above, and one of its objects is to provide an information terminal and a voice output method for an information terminal that can improve accessibility for users.
[0008] The above and other objects and novel features of the present invention will become apparent from the description of this specification and the accompanying drawings.
[0009] An information terminal according to one embodiment is configured to display a predetermined terminal screen, set an acoustic space corresponding to the terminal screen based on the user of the information terminal, and generate and output boundary sounds as stereophonic signals to allow the user to recognize the boundaries of the set acoustic space.
[0010] According to one embodiment, accessibility for users can be improved.
[0011] 1 is a block diagram showing an example of the configuration of an information terminal according to a first embodiment. FIG. 2 is a schematic diagram illustrating an example of an acoustic space set by the information terminal shown in FIG. 1. FIG. 3 is a schematic diagram illustrating a more detailed example of the terminal screen and acoustic space shown in FIG. 2, including their correspondence. FIG. 4 is a flowchart showing an example of processing content based on a stereophonic signal generation program in the information terminal shown in FIG. 1. FIG. 5 is a schematic diagram illustrating an example of processing content when a user operates a document in an information terminal according to a second embodiment. FIG. 6 is a schematic diagram illustrating an example of processing content different from that in FIG. 5. FIG. 7 is a flowchart showing an example of processing content based on the stereophonic signal generation program shown in FIG. 1 in an information terminal according to the second embodiment. FIG. 8 is a schematic diagram illustrating an example of an acoustic space different from that in FIG. 8, which is set by the information terminal shown in FIG. 1 in an information terminal according to a third embodiment. FIG. 9 is a schematic diagram illustrating an example of processing content based on the stereophonic signal generation program shown in FIG. 1 in an information terminal according to the third embodiment. FIG. 11 is a schematic diagram illustrating an example of processing content when a user operates a document in an information terminal according to a fourth embodiment. 14 is a block diagram showing an example of the configuration of an information terminal according to the fifth embodiment when the information terminal is a head-mounted display (HMD). FIG. 15 is a flowchart showing an example of processing based on the stereophonic signal generation program shown in FIG. 1 in an information terminal according to a fourth embodiment. FIG. 16 is a flowchart showing an example of processing based on the stereophonic signal generation program shown in FIG. 16 in an information terminal according to the fifth embodiment. FIG. 17 is a diagram showing an example of application of an information terminal according to a fifth embodiment to a head-mounted display (HMD), and is a schematic diagram showing an example of the relationship between the display space and acoustic space of the HMD. FIG. 18 is a schematic diagram showing an example of the relationship between the display space and acoustic space of the HMD and a camera image. FIG. 19 is a block diagram showing an example of the configuration of an information terminal according to the fifth embodiment when the information terminal is a head-mounted display (HMD). FIG. 19 is a flowchart showing an example of processing based on the stereophonic signal generation program shown in FIG. 16 in an information terminal according to the fifth embodiment.
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings for explaining the embodiments, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.
[0013] Furthermore, the audio playback technology according to the present invention can contribute to "9. Build resilient infrastructure, promote inclusive and sustainable industrialization, foster innovation and build resilient infrastructure" of the Sustainable Development Goals (SDGs) advocated by the United Nations.
[0014] (First embodiment) <Configuration of information terminal> Fig. 1 is a block diagram showing an example of the configuration of an information terminal according to a first embodiment. The information terminal 1 shown in Fig. 1 is, for example, a personal computer, a mobile terminal such as a smartphone, or a wearable display such as a head-mounted display. In this specification, a personal computer is also referred to as a PC, and a head-mounted display is also referred to as an HMD.
[0015] 1 , information terminal 1 includes a camera unit 10, a sensor unit 11, an image display unit 12, an operation input unit 13, an audio input unit 14, an audio output unit 15, a communication unit 16, a main processing unit 17, a memory unit 18, a storage unit 19, and an internal bus 20 connecting these units. Main processing unit 17 is configured with processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a DSP (Digital Signal Processor). Memory unit 18 is a storage device and is configured with, for example, a RAM or flash memory.
[0016] The memory unit 18 stores an OS (Operating System) and programs such as operation control applications used by the main calculation unit 17. The processor constituting the main calculation unit 17 controls each component and controls the operation of the information terminal 1 as a whole by executing programs such as the OS and operation control applications stored in a storage device, specifically, in the memory unit 18.
[0017] The memory unit 18 also stores various information data such as data of output audio signals generated by the information terminal 1, data of three-dimensional images in virtual reality (VR) and augmented reality (AR), and data of stereophonic content. In other words, the memory unit 18 also serves as a working RAM for the main processing unit 17.
[0018] The storage unit 19 is another type of storage device and is configured, for example, by a non-volatile storage medium such as a flash ROM (also referred to as FROM). The storage unit 19 stores a basic operation program 21, a stereophonic signal generation program 22, and data 23. The basic operation program 21 is an OS or the like, and the stereophonic signal generation program 22 is a program for generating a stereophonic signal and outputting the generated stereophonic signal to the audio output unit 15. The basic operation program 21 and the stereophonic signal generation program 22 are loaded into the memory unit 18 and executed by the main processing unit 17.
[0019] The data 23 includes various data required for executing the basic operation program 21 and the stereophonic signal generation program 22. The various data include acoustic space setting data 24, head-related transfer function data 25 used to generate the stereophonic signal, voice synthesis data 26 used for voice synthesis, and sound source data 27. The voice synthesis data 26 is, for example, data used when converting text data into voice data. The sound source data 27 is, for example, data for generating various sound effects, etc.
[0020] The camera unit 10, which includes a camera, captures images of the user of the information terminal 1 and the scenery in front of the information terminal 1. The camera unit 10 converts light input from a lens into an electrical signal using an electronic device such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) sensor. Then, the camera unit 10 generates visible image data of the surroundings and objects based on the electrical signal.
[0021] Furthermore, a TOF (Time Of Flight) sensor capable of acquiring the distance to an object being imaged as a distance image may be provided instead of or in combination with the camera unit 10. By using these visible images and distance images, it is possible to accurately detect pointing actions, etc., when making a selection in multiple virtual reality spaces displayed on the information terminal 1 such as an HMD. A pointing action is an action in which a user holds their hand over an image of the virtual reality space and pinches it with their thumb and index finger.
[0022] The sensor unit 11 is a group of sensors such as an illuminance sensor, a proximity sensor, or a distance sensor, and assists in the use of the information terminal 1. For example, by scanning the surroundings with the camera unit 10 and using the outputs of the sensor unit 11 and the camera unit 10, a three-dimensional space map can be generated. The sensor unit 11 may also include motion sensors such as an acceleration sensor, a gyro sensor, and a geomagnetic sensor.
[0023] The acceleration sensor detects acceleration, which is the change in speed per second, and detects movement, vibration, shock, etc. The gyro sensor detects angular velocity in the direction of rotation, and detects vertical, horizontal, and diagonal orientations. The geomagnetic sensor detects the Earth's magnetic force and detects the direction in which the information terminal 1 is facing. Therefore, by using the acceleration sensor, gyro sensor, or geomagnetic sensor, for example, an information terminal 1 such as an HMD can detect the movement of the head of a user wearing the HMD.
[0024] These sensors 11 make it possible to detect the position, tilt, direction, movement, etc. of the information terminal 1. Specifically, for example, if the information terminal 1 is a wearable display such as an HMD, a gyro sensor or a geomagnetic sensor can be used, and an acceleration sensor can also be used as needed, to detect the movement of the user's head. In particular, if a three-axis geomagnetic sensor that detects geomagnetism in the up and down directions in addition to the front and back and left and right directions is used, head movement can be detected with higher accuracy by detecting changes in geomagnetic field in response to head movement. Furthermore, if a device paired with these sensors is worn on the hand or arm, hand or arm movement can also be detected.
[0025] The image display unit 12 displays objects such as icons and menus, open documents, etc. The image display unit 12 may also display virtual reality (VR) information, real space information captured by the camera unit 10, etc. The image display unit 12 is configured with various displays such as an organic EL display or a liquid crystal display, and can display real space information, virtual reality (VR) information, augmented reality (AR) information, etc., as images.
[0026] Furthermore, in an information terminal 1 such as an HMD, the image display unit 12 may display, for example, notification information presented to the user through a menu screen or the like, or the operating status of the HMD, etc. Furthermore, the image display unit 12 may display information to notify the user when, for example, sound emission of an output audio signal is started, interrupted, or resumed. This allows the user to recognize that, for example, when sound emission is interrupted or resumed, the interruption or resumption is due to a normal control operation and not a malfunction.
[0027] The operation input unit 13 is an operation input interface that inputs signals based on user operations. In an information terminal 1 such as a PC, the operation input unit 13 is configured with a keyboard, a mouse, etc., and in an information terminal 1 such as a smartphone, the operation input unit 13 is configured with a touch panel or the like having a touch sensor incorporated into a flat display that is the image display unit 12. In an information terminal 1 such as a wearable display, the operation input unit 13 is realized by a virtual keyboard, virtual buttons, etc. that are displayed as virtual objects.
[0028] The audio input unit 14 is a microphone or the like, and collects the user's speech. For example, the audio input unit 14 collects the user's speech when the user is in a web conference, on a telephone or online call, or talking with others in real space. The audio output unit 15 is connected to a sound emitting device and includes an audio output interface that interfaces with the sound emitting device.
[0029] The sound emitting device is, for example, a speaker or headphones. The sound emitting device is provided inside or outside the audio output unit 15. For example, a sound emitting device such as a speaker is provided inside the audio output unit 15. On the other hand, if the sound emitting device is headphones or the like, the audio output unit 15 is provided with an audio output terminal that can be connected to external headphones via an audio output interface.
[0030] Here, when the speech sound to be output is sound using a stereophonic signal, the sound emitting device is preferably headphones. The headphones convert left and right output sound signals generated and input inside the information terminal 1 into left and right output sounds, respectively, and emit the sounds toward the user. When a user listens to sound through headphones, they may hear air-conducted sound, which enters the ear and is transmitted by air vibrations, or they may hear bone-conducted sound, which is transmitted by bone vibrations without passing through the ears. The headphones may be air-conducted or bone-conducted, in other words, bone-conducted.
[0031] The headphones may be semi-closed headphones that are worn in contact with the surface of the ear, or open headphones that do not completely block the ear, i.e., open-ear headphones. When using open-ear headphones, the user can hear ambient sounds that come through the headphones in addition to the original sound. In particular, among open-ear headphones, bone conduction headphones that do not block the ear at all are often used. Note that the audio output unit 15 is not limited to headphones, and may be earphones or the like.
[0032] The communication unit 15 includes various communication interfaces that handle processing based on multiple communication protocols, such as LAN, mobile communication, and near-field communication. The communication unit 16 selects a communication protocol suitable for the usage environment, etc., from the multiple communication protocols. For example, when the audio output unit 15 outputs a stereophonic signal to a sound emitting device such as headphones via wireless communication, near-field communication by the communication unit 15 is used. A representative communication protocol that handles near-field communication is Bluetooth (registered trademark).
[0033] However, the communication protocol for near field communication is not limited to this, and may be any protocol that enables short-range wireless communication, such as electronic tags, IrDA (Infrared Data Association), Zigbee (registered trademark), HomeRF (Home Radio Frequency, registered trademark), or wireless LAN (IEEE802.11a, IEEE802.11b, IEEE802.11g).
[0034] In addition, the communication unit 16 transmits and receives call data via wireless communication with a base station of a mobile wireless communication network (not shown), and also transmits and receives various data via a core network. For communication with base stations, etc., W-CDMA (Wideband Code Division Multiple Access, registered trademark) system, GSM (Global System for Mobile communications, registered trademark) system, LTE (Long Term Evolution) system, 5G (5th Generation) system, etc. are used. However, of course, other communication systems may also be used.
[0035] 1 includes components that are not essential to the embodiment. The information terminal 1 may not be provided with these components, or may further include other components depending on the application.
[0036] <Regarding Acoustic Space> Fig. 2 is a schematic diagram illustrating an example of an acoustic space set by the information terminal shown in Fig. 1. Fig. 2 shows the information terminal 1, an acoustic space 30 set by the information terminal 1, a user 31 of the information terminal 1, headphones 32 worn by the user 31, and a terminal screen 34 of the information terminal 1. The headphones 32 may be earphones. The acoustic space 30 is a space in which the terminal screen 34 is mapped, and has sound source positions that correspond to display positions on the terminal screen 34.
[0037] As a specific example, here, a character 35 is displayed at a certain display position on the terminal screen 34. In this case, a sound representing the character 35 is emitted from a sound source position associated with the display position in the acoustic space 30. In detail, the information terminal 1 generates a stereophonic signal 36 in which a sound image is localized at the sound source position in the acoustic space 30, and outputs the signal to the headphones 32. When the user 31 listens to the stereophonic signal 36 through the headphones 32, the sound sounds as if the sound representing the character 35 is being emitted from the sound source position in the acoustic space 30.
[0038] To set such an acoustic space 30, the information terminal 1, for example, photographs the user 31 with the camera unit 10 shown in Fig. 1 and detects the orientation of the user 31's face, and therefore the forward direction of the user's face. The information terminal 1 then sets a rectangular space directly in front of the user 31's face as the acoustic space 30. The setting contents of the acoustic space 30 are stored in the storage unit 19 as acoustic space setting data 24 shown in Fig. 1. The stored acoustic space setting data 24 can be reused by reading it out.
[0039] More specifically, the acoustic space 30 is defined by the three-dimensional relative coordinates of four vertices [1] to [4], with the center of the head of the user 31, more specifically the midpoint of the headphones 32, or the like, as the reference position 33, in other words the origin. The relative coordinates of the four vertices [1] to [4] are uniquely determined once the direction in which the acoustic space 30 is set relative to the reference position 33, the distance between the reference position 33 and the acoustic space 30, and the vertical and horizontal sizes of the acoustic space 30 are determined. The direction in which the acoustic space 30 is set is determined, for example, by the camera unit 10, as described above. On the other hand, the distance between the reference position 33 and the acoustic space 30 and the vertical and horizontal sizes of the acoustic space 30 may be arbitrarily set by the user 31, for example.
[0040] The acoustic space 30 thus set is fixed in three-dimensional space. Therefore, the positional relationship between the face of the user 31 and the acoustic space 30 remains fixed unless the user 31 changes the direction of their face. Furthermore, by sequentially detecting the direction of the user's face, it is also possible to make the acoustic space 30 follow the direction of the user's face. Methods for detecting the direction of the face include a method using the camera unit 10 described above, and, in the case of an information terminal 1 such as an HMD, a method using a motion sensor. Furthermore, even if the user 31 shakes their hand while holding the information terminal 1 and the information terminal 1 moves, the acoustic space 30 does not follow the movement of the information terminal 1.
[0041] 2, when using an information terminal 1 such as a smartphone, if the size of the acoustic space 30 is set to be equal to the size of the terminal screen 34, the user 31 may not be able to sufficiently distinguish the position of the sound source of the stereophonic signal resulting from sound image localization. Therefore, in such a case, it is desirable to set the size of the acoustic space 30 larger than the size of the terminal screen 34, as shown in FIG.
[0042] FIG. 3 is a schematic diagram illustrating a more detailed example of the terminal screen and acoustic space shown in FIG. 2, including the correspondence between them. Here, for convenience in explaining the correspondence, the acoustic space 30 and the terminal screen 34 are illustrated at the same size. In the acoustic space 30 shown in FIG. 3, boundary sounds 40-43 indicating the boundaries of the acoustic space 30 are emitted from the four corners of the acoustic space 30, i.e., vertices [1]-[4] shown in FIG. 2, respectively. In other words, the boundary sounds 40-43 are stereophonic signals with sound images localized at the four corners of the acoustic space 30.
[0043] The information terminal 1 generates such boundary sounds 40-43, i.e., stereophonic signals, as necessary, and allows the user 31 to listen to and hear them. This allows the user 31 to recognize the boundaries of the acoustic space 30, i.e., their position and size, and ultimately the outer frame of the entire terminal screen 34. The timbre of the boundary sounds 40-43 is selected from the sound source data 27 shown in FIG. 1 , and is selected to be easy to hear and distinguishable from the synthesized voice of the text, such as a musical instrument sound, a chime sound, or a monotone sound. Note that the sound source position at which the boundary sound is localized as a sound image may be any position at which the boundary of the acoustic space 30 can be recognized, and is not necessarily limited to the four corners of the acoustic space 30.
[0044] The terminal screen 34 shown in FIG. 3 displays various objects such as a folder icon 44, a document icon 45, a trash can icon 46, and a menu 47. For example, assume that the user 31 moves the cursor to the document icon 45 via the operation input unit 13 shown in FIG. 1. In this case, the information terminal 1 generates and outputs a synthesized voice representing the name of the document icon 45, "Document Name," as a stereophonic signal 48 in the acoustic space 30. The user 31 can recognize the display position of the document icon 45 in the acoustic space 30, and therefore in the terminal screen 34, based on the relative positional relationship between the sound source positions of the stereophonic signal 48 and the sound source positions of the boundary sounds 40-43.
[0045] Furthermore, the information terminal 1 may automatically rotate the cursor from folder icon 44 -> document icon 45 -> trash can icon 46 -> menu 47 on the terminal screen 34, while sequentially generating and outputting stereophonic signals corresponding to all icons and menus on the terminal screen 34. This allows the information terminal 1 to allow the user 31 to recognize the layout of the terminal screen 34 all at once.
[0046] 2 has been described as an example in which the acoustic space 30 is set by detecting the direction of the face of the user 31. On the other hand, in preparation for a case in which the direction of the face cannot be detected, the direction in which the acoustic space 30 is set may also be set arbitrarily by the user 31, in addition to the distance between the reference position 33 and the acoustic space 30 and the vertical and horizontal sizes of the acoustic space 30, as described in FIG. 2. In this case, the user 31 can adjust the direction in which the acoustic space 30 is set while, for example, having the information terminal 1 sequentially output the boundary sounds 40-42.
[0047] <Regarding the stereophonic signal generation program> Figure 4 is a flowchart showing an example of the processing content based on the stereophonic signal generation program in the information terminal shown in Figure 1. That is, this flowchart is realized by the main processing unit 17, which is a processor, executing the stereophonic signal generation program 22 loaded in the memory unit 18. Here, an example is taken of generating and outputting a stereophonic signal representing information about an object such as an icon or menu, as described in Figure 3.
[0048] 4, the main processing unit 17 starts processing (step S10) and first sets up an acoustic space 30 based on the position of the user 31 (step S11).The main processing unit 17 also selects boundary sounds 40-43 positioned at the four corners of the set acoustic space 30 from the sound source data 27 and outputs them as a stereophonic signal (step S12).
[0049] Specifically, in step S11, the main calculation unit 17 sets the acoustic space 30 by determining the three-dimensional relative coordinates of the four corners consisting of vertices [1] to [4], with the reference position 33 as the origin, as described in Fig. 2. At this time, for example, the size of the acoustic space 30 and the distance between the reference position 33 and the acoustic space 30 may be set according to the preferences of the user 31. Furthermore, the boundary sounds 40-43 in step S12 may also be selectable according to the preferences of the user 31.
[0050] 1. In step S11, the main calculation unit 17 sets the acoustic space 30 based on the acoustic space setting data 24. Similarly, in step S12, the main calculation unit 17 selects the boundary sounds 40-43 based on the acoustic space setting data 24.
[0051] Furthermore, when outputting the stereophonic signals in step S12, the main calculation unit 17 first determines head-related transfer functions corresponding to the relative coordinates of the four corners of the acoustic space 30, based on the head-related transfer function data 25. The main calculation unit 17 then converts the audio data of the boundary sounds 40-43 into stereophonic signals using the determined head-related transfer functions. As a result, the main calculation unit 17 can generate boundary sounds 40-43 with sound images localized at the four corners of the acoustic space 30, i.e., stereophonic signals, and output the generated stereophonic signals via the audio output unit 15.
[0052] After step S12, the main processing unit 17 repeatedly executes steps S13 to S19. In this repeated process, the main processing unit 17 waits for the user 31 to perform a predetermined cursor operation, causing the cursor to move onto an object displayed on the terminal screen 34, i.e., onto an icon or menu (step S13).
[0053] If a cursor operation is performed (step S13: YES), the main processing unit 17 identifies the icon or menu selected by the user 31 (step S14) and identifies the display position of the icon or menu on the terminal screen 34 (step S15).The main processing unit 17 also performs voice synthesis on information representing the icon or menu, such as text data such as a name, using the voice synthesis data 26 (step S16).
[0054] Next, the main processing unit 17 outputs the boundary sounds 40-43 selected in step S12 as a stereophonic signal (step S17). Following this, the main processing unit 17 outputs the synthetic sound generated in step S16 as a stereophonic signal (step S18). In step S18, the main processing unit 17 determines a sound source position in the acoustic space 30 corresponding to the display position on the terminal screen 34 identified in step S15, and determines a head-related transfer function based on the sound source position. The main processing unit 17 then converts the synthetic sound data using the determined head-related transfer function to generate a stereophonic signal with a sound image localized at the sound source position in the acoustic space 30, and outputs the signal via the audio output unit 15.
[0055] Thereafter, the main calculation unit 17 determines whether the process has ended based on, for example, a request from the user 31 (step S19), and if the process has not ended (if NO), the process returns to step S13, and if not (if YES), the process ends (step S20). Note that the flow shown in Fig. 4 may be a flow in which the setting of the acoustic space 30 and the selection of the boundary sound (steps S11 and S12) are performed at any timing in response to a request from the user 31.
[0056] 4, in response to cursor operation by user 31, boundary sounds 40-43 and synthesized sounds representing information about the object selected by the cursor, i.e., information about an icon, menu, etc., are output in sequence as stereophonic signals (steps S17 and S18). This allows user 31 to easily recognize the layout of icons and menus within acoustic space 30 and, ultimately, on terminal screen 34.
[0057] However, the output timing of the boundary sound may be changed as appropriate based on, for example, the settings of the user 31. As an example, the boundary sound may be output periodically at a predetermined time interval, or may be output each time the synthetic voice is output multiple times. Alternatively, the boundary sound may be output constantly at a low volume. Furthermore, the boundary sounds 40-43 at the four corners of the acoustic space 30 may be output simultaneously at all four corners, or may be output one by one in sequence, for example, in a clockwise or counterclockwise direction. Furthermore, more boundary sounds may be output than just at the four corners. For example, the number of boundary sounds may be eight in total by adding a sound at a position midway between each of the four corners.
[0058] <Major Effects of the First Embodiment> As described above, the information terminal and audio output method according to the first embodiment use stereophonic signals to set an acoustic space corresponding to the terminal screen, associate display positions on the terminal screen with sound source positions within the acoustic space, and allow the user to recognize the boundaries of the acoustic space using boundary sounds. This allows the user to recognize the layout on the terminal screen, making it easier to operate icons, menus, etc. As a result, user accessibility can be improved.
[0059] Second Embodiment <Regarding Document Operation> Fig. 5 is a schematic diagram illustrating an example of the processing content when a user operates a document in an information terminal according to a second embodiment. The information terminal according to the second embodiment is realized, for example, by the configuration shown in Fig. 1. Fig. 5 shows the correspondence between the acoustic space 30 and the terminal screen 34, as in Fig. 3. However, unlike Fig. 3, the terminal screen 34 shown in Fig. 5 displays a document 49 opened by the user 31.
[0060] Document 49 contains a sentence, and the cursor is placed at the beginning of the sentence "Regarding the budget proposal." In response to this, information terminal 1 performs speech synthesis on the text data of the sentence "Regarding the budget proposal," and reads the synthesized speech as a stereophonic signal 50 in acoustic space 30. Specifically, for example, stereophonic signal 50 with a sound image localized at the beginning of the sentence is output. Alternatively, stereophonic signal 50 with a sound image localized at the center of the sentence is output. Alternatively, instead of outputting stereophonic signal 50 with a sound image localized, stereophonic signal 50 with a sound image localized at each position of the characters in the sentence may be output. In this case, if the sentence is not folded within the layout, stereophonic signal 50 may be output beyond the layout.
[0061] Fig. 6 is a schematic diagram illustrating an example of processing content different from that of Fig. 5. In Fig. 6, a cursor is placed to select the words "Case B" in the same document 49 as in Fig. 5. In response to this, the information terminal 1 performs voice synthesis on the words "Case B" and reads the synthesized voice as a stereophonic signal 51 in the acoustic space 30. Specifically, for example, the stereophonic signal 51 is output with a sound image localized at the beginning or center of the word.
[0062] The difference between Figures 5 and 6 is whether a sentence or a word is read aloud. The user 31 can select from a variety of text reading methods, including reading the entire document, reading a paragraph, reading a sentence as shown in Figure 5, reading a word as shown in Figure 6, and even reading a single character. In this case, the user 31 can appropriately select the reading method by, for example, assigning each reading method to a key combination or typing pattern. This allows the user 31 to gradually narrow the reading range, for example, from the entire document to a paragraph, a sentence, a word, and a character, until they reach the part they want to edit. The type of synthetic voice may also be associated with the text reading method. For example, when reading a single word or a single character, it is recommended to associate a synthetic voice with a higher frequency to make it easier to understand the location of the sound image.
[0063] <Regarding the Stereophonic Signal Generation Program> Figure 7 is a flowchart showing an example of the processing content based on the stereophonic signal generation program shown in Figure 1 in an information terminal according to the second embodiment. The flow shown in Figure 7 differs from the flow shown in Figure 4 in that the readout target has been changed from icons and menu information to documents. In Figure 7, the processing of steps S10-S12 and steps S17-S20 is the same as in Figure 4. In Figure 7, after step S12, the processing of steps S21-S25 is executed instead of steps S13-S16 in Figure 4.
[0064] In steps S21-S25, the main processing unit 17 first waits for the user 31 to perform a predetermined cursor operation, thereby moving the cursor onto a sentence in the document 49 (step S21). If a cursor operation is performed (step S21: YES), the main processing unit 17 recognizes the reading unit based on the cursor operation (step S22) and identifies the reading unit based on the cursor operation (step S23).
[0065] Next, the main processing unit 17 identifies the display position of the identified reading portion on the terminal screen 34 (step S24). Furthermore, the main processing unit 17 performs voice synthesis on the text data of the identified reading portion using the voice synthesis data 26 (step S25). Thereafter, as in the case of Fig. 4, the main processing unit 17 outputs the boundary sound as a stereophonic signal (step S17) and outputs the synthesized voice generated in step S25 as a stereophonic signal (step S18).
[0066] In this flow, the user 31, for example, operates the cursor as appropriate to narrow down the editing points in the document 49. In response to this, the processes of steps S21 and S22 are repeatedly executed, allowing the user 31 to narrow down the editing points in the document 49. Furthermore, the execution timing of steps S11 and S12 and the execution timing of step S17 can be changed as appropriate, as described with reference to FIG.
[0067] <Major Effects of the Second Embodiment> As described above, the information terminal and audio output method according to the second embodiment can also achieve the same effects as those described in the first embodiment. Furthermore, by using the method according to the second embodiment, the user can recognize the layout of the document on the terminal screen, which makes it easier to perform document editing, for example. As a result, accessibility for users can be improved.
[0068] (Third embodiment) <Regarding acoustic space> Fig. 8 is a schematic diagram illustrating an example of an acoustic space different from that shown in Fig. 2, which is set by the information terminal shown in Fig. 1 in an information terminal according to a third embodiment. Fig. 9 is a schematic diagram illustrating an example of an acoustic space different from that shown in Fig. 8. The acoustic space in Figs. 8 and 9 differs in position and shape from that shown in Fig. 2. That is, the position of the acoustic space is not limited to a position directly facing the user 31 as shown in Fig. 2, and the shape of the acoustic space is not limited to a rectangular shape as shown in Fig. 2.
[0069] In Fig. 8, an acoustic space 30 is set above the head of a user 31. More specifically, the acoustic space 30 is set above the head of the user 31 so as to be parallel to the ground. For example, if there is environmental noise in front of the user 31, the position of the acoustic space 30 can be changed as appropriate, as shown in Fig. 8. This allows the user 31 to avoid the environmental noise and easily hear the stereophonic signal.
[0070] 9, a hemispherical space surrounding the user 31 is assumed, and the acoustic space 30 is formed by a portion of this hemispherical space. For example, in a rectangular acoustic space such as that shown in FIG. 2, the closer one gets to the boundary of the acoustic space, the farther the distance from the user 31 becomes. Therefore, the closer the user 31 gets to the boundary of the acoustic space, the more difficult it becomes for the user 31 to distinguish between different sound source positions. On the other hand, in the acoustic space 30 shown in FIG. 9, the distance from the user 31, i.e., the radius of the sphere, is the same at all points. Therefore, the user 31 can easily distinguish between different sound source positions throughout the entire acoustic space 30.
[0071] <Regarding the stereophonic signal generation program> Fig. 10 is a flowchart showing an example of the processing performed by an information terminal according to the third embodiment based on the stereophonic signal generation program shown in Fig. 1. The flow shown in Fig. 10 adds processing for resetting the acoustic space (steps S30 and S31) to the flow shown in Fig. 7. In addition, the return destination in step S19a has been changed accordingly.
[0072] 10 is a flow that allows the settings of the acoustic space to be changed not only at the time of initial setting (steps S11 and S12) but also at any timing desired by the user 31. Note that the processing associated with the resetting of the acoustic space (steps S30 and S31) is not limited to the flow shown in FIG. 7 but may be added to the flow shown in FIG. 4.
[0073] 10, the main processing unit 17 selects a boundary sound and outputs it as a stereophonic signal (step S12), and then detects the physical movement of the information terminal 1 using, for example, a motion sensor in the sensor unit 11 (step S30). That is, in this example, the acoustic space 30 can be reconfigured based on the physical movement of the information terminal 1, independently of the movement of the cursor on the terminal screen 34.
[0074] For example, the left and right movements, or the up and down movements of the information terminal 1, can be paired and assigned to expand and contract the acoustic space 30. Also, for example, the physical movement of the information terminal 1 after pressing a specific key by cursor operation can be assigned to the translation of the acoustic space 30. Alternatively, the acoustic space 30 may be configured to expand and contract in accordance with changes in the distance between the user 31 and the information terminal 1. With these operation examples as representative, the method by which the user 31 reconfigures the acoustic space 30 can be modified as appropriate.
[0075] If the main processing unit 17 does not detect a physical movement of the information terminal 1 (step S30: NO), it waits for a cursor operation by the user 31 (step S21), as in the case of Fig. 7. On the other hand, if the main processing unit 17 detects a physical movement of the information terminal 1 (step S30: YES), it resets the acoustic space 30 based on the movement and proceeds to step S12 (step S31). Furthermore, as in the case of step S19 in Fig. 7, the main processing unit 17 determines whether the process has ended (step S19a). If the process has not ended, it returns to step S30 and detects a physical movement of the information terminal 1.
[0076] <Major Effects of the Third Embodiment> As described above, the use of the information terminal and audio output method according to the third embodiment also provides the same effects as those described in the first and second embodiments. Furthermore, for example, an acoustic space suited to the surrounding environment can be set, thereby improving accessibility for users.
[0077] (Fourth embodiment) <Regarding document operation> Fig. 11 is a schematic diagram illustrating an example of the processing content when a user operates a document in an information terminal according to a fourth embodiment. The information terminal according to the fourth embodiment is realized, for example, by the configuration shown in Fig. 1. Fig. 11 shows the correspondence between the acoustic space 30 and the document 49 on the terminal screen 34, similar to Fig. 5. However, unlike Fig. 5, character attributes are assigned to the text in the document 49 in Fig. 11.
[0078] In Fig. 11, the sentence "Regarding the budget proposal..." is modified with bold font. When reading this sentence, information terminal 1 outputs character attribute sound 61 representing the bold font attribute as a stereophonic signal. More specifically, as in the case of Fig. 5, information terminal 1 outputs stereophonic signal 60 of synthesized speech for the text data of this sentence, and also outputs character attribute sound 61 as a stereophonic signal.
[0079] The character attribute sound 61 is selected from the sound source data 27 in Fig. 1, and in the example shown in Fig. 11, an animal cry is selected. The character attribute sound 61 may be output simultaneously with, for example, a stereophonic signal 60 of synthetic voice. However, the character attribute sound 61 is localized as a sound image outside the acoustic space 30, for example. This allows the user 31 to easily distinguish between the character attribute sound 61 and the stereophonic signal 60 of synthetic voice.
[0080] Fig. 12 is a schematic diagram illustrating an example of processing content different from that of Fig. 11. In Fig. 12, the sentence "Regarding the impact on our company..." is underlined. When reading this sentence, information terminal 1 outputs a stereophonic signal 62 of synthesized voice, as in the case of Fig. 11, and also outputs a character attribute sound 63 representing the underline attribute as a stereophonic signal.
[0081] The character attribute sound 63 is selected from the sound source data 27 in Fig. 1, and in the example shown in Fig. 12, the sound of waves is selected. As in the case of Fig. 11, the character attribute sound 63 is localized as a sound image outside the acoustic space 30, for example. Furthermore, the character attribute sound 63 representing the underline attribute may be localized as a sound image at a different position from the character attribute sound 61 representing the bold attribute. This allows the user 31 to recognize the difference in character attributes not only from the type of sound but also from the difference in position.
[0082] <Regarding the stereophonic signal generation program> Fig. 13 is a flowchart showing an example of the processing performed by an information terminal according to the fourth embodiment based on the stereophonic signal generation program shown in Fig. 1. The flow shown in Fig. 13 adds processing associated with character attribute sounds (steps S41 and S42) to the flow shown in Fig. 7. In addition, the processing of step S18 in Fig. 7 has been replaced with the processing of step S18a in Fig. 13.
[0083] In Fig. 13, the main processing unit 17 identifies a reading section (step S23) and then determines the character attribute of the identified reading section (step S41). Based on the result of this determination, the main processing unit 17 selects a character attribute sound corresponding to the character attribute (step S42). Thereafter, the processing from step S24 onward is performed, as in the case of Fig. 7. However, unlike the case of Fig. 7, the main processing unit 17 outputs a stereophonic signal of the synthesized speech in step S18a, and also outputs the character attribute sound as a stereophonic signal.
[0084] Here, the character attribute sound is output, for example, simultaneously with the synthesized voice. However, the timing of outputting the character attribute sound and the duration of the output may be determined as appropriate. For example, the main processing unit 17 may output the character attribute sound and then output the synthesized voice immediately after stopping the output. Alternatively, the main processing unit 17 may output the character attribute sound for a short time simultaneously with the synthesized voice, or may output the character attribute sound continuously throughout the output period of the synthesized voice.
[0085] <Major Effects of the Fourth Embodiment> As described above, the use of the information terminal and audio output method according to the fourth embodiment also provides the same effects as those described in the second and third embodiments. Furthermore, by outputting a stereophonic signal representing the character attributes in a document separately from the synthesized audio, the user can easily recognize the character attributes. As a result, user accessibility can be improved.
[0086] Fifth Embodiment <Application Example to a Head-Mounted Display (HMD)> Fig. 14 is a diagram showing an application example to a head-mounted display (HMD) of an information terminal according to a fifth embodiment, and is a schematic diagram showing an example of the relationship between the display space and acoustic space of the HMD. As in Fig. 2 and other cases, Fig. 14 shows a user 31 of the information terminal, headphones 32 worn by the user 31, and an acoustic space 30 set by the information terminal. As in Fig. 3 and other cases, boundary sounds 40-43 are emitted from the four corners of the acoustic space 30.
[0087] 14, the information terminal is an HMD 73 worn and used by a user 31. The HMD 73 generates a display space 74, which is a virtual space, and displays virtual objects 75-77 within the display space 74. The HMD 73 and the headphones 32 may be separate devices as shown in the figure, or may be integrated.
[0088] Now, let us consider a case where the user 31 selects a virtual object 77 by performing a predetermined operation on the HMD 73. In this case, the HMD 73 outputs to the headphones 32 a stereophonic signal 78 in which a sound image is localized at a sound source position corresponding to the display position of the virtual object 77, generating a synthesized voice representing information about the virtual object 77, such as its name. In this case, the display space 74 is usually set to a large space to achieve a sense of realism. Therefore, unlike the case of FIG. 2, the HMD 73 can set the acoustic space 30 to match the display space 74, as shown in FIG. 14 .
[0089] As a result, for the same virtual object, the display position in the display space 74 and the sound source position in the acoustic space 30 coincide with each other. As a result, the user 31 can easily recognize the layout of the display space 74 through the acoustic space 30. More specifically, the display space 74 is a three-dimensional space, and the virtual objects 75-77 are three-dimensional images having depth coordinates. The user 31 can recognize the layout of the display space 74, i.e., the three-dimensional space, including the depth coordinates of the virtual objects 75-77, through the acoustic space 30.
[0090] Fig. 15 is a schematic diagram showing an example of the relationship between the display space and acoustic space of the HMD shown in Fig. 14 and the relationship with the camera image. In Fig. 15, a camera capture range 79 and a physical object 80 captured within the camera capture range 79 are added to the display space 74 shown in Fig. 14. That is, the display space 74 is, for example, an augmented reality (AR) space. In this example, text is written on the physical object 80. The HMD 73 outputs information about the physical object 80, in this example, synthesized speech representing the text written on the physical object 80, as a stereophonic signal 81.
[0091] Specifically, the camera is built into the HMD 73 and captures an image in front of the HMD 73. The camera's capture range 79 may be narrower than the display space 74. If the HMD 73 is a video see-through type, the real object 80 is captured by the camera and displayed within the camera's capture range 79 in the display space 74. On the other hand, if the HMD 73 is an optical see-through type, the real object 80 passes through the optical system of the HMD 73 and is viewed directly by the user 31. At this time, the outer frame of the camera's capture range 79 may be displayed as a virtual object.
[0092] If text is written on the physical object 80 photographed by the camera, the HMD 73 performs character recognition on the text and converts it into text data. The HMD 73 then generates synthetic speech from the recognized text data and outputs the generated synthetic speech as a stereophonic signal 81 in which a sound image is localized at the position of the physical object 80. In this case, the HMD 73 may identify the type of the physical object 80 photographed by the camera by object recognition, and generate synthetic speech using the identified information as information on the physical object 80, and output the synthetic speech as a stereophonic signal.
[0093] As described above, when the information terminal is the HMD 73, unlike when the information terminal is a smartphone or the like, it is beneficial to match the acoustic space 30 with the display space 74. This allows the user 31 to perceive the display space 74 through the acoustic space 30 with the same sense of scale as the visual sense of scale. As a result, the user 31 can perceive the actual positions of the real objects displayed in the display space 74, in addition to the positions of the virtual objects displayed in the display space 74, through the acoustic space 30.
[0094] <Configuration of Head-Mounted Display (HMD)> Fig. 16 is a block diagram showing a configuration example of an information terminal according to the fifth embodiment when the information terminal is a head-mounted display (HMD). Compared to the configuration example shown in Fig. 1, the HMD 73 shown in Fig. 16 additionally includes a distance measurement unit 90 and virtual object data 92, which is one of the data 23, and furthermore, the image display unit 12 is replaced with a three-dimensional image display unit 91 as a display.
[0095] The distance measuring unit 90 measures the distance between the physical object photographed by the camera unit 10 and the distance measuring unit 90, and links the measurement result to the physical object 80 within the camera photographing range 79. The distance measuring unit 90 is, for example, a distance measuring sensor such as a TOF sensor using infrared or millimeter wave radar. For example, by scanning the surroundings of the user 31 with the distance measuring unit 90 and combining the measurement result by the distance measuring unit 90 with the image captured by the camera unit 10, a three-dimensional space map can be generated.
[0096] The three-dimensional image display unit 91 displays a left-eye image and a right-eye image, which are visually recognized by the left and right eyes of the user 31, thereby displaying a three-dimensional image. Data for the three-dimensional image is generated by a three-dimensional video generation processing unit (not shown). At this time, data for a virtual object, which is one of the three-dimensional images, is acquired from virtual object data 92.
[0097] Here, as described above, distance data from the distance measuring unit 90 is assigned to the real object photographed by the camera unit 10. Therefore, the HMD 73 can grasp not only the display position of the virtual object in the display space 74, but also the position of the real object. Therefore, the HMD 73 can localize the sound image of the synthesized sound representing the information of the virtual object and the synthesized sound representing the information of the real object in a three-dimensional acoustic space including depth, and can output each synthesized sound as a stereophonic signal.
[0098] 17A and 17B are flowcharts showing an example of processing based on the stereophonic signal generation program shown in Fig. 16 in an information terminal according to the fifth embodiment. In Fig. 17A, the main processing unit 17 starts processing (step S10) and first sets the acoustic space 30 so that it matches the display space 74 (step S50). Next, the main processing unit 17 selects a boundary sound and outputs the boundary sound as a stereophonic signal (step S12), as in the case of Fig. 4, etc.
[0099] Next, the main processing unit 17 uses the 3D image display unit 91 to display a 3D virtual object in the display space 74 (step S51). The main processing unit 17 also uses the camera unit 10 to capture an image in front of the HMD 73 (step S52). Next, in FIG. 17B , the main processing unit 17 waits for a cursor operation by the user 31 (step S53). In the HMD 73, cursor operation can also be performed by gestures or the like by the user 31.
[0100] If a cursor operation is performed (step S53: YES), the main processing unit 17 determines whether the cursor is pointing at a virtual object (step S54). If the cursor is pointing at a virtual object (step S54: YES), the main processing unit 17 identifies information about the virtual object, such as the type and name of the virtual object (step S55), and further identifies the display position of the virtual object in the display space 74 (step S56). Then, the main processing unit 17 performs voice synthesis on text data representing the identified information about the virtual object (step S57), and then proceeds to step S17 in FIG. 17A.
[0101] On the other hand, if the cursor is not pointing at a virtual object (step S54: NO), the main processing unit 17 determines whether the cursor is pointing inside the camera's image capture range 79 (step S58). If the cursor is pointing outside the camera's image capture range 79 (step S58: NO), the main processing unit 17 returns to step S53. If the cursor is pointing inside the camera's image capture range 79 (step S58: YES), the main processing unit 17 identifies the text portion written on the physical object 80 included in the camera's image capture range 79 and performs character recognition (step S59).
[0102] Next, the main processing unit 17 identifies the position of the text portion or the physical object 80 using the distance measurement unit 90 (step S60). Furthermore, the main processing unit 17 performs speech synthesis on the text data of the text portion recognized in step S59 (step S61), and then proceeds to step S17 in FIG. 17A. Note that in steps S59-S61, the main processing unit 17 performs character recognition on the text portion. However, in addition to or instead of this, the main processing unit 17 may perform object recognition and generate synthetic speech representing information about the recognized object. That is, the main processing unit 17 may generate synthetic speech representing information about the physical object.
[0103] 17A, the main calculation unit 17 outputs the boundary sound as a stereophonic signal (step S17), as in the case of Fig. 4 etc., and then outputs the synthetic sound generated in step S57 or step S61 as a stereophonic signal using the head-related transfer function data 25 (step S18). Thereafter, the main calculation unit 17 determines whether the process is complete, as in the case of Fig. 4 etc., and if the process is not complete, returns to step S51.
[0104] <Major Effects of the Fifth Embodiment> As described above, the use of the information terminal and audio output method according to the fifth embodiment can also achieve the same effects as those described in the first to fourth embodiments. Furthermore, by matching the acoustic space with the display space of the HMD, the user can easily recognize the layout of virtual objects and real objects in the display space. As a result, user accessibility can be improved. Furthermore, it becomes possible for the user to recognize the real space through audio.
[0105] Although the embodiments of the present invention have been described above, it goes without saying that the present invention can also be implemented in an information terminal that does not have a display unit such as the image display unit 12 in FIG. 1 or the three-dimensional image display unit 91 in FIG. 16. Furthermore, when implementing the present invention, the display unit may not display anything. This can extend the usable time of the information terminal when it is battery-powered.
[0106] An information terminal according to one embodiment includes a computer system including a processor and a memory, a storage unit, an audio input unit, an audio output unit, a screen display unit, and an operation input unit. The storage unit includes a program for processing the generation and output of a stereophonic signal, and the computer system executes the program.
[0107] The stereophonic signal generation and output process sets up an acoustic space corresponding to the device screen. The acoustic space can be set to any position, size, and shape based on the user of the device, and the position, size, and shape can be changed as needed during use by the user. The stereophonic signal generation and output process also outputs sounds that indicate the boundaries, including the four corners, of the set acoustic space as a stereophonic signal.
[0108] Furthermore, in the stereophonic signal generation and output process, when a user moves the cursor by inputting using a keyboard or voice recognition and places it on an icon or menu on the terminal screen, a synthetic voice is generated that represents the name of the icon or menu, etc. The generated synthetic voice is then localized as a sound image so that it can be heard from a position in the acoustic space corresponding to the position on the terminal screen, and is output as a stereophonic signal.
[0109] Alternatively, the stereophonic signal generation and output process may automatically select icons or menus while rotating them, and output information representing the selected icon or menu as a stereophonic signal that can be heard from a position in an acoustic space corresponding to the position on the terminal screen. The terminal user can recognize the position of the icon or menu on the terminal screen by comparing the stereophonic signal representing the icon or menu information with the stereophonic signal that notifies the user of the boundaries, including the four corners.
[0110] In addition, in the stereophonic signal generation and output process, when the user moves the cursor and places it on text in a document open on the terminal screen, the synthesized speech of that text data is output as a stereophonic signal from a position in the acoustic space corresponding to the position on the terminal screen. At this time, the entire text data may be read aloud, or only part of the text, such as the sentence, character, or word where the cursor is located, may be read aloud.
[0111] The user of the device compares the stereophonic signal of the text data with the stereophonic signal of the sounds indicating the boundaries including the four corners to recognize the position of the text data on the device screen. At this time, a process of reading out only part of the text of a sentence, character, word, etc., may be applied to allow the user to accurately identify the position.
[0112] In one embodiment, an audio output method for an information terminal sets an acoustic space corresponding to the terminal screen based on the user of the terminal, and outputs a sound indicating the boundaries including the four corners of the set acoustic space as a stereophonic signal.
[0113] In addition, synthesized voices corresponding to icons and menus on the terminal screen are output as stereophonic signals in the acoustic space, and synthesized voices corresponding to the text data of documents opened on the terminal screen are output as stereophonic signals in the acoustic space.
[0114] <Notes> [1] An information terminal that includes a computer unit consisting of a main calculation unit and memory, a storage unit, an audio output unit, a screen display unit, an input operation unit, etc., and that executes a stereophonic signal generation and output process, wherein the stereophonic signal generation and output process includes a process of setting an acoustic space based on the user of the information terminal, and a process of outputting a sound that notifies the boundaries including the four corners of the set acoustic space as a stereophonic signal.
[0115] [2] In the information terminal described in [1] above, the stereophonic signal generation and output process includes a process of obtaining synthetic speech from icons, menu names, etc., and a process of outputting the stereophonic signal of the synthetic speech to the audio output unit, and when the cursor is placed on an icon or menu on the terminal screen displayed on the screen display unit using the input operation unit, the stereophonic signal of the synthetic speech is localized so that it can be heard from a position in the acoustic space corresponding to the position of the icon or menu on the terminal screen.
[0116] [3] In the information terminal described in [1] above, the stereophonic signal generation and output process includes a process for obtaining synthetic speech of the text of a document opened on the terminal screen, a process for outputting a stereophonic signal of the synthetic speech to the audio output unit, and a process for controlling the range of the text from which the synthetic speech is obtained, and when the cursor is placed on the text of the open document, the sound image is localized so that the stereophonic signal of the synthetic speech of the range of the text determined by the process for controlling the range from which the synthetic speech is obtained is heard from a position in the acoustic space corresponding to the position of the document text on the terminal screen.
[0117] [4] In the information terminal described in [1] above, the stereophonic signal generation process includes a process of resetting the acoustic space.
[0118] [5] In the information terminal described in [3] above, the stereophonic signal generation process includes a process of obtaining the modification attributes of the text of a document opened on the terminal screen, and a process of outputting a sound corresponding to the modification attributes as a stereophonic signal, and localizing the stereophonic signal of the sound corresponding to the modification attributes outside the acoustic space.
[0119] [6] In the information terminal described in [1] above, the information terminal is a head-mounted display worn by a user, and includes a camera unit that photographs real objects in front of the information terminal, and a distance measurement unit that obtains the depth of the real objects. The screen display unit displays virtual objects, etc., in three dimensions. The stereophonic signal generation process includes a process for obtaining synthetic sounds related to the virtual objects, and a process for outputting a stereophonic signal of the synthetic sounds related to the virtual objects. When a cursor is placed on the virtual object, the stereophonic signal of the synthetic sounds related to the virtual object localizes a sound image at a position that reflects the depth information of the display space of the virtual object.
[0120] [7] In the information terminal described in [6] above, the stereophonic signal generation process includes a process of recognizing the text portion of a real object photographed by the camera unit, a process of obtaining synthetic speech of the recognized text portion, and a process of obtaining a stereophonic signal of the synthetic speech of the recognized text portion, and the stereophonic signal of the synthetic speech of the recognized text portion localizes a sound image at the position of the real object reflecting the depth information of the display space.
[0121] Although the embodiments of the present invention have been described above, it goes without saying that the configurations for realizing the technology of the present invention are not limited to the above-described embodiments, and various modifications are possible. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. All of these fall within the scope of the present invention. Furthermore, numerical values, messages, etc. appearing in the text and figures are merely examples, and the effects of the present invention will not be impaired even if different ones are used.
[0122] The programs described in each processing example may be independent programs, or multiple programs may constitute a single application program. The order in which each process is performed may also be changed.
[0123] Some or all of the functions of the present invention described above may be implemented in hardware, for example, by designing them as integrated circuits. They may also be implemented in software by a microprocessor unit, CPU, or the like interpreting and executing an operating program that implements each function. Furthermore, the scope of software implementation is not limited, and both hardware and software may be used. Some or all of the functions may also be implemented by a server.
[0124] The server may take any form as long as it can execute functions in cooperation with other components via communications, and may be, for example, a local server, a cloud server, an edge server, an internet service, etc. Information such as programs, tables, and files that implement each function may be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a non-transitory tangible computer-readable recording medium such as an IC card, an SD card, or a DVD, or may be stored in a device on a communications network.
[0125] Furthermore, the control lines and information lines shown in the diagram are those considered necessary for explanation, and do not necessarily represent all of the control lines and information lines on the product. In reality, it can be assumed that almost all components are interconnected.
[0126] 1: Information terminal, 10: Camera unit, 11: Sensor unit, 12: Image display unit, 13: Operation input unit, 14: Audio input unit, 15: Audio output unit, 16: Communication unit, 17: Main calculation unit, 18: Memory unit, 19: Storage unit, 20: Internal bus, 21: Basic program, 22: Stereophonic signal generation program, 23: Data, 24: Acoustic space setting data, 25: Head-related transfer function data, 26: Data for voice synthesis, 27: Sound source data, 30: Acoustic space, 31: User, 3 2: headphones, 33: reference position, 34: terminal screen, 36, 48, 50, 51, 60, 62, 78, 81: stereophonic signal, 40-43: boundary sound, 44-46: icon, 47: menu, 49: document, 61, 63: character attribute sound, 73: head-mounted display (HMD), 74: display space, 75-77: virtual object, 79: camera shooting range, 80: real object, 90: distance measurement unit, 91: three-dimensional image display unit, 92: virtual object data
Claims
1. An information terminal comprising: a processor; an audio output interface; a display; and an operation input interface, wherein the processor controls the audio output interface to be connected to a sound emitting device; controls the display to display a predetermined terminal screen; controls the operation input interface to input a signal based on a user's operation; sets an acoustic space corresponding to the terminal screen based on the user; generates boundary sounds as stereophonic signals to allow the user to recognize the boundaries of the set acoustic space; and controls the audio output interface to output the boundary sounds.
2. An information terminal according to claim 1, wherein the boundary sound is localized as sound images at a plurality of sound source positions including the four corners of the acoustic space.
3. An information terminal according to claim 1, wherein the size of the acoustic space is larger than the size of the terminal screen.
4. An information terminal according to claim 1, wherein the processor generates synthetic speech representing information about an object displayed on the terminal screen, and controls the processor to generate a stereophonic signal of the synthetic speech by localizing a sound image of the generated synthetic speech at a sound source position in the acoustic space corresponding to the display position of the object on the terminal screen, and to output the stereophonic signal to the audio output interface.
5. An information terminal according to claim 1, wherein the processor, based on the user's operation, identifies text data to be read out from among documents displayed on the terminal screen, generates synthetic speech that reads out the identified text data, and localizes a sound image of the generated synthetic speech at a sound source position in the acoustic space corresponding to the display position of the text data on the terminal screen, thereby generating a stereophonic signal of the synthetic speech and controlling the signal to be output to the audio output interface.
6. An information terminal according to claim 1, wherein the processor reconfigures the position or size of the acoustic space in response to a request from the user.
7. An information terminal according to claim 5, wherein the processor determines character attributes in the identified text data, generates character attribute sound representing the character attributes, localizes a sound image of the character attribute sound at a sound source position outside the acoustic space, thereby generating a stereophonic signal of the character attribute sound, and controls the generated stereophonic signal of the character attribute sound to be output to the audio output interface separately from the stereophonic signal of the synthesized voice.
8. An information terminal according to claim 1, wherein the information terminal is a head-mounted display worn on the user's head, the display displays a three-dimensional display space, and the processor sets the acoustic space to coincide with the display space, generates synthetic sound representing information about a virtual object displayed in the display space, and localizes a sound image of the generated synthetic sound at a sound source position in the acoustic space corresponding to the display position of the virtual object in the display space, thereby generating a stereophonic signal of the synthetic sound and controlling the signal to be output to the audio output interface.
9. An information terminal according to claim 8, comprising a camera for capturing images of a real object in front of the information terminal, and a distance sensor for measuring the distance between the information terminal and the real object, wherein the processor generates synthetic sound representing information about the real object displayed in the display space, identifies the position of the real object in the display space based on the measurement results of the distance sensor, and localizes a sound image of the generated synthetic sound at a sound source position in the acoustic space corresponding to the position of the real object in the display space, thereby generating a stereophonic signal of the synthetic sound and controlling the signal to be output to the audio output interface.
10. An audio output method for an information terminal that generates and outputs a stereophonic signal, comprising: displaying a predetermined terminal screen; setting an acoustic space corresponding to the terminal screen based on the user of the information terminal; and generating and outputting boundary sounds as stereophonic signals to allow the user to recognize the boundaries of the set acoustic space.
11. The audio output method for an information terminal according to claim 10, wherein the boundary sound is localized as sound images at a plurality of sound source positions including the four corners of the acoustic space.
12. The audio output method for an information terminal according to claim 10, wherein the size of the acoustic space is set to be larger than the size of the terminal screen.
13. An audio output method for an information terminal according to claim 10, comprising: generating synthetic audio representing information about an object displayed on the terminal screen; and localizing a sound image of the generated synthetic audio at a sound source position in the acoustic space corresponding to the display position of the object on the terminal screen, thereby generating and outputting a stereophonic signal of the synthetic audio.
14. A method for outputting audio from an information terminal as described in claim 10, comprising the steps of: identifying text data to be read out from a document displayed on the terminal screen based on the user's operation; generating synthetic speech that reads out the identified text data; and localizing a sound image of the generated synthetic speech at a sound source position in the acoustic space corresponding to the display position of the text data on the terminal screen, thereby generating and outputting a stereophonic signal of the synthetic speech.
15. The audio output method for an information terminal according to claim 10, wherein the position or size of the acoustic space is reset in response to a request from the user.
16. A method for outputting audio from an information terminal as defined in claim 14, comprising: determining character attributes in the identified text data; generating character attribute sound representing the character attributes; localizing a sound image of the character attribute sound at a sound source position outside the acoustic space to generate a stereophonic signal of the character attribute sound; and outputting the generated stereophonic signal of the character attribute sound separately from the stereophonic signal of the synthesized voice.
17. A method for outputting audio from an information terminal as defined in claim 10, wherein the information terminal is a head-mounted display worn on the user's head, and the method displays a three-dimensional display space as the specified terminal screen, sets the acoustic space to coincide with the display space, generates synthetic audio representing information about a virtual object displayed in the display space, and localizes a sound image of the generated synthetic audio at a sound source position in the acoustic space corresponding to the display position of the virtual object in the display space, thereby generating and outputting a stereophonic signal of the synthetic audio.
18. A sound output method for an information terminal as defined in claim 17, comprising: capturing an image of a real object in front of the information terminal; measuring the distance between the information terminal and the real object; generating synthetic sound representing information about the real object to be displayed in the display space; identifying the position of the real object in the display space based on the measurement result of the distance between the information terminal and the real object; and localizing a sound image of the generated synthetic sound at a sound source position in the acoustic space corresponding to the position of the real object in the display space, thereby generating and outputting a stereophonic signal of the synthetic sound.
Citation Information
Patent Citations
Text readout method
JP1996263260A
Information processing device and information processing method
WO2019225192A1
Acoustic reproduction method, acoustic reproduction device, and program
WO2021187335A1