Information terminal, operation input device, and method for operating information terminal

The information terminal uses motion and blink detection sensors with spatial audio output to provide hands-free input for visually impaired users, addressing the need for improved accessibility and usability in noisy environments.

WO2025210829A1PCT designated stage Publication Date: 2025-10-09MAXELL LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/013943
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing information terminals, particularly for visually impaired users, face challenges with hands-free input functions in noisy environments and the need for dedicated terminals, and there is a demand for improved accessibility features that are cost-effective and suitable for various settings.

Method used

An information terminal equipped with a motion sensor, blink detection sensor, and processor that utilizes motion and blink patterns for cursor control and application execution, combined with spatial audio output for hands-free operation.

Benefits of technology

Enables hands-free input functionality without dedicated terminals, providing accessibility for visually impaired users through eye-tracking and head movement inputs, enhancing usability in noisy environments and reducing reliance on visual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024013943_09102025_PF_FP_ABST
    Figure JP2024013943_09102025_PF_FP_ABST
Patent Text Reader

Abstract

An information terminal according to the present invention comprises: a motion sensor that detects a motion of a user and that outputs motion information; a blinking detection sensor that detects blinking of the user; a processor; and a display, wherein the processor moves, according to the motion information, a cursor displayed on a display screen of the display when the motion information is acquired, and executes processing indicated by the cursor when the blinking detection sensor detects blinking of the user in a prescribed pattern.
Need to check novelty before this filing date? Find Prior Art

Description

Information terminal, operation input device, and method for operating information terminal

[0001] The present invention relates to an information terminal, an operation input device, and an operation method for an information terminal.

[0002] Information terminals such as head-mounted displays (HMDs), smartphones, and tablet PCs are used for a wide range of purposes, including browsing web pages, creating documents, and playing videos. Furthermore, HMDs offer immersive experiences with computer-generated reality (CGR) content, such as virtual reality (VR) and mixed reality (MR). VR displays images of a virtual space generated by a computer or the like, while MR superimposes images of virtual objects onto images of real space.

[0003] Information terminals are equipped with a variety of input interfaces, such as keyboard input, touch input using a touchpad or touch-sensitive display, voice input, gesture input that recognizes finger movements, as well as eye-tracking input and head movement input, and an input interface suitable for the user and the usage environment is selected. For example, when experiencing CGR content using an HMD, it is desirable for the user to be hands-free, and the use of an eye-tracking input or head movement input interface is preferred. As an example, Patent Document 1 discloses an information terminal and input interface that utilizes eye-tracking input and head movement input.

[0004] Special Publication No. 2023-520345

[0005] Information terminals such as smartphones, PCs, and HMDs are being equipped with accessibility features designed for use by visually impaired people. For example, they have a text-to-speech function that reads web pages and documents aloud using synthesized speech. They also have spatial audio as an output interface, and offer experiential content such as games that do not rely on vision.

[0006] On the other hand, voice input and eye gaze input are well-known as hands-free input functions for information terminals, and although voice input can be the main input device used particularly by visually impaired people, there is a demand for further improvements in convenience when it comes to operation in places where the ambient noise is loud, or conversely, in places where voice input is not suitable, such as libraries and hospitals. On the other hand, from a cost perspective, it is desirable for information terminals equipped with accessibility functions to be based on information terminals that are widely used by able-bodied people, rather than using dedicated information terminals.

[0007] An object of the present invention is to provide a technique for realizing a hands-free input function that replaces voice input without using a dedicated terminal.

[0008] In order to solve the above problems, the present invention has the configurations described in the claims. As an example, the present invention is an information terminal comprising: a motion sensor that detects a user's motion and outputs motion information; a blink detection sensor that detects the user's blinks; a processor; and a display, wherein the processor, upon acquiring the motion information, moves a cursor displayed on a display screen of the display in accordance with the motion information; and, upon detecting a predetermined pattern of the user's blinks, executes a process indicated by the cursor.

[0009] According to the present invention, it is possible to provide a technology that realizes a hands-free input function that replaces voice input without using a dedicated terminal. The objects, configurations, and effects of the present invention other than those described above will be made clear in the following embodiments.

[0010] 10. A block diagram of an information terminal. An external view of an HMD. A diagram showing a state in which a user is wearing an information terminal and viewing the home screen of the information terminal superimposed on real space. A flowchart of transitioning to accessibility mode and icon selection. A diagram explaining reading a web page. A flowchart of reading a web page. A diagram explaining transition to linked content. A flowchart including transition to linked content. A diagram explaining voice input for a web page. A flowchart including voice input in a web application. A diagram showing cooperation with a wearable terminal. A flowchart of processing to reflect movement information of the wearable terminal in steps S16, S38, and S39 of confirming head movement detection in FIGS. 4, 6, 8, and 10.

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all drawings for explaining the embodiments, the same components are generally designated by the same reference numerals, and repeated description thereof will be omitted.

[0012] Furthermore, the operation technology according to the present invention is expected to contribute to "9. Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation" of the Sustainable Development Goals (SDGs) advocated by the United Nations.

[0013] First Embodiment A first embodiment will be described with reference to FIGS. 1 to 4. FIG.

[0014] 1 is a block diagram of an information terminal according to this embodiment, which is applicable to the second embodiment and subsequent embodiments in addition to the first embodiment. In the following description, an HMD will be used as an example of the information terminal 1.

[0015] 1 includes a front camera 10, a distance measurement sensor 11, an eye camera 12, a gesture camera 13, a sensor group 14, a motion sensor 15 which is part of the sensor group 14, a display 16, a touchpad 17, an audio input unit 18, an audio output unit 19, a communication interface 21, a main processor 22, a memory 23, and a storage 24, which are connected to one another by an internal bus 29. The communication interface 21, the main processor 22, the memory 23, and the storage 24 are unitized as a control unit 20 and are mounted on the information terminal 1 (see FIG. 2).

[0016] The storage 24 stores a basic operation program 25, an application program set 26, an input interface program set 27, and an output interface program set 28.

[0017] The application program set 26 includes a computer-generated reality processing program 30, a web browser processing program 31, a document processing program 32, a video playback processing program 33, a camera shooting processing program 34, and the like.

[0018] The input interface program set 27 includes an eye-tracking processing program 35, a blink detection processing program 36, a head movement detection processing program 37, a voice dictation processing program 38, a gesture recognition processing program 39, a touch input processing program 40, and the like.

[0019] The output interface program set 28 includes a two-dimensional image generation processing program 41, a three-dimensional image generation processing program 42, a synthetic voice generation processing program 43, a spatial audio generation processing program 44, and the like.

[0020] The main processor 22 is made up of a CPU and the like, and the memory 23 is made up of a RAM and the like.

[0021] The storage 24 is made up of a nonvolatile storage medium such as a flash ROM, and stores a basic operation program 25, an application program set 26, an input interface program set 27, and an output interface program set 28. The various programs in the basic operation program 25, the application program set 26, the input interface program set 27, and the output interface program set 28 are loaded into the memory 23 and executed by the main processor 22.

[0022] When a user wears information terminal 1 (HMD) and experiences viewing CGR content using computer-generated reality processing program 30, front camera 10 captures the real space in front of information terminal 1. Front camera 10 may be equipped with multiple cameras, or may capture a 360-degree surrounding area of ​​information terminal 1 using a 360-degree camera. Main processor 22 obtains images captured by front camera 10 from camera capture processing program 34 and detects real objects such as furniture and external people. Distance sensor 11 calculates the distance to the detected real objects and recognizes three-dimensional spatial relationships in the real space.

[0023] The image displayed on the display 16 is an image obtained by a two-dimensional image generation processing program 41 or a three-dimensional image generation processing program 42. The computer-generated reality processing program 30 generates an image of a virtual space in VR, and creates an image of a virtual object in MR, overlays it on real space, and displays it as a three-dimensional image on the display 16. Note that the display 16 may display a virtual image in the accessibility mode. In virtual image display, spatial position coordinates are associated with and stored for the display target of an icon, content, or virtual object, regardless of whether or not a physical display is performed on the display 16.

[0024] The audio output unit 19 outputs sounds of CGR content, etc. It also outputs synthetic audio generated by the synthetic audio generation processing program 43 as needed. Furthermore, the sound of CGR content or synthetic audio may be output as spatial audio by the spatial audio generation processing program 44. When synthetic audio is output as spatial audio, it is desirable to output it as spatial audio with sound image localization corresponding to the output source of the synthetic audio. This allows, for example, even a visually impaired person who cannot see the display screen to be notified by expressing the order of the web browser icon 54, the icon 55 for launching the document processing program 32, the icon 56 for launching a program that plays CGR content, and the icon 57 for launching a video playback processing program, as shown in FIG. 3 (described later), using sound image localization corresponding to the cursor position when the cursor points to the icon. Furthermore, the cursor position can be expressed as a sound image by using sound image localization of a notification sound indicating the movement of the cursor position.

[0025] The audio input unit 18 acquires the user's voice. The audio output unit 19 may be a speaker built into the information terminal 1, or may have an audio output terminal and send a spatial audio signal to headphones or the like.

[0026] The audio input unit 18 may be a built-in microphone, or may input an audio signal from an external microphone via an audio input terminal.

[0027] The sensor group 14 includes a GPS sensor, a biometric authentication sensor, an optical sensor, and a motion sensor 15 that combines an acceleration sensor, a gyro sensor, and a magnetic sensor to sense the user's up / down, left / right, forward / backward movements, and rotational movements accompanied by changes in the direction of the user's gaze.

[0028] The communication interface 21 is equipped with multiple communication protocols, such as Internet communication and near-field communication, and uses them depending on the purpose. The computer-generated reality processing program 30 of the information terminal 1 receives CGR content using Internet communication and allows the user to view the CGR content. Furthermore, near-field communication and the like connects headphones (not shown) to the audio output unit 19. Some of the arithmetic processing of the main processor 22 may be executed via the communication interface 21 on a server on the Internet or, for example, on a personal computer connected via direct communication.

[0029] The gaze tracking processing program 35 processes images of the user's eyes captured by the eye camera 12 and the camera image capturing processing program 34. The gaze tracking input interface is configured by including the eye camera 12 and the camera image capturing processing program 34. The gaze tracking processing program 35 also detects blinks of the user 2 using images from the eye camera 12, and therefore these correspond to blink detection sensors.

[0030] The head movement input interface includes a movement sensor 15 and a head movement detection processing program 37 .

[0031] The gesture input interface includes the gesture camera 13 and a gesture recognition processing program 39 .

[0032] The voice input interface includes a voice input unit 18 and a voice dictation processing program 38 .

[0033] The touch input interface includes a touch pad 17 and a touch input processing program 40 .

[0034] Furthermore, the eye camera 12 captures an image of the user's eyes, the camera photographing processing program 34 generates an eye image, and the blink detection processing program 36 analyzes the eye image to detect whether the user has blinked. That is, part of the input interface in the accessibility mode is configured to include the eye camera 12, the camera photographing processing program 34, and the blink detection processing program 36.

[0035] Fig. 2 is an external view of an HMD as the information terminal 1. In Fig. 2, the same components as those shown in the block diagram of the information terminal 1 in Fig. 1 are denoted by the same reference numerals.

[0036] The HMD 1a as the information terminal 1 includes a side-mounted housing 45a and a front-mounted housing 45b for mounting the HMD 1a on the user's head.

[0037] The front-mounted housing 45b is provided with a sensor group 14 and a voice input unit 18. A display 16 is disposed in front of the front-mounted housing 45b. Above the display 16, a front camera 10, a distance measurement sensor 11, a left eye camera 12a, a right eye camera 12b, and a gesture camera 13 are provided.

[0038] The side-mounted housing 45 a includes a touch pad 17 , a left speaker 19 a , a right speaker 19 b , and a control unit 20 .

[0039] The display 16 has a plurality of display elements for the left and right eyes as display elements, and displays images for the left and right eyes, allowing the user to perceive a three-dimensional image. The display 16 is a non-transmissive display panel in the case of a video see-through HMD, and is a transmissive display panel in the case of an optical see-through HMD. The display 16 may be of any other type as long as it is capable of three-dimensional display.

[0040] FIG. 3 is a diagram showing a state in which a user wears information terminal 1 and views the home screen of the information terminal superimposed on real space.

[0041] 3, a user 2 wears an information terminal 1 and headphones 3 on his / her head. A sofa 52 is placed in a real space 4 such as a living room where the user 2 is located.

[0042] The information terminal 1 is connected to a network 6 via an access point 5. The information terminal 1 is then connected to a CGR service server 7, a web service server 8, and a video service server 9 via the network 6. The information terminal 1 receives CGR content from the CGR service server 7 via the access point 5 and the network 6. In addition, the information terminal 1 receives and views web pages from the web service server 8. The information terminal 1 receives video services from the video service server 9. There are many other services in addition to those listed in Figure 3, and these services are provided by the respective service servers.

[0043] Information terminal 1 displays a home screen 51 of information terminal 1 superimposed on a real space image 50 captured from real space 4. Home screen 51 is a virtual reality screen generated by computer-generated reality processing program 30 of Fig. 1, and the size and depth can be set arbitrarily.

[0044] The real space image 50 covers the entire display screen of the information terminal 1, and the real space image is displayed regardless of whether the optical see-through or video see-through method is used.

[0045] The home screen 51 includes application icons 54 to 57 for launching applications executable on the information terminal 1, and a page-turn button 58 for the home screen 51. In Fig. 3, the positions of the cursors displayed at a certain point in time in the past are shown as black cursors 60, 61, and 62, and the current position of the cursor is shown as a white cursor 63. For ease of explanation, Fig. 3 illustrates the movement of the cursors on the home screen 51 by showing the past positions of the black cursors 60, 61, and 62 and the white cursor 63 in one diagram, but in reality, only one cursor is displayed on the home screen 51 at a given time.

[0046] 3 also illustrates spatial audio outputs 64, 65, 66, and 67. Spatial audio output 64 is an announcement that is output when cursor 60 is displayed. Spatial audio output 65 is a notification sound that is output when the position of cursor 60 transitions to cursor 61. Spatial audio output 66 is a notification sound that is output when the position of cursor 61 transitions to cursor 62, and is a double notification sound that indicates that the cursor has moved to the leftmost limit position on home screen 51. Spatial audio output 67 is an announcement that notifies users that the position of cursor 62 has transitioned to cursor 63 and that it is now overlapping with icon 54 of the web browser application.

[0047] Immediately after startup, the information terminal 1 is in default mode, with the cursor at an arbitrary position. If the user 2 blinks a predetermined number of times or more, the information terminal 1 detects the blink and transitions the information terminal 1 from default mode to accessibility mode. At this time, the cursor 60 is displayed at the center of the screen, and the home screen 51 is positioned in a coordinate system centered on the information terminal 1. This coordinate system is preferably a local coordinate system, but may also be a world coordinate system. Furthermore, when the information terminal 1 transitions from default mode to accessibility mode, even if the coordinate system of the default mode is a world coordinate system, it is preferable to change the coordinate system to a local coordinate system simultaneously with the transition to accessibility mode. In other words, even if the content of the real-space image changes as the user 2 moves, the home screen 51 is controlled to be at the center of the real-space image 50. Furthermore, a spatial audio output for confirmation, "Running in accessibility mode" 64, is output by localizing a sound image at the position of the cursor 60. When the default mode is switched to the accessibility mode, the cursor 60 is displayed at the center of the screen. However, the cursor may be positioned at the last cursor position when the home screen 51 was operated the previous time.

[0048] To find the web application icon 54, user 2 shakes his head to the left. Shaking his head here refers to moving his head left, right, up, or down and then returning it to its original position. When head movement detection simply detects head movement, it may detect movements of user 2 such as walking or turning around, which may differ from the input intended by the user. However, by detecting the head shaking motion, it is possible to avoid unintended input.

[0049] Furthermore, when shaking the head, if the action is performed within a predetermined time, the home screen 51 is not linked to the movement of the user 2's head, but is fixed as if the home screen 51 were in a coordinate system of real space.

[0050] When the head movement detection of the information terminal 1 detects a head shaking motion, cursor 60 moves to the position of left cursor 61 and a "beep, beep" sound is output as spatial audio. When the head is shaken to the left again, cursor 61 moves to the position of left cursor 62 and a "beep, beep" sound is output as spatial audio. The difference between "beep" and "beep, beep" is that "beep" indicates that the cursor has moved, while "beep, beep" indicates that the cursor cannot be moved any further in the same direction.

[0051] User 2 moves his / her head downward from the position of cursor 62 to the position of cursor 63. Cursor 63 is located on web application icon 54. A spatial audio output 67, which indicates the type of icon 54, "Web browser," is then output by localizing a sound image at the position of cursor 63. When launching the web application, user 2 blinks a predetermined number of times or more to launch the application.

[0052] Shaking the head left and right, up and down corresponds to the arrow keys on a keyboard, and blinking a predetermined number of times or more corresponds to the ENTER key. Shaking the head left and right, up and down, and blinking a predetermined number of times or more corresponds to a first action. After performing the first action and listening to the spatial audio output, the user can perform an undo process (UNDO function) of the first action by shaking the head in the opposite direction (second action).

[0053] FIG. 4 is a flowchart showing the transition to the accessibility mode and the selection of an icon.

[0054] When the process of FIG. 4 starts (S10), the information terminal 1 starts up in the default mode and executes the input interface program set 27 (S11).

[0055] The main processor 22 executes the blink detection processing program 36 (S12). If the main processor 22 executing the blink detection processing program 36 does not detect a blink based on the eye image from the eye camera 12 (S12: NO), the process returns to S11 and continues execution in the default mode.

[0056] If the main processor 22 does not detect a blink (S12: NO), the process returns to S11 and continues execution in the default mode.

[0057] If the main processor 22 detects a blink (S12: YES), the operation mode of the information terminal 1 is switched from the default mode to the accessibility mode (S13). Note that the blinks detected in S12 may be based on a predetermined pattern, such as multiple blinks within a predetermined time period or blinks in which the eyes are closed for a longer period than the normal blinking period, in order to distinguish them from normal blinks, i.e., blinks that are not intended to change the mode.

[0058] The main processor 22 displays the home screen 51 at the center of the display screen in real space in a coordinate system based on the information terminal 1 (S14), and the main processor 22 executing the synthetic voice generation processing program 43 generates synthetic voice saying "Running in accessibility mode," and the main processor 22 executing the spatial audio generation processing program 44 outputs spatial audio (S15).

[0059] The motion sensor 15 outputs motion information as a sensor output indicating head motion, and when the main processor 22 executing the head motion detection processing program 37 detects a head shaking motion based on the motion information (S16: YES), the main processor 22 moves the cursor in the direction of motion indicated by the motion information (S17).

[0060] If the main processor 22 cannot confirm the head shaking action (S16: YN), the process proceeds to before S25.

[0061] The main processor 22 checks whether the moved cursor is on an icon, and if it determines that it is on an icon (S18: YES), it outputs a synthesized voice explaining the icon, such as "This is a web browser," as spatial audio (S19), and transitions to S21.

[0062] If the main processor 22 determines that the moved cursor is not on an icon (S18: NO), a sound equivalent to a notification sound such as "beep" and "beep beep" is output as spatial audio (S20), and the process proceeds to before S16.

[0063] In S21, the main processor 22 again checks whether a blink has been detected, and if there is no blink (S21: NO), the process returns to S16 to further check whether head movement has been detected.

[0064] On the other hand, if the main processor 22 detects a blink again in S21 (S21: YES), it goes through point A (S22) and executes the application (S23). After that, when the execution of the application returns at point B (S24), it checks whether to end (S25). If not (S25: NO), it returns to before S16, and if it is to end (S25: YES), the series of processes ends (S26).

[0065] As described above, according to the first embodiment, it is possible to move the cursor and launch the application program indicated by the icon simply by blinking and moving the head, thereby providing an operation input device that allows hands-free operation input to the information terminal 1, an information terminal equipped with the same, and an operation input method. Furthermore, by realizing an accessibility function that applies an input interface for eye-tracking input, head movement input, and an output interface for spatial audio output, the position of an icon is notified to the user by sound image localization, and the cursor can be aligned to the position of the icon simply by blinking and moving the head, allowing operation input to the information terminal 1 without relying on visual information. Therefore, according to this embodiment, it is possible to provide an information terminal equipped with the same and an operation input method that are suitable for visually impaired people.

[0066] Second Embodiment A second embodiment will be described with reference to FIGS.

[0067] Figure 5 shows a first example of application execution and is a diagram illustrating the reading of a web page. In Figure 5, the same components as in Figure 3 are assigned the same numbers, and duplicated explanations will be omitted. Note that the access point 5, network 6, etc. are not shown.

[0068] In FIG. 5, a web screen 59 is superimposed on a real space image 50, and cursors 70, 71, 72, and 73 indicate positions to which the cursors move, and spatial audio outputs 74, 75, 76, and 77 are output at the positions of the cursors.

[0069] The web screen 59 is configured with a car image content on the top right, an airplane image content on the bottom left, a car text content on the top left, and an airplane text content on the bottom right.

[0070] At the position of the cursor 70, a synthesized voice explaining the image of a car, "This is a picture of a car," is output as spatial audio output 74.

[0071] When the user 2 shakes his / her head to the left at the position of the cursor 70, the cursor moves to the cursor 71, and the voice "This is a sentence about a car" is output as the spatial audio output 75 to notify that the cursor 71 is pointing to an explanation of the sentence about a car.

[0072] Furthermore, when user 2 shakes his / her head down at the position of cursor 71, cursor 72 moves to cursor 72, and a voice saying "This is a picture of an airplane" is output as spatial audio output 76 to notify that cursor 72 is pointing to the image of an airplane.

[0073] Furthermore, when user 2 shakes his / her head to the right at the position of cursor 72, cursor 72 moves to cursor 73, and the voice "This is the airplane sentence" is output as spatial audio output 77 to notify that cursor 73 is pointing to the airplane sentence.

[0074] Furthermore, if the user blinks a predetermined number of times or more at the position of the cursor 71 or cursor 73, the text about the car or the text about the airplane is read out.

[0075] FIG. 6 is a flowchart of a web page reading process of a web browser application.

[0076] The application starts to be executed at point A of S22 in FIG. 4, and ends to be executed at point B of S24.

[0077] The main processor 22 moves the cursor to the top of the web page or onto adjacent content (S30). In Figure 5, the top content of the web page is the position of cursor 71, and cursors 70, 72, and 73 correspond to positions on adjacent content.

[0078] The main processor 22 generates synthesized voice for the name, description, etc. of the content pointed to by the cursor (S31), and outputs spatial audio with a sound image localized at the position of the cursor (S32).

[0079] When the main processor 22 detects a blink of the user (S33: YES), it starts reading out the content (S34). In S34, the main processor 22 moves the cursor to the beginning of the sentence to be read out or to the beginning of an unread sentence. Furthermore, the main processor 22 generates a synthesized voice of the sentence (S35) and outputs it as spatial audio (S36).

[0080] The main processor 22 checks whether reading has been completed up to the end of the sentence of the content (S37). If the main processor 22 has not reached the end of the sentence of the content (S37: NO) and no head shaking movement is confirmed (S38: NO), the main processor 22 returns to S34 and moves to the beginning of the next unread sentence of the same content (S34).

[0081] On the other hand, if the end of the content has been reached (S37: YES), or if the end of the content has not been reached (S37: NO) but head shaking movement has been confirmed in S38 (S38: YES) and movement between contents has been instructed, the main processor 22 proceeds to S30.

[0082] If the main processor 22 does not detect a blink in S33 (S33: NO) and a head shaking movement is confirmed in S39 (S39: YES) and a move between contents is instructed, or if the main processor 22 does not detect a blink (S33: NO) and a head shaking movement is not confirmed in S39 (S39: NO) and the application execution cannot be terminated (S40: NO), the process proceeds to S30.

[0083] If no blinking is detected in S33 (S33: NO), and no head shaking movement is confirmed in S39 (S39: NO), and the application execution is to be terminated (S40: YES), the process exits from point B of S24 in FIG. 4 and returns to the home screen 51.

[0084] 7 is a diagram illustrating a second example of application execution, which explains transition to linked content. In FIG. 7, the same components as those in FIG. 3 and FIG. 5 are assigned the same numbers, and redundant explanations will be omitted.

[0085] In Figure 7, a web screen 59 is superimposed on a real space image 50, and the positions to which the cursor moves are indicated by cursors 80, 81, 82, and 83. The spatial audio outputs 84, 85, 86, and 87 output at each cursor position, the linked URL 89, and the linked content screen 90 are also shown.

[0086] In the upper left content of Fig. 7, each word in the sentence has a link destination. Move the cursor to select which linked content to display.

[0087] In FIG. 7, when the cursor 80 is positioned on the title of the content, a synthesized voice indicating the content, "This is a tourist guide to Japan," is output as spatial audio output 84.

[0088] When user 2 moves his / her head slowly downwards at the position of cursor 80, the cursor moves to cursor 81, and during the movement, a synthesized voice of several words, "Tokyo, Sapporo, Takayama," is output as spatial audio output 85. At this time, the spatial audio output may be configured to indicate how many words have been moved. Alternatively, the number of cursor movements may be displayed instead of the words.

[0089] In a normal head shake, only one word is moved at a time, but by slowly and for a long time moving the head, the cursor movement distance can be increased, and multiple words can be moved at once. Normally, one shake takes T seconds, and if the movement in one direction is approximately T / 2, then a slow and long head movement can move N words at once by continuing to move for NT / 2 seconds. If the head is returned at the normal speed of T / 2, the total movement time is (N+1) / 2, which is shorter than the movement time NT for repeating one shake movement N times.

[0090] When user 2 moves their head slowly and deeply to the right from the position of cursor 81, cursor 82 is reached and the synthesized speech "Takayama, Kanazawa, Tateyama" is output as spatial audio output 86. When the movement of cursor 82 is stopped, the synthesized speech "It is Tateyama" is output as spatial audio output 87 in accordance with the word at the position of cursor 82.

[0091] When the user blinks at the position of cursor 82, a content screen 90 of the link destination described in the word at the position of cursor 82 is displayed superimposed on the web screen 59. Cursor 82 then moves to cursor 83, and a synthesized voice saying "This is link content to the Tateyama Kurobe Alpine Route," indicating that cursor 83 is pointing to the content screen 90 of the link destination, is output as spatial audio 88, and the content screen 90 becomes ready to be read aloud. A URL 89 of the link destination is displayed at the top of the web screen 59.

[0092] Figure 8 is a flowchart that includes transition to linked content of a web application. The same steps as in the flowchart of Figure 6 are given the same numbers, and duplicate explanations will be omitted. The differences from the flowchart of Figure 6 are steps S50 to S57.

[0093] If no blinking is confirmed in S33 (S33: NO) and head movement is confirmed in S39 (S39: YES), the main processor 22 detects whether the head movement is slow and continuous (S50).

[0094] If the main processor 22 does not detect a slow, continuous movement (S50: NO), the process returns to S30, and the cursor moves to the next adjacent content (including a word linked to the linked content). For example, this is the case when the cursor moves one by one through the words (hereinafter sometimes referred to as words) linked to the linked content, such as "Takayama," "Kanazawa," and "Tateyama."

[0095] When the main processor 22 detects a slow, continuous movement of the user 2's head (S50: YES), it handles several words in succession according to the duration of the movement. For example, it moves the cursor over several words (S51), generates a synthesized voice of a series of multiple words, such as "Takayama Kanazawa Tateyama" (S52), and outputs it as spatial audio (S53).

[0096] When the user stops moving his / her head in response to the spatial audio output of "Tateyama..." and the main processor 22 detects a blink (S54: YES), a synthesized voice of the link name, "It's Tateyama," is generated (S55), output as spatial audio (S56), and the linked content is displayed (S57).

[0097] As described above, according to the second embodiment, as with the first embodiment, it is possible to provide an information terminal and an operating method for the information terminal that have accessibility functions that apply an input interface for eye-gaze tracking input and head movement input, and an output interface for spatial audio output. In particular, the second embodiment facilitates control of text read-out when an application is executed, and also facilitates access to linked content.

[0098] Third Embodiment A third embodiment will be described with reference to FIGS.

[0099] 9 is a third example of application execution, illustrating voice input for a web page. In FIG. 9, the same components as those in FIG. 3, FIG. 5, and FIG. 7 are assigned the same numbers, and redundant explanations will be omitted.

[0100] 9, a web screen 59 is superimposed on a real space image 50, and cursors 71 and 91 indicate the positions to which the cursors move, and also show spatial audio outputs 75 and 96 that are output at the positions of the cursors, and a voice input field 97. In the example of FIG. 9, the voice input field 97 is shown as a URL input field for a web page, but this is not limited to this and can be applied to the case of inputting text in all content.

[0101] In a situation where the main processor 22 outputs a synthesized voice saying "This is a sentence about a car" as spatial audio output 75 at the position of cursor 71 on the content in the upper left of Figure 9, user 2 shakes his head upward and moves cursor 71 to the position of cursor 95.

[0102] The cursor 91 is positioned at the voice input field 97, and the synthesized voice "This is the URL input field" is output as spatial audio output 96. After confirming the output of "This is the URL input field" in spatial audio output 96, user 2 inputs the URL text by voice. The main processor 22 recognizes the voice input and converts it into text by executing the voice dictation processing program 38 of FIG. 1.

[0103] Figure 10 is a flowchart that includes voice input in a web application. The same steps as in the flowchart in Figure 6 are given the same numbers, and duplicate explanations will be omitted. The differences from the flowchart in Figure 6 are steps S60 to S65.

[0104] If no blinking is confirmed in S33 (S33: NO) and a head shaking movement is confirmed in S39 (S39: YES), the main processor 22 moves the cursor in the direction of the head movement (S60).

[0105] When the main processor 22 confirms that the cursor has been moved to the voice input field (S61: YES), it generates a synthesized voice explaining that this is the voice input field 97, saying "This is the URL input field" (S62), and outputs this as spatial audio (S63).

[0106] When the user 2 confirms that the spatial audio indicates the voice input field 97, the user 2 speaks the voice to be input, and the main processor 22 performs voice dictation processing (S64). The main processor 22 displays a new web page based on the text obtained by the voice dictation processing (S65).

[0107] In the explanation of the second and third embodiments, the case where a web application is executed is described as an example of an application, but it goes without saying that the application of the present invention is not limited to web applications and can also be applied when other applications are executed.

[0108] As described above, according to the third embodiment, it is possible to provide an information terminal and an operating method for an information terminal that have accessibility functions that apply an input interface for eye tracking input and head movement input, as well as an audio input interface and an output interface for spatial audio output.

[0109] [Fourth embodiment] A fourth embodiment will be described with reference to Fig. 11 and Fig. 12. This embodiment is an example in which a mobile information terminal and a wearable terminal worn by a user are linked to perform motion detection input. Fig. 11 is a diagram showing the link with the wearable terminal.

[0110] 11, a user 2 wears a wearable terminal 100 in addition to an information terminal 1 and headphones 3. Examples of the wearable terminal 100 include a smart watch worn on the wrist and a smartphone placed in a pocket or the like.

[0111] The information terminal 1 and the wearable terminal 100 are connected by near field communication. The information terminal 1 receives movement information detected by the wearable terminal 100 and uses the information information in the head movement detection processing program 37 shown in FIG.

[0112] FIG. 12 is a flowchart of a process for reflecting movement information of the wearable device in steps S16, S38, and S39 for confirming head movement detection in FIGS. 4, 6, 8, and 10. In FIG.

[0113] The main processor 22 detects the movement information M of the head movement from the movement sensor 15 of the information terminal 1 (S70).

[0114] The main processor 22 also receives the movement information N from the wearable terminal 100 (S71).

[0115] The main processor 22 then compares the motion information M with the motion information N (S72) and, if the relationship M>>N holds, that is, if the motion information N of the wearable device 100 is much smaller than the motion information M of the information device 1 (corresponding to a case where the motion difference between the motion information N and the motion information M is greater than a predetermined motion difference) (S72: YES), it determines that the motion is of only the head of the user 2, and uses the motion information M as the motion detection result. If the main processor 22 determines that the relationship M>>N does not hold (S72: NO), it can be inferred that the motion information M is motion information associated with an action other than the movement of only the head of the user 2, such as walking or turning around, and therefore the motion information M is rejected and not used as the motion detection result. Therefore, the main processor 22 does not accept the operation input of cursor movement depending on the motion information M in this case.

[0116] As described above, according to the fourth embodiment, it is possible to provide an information terminal and an operating method for an information terminal that have accessibility functions that apply an input interface for eye-tracking input and head movement input, and an output interface for spatial audio output, and it is also possible to provide an information terminal and an operating method for an information terminal that further improves the accuracy of movement detection.

[0117] 1 to 12 illustrate an example of an input interface configuration using head movement detection, blink detection, and voice input. However, the present invention is not limited to these examples. It is also possible to use a combination of mouth movement using a face tracker, gesture input, voice commands, a digital crown, or touch input. In particular, by replacing blink detection, even users who have difficulty blinking multiple times can operate the device conveniently. Furthermore, in addition to the above, head movement and eye (iris) movement may be used instead of head movement detection.

[0118] In other words, the operation input device according to the present invention may be configured by combining movement information that detects slight movements of any limb, such as the fingers, wrists, or feet, instead of the movement of the user's head, with a blink detection sensor. This allows a user with limited head or neck movement to move a cursor using the movement of parts that the user can move, and execute a program indicated by the cursor by blink detection. In this case, the movement sensor may be a gyro sensor, acceleration sensor, direction sensor, or other motion sensor attached to the part to be moved, or the movement information may be obtained by analyzing images captured by a camera of the part to be moved by the user.

[0119] Although the above description concerns an operation input device for moving a cursor within a display screen controlled by the main processor 22, a hardware display and display screen are not required. The present invention may also be applied to an operation input device for moving a cursor placed within a virtual reality screen generated by the main processor 22. This allows the main processor 22 to generate and virtually display a virtual reality screen, rather than a display screen controlled by the display screen, particularly in information terminals used by visually impaired individuals. Since the display screen is not necessarily visible, the main processor 22 can virtually display the virtual reality screen, and move the cursor within the virtual reality screen. This reduces the weight of the information terminal 1 and frees it from size constraints imposed by hardware-based display and display screen configurations. It allows the user to move the cursor within a large virtual reality screen and execute desired processing. Furthermore, the power consumption of the information terminal 1 can be reduced, enabling extended use when using a battery. While the embodiments of the present invention described in Figures 1 to 12 use input interfaces that include head movement detection, blink detection, and voice input, these input interface configurations can also be used to zoom in and out on the screen. This can be achieved, for example, by switching to a zoom mode upon detection of a blink, changing the zoom ratio with a head shake, and ending the zoom mode upon detection of a blink.

[0120] Although the embodiments of the present invention have been described above, it goes without saying that the configurations for realizing the technology of the present invention are not limited to the above-described embodiments, and various modifications are possible. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. All of these fall within the scope of the present invention. Furthermore, numerical values, messages, etc. appearing in the text and figures are merely examples, and the effects of the present invention will not be impaired even if different ones are used.

[0121] Furthermore, the programs described in each processing example may be independent programs, or multiple programs may constitute a single application program. The order in which each process is performed may also be changed. Some or all of the functions of the present invention described above may be implemented in hardware, for example, by designing them as integrated circuits. Furthermore, they may also be implemented in software, with a microprocessor unit, CPU, or the like interpreting and executing operating programs that implement the respective functions.

[0122] Furthermore, the scope of software implementation is not limited, and hardware and software may be used together. Furthermore, some or all of the functions may be implemented by a server. The server may be, for example, a local server, a cloud server, an edge server, or an online service, as long as it can execute the functions in cooperation with other components via communications. Information such as programs, tables, and files that implement the functions may be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD, or may be stored in a device on a communications network.

[0123] Furthermore, the control lines and information lines shown in the diagram are those considered necessary for explanation, and do not necessarily represent all of the control lines and information lines on the product. In reality, it can be assumed that almost all components are interconnected.

[0124] As described above, the above embodiments include the following inventions. (Supplementary Note 1) An information terminal comprising: a motion sensor that detects user movement and outputs movement information; a blink detection sensor that detects user blinks; a processor; and a display, wherein upon acquiring the movement information, the processor moves a cursor displayed on a display screen of the display in accordance with the movement information, and executes processing indicated by the cursor when the blink detection sensor detects the user's blinks in a predetermined pattern. (Supplementary Note 2) An information terminal comprising: a motion sensor that detects user movement and outputs movement information; a blink detection sensor that detects the user's blinks; and a processor, wherein the processor generates a virtual reality screen and arranges at least one or more application icons and a cursor on the virtual reality screen, and upon acquiring the movement information, moves the cursor to match the icon in accordance with the movement information, and executes processing indicated by the cursor when the blink detection sensor detects the user's blinks in a predetermined pattern. (Supplementary Note 3) An operation input device comprising: a motion sensor that detects a user's motion and outputs motion information; a blink detection sensor that detects a blink of the user; and a processor, wherein the processor accepts an operation of moving a cursor displayed on a display screen whose display is controlled by the processor in accordance with the motion information, and an operation of instructing execution of a process indicated by the cursor when the blink detection sensor detects a blink of the user in a predetermined pattern. (Supplementary Note 4) An operation method for an information terminal, comprising: a step of moving a cursor displayed on a display screen whose display is controlled by a processor of the information terminal in accordance with motion information indicating the user's motion, and a step of executing the process indicated by the cursor when the blink detection sensor that detects a blink of the user detects a blink of the user in a predetermined pattern.

[0125] 1: Information terminal 1a: HMD 2: User 3: Headphones 4: Real space 5: Access point 6: Network 7: CGR service server 8: Web service server 9: Video service server 10: Front camera 11: Distance measurement sensor 12: Eye camera 12a: Left eye camera 12b: Right eye camera 13: Gesture camera 14: Sensor group 15: Motion sensor 16: Display 17: Touchpad 18: Audio input unit 19: Audio output unit 19a: Left speaker 19b: Right speaker 20: Control unit 21: Communication interface 22: Main processor 23: Memory 24: Storage 25: Basic operation program 26: Application program set 27: Input interface program set 28 : Output interface program set 29: Internal bus 30: Computer-generated reality processing program 31: Web browser processing program 32: Document processing program 33: Video playback processing program 34: Camera capture processing program 35: Eye tracking processing program 36: Detection processing program 37: Head movement detection processing program 38: Voice dictation processing program 39: Gesture recognition processing program 40: Touch input processing program 41: 2D image generation processing program 42: 3D image generation processing program 43: Synthetic voice generation processing program 44: Spatial audio generation processing program 45a: Side-mounted housing 45b: Front-mounted housing 50: Real space image 51: Home screen 52: Sofa 54: Icon 55: Icon 56: Icon 57: Icon 58: Button 59: Web screen 60 : Cursor 61 : Cursor 62 : Cursor 63 : Cursor 64 : Spatial audio output 65 : Spatial audio output 66 : Spatial audio output 67 : Spatial audio output 70 : Cursor 71 : Cursor 72 : Cursor 73 : Cursor 74 : Spatial audio output75: Spatial audio output 76: Spatial audio output 77: Spatial audio output 80: Cursor 81: Cursor 82: Cursor 83: Cursor 84: Spatial audio output 85: Spatial audio output 86: Spatial audio output 87: Spatial audio output 88: Spatial audio 89: URL 90: Content screen 91: Cursor 96: Spatial audio output 97: Audio input field 100: Wearable device

Claims

1. An information terminal comprising: a motion sensor that detects a user's motion and outputs motion information; a blink detection sensor that detects the user's blinks; a processor; and a display, wherein the processor, upon acquiring the motion information, moves a cursor displayed on the display screen in accordance with the motion information; and, upon detecting the user's blinks in a predetermined pattern, executes processing indicated by the cursor.

2. An information terminal as claimed in claim 1, wherein the processor performs the following processing as indicated by the cursor: launching and executing an application program corresponding to an icon located in the area of ​​the display screen indicated by the cursor; reading out content located in and including the area of ​​the display screen indicated by the cursor; or displaying linked content located in the area of ​​the display screen indicated by the cursor.

3. An information terminal according to claim 2, further comprising an audio output unit, wherein the processor, when moving the cursor in accordance with the movement information, generates synthetic audio corresponding to the icon or content at the cursor position, converts the synthetic audio into spatial audio with a sound image localized in accordance with the cursor position, and causes the audio output unit to output the synthetic audio as spatial audio.

4. An information terminal according to claim 1, wherein the movement information is information indicating a series of movements of the user's head in any of the left, right, up, and down directions and the head returning to its original position.

5. An information terminal according to claim 4, further comprising an audio output unit, wherein the movement information is information indicating continuous movements of the user's head in the same direction within a predetermined time period, and wherein the processor, when detecting continuous movements of the user's head in the same direction within a predetermined time period based on the movement information, continuously moves a plurality of icons or content displayed on the display screen, generates continuous synthetic audio of the continuously moved icons or content, converts the continuous synthetic audio into spatial audio with sound images localized in accordance with the position of the cursor, and causes the audio output unit to execute spatial audio output of the synthetic audio.

6. An information terminal as claimed in claim 4, wherein the movement information is information indicating that a first series of movements of the head in either the left, right, or up or down direction and returning the head to its original position has been performed, and information indicating that after the first movement, a second series of movements of the head in the opposite direction of the left, right, or up or down direction and returning the head to its original position has been performed, and the processor performs processing to cancel the processing performed in the first movement based on the movement information.

7. An information terminal as claimed in claim 1, further comprising a communication interface for communicating with a wearable terminal worn by a user, wherein the motion sensor is provided integrally with the information terminal, and wherein the processor compares the motion information acquired from the motion sensor with the motion information received from the wearable terminal, and when the amount of motion difference between the motion information of the wearable terminal and the motion information of the information terminal is greater than a predetermined amount of motion difference, validates the motion information of the motion sensor.

8. An information terminal according to claim 1, further comprising: a microphone; and a voice output unit; wherein a voice input field is displayed on the display; and wherein, when the cursor is located in the voice input field, the processor generates and outputs synthesized voice to notify the user that the cursor is located in the voice input field, and converts the voice collected by the microphone into text and inputs the text into the voice input field.

9. An information terminal comprising: a motion sensor that detects a user's motion and outputs motion information; a blink detection sensor that detects the user's blinks; and a processor, wherein the processor generates a virtual reality screen, arranges at least one application icon and a cursor on the virtual reality screen, moves the cursor in accordance with the icon in accordance with the motion information when the motion information is acquired, and executes processing indicated by the cursor when the blink detection sensor detects the user's blinks in a predetermined pattern.

10. An operation input device comprising: a motion sensor that detects a user's motion and outputs motion information; a blink detection sensor that detects the user's blinks; and a processor, wherein the processor accepts an operation to move a cursor displayed on a display screen controlled by the processor in accordance with the motion information; and an operation to instruct the execution of a process indicated by the cursor when the blink detection sensor detects the user's blinks in a predetermined pattern.

11. A method for operating an information terminal, comprising the steps of: moving a cursor displayed on a display screen controlled by a processor of the information terminal in accordance with movement information indicating user movement; and executing processing indicated by the cursor when a blink detection sensor that detects the user's blinks detects a predetermined pattern of the user's blinks.

Citation Information

Patent Citations

  • Head mounted video display device system

    JP1996202281A

  • Device and method for pointer control signal generation

    JP2005352580A

  • Computer-implemented method for enabling purchase related to augmented reality environment, computer-readable medium, ar device, and system for enabling purchase related to augmented reality environment

    JP2023099344A

  • Head-mounted display device and control method for head-mounted display device

    JP6277673B2

  • Modifying visual content to facilitate improved speech recognition

    JP6545716B2