Information processing method, information processing device, and program
Patent Information
- Application Number
- PCT/JP2025/003062
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-01-30
- Publication Date
- 2025-10-02
Smart Images

Figure JP2025003062_02102025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, and program
[0001] The present disclosure relates to an information processing method, an information processing device, and a program.
[0002] Patent Document 1 discloses a glasses-type or head-mounted display device according to the background art. The display device includes a display means, a user information acquisition means, and a control means. The display means is positioned in front of the user's eyes and displays images. The user information acquisition means acquires information about the user's condition from a visual sensor that detects the user's visual information or a biosensor that detects the user's biometric information. The control means determines whether the user is in a comfortable or uncomfortable state from the information acquired by the user information acquisition means, and controls the display operation of the display means based on the determination result.
[0003] The visual sensor has an imaging unit that captures an image of the user's eye. Analysis of the captured image detects a specific blinking or eye movement, which initiates the display of an image on the display means. Analysis of the captured image also detects a line of sight movement in a specific direction, which controls the forward / backward movement of the displayed image or executes scrolling of the display screen.
[0004] According to the display device disclosed in Patent Document 1, eye gaze control is performed simply by analyzing an image of the user's eye, so the accuracy of eye gaze control is low and erroneous operation is likely to occur.
[0005] JP 2013-77013 A
[0006] An object of the present disclosure is to provide an information processing method, an information processing device, and a program that can improve the operation accuracy of eye-gaze operations.
[0007] An information processing method according to one aspect of the present disclosure includes an information processing device that acquires gaze information of a user, identifies a focal length of the user's gaze based on the gaze information, identifies an object to be operated from one or more objects included in the user's field of view based on the focal length, acquires biometric information of the user, identifies an intention of the user to perform a gaze operation based on the biometric information, and performs a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to perform the operation.
[0008] FIG. 1 is a simplified diagram showing the overall configuration of AR glasses according to a first embodiment. FIG. 2 is a schematic diagram of a left eye unit of the AR glasses viewed from the side of the user's face. FIG. 3 is a simplified diagram showing the functional configuration of a processing unit according to the first embodiment. FIG. 4 is a flowchart showing a first example of processing executed by the processing unit according to the first embodiment. FIG. 5 is a flowchart showing a second example of processing executed by the processing unit according to the first embodiment. FIG. 6 is a flowchart showing a third example of processing executed by the processing unit according to the first embodiment. FIG. 7 is a timing chart showing the relationship between the timing of specifying an operation object and the period for acquiring biometric information. FIG. 8 is a timing chart showing the relationship between the timing of specifying an intention and the period for acquiring gaze information. FIG. 1 is a simplified diagram showing the overall configuration of AR glasses according to a second embodiment. FIG. 1 is a simplified diagram showing the functional configuration of a processing unit according to the second embodiment. FIG. 2 is a flowchart showing a first example of processing executed by the processing unit according to the second embodiment. FIG. 3 is a schematic diagram of the left eye unit of the AR glasses viewed from the side of the user's face. FIG. 4 is a schematic diagram of the left eye unit of the AR glasses viewed from the side of the user's face. FIG. 5 is a schematic diagram of the left eye unit of the AR glasses viewed from the side of the user's face. Fig. 1 is a schematic diagram of the left eye unit of the AR glasses viewed from the user's face side. Fig. 2 is a schematic diagram of the left eye unit of the AR glasses viewed from the user's face side. Fig. 3 is a flowchart showing details of the execution process of gaze operation. Fig. 4 is a schematic diagram of the left eye unit of the AR glasses viewed from the user's face side. Fig. 5 is a simplified diagram showing the overall configuration of AR glasses according to a modified example. Fig. 6 is a simplified diagram showing the functional configuration of a processing unit according to a modified example. Fig. 7 is a flowchart showing an example of processing executed by a processing unit according to a modified example.
[0009] (Findings forming the basis of the present disclosure) Research is underway into user interfaces for operating objects displayed on a screen in AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, etc. In some products, a user moves their line of sight onto a desired object and then performs a predetermined hand gesture to operate the object.
[0010] However, when AR glasses are used to assist work in a factory, for example, it is difficult for a user to perform hand gestures while working with both hands. Therefore, it is desirable to realize a hands-free user interface that does not rely on hand gestures.
[0011] For example, the display device disclosed in Patent Document 1 includes a visual sensor that detects visual information of a user. The visual sensor has an imaging unit that captures an image of the user's eye. Analysis of the captured image detects a specific blinking action or a specific eye movement, and then the display of an image on the display means begins. Analysis of the captured image also detects a line of sight movement in a specific direction, and then the display image is advanced / rewinded, or a scrolling process of the display screen is executed.
[0012] However, with the display device disclosed in Patent Document 1, gaze control is performed simply by analyzing an image of the user's eye, so the accuracy of gaze control is low and erroneous operation is likely to occur.
[0013] In order to solve this problem, the inventor discovered that the accuracy of gaze operations can be improved by identifying the object to be operated based on the focal length of the user's gaze, identifying the user's intention to perform gaze operations based on the user's biometric information, and performing gaze operations based on the results of identifying the object to be operated and the results of identifying the intention to perform, and thus came up with the present disclosure.
[0014] Next, each aspect of the present disclosure will be described.
[0015] An information processing method according to a first aspect of the present disclosure includes an information processing device that acquires gaze information of a user, identifies a focal length of the user's gaze based on the gaze information, identifies an object to be operated from one or more objects included in the user's field of view based on the focal length, acquires biometric information of the user, identifies an intention of the user to perform a gaze operation based on the biometric information, and performs a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to perform the operation.
[0016] According to the first aspect, it is possible to improve the accuracy of identifying the operation target object, and as a result, it is possible to improve the operation accuracy of the gaze operation on the operation target object.
[0017] In the information processing method according to the second aspect of the present disclosure, in the first aspect, the process of acquiring the biometric information may be started after the operation target object is identified.
[0018] According to the second aspect, the processing load on the information processing device can be reduced.
[0019] In the information processing method according to a third aspect of the present disclosure, in the first aspect, the process of acquiring the line of sight information may be started after the intention to execute is identified.
[0020] According to the third aspect, the processing load on the information processing device can be reduced.
[0021] In the information processing method according to a fourth aspect of the present disclosure, in the first aspect, the process of acquiring the line of sight information and the process of acquiring the biological information may be executed in parallel.
[0022] According to the fourth aspect, the operation accuracy of the eye-gaze operation can be further improved.
[0023] In the information processing method according to the fifth aspect of the present disclosure, in the fourth aspect, the acquired biometric information may further be stored, and in identifying the execution intention, the execution intention may be identified based on the biometric information acquired within a predetermined period including before and after the time when the target object is identified.
[0024] According to the fifth aspect, the operation accuracy of the eye-gaze operation can be further improved.
[0025] In the information processing method according to the sixth aspect of the present disclosure, in the fourth aspect, the acquired gaze information may further be stored, and in identifying the object to be operated, the object to be operated may be identified based on the gaze information acquired within a predetermined period including before and after the time when the execution intention is identified.
[0026] According to the sixth aspect, the operation accuracy of the eye-gaze operation can be further improved.
[0027] An information processing method according to a seventh aspect of the present disclosure, in any one of the first to sixth aspects, further comprises displaying one or more virtual objects within the user's field of view, the one or more objects including the one or more virtual objects, identifying a gaze object that the user is presumed to be gazing at from among the one or more virtual objects, changing a display distance that is a distance in a depth direction of the field of view from the eyeball of the user to the gaze object, and identifying the operation target object from among the one or more virtual objects based on the change in the display distance and the change in the focal length.
[0028] According to the seventh aspect, it is possible to further improve the accuracy of specifying the operation target object, and as a result, it is possible to further improve the operation accuracy of the gaze operation.
[0029] An information processing method according to an eighth aspect of the present disclosure is any one of the first to seventh aspects, further comprising: displaying a virtual object within the user's field of view; the one or more objects include the virtual object; and, in identifying the object to be operated, if multiple objects including the virtual object overlap within the user's field of view, changing a display distance which is the distance in the depth direction of the field of view from the user's eyeball to the virtual object, and identifying the object to be operated from among the multiple objects based on the change in the display distance and the change in the focal length.
[0030] According to the eighth aspect, it is possible to improve the accuracy of specifying the operation target object, and as a result, it is possible to further improve the operation accuracy of the eye-gaze operation.
[0031] In the information processing method according to the ninth aspect of the present disclosure, in the seventh or eighth aspect, it is preferable that personal characteristic information regarding the user's focus adjustment function is acquired, and when displaying the virtual object, a display distance at which the virtual object is displayed is set based on the personal characteristic information.
[0032] According to the ninth aspect, the operation accuracy of the eye-gaze operation can be further improved. In addition, the burden on the user is reduced, thereby improving operability.
[0033] In the information processing method according to the tenth aspect of the present disclosure, in any one of the seventh to ninth aspects, it is preferable to further acquire personal characteristic information regarding the user's focus adjustment function, and, when changing the display distance, set a change range of the display distance based on the personal characteristic information.
[0034] According to the tenth aspect, the operation accuracy of the eye-gaze operation can be further improved. Also, the burden on the user can be reduced, thereby improving operability.
[0035] In an information processing method according to an eleventh aspect of the present disclosure, in any one of the first to tenth aspects, it is preferable that the information processing method further identifies the user's gaze direction based on the gaze information, and in identifying the object to be operated, identifies the object to be operated based on the focal length and the gaze direction.
[0036] According to the eleventh aspect, it is possible to improve the accuracy of specifying the operation target object, and as a result, it is possible to further improve the operation accuracy of the eye-gaze operation.
[0037] The information processing method according to the twelfth aspect of the present disclosure, in any one of the first to eleventh aspects, may further include acquiring an image of a real space included in the user's field of view, and the one or more objects may include real objects included in the image.
[0038] According to the twelfth aspect, it is possible to perform a gaze operation not only on a virtual object but also on a real object.
[0039] In the information processing method according to a thirteenth aspect of the present disclosure, in any one of the first to twelfth aspects, it is preferable that the display mode of the specified operation target object is made different from the display mode of other objects.
[0040] According to the thirteenth aspect, the user can recognize the identified operation target object at a glance, thereby improving operability.
[0041] In an information processing method according to a fourteenth aspect of the present disclosure, in any one of the first to thirteenth aspects, the biometric information may include at least one of gaze, focal length, pupil diameter, blood pressure, electroencephalogram, electromyogram, electrooculography, voice, and gesture.
[0042] According to the fourteenth aspect, it is possible to improve the accuracy of identifying the intention of performing a gaze operation, and as a result, it is possible to further improve the operation accuracy of the gaze operation.
[0043] In an information processing method according to a fifteenth aspect of the present disclosure, in any one of the first to fourteenth aspects, when performing the gaze operation, gaze information of the user is acquired, a change in the focal length of the user's gaze is identified based on the gaze information, and the object to be operated is moved in the depth direction of the field of view based on the change in focal length.
[0044] According to the fifteenth aspect, the operation target object can be moved in the depth direction of the field of view, thereby improving operability.
[0045] In an information processing method according to a sixteenth aspect of the present disclosure, in any one of the first to fifteenth aspects, the information processing device may be mounted on AR glasses or VR glasses.
[0046] According to the sixteenth aspect, it is possible to realize AR glasses or VR glasses equipped with a highly accurate gaze control function.
[0047] An information processing device according to a seventeenth aspect of the present disclosure is an information processing device including a processing unit, which acquires gaze information of a user, identifies a focal length of the user's gaze based on the gaze information, identifies an object to be operated from one or more objects included in the user's field of view based on the focal length, acquires biometric information of the user, identifies an intention of the user to perform a gaze operation based on the biometric information, and performs a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to execute.
[0048] According to the seventeenth aspect, it is possible to improve the accuracy of specifying the operation target object, and as a result, it is possible to improve the operation accuracy of the gaze operation on the operation target object.
[0049] A program according to an eighteenth aspect of the present disclosure is a program for causing an information processing device to execute a process, the process acquiring a user's gaze information, identifying a focal length of the user's gaze based on the gaze information, identifying an object to be operated from one or more objects included in the user's field of view based on the focal length, acquiring the user's biometric information, identifying the user's intention to perform a gaze operation based on the biometric information, and performing a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to execute.
[0050] According to the eighteenth aspect, it is possible to improve the accuracy of specifying the operation target object, and as a result, it is possible to improve the operation accuracy of the gaze operation on the operation target object.
[0051] The present disclosure can also be realized as a program that causes a computer to execute each characteristic configuration included in such a method or apparatus, or as a system operated by this program. Needless to say, such a computer program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.
[0052] (Embodiments of the Present Disclosure) Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Elements with the same reference numerals in different drawings indicate the same or corresponding elements. Furthermore, the components, the arrangement positions of the components, the connection forms, the order of operations, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. The present disclosure is limited only by the claims. Therefore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept of the present disclosure are not necessarily required to achieve the objectives of the present disclosure, but are described as constituting more preferred forms.
[0053] 1 is a simplified diagram showing the overall configuration of AR glasses 1 according to a first embodiment of the present disclosure. The AR glasses 1 include a display unit 11, a processing unit 12, an input unit 13, a biometric information sensor 14, a storage unit 15, and an internal camera 16.
[0054] The display unit 11 is configured using a liquid crystal display, an organic EL display, etc. The processing unit 12 is configured using a processor (information processing device) such as a CPU, etc. The input unit 13 is configured using physical buttons, a microphone, etc.
[0055] The biometric information sensor 14 detects biometric information of the user wearing the AR glasses 1. The biometric information includes at least one of gaze, focal length, pupil diameter, blood pressure, electroencephalogram (EEG), electromyogram (EMG), electrooculogram (EOG), voice, and gesture. Gestures include predetermined movements by moving the user's eyelids or eyeballs, and predetermined movements by moving the user's fingers. Eyelid movements include lifting the eyelids and widening the eyes, or blinking repeatedly. If the biometric information is, for example, a gesture, the biometric information sensor 14 includes an internal camera or an external camera mounted on the AR glasses 1. If the biometric information is, for example, a brain wave, the biometric information sensor 14 includes an EEG sensor mounted on the AR glasses 1.
[0056] The storage unit 15 is configured using a semiconductor memory or the like. A program 21 and graphic data 22 are stored in the storage unit 15. The storage unit 15 includes a computer-readable non-volatile storage medium. The program 21 is stored in the storage medium. The graphic data 22 includes image data of virtual objects to be displayed on the display unit 11, etc.
[0057] The internal camera 16 is configured using an optical system, a CMOS image sensor, etc. The internal camera 16 captures moving images of both eyes of the user wearing the AR glasses 1.
[0058] 2 is a schematic diagram of the left eye unit of the AR glasses 1 viewed from the user's face side. Of the front and back surfaces of the left eye unit, the side facing the user's face (back side) corresponds to the internal side, and the side opposite the user's face (front side) corresponds to the external side. Note that the configuration of the right eye unit is the same as the configuration of the left eye unit, so the right eye unit is not shown in the figure.
[0059] The left-eye unit has a frame 25. An internal camera 16 is disposed inside the frame 25. A display unit 11 that virtually displays virtual objects 41, 42, such as icons, is disposed inside the frame 30. A transparent or semi-transparent lens is disposed inside the frame 30, and the user can view the scenery in real space included in the field of view through the lens. Note that the number of displayed virtual objects 41, 42 is not limited to two, and may be one or more. Furthermore, the display method for the virtual objects 41, 42 is not limited to projection onto the display unit 11, and direct retinal imaging, etc., may also be used.
[0060] Fig. 3 is a simplified diagram showing the functional configuration of the processing unit 12 according to the first embodiment. The functional configuration shown in Fig. 3 is realized by the processing unit 12 executing the program 21 read out from the storage unit 15. The processing unit 12 includes a gaze information acquisition unit 31, a focal length identification unit 32, a gaze direction identification unit 33, an operation target object identification unit 34, a display control unit 35, a biometric information acquisition unit 36, an execution intention identification unit 37, and a gaze operation execution unit 38. Details of the processing content of each of these units will be described later.
[0061] FIG. 4 is a flowchart showing a first example of processing executed by the processing unit 12 according to the first embodiment.
[0062] First, in step S01, the gaze information acquisition unit 31 acquires gaze information relating to the gaze of the user wearing the AR glasses 1. The gaze information includes image data D1 of both eyes of the user input from the internal cameras 16 of both eyes of the AR glasses 1. The gaze information acquisition unit 31 inputs data D2 indicating the gaze information to the focal length identification unit 32 and the gaze direction identification unit 33.
[0063] Next, in step S02, the focal length identification unit 32 calculates the angle of convergence of the user's eyes based on the data D2, thereby identifying the focal length of the user's line of sight. The focal length identification unit 32 may identify the focal length by calculating the thickness of the user's eye lenses instead of the angle of convergence. The focal length identification unit 32 inputs data D3 indicating the identified focal length to the operation target object identification unit 34. Also in step S02, the gaze direction identification unit 33 calculates the direction of the user's eyeballs based on the data D2, thereby identifying the gaze direction of the user's line of sight. The gaze direction identification unit 33 inputs data D4 indicating the identified gaze direction to the operation target object identification unit 34.
[0064] Next, in step S03, the operation target object identification unit 34 identifies an operation target object from among one or more virtual objects 41, 42 included in the user's field of view. The operation target object is an object on which the user is about to perform gaze operation. Data D5 indicating the display positions of the virtual objects 41, 42 on the display unit 11 is input from the display control unit 35 to the operation target object identification unit 34. The display positions include position coordinates of each virtual object 41, 42 in the in-plane direction of the display screen of the display unit 11. The display positions also include a display distance, which is the distance from the user's eyeball to each virtual object 41, 42 in the depth direction of the user's field of view. The display positions of each virtual object 41, 42 can be identified three-dimensionally based on the position coordinates and the display distance. If the user's gaze remains at the display position of a specific object for a predetermined period of time or more, that is, if there is a gaze object that the user is gazing at, the operation target object identification unit 34 identifies the gaze object as the operation target object based on the data D3 to D5. Specifically, when there is a gazed-up object whose focal length indicated in data D3 matches the display distance included in data D5 and whose line-of-sight direction indicated in data D4 matches the position coordinates included in data D5, the operation target object identification unit 34 identifies the gazed-up object as the operation target object. However, when there is only one object displayed and its position coordinates are fixed, the determination of whether the line-of-sight direction and position coordinates match may be omitted. The operation target object identification unit 34 inputs data D6 indicating the identified operation target object to the gaze operation execution unit 38.
[0065] If the operation target object specification unit 34 does not specify an operation target object (step S03: NO), the process ends.
[0066] If the operation target object identification unit 34 has identified the operation target object (step S03: YES), then in step S04, the biometric information acquisition unit 36 starts a process of acquiring output data D7 from the biometric information sensor 14. The biometric information acquisition unit 36 acquires biometric information of the user based on the output data D7, and inputs data D8 indicating the acquired biometric information to the execution intention identification unit 37.
[0067] Next, in step S05, the intention identification unit 37 identifies the intention of the user to perform a gaze operation based on the data D8. The biometric information includes at least one of gaze, focal length, pupil diameter, blood pressure, electroencephalogram, electromyogram, electrooculography, voice, and gesture, and an intention estimation rule is set in advance for each piece of biometric information. The intention identification unit 37 determines whether or not the biometric information indicated by the data D8 satisfies the estimation rule corresponding to the biometric information. If the intention identification unit 37 determines that an intention is present, it inputs data D9 indicating that the intention has been identified to the gaze operation execution unit 38. The data D9 includes the execution content of the gaze operation according to the biometric information.
[0068] If the intention identifying unit 37 does not identify an intention (step S05: NO), the process ends.
[0069] If the execution intention identification unit 37 identifies an execution intention (step S05: YES), then in step S06, the gaze operation execution unit 38 executes a gaze operation with the execution content indicated by data D9 on the object to be operated indicated by data D6.
[0070] FIG. 5 is a flowchart showing a second example of the process executed by the processing unit 12 according to the first embodiment.
[0071] In the first example, the processing unit 12 starts the process of acquiring biometric information after identifying the operation target object. Conversely, in the second example, the processing unit 12 starts the process of acquiring gaze information after identifying the intention to perform the gaze operation.
[0072] First, in step S04, the biometric information acquisition unit 36 acquires the biometric information of the user.
[0073] Next, in step S05, the intention identifying unit 37 identifies the intention of the user's eye gaze operation.
[0074] If the intention identifying unit 37 does not identify an intention (step S05: NO), the process ends.
[0075] If the intention identifying unit 37 identifies an intention (step S05: YES), the gaze information acquiring unit 31 starts a process of acquiring the user's gaze information in step S01.
[0076] Next, in step S02, focal length specifying unit 32 specifies the focal length of the user's line of sight. Also in step S02, line of sight direction specifying unit 33 specifies the line of sight direction of the user's line of sight.
[0077] Next, in step S03, the operation target object specification unit 34 specifies an operation target object from among one or more virtual objects 41, 42 included in the user's field of view.
[0078] If the operation target object specification unit 34 does not specify an operation target object (step S03: NO), the process ends.
[0079] If the operation target object specifying unit 34 specifies the operation target object (step S03: YES), then in step S06, the line-of-sight operation executing unit 38 executes a line-of-sight operation on the operation target object.
[0080] FIG. 6 is a flowchart showing a third example of the process executed by the processing unit 12 according to the first embodiment.
[0081] In the third example, the processing unit 12 executes the process of acquiring gaze information and the process of acquiring biological information in parallel.
[0082] When the process starts, the line-of-sight information acquisition unit 31 acquires line-of-sight information of the user in step S01. In parallel with step S01, the biometric information acquisition unit 36 acquires biometric information of the user in step S04.
[0083] After step S01, in step S02, the line-of-sight direction identification unit 33 identifies the line-of-sight direction of the user's line of sight. Next, in step S03, the operation target object identification unit 34 identifies an operation target object from one or more virtual objects 41, 42 included in the user's field of view. If the operation target object identification unit 34 does not identify an operation target object (step S03: NO), the process ends.
[0084] In step S05, the intention identifying unit 37 identifies the intention of the eye gaze operation by the user. If the intention identifying unit 37 does not identify the intention (step S05: NO), the process ends.
[0085] If the operation target object identification unit 34 identifies the operation target object and the execution intention identification unit 37 identifies the execution intention (steps S03, S05: YES), then in step S06, the gaze operation execution unit 38 executes gaze operation on the operation target object.
[0086] In the third example, the processing unit 12 may adjust the execution period of the process of identifying the operation target object based on the gaze information and the execution period of the process of identifying the execution intention based on the biometric information as follows.
[0087] 7 is a timing chart showing the relationship between the timing of specifying the operation target object and the period of time during which biometric information is acquired. The biometric information acquisition unit 36 stores data D8 indicating the acquired biometric information in the storage unit 15.
[0088] When the operation target object identification unit 34 identifies the operation target object at time T11, the intention identification unit 37 reads data D8 acquired within a period P1 from time T10 to time T12, including the periods before and after time T11, from the storage unit 15. The intention identification unit 37 identifies the intention of the user's gaze operation based on the biometric information within the period P1 indicated by the data D8 read from the storage unit 15.
[0089] 8 is a timing chart showing the relationship between the timing of identifying an intention and the period of acquiring gaze information. The gaze information acquiring unit 31 stores data D2 indicating the acquired gaze information in the storage unit 15.
[0090] When the intention identification unit 37 identifies the intention at time T21, the focal length identification unit 32 and the gaze direction identification unit 33 read data D2 acquired during a period P2 from time T20 to time T22, including the periods before and after time T21, from the storage unit 15. The focal length identification unit 32 and the gaze direction identification unit 33 identify the focal length and gaze direction of the user's gaze based on the gaze information during the period P2 indicated by the data D2 read from the storage unit 15. The operation target object identification unit 34 identifies the operation target object based on the focal length and gaze direction identified based on the gaze information during the period P2.
[0091] According to this embodiment, by using the focal length, it is possible to improve the accuracy of identifying the operation target object, and as a result, it is possible to improve the operation accuracy of the gaze operation on the operation target object.
[0092] Furthermore, according to the example shown in FIG. 4, the processing load on the processing unit 12 can be reduced by starting the process of acquiring biometric information after identifying the operation target object.
[0093] Furthermore, according to the example shown in FIG. 5, the processing load on the processing unit 12 can be reduced by starting the process of acquiring line-of-sight information after the intention to execute is identified.
[0094] Furthermore, according to the example shown in FIG. 6, the operation accuracy of the eye-gaze operation can be further improved by executing the process of acquiring the eye-gaze information and the process of acquiring the biometric information in parallel.
[0095] Furthermore, according to the example shown in FIG. 7, the intention to execute can be identified based on biometric information acquired within a predetermined period including before and after the point in time when the object to be operated is identified, thereby further improving the accuracy of gaze operation.
[0096] Furthermore, according to the example shown in FIG. 8, the operation accuracy of gaze operations can be further improved by identifying the object to be operated based on gaze information acquired within a predetermined period including before and after the time when the execution intention is identified.
[0097] Second Embodiment Fig. 9 is a simplified diagram showing the overall configuration of AR glasses 1 according to a second embodiment of the present disclosure. The AR glasses 1 further include an external camera 17 and a distance measurement sensor 18 in addition to the configuration shown in Fig. 1 .
[0098] The external camera 17 is configured using an optical system, a CMOS image sensor, etc. The external camera 17 captures an image of real space included in the field of view of the user wearing the AR glasses 1. The image of real space may include video data of the real space captured by the external camera 17.
[0099] The distance measurement sensor 18 is configured using a light detection and ranging (LiDAR) sensor, a time of flight (ToF) sensor, etc. The distance measurement sensor 18 measures the distance to a real object included in the real space.
[0100] FIG. 10 is a simplified diagram showing the functional configuration of the processing unit 12 according to the second embodiment. The processing unit 12 further includes an individual characteristic learning unit 61 in addition to the configuration shown in FIG. 3. The focusable depth distance of the field of view varies from person to person depending on visual acuity, etc. The individual characteristic learning unit 61 learns the user's individual characteristics related to the focus adjustment function by machine learning or the like performed during calibration. The individual characteristic learning unit 61 inputs data D11 indicating the learned individual characteristics to the display control unit 35.
[0101] Fig. 11 is a flowchart showing a first example of processing executed by the processing unit 12 according to the second embodiment. Fig. 11 discloses an example in which this embodiment is applied to the example shown in Fig. 6, but this embodiment can also be applied to the example shown in Fig. 4 or the example shown in Fig. 5.
[0102] FIG. 12 is a schematic diagram of the left eye unit of the AR glasses 1 viewed from the user's face side. An external camera 17 and a distance measurement sensor 18 are disposed on the exterior side of the frame 25. In the example shown in FIG. 12, virtual objects 41 and 42 are displayed on the display unit 11, and a real object 43 (a book in this example) existing in real space is included in the user's field of view. Here, the display control unit 35 sets the display distance of the virtual objects 41 and 42 within the user's focusable range based on the data D11. The virtual objects 41 and 42 and the real object 43 are disposed side by side in close proximity to each other.
[0103] Referring to FIG. 11, in step S11 following step S02, the operation target object identification unit 34 identifies a gaze object that is estimated to be gazed at by the user from among a plurality of objects including virtual objects 41, 42 and real object 43, based on data D3 to D5.
[0104] If the operation target object specifying unit 34 does not specify a gaze object (step S11: NO), the process ends.
[0105] In the following example, it is assumed that the virtual object 42 is identified as the gaze object. If the operation target object identification unit 34 identifies the gaze object (step S11: YES), then in step S12, the display control unit 35 changes the display distance of the gaze object.
[0106] FIG. 13 is a schematic diagram of the left eye unit of the AR glasses 1 as viewed from the user's face. By shortening the display distance of the gazed-up object by the display control unit 35, the virtual object 42, which is the gazed-up object, is displayed larger and closer to the user than the display position shown in FIG. 12 . Here, the display control unit 35 sets the changed display distance of the virtual object 42 within the user's focusable range based on data D11. The display control unit 35 may also lengthen the display distance of the gazed-up object to display the virtual object 42 smaller and further back than the display position shown in FIG. 12 . Furthermore, when a real object 43 is the gazed-up object, the display control unit 35 may create a figure simulating the real object 43 based on an image captured by the external camera 17, and change the display distance of the figure based on the distance to the real object 43 measured by the distance measurement sensor 18.
[0107] Next, in step S13, the operation target object identification unit 34 acquires data D3 again after changing the display distance, and identifies the operation target object based on the change in display distance and the change in focal length. If the focal length has changed in response to the change in display distance, the operation target object identification unit 34 identifies the gazed-up object as the operation target object. If the focal length has not changed in response to the change in display distance, the operation target object identification unit 34 does not identify the gazed-up object as the operation target object. In this case, the operation target object identification unit 34 may identify the virtual object 41 or real object 50 that is close to the virtual object 42 as the operation target object.
[0108] If the operation target object specifying unit 34 does not specify an operation target object (step S13: NO), the process ends.
[0109] If the operation target object specifying unit 34 specifies the operation target object (step S13: YES), the process of step S06 is executed.
[0110] 14 is a schematic diagram of the left eye unit of the AR glasses 1 viewed from the user's face side. When the operation target object identification unit 34 identifies the operation target object (step S13: YES), the display control unit 35 may differentiate the display mode of the identified operation target object from the display modes of other objects. In the example shown in FIG. 14 , the virtual object 42, which is the operation target object, is displayed brightly with high brightness, while the other objects, the virtual object 41 and the real object 43, are displayed darkly with low brightness. Note that the method of differentiating the display modes is not limited to changing the brightness, and may also be changing the clarity, etc.
[0111] Fig. 15 is a flowchart showing a second example of processing executed by the processing unit 12 according to the second embodiment. Fig. 15 discloses an example in which this embodiment is applied to the example shown in Fig. 6, but this embodiment can also be applied to the example shown in Fig. 4 or the example shown in Fig. 5.
[0112] Fig. 16 is a schematic diagram of the left eye unit of the AR glasses 1 viewed from the user's face. In the example shown in Fig. 16, in addition to a virtual object 42 being displayed on the display unit 11, a real object 50 (a person in this example) existing in real space is included in the user's field of view. Here, the display control unit 35 sets the display distance of the virtual object 42 within the user's focusable range based on data D11. At least a portion of the virtual object 42 and the real object 50 overlap within the user's field of view.
[0113] 15 , in step S21 after step S02, operation target object identification unit 34 determines whether virtual object 42 estimated to be the gaze object overlaps with real object 50. Operation target object identification unit 34 identifies the display position of virtual object 42 based on data D5, and identifies the position of real object 50 based on the image captured by external camera 17.
[0114] If the virtual object 42 and the real object 50 overlap (step S21: YES), the display control unit 35 changes the display distance of the virtual object 42 in step S22.
[0115] FIG. 17 is a schematic diagram of the left eye unit of the AR glasses 1 as viewed from the user's face. By shortening the display distance of the virtual object 42, the virtual object 42 is displayed larger and closer to the user than the display position in FIG. 16 . Here, the display control unit 35 sets the changed display distance of the virtual object 42 within the user's focusable range based on data D11. The display control unit 35 may increase the display distance of the virtual object 42, thereby displaying the virtual object 42 smaller and further back than the display position in FIG. 16 . The display control unit 35 may also create a figure simulating a real object 50 based on an image captured by the external camera 17, and change the display distance of the figure based on the distance to the real object 50 measured by the ranging sensor 18.
[0116] Next, in step S23, the operation target object identification unit 34 acquires data D3 again after changing the display distance, and identifies the operation target object based on the change in display distance and the change in focal length. If the focal length has changed in response to the change in display distance, the operation target object identification unit 34 identifies the virtual object 42 as the operation target object. If the focal length has not changed in response to the change in display distance, the operation target object identification unit 34 does not identify the virtual object 42 as the operation target object. In this case, the operation target object identification unit 34 may identify a real object 50 that overlaps with the virtual object 42 as the operation target object.
[0117] If the virtual object 42 and the real object 50 do not overlap (step S21: NO), then in step S23, the operation target object identification unit 34 identifies the virtual object 42 as the operation target object based on the data D3 to D5 if the virtual object 42 is the gaze object. If the virtual object 42 is not the gaze object, the operation target object identification unit 34 does not identify the virtual object 42 as the operation target object.
[0118] If the operation target object specification unit 34 does not specify an operation target object (step S23: NO), the process ends.
[0119] If the operation target object specifying unit 34 specifies the operation target object (step S23: YES), the process of step S06 is executed.
[0120] 18 is a flowchart showing the details of the eye-gaze operation execution process in step S06. In this example, it is assumed that the execution content of the eye-gaze operation indicated by data D9 is movement (dragging) of an object.
[0121] In step S061 , the line-of-sight information acquisition unit 31 acquires line-of-sight information relating to the line of sight of the user wearing the AR glasses 1 .
[0122] Next, in step S062, focal length specifying unit 32 specifies the focal length of the user's line of sight based on the line of sight information. Also in step S062, line of sight direction specifying unit 33 specifies the line of sight direction of the user's line of sight based on the line of sight information.
[0123] Next, in step S063, the line-of-sight operation execution unit 38 moves the operation target object along the trajectory of the focal length and line-of-sight direction of the user's line of sight.
[0124] 19 is a schematic diagram of the left eye unit of the AR glasses 1 as viewed from the user's face. By a drag operation using eye gaze control, the virtual object 42 has been moved to the far left from the display position shown in FIG.
[0125] According to this embodiment, by identifying the object to be operated based on changes in the display distance and the focal length, the accuracy of identifying the object to be operated can be further improved, and as a result, the accuracy of gaze operation can be further improved.
[0126] Furthermore, according to this embodiment, the accuracy of eye-gaze operation can be further improved by setting the display distance for displaying the virtual objects 41 and 42 based on the personal characteristic information indicated by the data D11. This also reduces the burden on the user, thereby improving operability.
[0127] Furthermore, according to this embodiment, the range of change in the display distance is set based on personal characteristic information, thereby further improving the accuracy of eye-gaze operation. Also, the burden on the user is reduced, thereby improving operability.
[0128] Furthermore, according to this embodiment, gaze operations can be performed not only on the virtual objects 41 and 42 but also on the real objects 43 and 50 .
[0129] Furthermore, according to this embodiment, by making the display mode of the identified object to be operated different from the display mode of other objects, the user can understand the identified object to be operated at a glance, thereby improving operability.
[0130] Furthermore, according to this embodiment, the operation target object can be moved in the depth direction of the field of view by eye-gaze operation, thereby improving operability.
[0131] 20 is a simplified diagram showing the overall configuration of AR glasses 1 according to a modification of the present disclosure, in which the biometric sensor 14 is omitted from the configuration shown in FIG.
[0132] 21 is a simplified diagram showing the functional configuration of the processing unit 12 according to the modified example, in which the biometric information acquisition unit 36 and the intention identification unit 37 are omitted from the configuration shown in FIG.
[0133] 22 is a flowchart showing an example of processing executed by the processing unit 12 according to the modified example, in which steps S04 and S05 are omitted from the flowchart shown in FIG.
[0134] In cases where the intention to be executed by gaze operation is simple, such as selecting an object, the processing unit 12 may omit identifying the intention to be executed based on biometric information and perform the gaze operation based only on the result of identifying the object to be operated based on gaze information.
[0135] Note that the present disclosure is not limited to the display of objects on a display unit such as AR glasses or VR glasses, but can also be applied to the display of objects on a two-dimensional display or three-dimensional display positioned facing the user.
[0136] The present disclosure is widely applicable to AR glasses, VR glasses, smart glasses, head-mounted displays, computer user interfaces, digital assistants, information terminals, wearable devices, and the like.
Claims
1. An information processing method in which an information processing device acquires gaze information of a user, identifies a focal length of the user's gaze based on the gaze information, identifies an object to be operated from one or more objects included in the user's field of view based on the focal length, acquires biometric information of the user, identifies an intention of the user to perform a gaze operation based on the biometric information, and performs a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to perform the operation.
2. The information processing method according to claim 1, further comprising starting a process of acquiring the biometric information after the object to be operated is identified.
3. The information processing method according to claim 1, further comprising: starting a process of acquiring the gaze information after the execution intention is identified.
4. The information processing method according to claim 1, wherein the process of acquiring the line-of-sight information and the process of acquiring the biological information are executed in parallel.
5. The information processing method according to claim 4, further comprising storing the acquired biometric information, and in identifying the execution intention, identifying the execution intention based on the biometric information acquired within a predetermined period including before and after the time when the object to be operated is identified.
6. The information processing method of claim 4, further comprising storing the acquired gaze information, and in identifying the object to be operated, identifying the object to be operated based on the gaze information acquired within a predetermined period including before and after the time when the execution intention is identified.
7. The information processing method according to claim 1, further comprising: displaying one or more virtual objects within the user's field of view, the one or more objects including the one or more virtual objects; identifying a gaze object that the user is presumed to be gazing at from the one or more virtual objects; changing a display distance, which is the distance in the depth direction of the field of view from the user's eyeball to the gaze object, in identifying the object to be operated; and identifying the object to be operated from the one or more virtual objects based on the change in the display distance and the change in the focal length.
8. The information processing method according to claim 1, further comprising: displaying a virtual object within the user's field of view; the one or more objects including the virtual object; and, in identifying the object to be operated, if multiple objects including the virtual object overlap within the user's field of view, changing a display distance which is the distance in the depth direction of the field of view from the user's eyeball to the virtual object; and identifying the object to be operated from among the multiple objects based on the change in display distance and the change in focal length.
9. The information processing method according to claim 7 or 8, further comprising acquiring personal characteristic information relating to the user's focus adjustment function, and setting a display distance at which the virtual object is displayed based on the personal characteristic information when displaying the virtual object.
10. An information processing method according to claim 7 or 8, further comprising acquiring personal characteristic information relating to the user's focus adjustment function, and setting a range of change in the display distance based on the personal characteristic information when changing the display distance.
11. The information processing method of claim 1, further comprising: identifying the user's gaze direction based on the gaze information; and, in identifying the object to be operated, identifying the object to be operated based on the focal length and the gaze direction.
12. The information processing method according to claim 1, further comprising acquiring an image of a real space included in the user's field of view, and the one or more objects including a real object included in the image.
13. The information processing method according to claim 1, further comprising: making the display mode of the identified object to be operated different from the display mode of other objects.
14. The information processing method according to claim 1, wherein the biological information includes at least one of gaze, focal length, pupil diameter, blood pressure, electroencephalogram, electromyogram, electrooculography, voice, and gesture.
15. The information processing method of claim 1, wherein, in performing the gaze operation, gaze information of the user is acquired, a change in the focal length of the user's gaze is identified based on the gaze information, and the object to be operated is moved in the depth direction of the field of view based on the change in focal length.
16. The information processing method according to claim 1, wherein the information processing device is mounted on AR glasses or VR glasses.
17. An information processing device having a processing unit, wherein the processing unit: acquires gaze information of a user; identifies a focal length of the user's gaze based on the gaze information; identifies an object to be operated from one or more objects included in the user's field of view based on the focal length; acquires biometric information of the user; identifies an intention of the user to perform a gaze operation based on the biometric information; and performs a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to perform the operation.
18. A program for causing an information processing device to execute a process, the process comprising: acquiring a user's gaze information; identifying the focal length of the user's gaze based on the gaze information; identifying an object to be operated from one or more objects included in the user's field of view based on the focal length; acquiring the user's biometric information; identifying the user's intention to perform a gaze operation based on the biometric information; and performing a gaze operation on the object to be operated based on the result of identifying the object to be operated and the result of identifying the intention to perform the operation.