Information providing device, information providing system, information providing method, and information providing program

The system addresses the limitation of requiring special tools by using gaze tracking and environmental sensors to deliver personalized information, improving accessibility and accuracy.

WO2025163691A1PCT designated stage Publication Date: 2025-08-07MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002543
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing information provision systems require users to have a printed piece of paper or a mobile terminal to operate, limiting accessibility for those without these tools.

Method used

An information provision system that estimates user intentions based on gaze tracking and environmental factors to provide information without the need for special tools, utilizing a camera, microphone, and display devices to determine and deliver relevant content.

Benefits of technology

Enables information delivery tailored to user intentions and environments, enhancing accessibility and accuracy by eliminating the need for paper or mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002543_07082025_PF_FP_ABST
    Figure JP2024002543_07082025_PF_FP_ABST
Patent Text Reader

Abstract

An information providing device according to the present invention comprises: an intent inference unit (23) that, on the basis of a plurality of images obtained by capturing an image of a person and display information displayed on a display device at the time of capturing the plurality of images, infers an intent of the person; a selection unit (24) that selects, in accordance with the inferred intent, a method or the content of information for provision; and an output control unit (25) that controls an output device so as to provide the information in accordance with the selection.
Need to check novelty before this filing date? Find Prior Art

Description

Information provision device, information provision system, information provision method, and information provision program

[0001] The present disclosure relates to an information providing technology that is a technology for providing information.

[0002] In public transportation service locations such as train stations, bus stops, and airport terminals, or in facilities such as department stores and shopping malls (hereinafter, these locations or facilities may be collectively referred to as "service spots"), information providing devices may be installed to provide various information related to the use of the service spots. For example, Patent Document 1 discloses a technology related to "a stationary guidance information display device, comprising: a code reading means for reading code information from a visible code; a storage means for storing a group of destination information in which the code information corresponds to destination information; a route guidance generating means for searching the group of destination information for a destination associated with the read code information and generating route guidance information to the destination; and a display means for displaying the generated route guidance information in the form of a three-dimensional map" (claim 1 of Patent Document 1), and it is envisioned that such a guidance information display device will be installed and used within train stations (paragraph 0022 of Patent Document 1). According to Patent Document 1, a user holds a piece of paper on which a visible code such as an AR marker is printed or a mobile terminal on which the code is displayed over a code reader (code reading means), and the code reader reads the visible code and the specified information is displayed (paragraphs 0024 to 0025 of Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2017-3444

[0004] However, according to the technology of Patent Document 1, the user of the guidance information display device must have a piece of paper on which the visible code is printed or a mobile terminal on which it is displayed, and there is a problem in that if the user does not have such a piece of paper or a mobile terminal, the user cannot operate the guidance information display device.

[0005] The present disclosure has been made to solve such problems, and aims to provide an information provision technology that can provide information without requiring special tools such as paper or a mobile terminal.

[0006] One aspect of an information providing device according to an embodiment of the present disclosure includes an intention estimation unit that estimates the intention of a person based on multiple images of the person and display information displayed on a display device at the time the multiple images are captured, a selection unit that selects the content or method of information to be provided in accordance with the estimated intention, and an output control unit that controls an output device to provide information in accordance with the result of the selection.

[0007] According to the information providing device according to the embodiment of the present disclosure, information can be provided without requiring special tools such as paper or a mobile terminal.

[0008] 1 is a diagram illustrating an example of the configuration of an information providing device and an information providing system. FIG. 2 is a diagram illustrating an example of the configuration of the hardware of the information providing device. FIG. 3 is a diagram illustrating an example of the configuration of the hardware of the information providing device. FIG. 4 is a diagram illustrating a flowchart of an information providing method by the information providing device. FIG. 5 is a diagram illustrating an example of information display by the information providing device. FIG. 6 is a diagram illustrating an example of gaze detection. FIG. 7 is a diagram illustrating an example of information display by the information providing device. FIG. 8 is a diagram illustrating an example of gaze detection. FIG. 9 is a diagram illustrating an example of information display after gaze detection. FIG. 10 is a diagram illustrating an example of information display after gaze detection. FIG. 11 is a diagram illustrating an example of gaze detection. FIG. 12 is a diagram illustrating an example of information display after gaze detection. FIG. 13 is a diagram illustrating an example of a correspondence pattern table (correspondence pattern table 1). FIG. 14 is a diagram illustrating an example of a correspondence pattern table (correspondence pattern table 2).

[0009] Various embodiments of the present disclosure will be described in detail below with reference to the drawings. In the drawings, identical or similar parts are designated by identical or similar reference numerals, and redundant explanations of such parts will be omitted. In addition, in this disclosure, the term "or" is used to mean an inclusive logical OR unless otherwise specified.

[0010] Embodiment 1. <Configuration> An information provision system Sys. according to embodiment 1 of the present disclosure and an information provision device 20 included in the information provision system Sys. will be described with reference to FIG. 1 . As an example, as shown in FIG. 1 , the information provision system Sys. includes a camera 11, a microphone 12, a storage device 13, a storage device 14, a storage device 15, an information provision device 20, a display device 31, and a speaker 32. The display device 31 and the speaker 32 are both examples of output devices in the present disclosure. The output devices including the display device 31 and the speaker 32 provide various information under the control of the information provision device 20. As an example, the display device 31 displays train operation information on a screen included in the display device 31 under the control of the information provision device 20. As another example, the speaker 32 provides the train operation information by voice under the control of the information provision device 20.

[0011] The information provision system Sys. is installed at service provision spots such as railway stations and is used as a digital signage system. The entire information provision system Sys. does not need to be installed in the same location. For example, the camera 11, microphone 12, display device 31, and speaker 32 may be installed on a station platform as an integrated device or separately, and the storage devices 13, 14, 15, and information provision device 20 may be installed in the train crew room of the station.

[0012] (Camera) The camera 11 typically captures multiple time-series images of a person (hereinafter, sometimes referred to as a "user") standing in front of the display device 31, and outputs the captured multiple images to the information providing device 20. The camera 11 is installed so as to be able to capture images of the person standing in front of the display device 31. The camera 11 may be provided integrally with the display device 31. The camera 11, the display device 31, and the information providing device 20 may also be provided integrally. The camera 11 captures images at a frame rate of, for example, 20 fps (frames per second) or 30 fps (frames per second).

[0013] (Microphone; Environmental Information Acquisition Device) The microphone 12 is an example of an environmental information acquisition device that acquires information about the user's environment. When the environmental information acquisition device is the microphone 12, the microphone 12 is installed near the camera 11 to collect sounds around the user. The microphone 12 collects sounds and outputs the collected sounds to the information providing device 20. The environmental information acquisition device is not limited to the microphone 12, and may be, for example, a processor that acquires weather information for the area in which the information providing device 20 is installed. When the environmental information acquisition device is a processor that acquires weather information, the environmental information acquisition device acquires the weather information via a communication network (not shown).

[0014] (Storage Device) The storage device 13 is a storage device that holds a personal information database that stores personal information. Examples of personal information include multiple ID numbers that identify users, and facial images and attributes of the users that are associated with each ID number. The facial images of the users may include data that represent the features of the facial images in addition to the facial image data.

[0015] The storage device 14 is a storage device that stores correspondence patterns. For example, the correspondence patterns are stored in the form of a table as shown in FIG. 8 , which associates "user's eye movement," "characteristics of displayed information," "user's intention," and "content of information to be provided or method of providing information." The correspondence patterns may also be associated with "user attributes," "the user's current environment," and "content of information to be provided or method of providing information," as shown in FIG. 9 . For convenience of explanation, the table in FIG. 8 will be referred to as "correspondence pattern table 1," and the table in FIG. 9 will be referred to as "correspondence pattern table 2." In FIG. 8 , nine correspondence patterns are defined in correspondence pattern table 1 as an example. The correspondence patterns defined in correspondence pattern table 1 may be collectively referred to as "correspondence pattern 1." In FIG. 9 , seven correspondence patterns are defined in correspondence pattern table 2 as an example. The correspondence patterns defined in correspondence pattern table 2 may be collectively referred to as "correspondence pattern 2."

[0016] Correspondence pattern 2 in Fig. 9 may be used in combination with correspondence pattern 1 in Fig. 8. For example, if the content defined as "content of information to be provided" in correspondence pattern 1 in Fig. 8 can be provided by both display and audio, it may be defined that the information is provided by either display or audio depending on the "user attributes" or the "environment in which the user is currently located" as in correspondence pattern 2 in Fig. 9.

[0017] The storage device 15 is a storage device that stores output device information related to an output device used for providing information by the information providing device 20. Examples of the output device information include information that identifies the display device 31, the speaker 32, or a lighting device (not shown).

[0018] 1, the information providing device 20 includes an eye gaze tracking unit 21, a display information acquisition unit 22, an intention estimation unit 23, a selection unit 24, and an output control unit 25. The information providing device 20 may also include, as optional additional functional units, an attribute determination unit 26 and an environment determination unit 27. These functional units included in the information providing device 20 will be described below.

[0019] (Gaze Tracking Unit) The gaze tracking unit 21 tracks the gaze of the user captured by the camera 11 from multiple images captured by the camera 11 using known gaze tracking technology. The gaze tracking unit 21 outputs tracked gaze information to the display information acquisition unit 22 as gaze information. The gaze information includes position information indicating which position on the screen of the display device 31 the tracked gaze is looking at, and characteristics of the tracked gaze movement. The gaze movement characteristics are, for example, a characteristic that the gaze moves back and forth between two or more areas, or a characteristic that the gaze does not move from the same area. Note that when multiple people are captured by the camera 11, the gaze tracking unit 21 may track the gaze of each person, or may include people whose gaze is tracked and people whose gaze is not tracked. Hereinafter, the user whose gaze is tracked may be referred to as the "target person."

[0020] (Display Information Acquisition Unit) The display information acquisition unit 22 acquires display information displayed in the line of sight indicated by the line of sight information. That is, the display information acquisition unit 22 acquires display information displayed at the position indicated by the position information included in the line of sight information. The display information acquisition unit 22 accesses a storage device (not shown) that stores display information (display content) to acquire the display content displayed on the display device 31, and acquires the display information displayed in the line of sight indicated by the line of sight information. From the relative positional relationship between the user's line of sight and the display position where the display content is displayed on the display device 31, the display information displayed in the line of sight indicated by the line of sight information can be identified.

[0021] The display information acquisition unit 22 analyzes the characteristics of the acquired display information. For example, it compares a display of "A" with a display of "B" and analyzes that "different displays are made." As another example, it compares a display of "A" with a text display including "A" and analyzes that "the same A is displayed." The display information acquisition unit 22 outputs the analysis result to the intention estimation unit 23.

[0022] (Intention Estimation Unit) The intention estimation unit 23 estimates the user's intention based on multiple images of the user and display information displayed on the display device 31 at the time the multiple images were captured. To perform this estimation, the intention estimation unit 23 estimates the user's intention from characteristics of the gaze movement tracked by the gaze tracking unit 21 and characteristics of the information displayed in the line of sight acquired by the display information acquisition unit 22. As an example, the user's intention may be estimated rule-based. For example, a correspondence pattern table 1 as shown in FIG. 8 may be created in advance and the intention estimation unit 23 may refer to the correspondence pattern table 1. The correspondence pattern 1 is stored in advance in the storage device 14, and the intention estimation unit 23 accesses the storage device 14 to refer to the correspondence pattern 1.

[0023] The intention estimation unit 23 may estimate the user's intention using a machine learning model that has learned the user's intention. Such a machine learning model may be generated by supervised learning or reinforcement learning, for example, using as input an image obtained by tracking the user's gaze for each piece of display information displayed on the display device 31 and the user's intention as training data. The machine learning model may also be generated using other methods. Details of the operation of the intention estimation unit 23 will be described later.

[0024] (Selection Unit) The selection unit 24 selects the content of information to be provided or the method of providing information according to the estimation result estimated by the intention estimation unit 23. This selection is performed by referring to, for example, a correspondence pattern table 1 as shown in Fig. 8 or a correspondence pattern table 2 as shown in Fig. 9.

[0025] The selection unit 24 may use the determination result by the attribute determination unit 26 or the determination result by the environment determination unit 27 instead of the estimation result by the intention estimation unit 23. Furthermore, the selection unit 24 may use the determination result by the attribute determination unit 26 or the determination result by the environment determination unit 27 in addition to the estimation result by the intention estimation unit 23. The selection unit 24 outputs the selection result to the output control unit 25. Details of the operation of the selection unit 24 will be described later.

[0026] (Output control unit) The output control unit 25 controls the provision of information by an output device associated with the selected content or method of information to be provided, in accordance with the selection result output from the selection unit 24. Information on the associated output device is stored in the storage device 15, and the output control unit 25 accesses the storage device 15 to control the output device.

[0027] (Attribute Determination Unit) The attribute determination unit 26 determines the attributes of the user captured by the camera 11. In the present disclosure, attributes refer to a person's profile including at least one of the person's name, age, gender, whether or not the person is healthy, whether or not the person has a visual or hearing impairment, whether or not the person uses a wheelchair, assistance history, whether or not the person has a commuter pass, or hobbies. The assistance history is a term that refers to a history of receiving assistance. The attribute determination unit 26 analyzes the facial image of the person captured by the camera 11 using known technology and compares information representing the facial image of the person captured by the camera 11 with personal information acquired by accessing a personal information database stored in the storage device 13 to determine the attributes of the person captured by the camera 11. The attribute determination unit 26 outputs the determined attributes to the selection unit 24.

[0028] (Environment Determination Unit) The environment determination unit 27 determines the environment of the target person. To determine the environment of the target person, the environment determination unit 27 determines the environment of the target person, such as whether the sound around the target person is loud or quiet, whether the weather at the location where the target person is located is sunny or rainy, and whether the target person is alone, based on information acquired by an environment information acquisition device such as the microphone 12.

[0029] As one example, if the display device 31 is installed on a station platform, the environment determination unit 27 determines that the surrounding sound is loud if the noise level is 70 dB or higher, and determines that the surrounding sound is quiet if the noise level is less than 70 dB. As another example, if the display device 31 is installed in a station concourse, the environment determination unit 27 determines that the surrounding sound is loud if the noise level is 60 dB or higher, and determines that the surrounding sound is quiet if the noise level is less than 60 dB.

[0030] As another example, the environment determination unit 27 determines that the weather is rainy if the amount of rainfall in the area where the display device 31 is installed is 1 mm or more, and determines that the weather is sunny if the amount of rainfall in the area is less than 1 mm. If the environment determination unit 27 can acquire information indicating whether the area where the display device 31 is installed is sunny or rainy from a server (not shown), the environment determination unit 27 may use the acquired information as is.

[0031] As another example, if an image containing a target person also contains an image of a person whose gaze cannot be detected in addition to the target person, the environment determination unit 27 may determine that there is only one target person.

[0032] Next, an example of the hardware configuration of the information providing device 20 will be described with reference to Figures 2A and 2B. Each function of the information providing device 20 is realized by a processing circuitry. The processing circuitry may be a dedicated processing circuit 100a as shown in Figure 2A, or a processor 100b as a computer that executes a program stored in a memory 100c as shown in Figure 2B.

[0033] When the processing circuitry is a dedicated processing circuit 100a, the dedicated processing circuit 100a may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination thereof. The functions of the information providing device 20 may be realized by multiple separate processing circuits, or the functions of the information providing device 20 may be realized together in a single processing circuit.

[0034] When the processing circuitry is a processor 100b, the functions of the information providing device 20 are realized by software, firmware, or a combination of software and firmware. The software and firmware are written as programs and stored in the memory 100c. The processor 100b realizes the functions of the information providing device 20 by reading and executing the programs stored in the memory 100c. Here, examples of the memory 100c include non-volatile or volatile semiconductor memories such as random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM), as well as magnetic disks, flexible disks, optical disks, compact disks, minidisks, and DVDs.

[0035] It is also possible to implement some of the functions of the information providing device 20 using dedicated hardware, and other functions using software or firmware. In this way, the processing circuit can implement the functions of the information providing device 20 using hardware, software, firmware, or a combination of these.

[0036] <Operation: Overview> Next, an overview of the information providing method by the information providing device 20 will be described with reference to FIG.

[0037] (Step ST1) In step ST1, the gaze tracking unit 21 detects the gaze of the user from a plurality of images captured of the user standing in front of the display device 31, tracks the trajectory of the gaze, and analyzes the characteristics of the gaze movement. The gaze tracking unit 21 outputs the tracked gaze information to the display information acquisition unit 22 as gaze information.

[0038] (Step ST2) In step ST2, the display information acquisition unit 22 acquires display information displayed ahead of the line of sight being tracked, and analyzes the characteristics of the acquired display information.

[0039] (Step ST3) In step ST3, the intention estimation unit 23 estimates the user's intention based on, for example, characteristics of eye movement and characteristics of the displayed information. The intention estimation unit 23 estimates the user's intention by referring to a pre-created pattern that associates eye movement, displayed information, and the user's intention. As described above, the intention estimation unit 23 may estimate the user's intention using a machine learning model.

[0040] (Step ST4) In step ST4, the selection unit 24 selects the content or method of the information to be provided in accordance with the estimated intention.

[0041] (Step ST5) In step ST5, the output control unit 25 controls the output device to provide information according to the result of the selection.

[0042] (Step ST6) If the information providing device 20 additionally includes the attribute determining unit 26, in step ST6, the attribute determining unit 26 determines the attribute of the user.

[0043] In this case, in step ST4, the selection unit 24 selects the content or method of the information to be provided based on the determined attribute in addition to the estimated intention.

[0044] (Step ST7) If the information providing device 20 additionally includes the environment determining unit 27, in step ST7, the environment determining unit 27 determines the environment of the user.

[0045] In this case, in step ST4, the selection unit 24 selects the content or method of information to be provided based on the estimated intention as well as the determined environment. In step ST4, the selection unit 24 may select the content or method of information to be provided based on the estimated intention, the determined attribute, and the determined environment.

[0046] <Operation; Details> <Example 1; Figures 4A and 4B> A detailed example of the operation will be described below with reference to Figures 4A and 4B. Figure 4A is a diagram showing an example of information displayed on display device 31. More specifically, Figure 4A shows a display example assuming that display device 31 is installed on the platform of "Citizens' Hall" station. Display device 31 displays train operation information and stop stations, as well as exits and nearby information for Civic Hall station.

[0047] The train operation information indicates that a train bound for "AB Street" will arrive in two minutes. The stops of the train bound for "AB Street" are "AB Street" station and "CD Square" station, which are indicated by solid lines in Figure 4A. "EF Street" station and "GH Square" station are stations located in the opposite direction from the direction of travel of the train bound for "AB Street", and the train bound for "AB Street" does not stop at these stations. The dashed lines in Figure 4A indicate that the train bound for "AB Street" does not stop at "EF Street" station or "GH Square" station.

[0048] 4A, it is shown that if you go in the direction of arrow A1, you will reach Exit E1, and that the "Citizens Hall" and "LM Tower" are located near Exit E1. Also, in the display example of Fig. 4A, it is shown that if you go in the direction of arrow A2, you will reach Exit E2, and that the "riverside" and "bus stop" are located near Exit E2.

[0049] 4B is a diagram showing the movement of the user's gaze tracked by the gaze tracking unit 21 in the display example of FIG. 4A. In FIG. 4B, the trajectory of the gaze gazing at the area where "EF Street" or "GH Square" is displayed is represented by LOS1, the trajectory of the gaze gazing at the area where the destination "AB Street" is displayed is represented by LOS2, and the trajectory of the gaze from the destination "AB Street" to the area where "EF Street" or "GH Square" is displayed is represented by LOS3. Note that the trajectory of LOS3 is indicated by an arrow, showing that the gaze is moving from the area where the destination "AB Street" is displayed toward the area where "EF Street" or "GH Square" is displayed.

[0050] The intention estimation unit 23 acquires these gaze trajectories LOS1 to LOS3 and the display information displayed at the ends of these gazes, i.e., at the positions of the trajectories LOS1 to LOS3. Based on the acquired information, when the intention estimation unit 23 detects a movement of the user's gaze going back and forth between the area where "EF Street" or "GH Square" is displayed and the area where the destination "AB Street" is displayed, the intention estimation unit 23 determines that the gaze movement is going back and forth between the same multiple areas. Note that the "back and forth movement" may be a movement that attempts to go back and forth, and does not necessarily have to be a movement that has already gone back and forth. For example, it may be determined that there is a "back and forth movement" when LOS2 is detected following LOS1, and then LOS3 is detected, and the gaze does not necessarily have to completely reach the area indicated by LOS1.

[0051] When such a determination is made, the intention estimation unit 23 analyzes the characteristics of the information displayed at the line of sight based on the information displayed at the line of sight acquired by the display information acquisition unit 22. In the example of FIG. 4B , the information displayed at the position of the trajectory LOS1 is "EF Street" or "GH Square," and the information displayed at the position of the trajectory LOS2 is "AB Street," and the information at the line of sight relates to a destination in the opposite direction. In this way, when the line of sight moves back and forth between multiple identical areas and the information at the line of sight relates to a destination in the opposite direction, the intention estimation unit 23 refers to the correspondence pattern table 1 in FIG. 8 and determines that the correspondence pattern corresponds to the correspondence pattern defined in the first row of the correspondence pattern table 1, and estimates that the station the user wants to go to is on the opposite line. The intention estimation unit 23 outputs the result of the estimation to the selection unit 24.

[0052] Upon receiving the estimation result estimated by the intention estimation unit 23, the selection unit 24 selects "track number information for the direction of the station the user wishes to visit" as information to be provided to the user by referring to the correspondence pattern table 1 in Fig. 8. The selection unit 24 outputs the selection result to the output control unit 25.

[0053] When the output control unit 25 receives the selection result selected by the selection unit 24, the output control unit 25 controls an output device such as the display device 31 or the speaker 32 to provide information according to the selection result. When "track number information for the direction of the station you want to go to" is selected, the output control unit 25 may control the display device 31 to display text information such as "Trains bound for GH Square are on the opposite line," or may control the speaker 32 to issue audio information such as "Trains bound for GH Square are on the opposite line."

[0054] <Example 2; Figs. 5A to 5C> A detailed example of the operation will be described below with reference to Figs. 5A to 5C. Fig. 5A shows a display example in which a display area DA2 is added to the display area DA1 in which the display example described with reference to Fig. 4A is displayed. An information button 41 and a guidance message 42 are displayed in the display area DA2. A touch sensor such as a touch panel is provided in the display area DA2 so that the information button 41 can be operated by touch. The guidance message 42 displays a message saying, "Please press this button if you have an inquiry."

[0055] Fig. 5B is a diagram showing the movement of the user's gaze tracked by the gaze tracking unit 21 in the display example of Fig. 5A. In Fig. 5B, the trajectory of the gaze gazing at the area where the information button 41 is displayed is represented by LOS4, the trajectory of the gaze gazing at the area where the guidance message 42 is displayed is represented by LOS5, and the trajectory of the gaze from the area where the guidance message 42 is displayed to the area where the information button 41 is displayed is represented by LOS6. Note that the trajectory of LOS6 is indicated by an arrow, showing that the gaze is moving from the area where the guidance message 42 is displayed to the area where the information button 41 is displayed.

[0056] The intention estimation unit 23 acquires these line-of-sight trajectories LOS4 to LOS6 and the display information displayed at the ends of these line-of-sight trajectories, i.e., at the positions of the trajectories LOS4 to LOS6. Based on the acquired information, when the intention estimation unit 23 detects that the user's line of sight moves back and forth between the area where the information button 41 is displayed and the area where the guidance message 42 is displayed, it determines that the line of sight moves back and forth between the same multiple areas.

[0057] When such a determination is made, the intention estimation unit 23 analyzes the characteristics of the information displayed at the gaze point based on the information displayed at the gaze point acquired by the display information acquisition unit 22. In the example of FIG. 5B , the information displayed at the position of the trajectory LOS4 is the information button 41, and the guidance message 42 displayed at the position of the trajectory LOS5 is the guidance message for the information button 41, and all of the information at the gaze point relates to the information button 41. In this way, when the gaze moves back and forth between multiple identical areas and the information at the gaze point relates to the same target, the information button, the intention estimation unit 23 refers to the correspondence pattern table 1 in FIG. 8 and determines that the information corresponds to the correspondence pattern defined in the second row of the correspondence pattern table 1, and estimates that "I don't know what the information button is." The intention estimation unit 23 outputs the result of the estimation to the selection unit 24.

[0058] As can be seen from the examples of Figures 4B and 5B, even if the gaze movement is the same movement of moving back and forth between the same multiple areas, the intention estimation unit 23 estimates the user's intention differently by taking into account what is displayed in front of the gaze.

[0059] Upon receiving the estimation result estimated by the intention estimation unit 23, the selection unit 24 selects "detailed explanation regarding the information button" as information to be provided to the user by referring to the correspondence pattern table 1 in Fig. 8. The selection unit 24 outputs the selection result to the output control unit 25.

[0060] When the output control unit 25 receives the selection result selected by the selection unit 24, it controls an output device such as the display device 31 or the speaker 32 to provide information according to the selection result. If "Detailed explanation about the information button" is selected, for example, as shown in FIG. 5C , the output control unit 25 may control the display device 31 to display a text message such as "Touch the i button to speak to a station staff member. Please contact us if you need assistance or in an emergency." The output control unit 25 may also control the speaker 32 to provide a similar message by voice.

[0061] <Example 3; Figures 6A and 6B> A detailed example of the operation will be described below with reference to Figures 6A and 6B. Figure 6A is a diagram showing the movement of the user's gaze tracked by the gaze tracking unit 21 in the display example of Figure 4A. In Figure 6A, the trajectory of the gaze gazing at the area where "2" is displayed is represented by LOS7, the trajectory of the gaze gazing at the area where "min" is displayed is represented by LOS8, and the trajectory of the gaze from the area where "min" is displayed to the area where "2" is displayed is represented by LOS9. Note that the trajectory of LOS9 is indicated by an arrow, indicating that the gaze is moving from the area where "min" is displayed toward the area where "2" is displayed.

[0062] The intention estimation unit 23 acquires these gaze trajectories LOS7 to LOS9 and the display information displayed at the ends of these gazes, i.e., at the positions of the trajectories LOS7 to LOS9. Based on the acquired information, when the intention estimation unit 23 detects that the user's gaze moves back and forth between the area where "2" is displayed and the area where "min" is displayed, it determines that the gaze movement is moving back and forth between the same multiple areas.

[0063] When such a determination is made, the intention estimation unit 23 analyzes the characteristics of the information displayed at the line of sight based on the information displayed at the line of sight acquired by the display information acquisition unit 22. In the example of FIG. 6A , the information displayed at the position of the trajectory LOS7 is “2,” and the information displayed at the position of the trajectory LOS8 is “min.” Both pieces of information at the line of sight relate to arrival times. In this way, when the line of sight moves back and forth between multiple identical regions, and the information at the line of sight relates to the same target, the arrival time, the intention estimation unit 23 refers to the correspondence pattern table 1 in FIG. 8 and determines that the information corresponds to the correspondence pattern defined in the third row of the correspondence pattern table 1, and estimates that the user “wants to see the remaining time in seconds, not in minutes.” The intention estimation unit 23 outputs the estimation result to the selection unit 24.

[0064] Upon receiving the estimation result estimated by the intention estimation unit 23, the selection unit 24 refers to the correspondence pattern table 1 in Fig. 8 and selects "display of a countdown ring" as information to be provided to the user. The selection unit 24 outputs the selection result to the output control unit 25.

[0065] When the selection result selected by the selection unit 24 is "display of a countdown ring," the output control unit 25 controls the display device 31 to provide information according to the selection result. For example, as shown in Fig. 6B, the output control unit 25 controls the display of the display device 31 to display a countdown ring. The display of this ring may change, for example, every second.

[0066] <Example 4; Figures 7A and 7B> A detailed example of the operation will be described below with reference to Figures 7A and 7B. Figure 7A is a diagram showing the movement of the user's gaze tracked by the gaze tracking unit 21 in the display example of Figure 4A. In Figure 7A, the trajectory of the gaze gazing at the area where "Citizens Hall" and "LM Tower" are displayed is represented by LOS10, the trajectory of the gaze gazing at the area where arrow A1 is displayed is represented by LOS11, and the trajectory of the gaze from the area where arrow A1 is displayed to the area where "Citizens Hall" and "LM Tower" are displayed is represented by LOS12. Note that the trajectory of LOS12 indicates the movement of the gaze from the area where arrow A1 is displayed toward the area where "Citizens Hall" and "LM Tower" are displayed.

[0067] The intention estimation unit 23 acquires the trajectories LOS10 to LOS12 of these lines of sight and the display information displayed at the ends of these lines of sight, i.e., at the positions of the trajectories LOS10 to LOS12. Based on the acquired information, when the intention estimation unit 23 detects that the user's line of sight moves back and forth between the area where "Citizens Hall" and "LM Tower" are displayed and the area where arrow A1 is displayed, it determines that the line of sight moves back and forth between the same multiple areas.

[0068] When such a determination is made, the intention estimation unit 23 analyzes the characteristics of the information displayed at the line of sight based on the information displayed at the line of sight acquired by the display information acquisition unit 22. In the example of FIG. 7A , the information displayed at the position of the trajectory LOS10 is "Citizens Hall" and "LM Tower," and the information displayed at the position of the trajectory LOS11 is an arrow A1. All of the information at the line of sight relates to an exit. In this way, when the line of sight moves back and forth between multiple identical areas, and the information at the line of sight relates to the same object, an exit, the intention estimation unit 23 refers to the correspondence pattern table 1 in FIG. 8 and determines that the information corresponds to the correspondence pattern defined in the fourth row of the correspondence pattern table 1, and infers that "the destination is not displayed, so it is unclear which exit the destination is." The intention estimation unit 23 outputs the result of the estimation to the selection unit 24.

[0069] Upon receiving the estimation result estimated by the intention estimation unit 23, the selection unit 24 selects “additional display related to destination” as information to be provided to the user by referring to the table in Fig. 8. The selection unit 24 outputs the selection result to the output control unit 25.

[0070] When the output control unit 25 receives the selection result selected by the selection unit 24, it controls an output device such as the display device 31 or the speaker 32 to provide information according to the selection result. When "additional display related to destination" is selected, the output control unit 25 controls the display of the display device 31 to display "MN Hotel" and "taxi stand", for example, as shown in Fig. 7B. The specific content of the additional display may be determined based on information about "Citizens Hall", "LM Tower", or the position where the arrow A1 is displayed.

[0071] <Consideration of Attributes or Environment; Fig. 9> The selection unit 24 may determine the content of information to be provided or the method of providing information based on the attributes or environment of the user. This point will be described with reference to Fig. 9. As described above, Fig. 9 is a diagram showing an example of the correspondence pattern table 2.

[0072] The selection unit 24 may determine a method of providing information by referring to the correspondence pattern table 2 based on the attributes of the user (target person) determined by the attribute determination unit 26 or the environment of the target person determined by the environment determination unit 27.

[0073] As an example, if the determination result by the attribute determination unit 26 indicates that the user is "hearing impaired," the selection unit 24 may refer to the corresponding pattern table 2 based on the determination result and determine that the information should be "displayed in text on the screen."

[0074] As an example, if the judgment result by the environment judgment unit 27 indicates that "surrounding sounds are low," the selection unit 24 may refer to the corresponding pattern table 2 based on the judgment result and determine that the information should be "transmitted as audio information from broadcasting equipment near the person."

[0075] As an example, if the judgment result by the environment judgment unit 27 indicates that "surrounding noise is loud," the selection unit 24 may refer to the corresponding pattern table 2 based on the judgment result and determine that the information should be "displayed in text on the screen."

[0076] As an example, if the determination result by the attribute determination unit 26 indicates "wheelchair user" and the determination result by the environment determination unit 27 indicates that there is "one user," the selection unit 24 may determine, based on these determination results, to "display a wheelchair route button" by referring to the correspondence pattern table 2. Note that the "wheelchair route button" refers to a button that indicates a route that can be traveled by wheelchair.

[0077] As an example, if the determination result by the attribute determination unit 26 indicates a "wheelchair user" and the determination result by the environment determination unit 27 indicates that there is "one user" and the weather is "rainy," the selection unit 24 may refer to the corresponding pattern table 2 based on these determination results and determine to "display the boarding location and contact information for secondary transportation (taxi, bus)."

[0078] As an example, if the determination result by the attribute determination unit 26 indicates that the user is "visually impaired," the selection unit 24 may determine, based on the determination result, to "broadcast using left and right directional speakers what is on the right or left side." In this case, the selection unit 24 may determine, along with or instead of this determination, to "broadcast an announcement informing the user of the presence of signage."

[0079] The information providing device 20 or the information providing system Sys. described above detects the user's line of sight, infers the user's intention based on the detected line of sight, and provides information according to the inferred intention. Therefore, it is possible to provide information without requiring special tools such as paper or a mobile terminal as in the prior art.

[0080] Furthermore, the information providing device 20 or the information providing system Sys. estimates the user's intention based on what information is displayed in the detected line of sight of the user, and thus can accurately estimate the user's intention. When making the estimation, the accuracy of the intention estimation can be improved by acquiring and analyzing the display information displayed on the display device 31 at the time the image of the user is captured.

[0081] Furthermore, the information providing device 20 or the information providing system Sys. may select the content of the information to be provided or the method of providing the information, taking into consideration the attributes or environment of the user. By taking into consideration the attributes or environment, it becomes possible to provide accurate information.

[0082] It is possible to combine the embodiments, and to modify or omit each embodiment as appropriate.

[0083] The information providing device or information providing system of the present disclosure can be used as a digital signage system that presents information to station users and the like.

[0084] 11 camera, 12 microphone, 13 storage device, 14 storage device, 15 storage device, 20 information providing device, 21 gaze tracking unit, 22 display information acquisition unit, 23 intention estimation unit, 24 selection unit, 25 output control unit, 26 attribute determination unit, 27 environment determination unit, 31 display device, 32 speaker, 41 information button, 42 guidance message, 100a processing circuit, 100b processor, 100c memory.

Claims

1. An information providing device comprising: an intention estimation unit that estimates the intention of a person based on multiple images of the person and display information displayed on a display device at the time the multiple images were captured; a selection unit that selects the content or method of information to be provided in accordance with the estimated intention; and an output control unit that controls an output device to provide information in accordance with the result of the selection.

2. An information providing device as described in claim 1, further comprising: an eye tracking unit that detects the person's eye gaze, tracks the detected eye gaze trajectory, and analyzes the characteristics of the eye gaze movement; and a display information acquisition unit that acquires the display information and analyzes the characteristics of the acquired display information, wherein the intention estimation unit estimates the person's intention based on the analyzed eye gaze movement characteristics and the analyzed display information characteristics by referring to a pre-created correspondence pattern that associates eye gaze movement, display information, and the user's intention.

3. An information providing device according to claim 2, wherein the eye movement is characterized by a movement back and forth between a plurality of areas on the display device.

4. An information providing device according to any one of claims 1 to 3, further comprising an attribute determination unit that determines an attribute of the person, wherein the selection unit makes the selection further based on the determined attribute.

5. An information providing device according to any one of claims 1 to 4, further comprising an environment determination unit that determines the environment of the person, wherein the selection unit makes the selection further based on the determined environment.

6. An information provision system comprising: an information provision device according to any one of claims 1 to 5; a camera that captures the plurality of images; and the display device.

7. The information providing system according to claim 6, wherein the output device is the display device.

8. An information provision method performed by an information provision device comprising an eye-gaze tracking unit, a display information acquisition unit, an intention estimation unit, a selection unit, and an output control unit, comprising: a step in which the eye-gaze tracking unit detects the eye gaze of a person from multiple images capturing the person, tracks the detected eye gaze trajectory, and analyzes characteristics of the eye gaze movement; a step in which the display information acquisition unit acquires display information that is displayed in the direction of the tracked eye gaze at the time the multiple images were captured, and analyzes characteristics of the acquired display information; a step in which the intention estimation unit estimates the person's intention based on the analyzed eye gaze movement characteristics and the analyzed display information characteristics, by referring to a pre-created correspondence pattern that associates eye gaze movement, display information, and the user's intention; a step in which the selection unit selects information to be provided in accordance with the estimated intention; and a step in which the output control unit controls the output device to provide the selected information in accordance with the selection result.

9. An information provision program that causes a computer to execute the following functions: a function to detect a person's gaze from multiple images of the person, track the detected gaze trajectory, and analyze the characteristics of the gaze movement; a function to acquire display information that is displayed in front of the tracked gaze at the time the multiple images were captured, and analyze the characteristics of the acquired display information; a function to infer the person's intention by referring to a pre-created correspondence pattern that associates gaze movement, display information, and the user's intention based on the analyzed gaze movement characteristics and the analyzed display information characteristics; a function to select information to be provided according to the inferred intention; and a function to control an output device to provide the selected information according to the selection results.

Citation Information

Patent Citations

  • Recognition evaluation system and method for advertisement

    JP2007299023A

  • Method and apparatus for controlling image display in response to viewer factors and reactions

    JP2012532366A

  • Advertisement linkage service providing system and advertisement linkage service providing method

    JP2020166768A

  • Advertisement interest level evaluation system and advertisement interest level evaluation method

    JP2022087982A