Information provision system, method, and program
The information provision system dynamically adjusts audio output based on user intentions, addressing the limitations of existing systems by providing personalized and interactive information delivery.
Patent Information
- Application Number
- JP2022021703
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-02-16
AI Technical Summary
Existing product display systems fail to consider customer intentions, providing unwanted information and requiring customers to wait for desired information, with no option to re-listen to missed parts.
An information provision system that uses position and gaze direction acquisition, head movement detection, and intention estimation to dynamically adjust audio output based on user preferences, including selection of explanatory information type, volume, and continuation of output.
Enables personalized audio information delivery, allowing users to engage interactively and receive relevant information at appropriate volumes and formats, enhancing user experience.
Smart Images

Figure 0007821430000001 
Figure 0007821430000002 
Figure 0007821430000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information providing system, method, and program. [Background technology]
[0002] Patent Document 1 describes a product display shelf equipped with a CD player and a speaker that provides customers with information explaining products. In this product display shelf, CDs on which descriptions of each of the multiple products displayed are recorded are played by the CD player, and the played audio is output from the speaker. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 8-160897 Summary of the Invention [Problem to be solved by the invention]
[0004] In the display shelf described in Patent Document 1, descriptions of multiple products are played in a predetermined order. If a customer approaches a display shelf and a description of a product that the customer is not interested in is being played, the customer will be provided with information that the customer does not want. Furthermore, if the customer wants to hear a description of a product that interests them, the customer must wait near the display shelf for a while. Furthermore, because the product descriptions are only played in a predetermined order, even if the customer misses part of the description, they cannot immediately listen to that part again.
[0005] As described above, the configuration according to Patent Document 1 has a problem in that it is not possible to provide information by voice in consideration of the customer's intentions. [Means for solving the problem]
[0006] The present disclosure can be realized in the following forms.
[0007] (1) According to an embodiment of the present disclosure, an information provision system is provided. The information provision system provides information through sound. The information provision system includes: a position / direction acquisition unit that acquires position information indicating a user's location and gaze direction information indicating a gaze direction, i.e., the direction in which the user's face is facing; a storage unit that stores target position information indicating the positions of multiple targets that the user may be viewing; explanatory information describing each target; and information indicating settings related to information provision; a target estimation unit that estimates the target the user is viewing based on the position information, gaze direction information, and target position information; an information output unit that outputs explanatory information about the estimated target by sound in accordance with the settings related to information provision; a head movement detection unit that detects a user's head movement; and an intention estimation unit that estimates the user's intention from the user's head movement while the explanatory information is being output and selects settings related to information provision in accordance with the user's intention. When the settings related to information provision are changed, the information output unit outputs the explanatory information in accordance with the changed settings related to information provision. According to this embodiment, settings related to information provision are selected in accordance with the user's intention estimated while the explanatory information is being output. The explanatory information is provided to the user in accordance with the settings related to information provision. Therefore, the settings related to information provision can be dynamically changed in accordance with the user's intention. This makes it possible to provide audio information taking the user's intention into consideration. (2) In the information provision system of the above aspect, the storage unit may store, as explanatory information for explaining the object, first explanatory information that is an explanation about the object and second explanatory information that is an explanation about the object that is different from the first explanatory information. The setting related to information provision may include information indicating which of the first explanatory information and the second explanatory information is selected as the explanatory information. According to this embodiment, either the first explanatory information or the second explanatory information different from the first explanatory information is selected in accordance with the user's intention estimated while the explanatory information is being output. Thus, audio information can be provided taking the user's intention into consideration. (3) In the information provision system of the above form, the memory unit further stores third explanatory information, which is an explanation about the object that is different from the first explanatory information and the second explanatory information, as explanatory information that explains the object, the first explanatory information is a normal explanation about the object, the second explanatory information is a more detailed explanation than the first explanatory information, and the third explanatory information is a simpler explanation than the first explanatory information, and the settings regarding information provision may include information indicating which of the first explanatory information, the second explanatory information, and the third explanatory information has been selected as the explanatory information. According to this embodiment, a normal explanation, a detailed explanation, or a simple explanation is selected according to the user's intention estimated while the explanation information is being output. Therefore, for example, if it is estimated that the user desires a simple explanation while the normal explanation is being output, the audio output is switched to the simple explanation. In this way, audio information can be provided taking the user's intention into consideration. (4) In the information providing system of the above aspect, the settings related to information provision may include setting information related to audio output. According to this embodiment, settings related to audio output are selected according to the user's intention estimated while the explanatory information is being output. For example, if it is estimated that the user finds it difficult to hear, the settings are changed to increase the volume. Therefore, since the volume is increased while the explanatory information is being output, the user can hear the explanatory information at a volume that is easy to hear. In this way, audio information can be provided taking into account the user's intention. (5) In the information providing system of the above aspect, the settings related to information provision may include information indicating whether or not to continue outputting the explanatory information. According to this embodiment, whether or not to continue outputting the explanatory information is selected according to the user's intention estimated while the explanatory information is being output. For example, if it is estimated that the user does not need the explanatory information to be output, the setting is changed to not continue outputting the explanatory information. Thus, explanatory information that the user does not want to be provided to the user is not provided to the user. (6) In the information providing system of the above aspect, the information output unit may output a question to the user by voice, and the intention estimation unit may estimate an answer to the user's question from a movement of the user's head. According to this embodiment, it is possible to provide a participatory information providing system in which the user can receive information while participating in the process, rather than simply passively receiving the information. (7) In the information providing system of the above aspect, the target includes a moving object. The target estimation unit may estimate that the moving object is the target being viewed by the user when the moving object remains within a range visible to the user's eyes for a predetermined period of time. According to this embodiment, explanatory information about not only stationary objects but also moving objects can be provided to the user. (8) The information providing system of the above aspect may further include a sound source position acquisition unit that acquires a virtual position of a sound source corresponding to each target. The information output unit may output, to a portable audio output device worn on the user's head, audio that has been subjected to stereophonic processing on audio representing the explanatory information, according to the virtual position of the sound source as seen from the user's current position. According to this embodiment, it is possible to provide the user with information about the object being viewed while giving the user a sense of realism. (9) In the information providing system of the above aspect, the storage unit may store intention definition data in which non-verbal actions based on the culture to which the language used by the user belongs are defined. The intention estimation unit may estimate the intention of the user from the intention definition data and the user's head movement. According to this embodiment, even if users use different languages, it is possible to infer the intention of the users from the movement of their heads. (10) In the information provision system of the above form, the intention estimation unit may estimate the user's intention by inputting parameters representing the user's head movement, the user's movement speed, the distance between the user and the object, and the user's relative angle with respect to the object into a trained machine learning model. According to this embodiment, the user's intention can be estimated with high accuracy. The present disclosure can be realized in various forms other than an information providing system, for example, a method in which a computer carried by a user provides information by voice, and a computer program for implementing the method. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a diagram showing a schematic configuration of an information providing system according to an embodiment; [Figure 2] FIG. 10 is a diagram for explaining a method of expressing the movement of the user's head by a rotation angle. [Figure 3] FIG. 1 is a diagram showing the positional relationship between a user and a virtually placed sound source. [Figure 4] 10 is a flowchart of an information providing process. [Figure 5] 10 is a flowchart of an explanation information output process. [Figure 6] 10 is a flowchart of a motion detection process. [Figure 7] 10 is a flowchart of an intention estimation process. DETAILED DESCRIPTION OF THE INVENTION
[0009] A. Embodiment FIG. 1 is a diagram showing the configuration of an information provision system 1000 according to an embodiment. The information provision system 1000 provides the user with explanatory information describing an object that the user is viewing by voice. The information provision system 1000 also provides information according to the estimated intention of the user. In this embodiment, an example will be described in which the information provision system 1000 provides information about tourist spots to a user visiting a tourist spot. The information provision system 1000 includes a mobile terminal 100 and earphones 200.
[0010] The mobile terminal 100 is a communication terminal carried by a user. In this embodiment, the mobile terminal 100 is a smartphone owned by the user. It is assumed that application software for providing the user with information about tourist spots is installed on the mobile terminal 100. Hereinafter, this application software will be referred to as a guidance application. By executing the guidance application, the user can receive information about tourist spots from the information providing system 1000. It is assumed that the user travels around tourist spots while carrying the mobile terminal 100. The guidance application has a function of estimating the user's current location and an object the user is viewing, and providing the user with information about tourist spots. The mobile terminal 100 will also be referred to as a computer carried by the user.
[0011] The earphones 200 are portable audio output devices that are worn on the head of a user. The earphones 200 are portable audio output devices that output audio representing signals received from the mobile terminal 100. In this embodiment, the earphones 200 are wireless earphones that the user owns. The user wears the earphones 200 in their ears and travels around tourist spots.
[0012] The mobile terminal 100 has, as its hardware configuration, a CPU (Central Processing Unit) 101, a memory 102, and a communication unit 103. The memory 102 and the communication unit 103 are connected to the CPU 101 via an internal bus 109.
[0013] The CPU 101 executes various programs stored in the memory 102 to realize various functions of the mobile terminal 100. The memory 102 stores the programs executed by the CPU 101 and various data used to execute the programs. The memory 102 is also used as a work memory for the CPU 101.
[0014] The communication unit 103 includes a network interface circuit and communicates with external devices under the control of the CPU 101. In this embodiment, the communication unit 103 is assumed to be able to communicate with external devices in accordance with the Wi-Fi (registered trademark) communication standard. Furthermore, the communication unit 103 includes a GNSS (Global Navigation Satellite System) receiver and receives signals from positioning satellites under the control of the CPU 101. In the information providing system 1000, the GPS (Global Positioning System) is used as the GNSS.
[0015] The earphones 200 output sound representing a signal supplied from the mobile terminal 100. The earphones 200 include a DSP (Digital Signal Processor) 201, a communication unit 202, a sensor 203, and a driver unit 204. The communication unit 202, the sensor 203, and the driver unit 204 are connected to the DSP 201 via an internal bus 209.
[0016] The DSP 201 controls the communication unit 202, the sensor 203, and the driver unit 204. The DSP 201 outputs an audio signal received from the mobile terminal 100 to the driver unit 204. Furthermore, every time a measurement value is supplied from the sensor 203, the DSP 201 transmits the measurement value to the mobile terminal 100. The communication unit 202 includes a network interface circuit, and communicates with an external device under the control of the DSP 201. The communication unit 202 wirelessly communicates with the mobile terminal 100 in accordance with, for example, the Bluetooth (registered trademark) standard.
[0017] The sensor 203 includes an acceleration sensor, an angle sensor, and an angular velocity sensor. For example, a triaxial acceleration sensor is used as the acceleration sensor. A triaxial angular velocity sensor is used as the angular velocity sensor. The sensor 203 performs measurements at predetermined time intervals and outputs the measured acceleration and angular velocity values to the DSP 201. The driver unit 204 converts the audio signal supplied from the DSP 201 into sound waves and outputs them.
[0018] The mobile terminal 100 functionally comprises a storage unit 110, a position and direction acquisition unit 120, an object estimation unit 130, a head movement detection unit 140, an intention estimation unit 150, and an information output unit 160.
[0019] The storage unit 110 stores location information of places that the user may visit, such as location coordinates representing the locations of art museums, parks, observation decks, etc. The location information of places that the user may visit is also referred to as location location information. The storage unit 110 also stores location coordinates representing the locations of exhibits in art museums, for example, as location information of objects that may be visually recognized by the user. The location information of objects that may be visually recognized by the user is also referred to as object location information. Furthermore, the storage unit 110 stores sound source data having an audio signal that reads out information describing exhibits in art museums, for example, as explanatory information describing objects that may be visually recognized by the user. Furthermore, the storage unit 110 stores information representing the location where a sound source, which will be described later, is virtually located for each object that may be visually recognized.
[0020] The storage unit 110 also stores intention definition data that associates the user's head movements with the user's intentions. Examples of the association between the user's head movements and intentions defined in the intention definition data are described below. The user's head tilting movement indicates that the user feels that they do not understand. The user's repeated head tilting movement indicates that the user feels that they cannot hear well. The user's nodding movement indicates that the user has positive emotions. The user's head shaking movement indicates that the user has negative emotions. The user's repeated head shaking movement indicates that the user has more negative emotions.
[0021] The storage unit 110 stores setting data representing settings related to information provision. The settings related to information provision represent settings for outputting explanatory information by audio. In an embodiment, the settings related to information provision include information indicating a selection of a type of explanatory information, information indicating a volume of the audio for outputting the explanatory information, information indicating whether or not to perform frame rewinding of the explanatory information, and information indicating whether or not to continue outputting the explanatory information.
[0022] In the information providing system 1000, the explanatory information provided to the user is one of three types of explanatory information: normal explanatory information, detailed explanatory information, and simplified explanatory information. For example, assume that explanatory information about target T1 is provided to the user. The normal explanatory information is information that explains target T1 that is normally planned to be provided to the user. The detailed explanatory information is information that explains target T1 in more detail than the normal explanatory information. The simplified explanatory information is information that explains target T1 more simply than the normal explanatory information. The normal explanatory information is also referred to as first explanatory information. The detailed explanatory information is also referred to as second explanatory information, and the simplified explanatory information is also referred to as third explanatory information. The detailed explanatory information is also referred to as third explanatory information, and the simplified explanatory information is also referred to as second explanatory information. The information indicating the selection of the type of explanatory information indicates whether the normal explanatory information, the detailed explanatory information, or the simplified explanatory information has been selected.
[0023] The information indicating the volume of the audio for outputting explanatory information represents the volume of the audio output from the earphone 200. The setting for whether to perform frame rewind of explanatory information refers to whether to perform frame rewind of a portion of the explanatory information that was previously output as audio. Frame rewind means to output a portion of the explanatory information that was output as audio again. The information indicating whether to continue outputting explanatory information indicates whether to continue outputting the audio of explanatory information or to stop it midway. The information indicating the volume of the audio for outputting explanatory information is also called setting information related to audio output.
[0024] The functions of storage unit 110 are realized by memory 102. The location location information, target location information, explanation information, and information representing the location of the sound source are assumed to be stored in memory 102 as part of the data for executing the guidance application when the guidance application is installed in mobile terminal 100.
[0025] The position / direction acquiring unit 120 acquires information indicating the current position of the mobile terminal 100 as information indicating the current position of the user. Furthermore, the position / direction acquiring unit 120 acquires information indicating the line of sight of the user from the measurement value by the sensor 203. The function of the position / direction acquiring unit 120 is realized by the CPU 101.
[0026] The object estimation unit 130 estimates the object that the user is gazing at. A method for estimating the object that the user is gazing at will be described later. The function of the object estimation unit 130 is realized by the CPU 101.
[0027] FIG. 2 is a diagram illustrating a method for detecting a user's head movement. Head movement detection unit 140 detects the head movement of a user wearing earphones 200. In this embodiment, the user's head movement is expressed as a rotation angle. The rotation axis along the front-to-back direction of the user is defined as the roll axis, the rotation axis along the left-to-right direction of the user is defined as the pitch axis, and the rotation axis along the direction of gravity is defined as the yaw axis. The movement of the user tilting their head can be expressed as a rotation around the roll axis. The movement of the user nodding can be expressed as a rotation around the pitch axis. The movement of the user turning their head can be expressed as a rotation around the yaw axis.
[0028] Hereinafter, the amount of rotational angle displacement around the roll axis will be referred to as the roll angle, the amount of angular displacement around the pitch axis as the pitch angle, and the amount of angular displacement around the yaw axis as the yaw angle. The movement of the user's head is represented by the roll angle, pitch angle, and yaw angle. The roll angle ranges from +30 degrees to -30 degrees, assuming that 0 degrees is when the user is facing forward. The pitch angle ranges from +45 degrees to -45 degrees, assuming that 0 degrees is when the user is facing forward. The yaw angle ranges from +60 degrees to -60 degrees, assuming that 0 degrees is when the user is facing forward.
[0029] Head movement detection unit 140 detects the roll angle, pitch angle, and yaw angle from the acceleration measurement values and angular velocity measurement values measured by sensor 203. Head movement detection unit 140 supplies information representing the detection results for the roll angle, pitch angle, and yaw angle to intention estimation unit 150. The functions of head movement detection unit 140 are realized by CPU 101.
[0030] The intention estimation unit 150 identifies the user's head movement from the roll angle, pitch angle, and yaw angle detected by the head movement detection unit 140. Then, the intention estimation unit 150 estimates the user's intention from the identified user's head movement and the intention definition data. Furthermore, the intention estimation unit 150 selects settings related to information provision according to the estimated user's intention. Note that there are cases where the settings related to information provision are not changed according to the estimated user's intention. In such cases, the intention estimation unit 150 selects to maintain the current settings. The functions of the intention estimation unit 150 are realized by the CPU 101.
[0031] When the object estimation unit 130 estimates the user's visual target, the information output unit 160 causes the earphones 200 to output explanatory information explaining the estimated target in accordance with the settings related to information provision stored in the storage unit 110. Specifically, the information output unit 160 causes the earphones 200 to output explanatory information of the selected type at a volume specified in the settings related to information provision.
[0032] Assume that after the output of the explanatory information has started, the settings related to information provision are changed in accordance with the estimated user's intention. In this case, the information output unit 160 causes the earphone 200 to output the explanatory information in accordance with the settings related to information provision after the change.
[0033] FIG. 3 is a diagram showing the positional relationship between a user P and a virtually placed sound source SS. FIG. 3 shows the user P and the sound source SS as viewed from above. In the embodiment, the information output unit 160 outputs audio from the earphones 200 that reads out explanatory information using stereophonic sound. The position of the sound source SS is set to the same position as the visual target. First, the information output unit 160 reads information on the virtually placed position of the sound source SS for the estimated visual target from the storage unit 110. The information output unit 160 acquires the virtual position of the sound source by reading information on the virtually placed position of the sound source for the visual target from the storage unit 110. The information output unit 160 is also referred to as a sound source position acquisition unit.
[0034] Furthermore, the information output unit 160 calculates the relative angle of the direction in which the sound source SS is located as seen from the user P with respect to the line of sight D of the user P. In the horizontal plane, the magnitude of the angle that the line of sight D forms with respect to a reference direction N is angle r1. The reference direction N is, for example, a direction facing north. The magnitude of the angle that the direction in which the sound source SS is located as seen from the user P with respect to the reference direction N is angle r2. The information output unit 160 calculates the angle r1 from the line of sight D and the reference direction N. The information output unit 160 calculates the angle r2 from the position of the sound source SS, the position of the user P, and the reference direction N. The information output unit 160 calculates angle r3, which is the difference between angle r1 and angle r2, as the relative angle of the direction in which the sound source SS is located with respect to the line of sight D of the user P.
[0035] Next, the information output unit 160 calculates the distance between the user P and the sound source SS from the positions of the user P and the sound source SS. Based on the calculated angle and distance, the information output unit 160 outputs audio that has been subjected to stereophonic processing to the earphones 200. For example, an existing algorithm for generating stereophonic sound is used for the stereophonic processing. The functions of the information output unit 160 are implemented by the CPU 101.
[0036] For example, suppose the center of a painting on display in a museum is set as the location of a virtual sound source. In this case, a user looking at the painting can feel as if the audio of explanatory information is being output from the center of the painting. In this way, this embodiment can provide the user with information about the object they are viewing while giving them a sense of realism.
[0037] 4 is a flowchart of an information provision process in which the information provision system 1000 provides information to a user via the mobile terminal 100. The information provision process is started at a predetermined time interval. The predetermined time interval is, for example, 0.5 seconds. Even if the predetermined time has elapsed, if the information provision process started immediately before has not ended on the same mobile terminal 100, a new information provision process will not be started. Furthermore, at the time the information provision process is started, the information indicating the settings related to information provision stored in the storage unit 110 will be the initial setting information.
[0038] In step S10, the position and direction acquisition unit 120 acquires position information of the mobile terminal 100. Specifically, first, the position and direction acquisition unit 120 acquires position coordinates indicating the current position of the mobile terminal 100 based on a GPS signal received from a GPS satellite. If the position and direction acquisition unit 120 is unable to receive a GPS signal, it acquires position coordinates indicating the current position of the mobile terminal 100 based on the radio wave intensities received from multiple Wi-Fi (registered trademark) base stations. The position and direction acquisition unit 120 supplies the position coordinates of the mobile terminal 100 to the target estimation unit 130.
[0039] In step S20, the position / direction acquisition unit 120 identifies the user's line of sight. The position / direction acquisition unit 120 determines whether the user is gazing at something based on the acceleration measurement value and the angular velocity measurement value measured by the sensor 203. For example, the position / direction acquisition unit 120 determines that the user is gazing at something when the acceleration measurement value and the angular velocity measurement value satisfy predetermined conditions. If it is determined that the user is gazing at something, the position / direction acquisition unit 120 identifies the direction in which the user's face is facing based on the acceleration and angular velocity.
[0040] The direction in which the user's face is facing can be expressed by an azimuth angle and an elevation angle or depression angle. Here, the azimuth angle refers to the angle that the direction in which the user's face is facing makes with respect to a reference direction. The elevation angle refers to the angle that the direction of the user's line of sight when looking at an object above makes with respect to the horizontal plane. The depression angle refers to the angle that the direction of the user's line of sight when looking at an object below makes with respect to the horizontal plane. In this embodiment, the direction in which the user's face is facing is referred to as the user's line of sight. Information indicating the user's line of sight is also referred to as line of sight direction information. The position and direction acquisition unit 120 supplies line of sight direction information indicating the user's line of sight to the object estimation unit 130.
[0041] On the other hand, when the position / direction acquiring unit 120 determines that the user is not gazing at anything, it notifies the object estimating unit 130 that the line of sight direction cannot be identified.
[0042] In step S30, the object estimation unit 130 determines whether or not there is an object being gazed at by the user. Specifically, first, the object estimation unit 130 reads, from the storage unit 110, position information about objects within a preset range centered on the user's current position indicated by the position information supplied from the position and direction acquisition unit 120, as information about candidate gazed objects. The object estimation unit 130 determines whether any of the candidate gazed objects is within the user's field of view, based on the position information of the objects within the preset range and the position information and gaze direction information supplied from the position and direction acquisition unit 120. It is assumed that ranges are set in advance for each of the azimuth angle, elevation angle, and depression angle as the range of the user's field of view.
[0043] For example, it is assumed that the object estimation unit 130 determines that an object T1 is in the user's field of view. In this case, the object estimation unit 130 determines whether the state in which the object T1 is in the user's field of view continues for a preset period. The preset period is, for example, one second. When the state in which the object T1 is in the user's field of view continues for a preset period, the object estimation unit 130 determines that the user is viewing the object T1. If it is determined that a viewed object exists (step S30; YES), the object estimation unit 130 supplies information indicating the determined object to the information output unit 160.
[0044] On the other hand, if the object estimation unit 130 determines that the gaze target cannot be estimated (step S30; NO), the information provision process is terminated. For example, if the object estimation unit 130 is notified by the position / direction acquisition unit 120 that the user's gaze direction cannot be identified, the object estimation unit 130 determines that the gaze target cannot be estimated. Furthermore, if the state in which the object T1 is in the user's field of view has not continued for a predetermined period, the object estimation unit 130 determines that the gaze target cannot be estimated. Furthermore, if there is no object that can be the gaze target within a predetermined range centered on the user's current position, the object estimation unit 130 determines that the gaze target cannot be estimated.
[0045] In step S40, an explanation information output process is executed to output explanation information of the estimated target by voice, after which the process shown in FIG.
[0046] Fig. 5 is a flowchart of the explanation information output process in step S40 in Fig. 4. In step S41, the information output unit 160 reads out setting data related to information provision stored in the storage unit 110.
[0047] In step S42, the information output unit 160 reads out explanatory information about the estimated visually recognized object from the storage unit 110, and starts outputting the explanatory information via the earphones 200 as audio.
[0048] In step S43, the information output unit 160 determines whether the explanatory information has been output to the end. If the explanatory information has not been output to the end (step S43; NO), the process of step S44 is executed. On the other hand, if the explanatory information has been output to the end (step S43; YES), the explanatory information output process is terminated.
[0049] In step S44, a movement detection process is executed by the head movement detection section 140. In the movement detection process, the movement of the user's head for a preset period is detected.
[0050] In step S45, an intention estimation process is executed by the intention estimation unit 150. In the intention estimation process, the intention of the user is estimated from the movement of the user's head. Furthermore, settings related to information provision are selected according to the intention of the user.
[0051] In step S46, the information output unit 160 determines whether or not the setting data related to information provision has been updated based on the notification from the intention estimation unit 150. If the setting data related to information provision has been updated (step S46; YES), the information output unit 160 executes the process of step S47. On the other hand, if the setting data related to information provision has not been updated (step S46; NO), the process of step S43 is executed.
[0052] In step S47, the information output unit 160 suspends the output of the explanatory information. In step S48, the information output unit 160 reads the setting data related to information provision from the storage unit 110. In step S49, the information output unit 160 resumes output of the explanatory information in accordance with the updated setting data related to information provision. Thereafter, the process of step S43 is executed again.
[0053] FIG. 6 is a flowchart of the movement detection process shown in step S44 of FIG. 5. In step S101, head movement detection unit 140 starts a timer to begin measuring time. In this embodiment, to estimate the user's intention, the user's head movement is observed for a set period of time. The set period is, for example, 0.5 seconds. The timer is used to measure the set period.
[0054] In step S102, head movement detection unit 140 acquires the roll angle, pitch angle, and yaw angle that represent the movement of the user's head. Specifically, head movement detection unit 140 calculates the roll angle, pitch angle, and yaw angle that represent the movement of the user's head from the acceleration measurement value and angular velocity measurement value measured by sensor 203.
[0055] In step S103, head motion detection unit 140 determines whether rotation about the roll axis has been detected. For example, head motion detection unit 140 determines that rotation about the roll axis has been detected when the roll angle is equal to or greater than a predetermined rotation angle. If head motion detection unit 140 detects rotation about the roll axis (step S103; YES), it executes the process of step S106. On the other hand, if head motion detection unit 140 determines in step S103 that rotation about the roll axis has not been detected (step S103; NO), it executes the process of step S104.
[0056] In step S104, head motion detection unit 140 determines whether rotation about the yaw axis has been detected. For example, if the yaw angle is equal to or greater than a predetermined rotation angle, head motion detection unit 140 determines that rotation about the yaw axis has been detected. If head motion detection unit 140 detects rotation about the yaw axis (step S104; YES), it executes the process of step S107. On the other hand, if head motion detection unit 140 determines in step S104 that rotation about the yaw axis has not been detected (step S104; NO), it executes the process of step S105.
[0057] In step S105, head motion detection unit 140 determines whether rotation around the pitch axis has been detected. For example, if the pitch angle is equal to or greater than a predetermined rotation angle, head motion detection unit 140 determines that rotation around the pitch axis has been detected. If head motion detection unit 140 detects rotation around the pitch axis (step S105; YES), it executes the process of step S108. On the other hand, if head motion detection unit 140 determines in step S105 that rotation around the pitch axis has not been detected (step S105; NO), it executes the process of step S109.
[0058] In step S106, head motion detection unit 140 increments roll axis counter Cr by 1. Head motion detection unit 140 also resets yaw axis counter Cy and pitch axis counter Cp. After that, head motion detection unit 140 executes the process of step S109. Roll axis counter Cr is a counter that indicates the number of times rotation around the roll axis has been detected. Yaw axis counter Cy is a counter that indicates the number of times rotation around the yaw axis has been detected. Pitch axis counter Cp is a counter that indicates the number of times rotation around the pitch axis has been detected.
[0059] In step S107, head motion detection unit 140 increments yaw axis counter Cy by 1. Head motion detection unit 140 also resets roll axis counter Cr and pitch axis counter Cp. After that, head motion detection unit 140 performs the process of step S109.
[0060] In step S108, head motion detection unit 140 increments pitch axis counter Cp by 1. Head motion detection unit 140 also resets roll axis counter Cr and yaw axis counter Cy. After that, head motion detection unit 140 performs the process of step S109.
[0061] In step S109, head motion detection unit 140 determines whether a preset time has elapsed since the timer was started. If the preset time has elapsed (step S109; YES), head motion detection unit 140 stops the timer and ends the motion detection process. On the other hand, if the preset time has not elapsed (step S109; NO), the process of step S102 is executed again.
[0062] Fig. 7 is a flowchart of the intention estimation process in step S45 of Fig. 5. In step S201, the intention estimation unit 150 determines whether the value of the counter Cr for the roll axis is 1 or greater. If the value of the counter Cr for the roll axis is 1 or greater (step S201; YES), the intention estimation unit 150 executes the process of step S205. On the other hand, if the value of the counter Cr for the roll axis is not 1 or greater (step S201; NO), the intention estimation unit 150 executes the process of step S202.
[0063] In step S202, the intention estimation unit 150 determines whether the value of the yaw axis counter Cy is equal to or greater than 1. If the value of the yaw axis counter Cy is equal to or greater than 1 (step S202; YES), the intention estimation unit 150 executes the process of step S208. On the other hand, if the value of the yaw axis counter Cy is not equal to or greater than 1 (step S202; NO), the intention estimation unit 150 executes the process of step S203.
[0064] In step S203, the intention estimation unit 150 determines whether the value of the counter Cp for the pitch axis is equal to or greater than 1. If the value of the counter Cp for the pitch axis is equal to or greater than 1 (step S203; YES), the intention estimation unit 150 executes the process of step S204. On the other hand, if the value of the counter Cp for the pitch axis is not equal to or greater than 1 (step S203; NO), the intention estimation unit 150 executes the process of step S211.
[0065] In step S204, the intention estimation unit 150 selects the detailed version of the explanatory information as the explanatory information. The intention estimation unit 150 updates the setting data related to information provision stored in the storage unit 110 with the selected content. Thereafter, the intention estimation unit 150 executes the process of step S211.
[0066] In step S205, the intention inference unit 150 selects to rewind the explanatory information frame by frame. The intention inference unit 150 updates the setting data related to information provision stored in the storage unit 110 with the selected content. Thereafter, the intention inference unit 150 executes the process of step S206.
[0067] In step S206, if the value of the counter Cr is 2 or more (step S206; YES), the intention estimation unit 150 executes the process of step S207. On the other hand, if the value of the counter Cr is not 2 or more (step S206; NO), the intention estimation unit 150 executes the process of step S211.
[0068] In step S207, the intention estimation unit 150 updates the setting data related to information provision stored in the storage unit 110 so as to increase the volume value of the output voice by a preset value. After that, the intention estimation unit 150 executes the process of step S211.
[0069] In step S208, the intention estimation unit 150 selects the simplified version of the explanatory information as the explanatory information. The intention estimation unit 150 updates the setting data related to information provision with the selected content. Thereafter, the intention estimation unit 150 executes the process of step S209.
[0070] In step S209, if the value of the counter Cy is 2 or more (step S209; YES), the intention estimation unit 150 executes the process of step S210. On the other hand, if the value of the counter Cy is not 2 or more (step S209; NO), the intention estimation unit 150 executes the process of step S211.
[0071] In step S210, the intention inference unit 150 selects to stop outputting the explanation information midway. The intention inference unit 150 updates the setting data related to information provision with the selected content. Thereafter, the intention inference unit 150 executes the process of step S211.
[0072] In step S211, the intention estimation unit 150 notifies the information output unit 160 whether the setting data regarding information provision has been updated. Then, the intention estimation process ends. After that, the process of step S46 shown in FIG. 5 is executed.
[0073] If detailed explanatory information is selected in the updated setting data related to information provision, the information output unit 160 reads out the detailed explanatory information about the visually recognized object from the storage unit 110. The information output unit 160 resumes outputting the detailed explanatory information to the earphones 200. The information output unit 160 outputs the explanatory information from a position in the detailed information that corresponds to the position where it was most recently interrupted. In response to this, the earphones 200 resume outputting the detailed explanatory information from the interrupted point.
[0074] For example, if the user nods when the standard version of the explanatory information is provided to the user, it is considered that the user has a positive feeling toward the explanatory information. In this case, it is considered that the user wants to hear a more detailed explanation. With the configuration according to the embodiment, it is possible to switch to providing the detailed version of the explanatory information in accordance with the estimated user's intention. In this way, it is possible to provide audio information taking the user's intention into consideration.
[0075] If frame rewinding of explanatory information is selected in the updated setting data related to information provision, the information output unit 160 causes the earphones 200 to output again part of the explanatory information that was output immediately before. In response to this, the earphones 200 output, by voice, for example, the sentence that was output immediately before. Thereafter, the information output unit 160 resumes output of the explanatory information from the position where it was previously interrupted. In response to this, the earphones 200 resume output of the explanatory information from the point where it was interrupted.
[0076] For example, if the user tilts their head, it is assumed that the user missed the immediately preceding explanatory information. In this case, part of the immediately preceding explanatory information is output again. This allows the user to relisten to the part they missed. In this way, audio information can be provided taking the user's intentions into consideration.
[0077] If the value of the volume of the output sound is increased in the updated setting data related to information provision, the information output unit 160 issues an instruction specifying the updated volume and resumes outputting the explanatory information to the earphone 200. In response to this, the earphone 200 resumes outputting the explanatory information at the updated volume.
[0078] For example, if the user repeatedly tilts their head, it is considered that the user feels that they cannot hear the explanatory information well. In this case, in the configuration according to the embodiment, the setting is changed to increase the volume. Therefore, since the volume is increased while the explanatory information is being output, the user can easily hear the explanatory information. In this way, audio information can be provided taking into consideration the user's intention.
[0079] If simplified explanatory information is selected in the updated setting data related to information provision, the information output unit 160 reads the simplified explanatory information about the visually recognized object from the storage unit 110. The information output unit 160 resumes outputting the simplified explanatory information to the earphones 200. The information output unit 160 outputs the explanatory information from a position in the simplified version that corresponds to the position where it was most recently interrupted. In response to this, the earphones 200 resume outputting the simplified explanatory information from the point where it was interrupted.
[0080] For example, if a user shakes their head when regular explanatory information is provided to the user, it is considered that the user has a negative feeling toward the explanatory information. In this case, it is considered that the user desires a simplified explanation. With the configuration according to the embodiment, it is possible to switch to providing simplified explanatory information in accordance with the estimated user's intention. In this way, it is possible to provide audio information taking into account the user's intention.
[0081] If the updated setting data related to information provision has selected to stop outputting the explanatory information, the information output unit 160 stops outputting the explanatory information, and as a result, output of the explanatory information from the earphones 200 is not resumed.
[0082] For example, if a user repeatedly shakes their head, it is considered that the user has negative feelings about the explanatory information. In this case, it is considered that the user does not want the explanatory information to be provided. With the configuration according to the embodiment, it is possible to switch the setting to stop the provision of explanatory information according to the estimated user's intention. Therefore, explanatory information that the user does not want is not provided to the user.
[0083] As described above, in the information provision system 1000, settings related to information provision are selected according to the user's intention estimated while the explanatory information is being output. The explanatory information is provided to the user according to the settings related to information provision. Therefore, the settings related to information provision can be dynamically changed according to the user's intention. This makes it possible to provide audio information taking the user's intention into consideration.
[0084] B1. Other embodiment 1 In the embodiment, an example has been described in which a user visually recognizes an object whose position is fixed. However, the object visually recognized by the user may be a moving object. The moving object may be, for example, a ship or an airplane. For example, when a user is watching a ship sailing on the sea from an observation deck in a park with an observation deck, the information provision system 1000 can output explanatory information about the ship by voice. Furthermore, for example, when a user is watching an airplane take off or land from an observation deck at an airport, the information provision system 1000 can output explanatory information about the airplane by voice. The following description will focus on configurations different from the embodiment.
[0085] In another embodiment 1, specific area information indicating the range of a specific area in which a user may view a moving object is pre-stored in the storage unit 110. The specific area may be, for example, an observation deck in a park or an observation deck in an airport.
[0086] For example, suppose a user is in a park with an observation deck and is watching ships sailing on the sea from the observation deck. The position / direction acquisition unit 120 acquires information indicating the current location of the mobile terminal 100 as information indicating the user's current location. The position / direction acquisition unit 120 also acquires information indicating the user's line of sight. Based on the acceleration measurement values and angular velocity measurement values received from the earphones 200, the position / direction acquisition unit 120 identifies the direction in which the user's face is facing as the user's line of sight.
[0087] The target estimation unit 130 estimates the target that the user is viewing. Specifically, the target estimation unit 130 first determines whether the user is within the range of the specific area based on the location information provided by the position / direction acquisition unit 120 and the specific area information stored in the storage unit 110. If the target estimation unit 130 determines that the user is within the range of the specific area, it determines candidate targets that the user may view based on the user's current location, date and time, flight schedule, and route information. Furthermore, the target estimation unit 130 determines whether the user is viewing the candidate target. If the target determined as a candidate target for viewing remains within the user's field of view for a predetermined period of time, the target estimation unit 130 determines that the user is viewing the target determined as the candidate target for viewing. The user's field of view is also referred to as the range that the user's eyes can see.
[0088] When the object estimation unit 130 estimates the user's gazed target, the information output unit 160 outputs explanatory information describing the estimated target from the earphones 200. The information output unit 160 acquires the position of a virtual sound source as follows. The information output unit 160 outputs audio that has been subjected to stereophonic processing from the earphones 200 based on the distance between the user and the gazed target and the relative angle of the direction of the gazed target as seen from the user. Because the gazed target is moving, the information output unit 160 may calculate the target position as the position of a virtual sound source at predetermined intervals. The predetermined interval is, for example, five seconds. The information output unit 160 may output audio using stereophonic sound based on the newly calculated distance between the sound source and the user and the relative angle of the direction of the sound source as seen from the user with respect to the user's line of sight. In this case, the user can also feel as if explanatory information is being output from the object being gazed.
[0089] Furthermore, when a plurality of objects are in the user's field of view, the information output unit 160 may output explanatory information in order from the object closest to the user to the object furthest from the user, for example.
[0090] The intention estimation unit 150 identifies the user's head movement from the detection result of the head movement detection unit 140, and estimates the user's intention from the identified user's head movement and the intention definition data. The intention estimation unit 150 selects settings related to information provision according to the user's intention estimated while outputting the explanation information.
[0091] On the other hand, it is assumed that the object estimation unit 130 determines that the user is not within the range of the specific area based on the location information supplied from the position / orientation acquisition unit 120 and the specific area information stored in the storage unit 110. In this case, in the information provision system 1000, explanatory information about the object whose position is fixed is provided to the user, as in the first embodiment.
[0092] B2. Other embodiment 2 The object that the user is viewing may also be a star. For example, when the user is outdoors during the nighttime hours and the elevation angle representing the user's line of sight is within a preset range, the information provision system 1000 can output explanatory information about the constellations in audio. In this case, the object estimation unit 130 can determine the object that the user is viewing based on the user's current location, date and time, the user's line of sight, and a star chart associated with the direction, date, and time. The object estimation unit 130 may read star chart data stored in advance in the storage unit 110. Alternatively, the object estimation unit 130 may read star chart data stored in a cloud server.
[0093] B3. Other embodiment 3 In the embodiment, the user simply listens to explanatory information about the object being viewed. However, the explanatory information may include a question for the user. For example, the information output unit 160 of the mobile terminal 100 outputs a quiz about the viewed object by voice. Furthermore, the information output unit 160 sequentially outputs answer options by voice along with numbers indicating the options. If the user nods after a number indicating an option is output, the intention estimation unit 150 may determine that the option selected by the user is the option indicated by that number.
[0094] According to this embodiment, it is possible to provide a participatory information providing system in which the user can receive information while participating in the process, rather than simply passively receiving the information.
[0095] B4. Other embodiment 4 In the embodiment, the mobile terminal 100 determines that the user is affirmative when the user nods. However, depending on the culture to which the language used by the user belongs, the non-verbal behavior that signifies affirmative may differ. The non-verbal behavior is a so-called gesture. For example, depending on the culture to which the language used by the user belongs, nodding the head may signify negation.
[0096] Therefore, the storage unit 110 of the mobile terminal 100 may store in advance intention definition data defined for each language used. The intention estimation unit 150 may estimate the user's intention implied by the user's head movement based on the intention definition data corresponding to the language used by the user. Note that the intention estimation unit 150 can acquire information about the language used by the user from, for example, setting information about the language set in the mobile terminal 100. In this way, even if users use different languages, it is possible to estimate the user's intention from the head movement.
[0097] B5. Other Embodiment 5 In the embodiment, the intention estimation unit 150 estimates the user's intention from the identified user's head movement and the intention definition data. Alternatively, the intention estimation unit 150 may estimate the user's intention using a machine learning model that has undergone machine learning. When parameters representing the user's head movement, the user's movement speed, the distance between the user and an object, and the user's relative angle with respect to the object are input, the machine learning model outputs a result of estimating the user's intention. According to this embodiment, the user's intention can be estimated with high accuracy.
[0098] B6. Other Embodiment 6 In the embodiment, the intention inference unit 150 determines that rotation around a certain rotation axis has been detected when the rotation angle of that rotation axis is equal to or greater than a predetermined rotation angle. However, there are cases where rotation around two rotation axes is detected at the same time. In such cases, the intention inference unit 150 may use the rotation around the rotation axis with the larger rotation angle.
[0099] B7. Other Embodiments 7 The settings related to information provision stored in the storage unit 110 may include information indicating the speed at which the explanatory information is read out in addition to the information described in the first embodiment. The information indicating the speed at which the explanatory information is read out indicates the speed at which the explanatory information is read out in voice output from the earphones 200. The information indicating the speed at which the explanatory information is read out is also referred to as setting information related to voice output.
[0100] For example, when the intention estimation unit 150 estimates that the user finds it difficult to hear the explanatory information, the intention estimation unit 150 may update the information indicating the speed at which the explanatory information is read out so as to slow down the speed at which the explanatory information is read out.
[0101] B8. Other Embodiment 8 In the embodiment, an example has been described in which the position and direction acquisition unit 120 acquires information indicating the current position of the mobile terminal 100 indoors based on the radio wave intensity received from a plurality of Wi-Fi (registered trademark) base stations. Alternatively, the position information of the mobile terminal 100 indoors may be acquired as follows. The mobile terminal 100 is assumed to be equipped with a geomagnetic sensor. In this case, the position and direction acquisition unit 120 may acquire the position information of the mobile terminal 100 using the geomagnetic sensor.
[0102] Alternatively, the position and direction acquisition unit 120 first acquires the position information of the mobile terminal 100 based on the radio wave intensity received from a Wi-Fi (registered trademark) base station. If the position and direction acquisition unit 120 is unable to acquire the position information, the position and direction acquisition unit 120 may acquire the position information of the mobile terminal 100 using a geomagnetic sensor.
[0103] In the embodiment, an example has been described in which the position and direction acquisition unit 120 uses GPS to acquire the current position of the mobile terminal 100 outdoors. Alternatively, the position and direction acquisition unit 120 may use another satellite positioning system such as the Quasi-Zenith Satellite System. Furthermore, alternatively, the position and direction acquisition unit 120 may acquire the current position of the mobile terminal 100 using GPS and the Quasi-Zenith Satellite System.
[0104] B9. Other Embodiments 9 In the embodiment, the storage unit 110 stores sound source data having an audio signal that reads out explanatory information about an object that may be a visual target of the user. However, the sound source data does not have to be stored in the storage unit 110. The information output unit 160 may access the sound source data stored in a cloud server and transmit the audio signal included in the sound source data to the earphones 200. In this case, it is sufficient that the storage unit 110 stores a URL (Uniform Resource Locator) that identifies the location of the sound source data stored in the cloud server.
[0105] B10. Other Embodiments 10 In addition, in the embodiment, an example has been described in which the explanatory information provided to the user is one of three types of explanatory information: normal explanatory information, detailed explanatory information, and simplified explanatory information. However, the types of explanatory information are not limited to three. Alternatively, one of two types of explanatory information, normal explanatory information and simplified explanatory information, may be provided to the user. Alternatively, the types of explanatory information may be four or more.
[0106] In the embodiment, an example has been described in which the three types of explanatory information are normal explanatory information, detailed explanatory information, and simplified explanatory information. Different types of explanatory information may be provided depending on the user's age. For example, explanatory information of one type provided to elementary school students, explanatory information of one type provided to junior high and high school students, or explanatory information of one type provided to university students and working adults may be provided depending on the user's age. For example, the information providing system 1000 determines the user's age group based on age information entered by the user when installing the guidance application. Each type of explanatory information has content that can be understood by users according to their age. Furthermore, normal explanatory information, detailed explanatory information, and simplified explanatory information are provided for each age type.
[0107] Alternatively, for a particular object, one of the three types of explanatory information may be provided to the user, and for other objects, one of the two types of explanatory information may be provided to the user.
[0108] Furthermore, in the embodiment, the earphone 200 is given as an example of an audio output device, but the audio output device may be headphones or a bone conduction headset.
[0109] In the embodiment, an example has been described in which the communication unit 103 communicates with an external device in accordance with the Wi-Fi (registered trademark) communication standard. However, the communication unit 103 may communicate with an external device in accordance with another communication standard, such as Bluetooth (registered trademark). The communication unit 103 may be compatible with multiple communication standards.
[0110] Furthermore, the means for realizing the functions of the mobile terminal 100 is not limited to software, and some or all of the functions may be realized by dedicated hardware. For example, the dedicated hardware may be a circuit such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).
[0111] In the embodiment, an example has been described in which the mobile terminal 100, which is a computer carried by a user, is a smartphone. Alternatively, the mobile terminal 100 may be a mobile phone, a tablet terminal, or the like. Furthermore, the mobile terminal 100 may be a wearable computer. Examples of wearable computers include a smart watch and a head-mounted display.
[0112] In the embodiment, when the information output unit 160 determines from the notification from the intention estimation unit 150 that the setting data related to information provision has been updated, the information output unit 160 suspends the output of the explanatory information. However, the information output unit 160 does not necessarily have to suspend the output of the explanatory information. For example, the information output unit 160 may read the updated setting data while continuing to output the explanatory information by voice, and then output the explanatory information according to the updated setting data related to information provision.
[0113] Furthermore, when rotation about the roll axis is detected, the information output unit 160 may suspend output of the explanatory information and re-output part of the explanatory information that was previously output in accordance with the updated setting data for information provision. When rotation about the yaw axis or the pitch axis is detected, the information output unit 160 may switch the explanatory information to be provided, for example, to detailed explanatory information or simplified explanatory information in accordance with the updated setting data for information provision without suspending output of the explanatory information.
[0114] Furthermore, the system may select which of the three types of explanatory information to provide regardless of the user's intention, which is estimated from the user's head movement. For example, outputting explanatory information by voice for a long period of time outdoors during hot or cold weather may cause the user to stay outdoors. In such a case, the system may select to provide simplified explanatory information based on, for example, date, time, and location information.
[0115] The head motion detection unit 140 may detect the roll angle, pitch angle, and yaw angle from the acceleration measurement value, angular velocity measurement value, and geomagnetic field strength measurement value. In this case, the sensor 203 includes a geomagnetic sensor in addition to the acceleration sensor, angle sensor, and angular velocity sensor.
[0116] The present disclosure is not limited to the above-described embodiments and can be realized in various configurations without departing from the spirit thereof. For example, the technical features in the embodiments corresponding to the technical features in each aspect described in the Summary of the Invention section can be appropriately replaced or combined to solve some or all of the above-described problems or achieve some or all of the above-described effects. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted. [Explanation of symbols]
[0117] 100...Mobile terminal, 101...CPU, 102...memory, 103...communication unit, 109...internal bus, 110...storage unit, 120...position and direction acquisition unit, 130...object estimation unit, 140...head movement detection unit, 150...intention estimation unit, 160...information output unit, 200...earphone, 201...DSP, 202...communication unit, 203...sensor, 204...driver unit, 209...internal bus, 1000...information provision system, Cr...counter, Cy...counter, D...gaze direction, N...reference orientation, P...user, SS...sound source, T1...object, r1...angle, r2...angle, r3...angle
Claims
1. An information providing system that provides information by sound, a position and direction acquisition unit that acquires position information indicating a position where a user is located and gaze direction information indicating a gaze direction that is a direction in which the user's face is facing; a storage unit configured to store target position information indicating the position of each of a plurality of targets that may be visually recognized by the user, explanatory information that explains each of the targets, and information indicating settings related to information provision; an object estimation unit that estimates the object being gazed at by the user based on the position information, the gaze direction information, and the object position information; an information output unit that outputs the explanatory information about the estimated target by voice in accordance with settings related to the information provision; a head movement detection unit that detects a movement of the user's head; an intention estimation unit that estimates an intention of the user from a head movement of the user while the explanation information is being output, and selects a setting regarding the information provision in accordance with the intention of the user; Equipped with when the setting regarding the information provision is changed, the information output unit outputs the explanation information in accordance with the setting regarding the information provision after the change. Information provision system.
2. 2. The information providing system according to claim 1, the storage unit stores, as the explanatory information that explains the object, first explanatory information that is an explanation about the object, and second explanatory information that is an explanation about the object that is different from the first explanatory information; the setting regarding information provision includes information indicating which of the first explanatory information and the second explanatory information is selected as the explanatory information; Information provision system.
3. 3. The information providing system according to claim 2, the storage unit further stores, as the explanation information explaining the object, third explanation information that is an explanation for the object different from the first explanation information and the second explanation information; The first description information is a general description of the object, the second description information is a more detailed description than the first description information, and the third description information is a more brief description than the first description information, the setting regarding the information provision includes information indicating which of the first explanatory information, the second explanatory information, and the third explanatory information is selected as the explanatory information; Information provision system.
4. 4. The information providing system according to claim 1, the information provision settings include setting information related to the audio output; Information provision system.
5. 5. The information providing system according to claim 1, the information provision setting includes information indicating whether to continue outputting the explanation information; Information provision system.
6. 6. The information providing system according to claim 1, the information output unit outputs a question to the user by voice; the intention estimation unit estimates an answer of the user to the question from a head movement of the user. Information provision system.
7. 7. The information providing system according to claim 1, the target includes a moving object; the object estimation unit estimates that the moving object is the object being visually recognized by the user when a state in which the moving object is within a range visible to the user's eyes continues for a predetermined period of time; Information provision system.
8. 8. The information providing system according to claim 1, a sound source position acquisition unit that acquires a virtual position of a sound source corresponding to each of the targets; Furthermore, the information output unit outputs, to a portable audio output device worn on the head of the user, audio obtained by performing stereophonic processing on the audio representing the explanatory information, in accordance with the virtual position of the sound source as seen from the current position of the user. Information provision system.
9. 9. The information providing system according to claim 1, the storage unit stores intention definition data in which non-verbal actions corresponding to a culture to which the language used by the user belongs are defined; the intention estimation unit estimates the intention of the user from the intention definition data and a head movement of the user. Information provision system.
10. 10. The information providing system according to claim 1, the intention estimation unit estimates the intention of the user by inputting parameters representing a head movement of the user, a moving speed of the user, a distance between the user and the object, and a relative angle of the user with respect to the object into a trained machine learning model. Information provision system.
11. A method for providing information by sound in a computer carried by a user, comprising: The computer acquiring position information indicating a position where the user is located and gaze direction information indicating a gaze direction in which the face of the user is facing; estimating an object being gazed at by the user from the position information, the gaze direction information, and target position information set in advance for each of a plurality of objects that may be gazed at by the user; outputting explanatory information about the estimated target by voice in accordance with settings related to information provision; detecting a head movement of the user; a step of estimating an intention of the user from a head movement of the user while the explanation information is being output; selecting a setting regarding the information provision in accordance with the estimated intention of the user; When the setting regarding the information provision is changed, outputting the explanatory information by voice in accordance with the changed setting regarding the information provision; A method comprising:
12. A program executed by a computer carried by a user, The computer, a function of acquiring position information indicating a position where the user is located and gaze direction information indicating a gaze direction, which is a direction in which the face of the user is facing; a function of estimating the target being gazed at by the user based on the position information, the gaze direction information, and target position information previously set for each of a plurality of targets that may be gazed at by the user; a function of outputting explanatory information about the estimated target by voice in accordance with settings related to information provision; a function of detecting a head movement of the user; a function of estimating the user's intention from a head movement of the user while the explanation information is being output; a function of selecting a setting regarding the provision of information in accordance with the estimated intention of the user; a function of outputting the explanatory information by voice in accordance with the changed setting regarding the information provision when the setting regarding the information provision is changed; A program to make this happen.
Citation Information
Patent Citations
Merchandise introducing device
JP1996160897A
Guiding device
JP1997170929A
Interactive sign system
JP2009104426A
Action support device
JP2016167208A
Automatic guide service system using virtual realityand service method thereof
KR1020050061856A