Information providing device, information providing method, information providing program, and storage medium
The information providing device improves convenience by identifying objects based on gaze concentration and providing relevant information without manual pointing, utilizing learning models for accurate object recognition, addressing the limitations of conventional systems.
Patent Information
- Application Number
- JP2025119166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-11
AI Technical Summary
Conventional object identification systems require vehicle occupants to point at objects with their hands or fingers to obtain information, which hinders convenience.
An information providing device that includes an image acquisition unit, information acquisition unit, identification unit, area extraction unit, and object recognition unit to identify objects based on gaze concentration and provide relevant information without manual pointing, utilizing learning models for accurate object recognition.
Enhances convenience by allowing occupants to receive object information without manual pointing, reduces processing load through statistical gaze identification, and ensures accurate object recognition using visual saliency technology when statistical methods fail.
Smart Images

Figure 2025133991000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information providing device, an information providing method, an information providing program, and a storage medium. [Background technology]
[0002] Conventionally, there is known an object identification device that identifies an object present around a vehicle and reads out information about the object, such as its name, by voice (see, for example, Patent Document 1). The object identification device described in Patent Document 1 identifies, as an object, a facility or the like on a map that is present in the direction indicated by the vehicle occupant pointing with their hand or finger. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-80060 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the technology described in Patent Document 1 has the problem that it requires vehicle occupants who wish to obtain information about an object to point to the object with their hand or finger, which makes it difficult to improve convenience.
[0005] The present invention has been made in view of the above, and has an object to provide an information providing device, an information providing method, an information providing program, and a storage medium that can improve convenience, for example. [Means for solving the problem]
[0006] The information provision device described in claim 1 is characterized by comprising an image acquisition unit that acquires captured images of the surroundings of a moving body; an information acquisition unit that acquires position information of the moving body when a specified request is made by an occupant; an identification unit that uses the position information to identify an object on which gazes are statistically concentrated; an area extraction unit that, if the identification unit cannot identify the object on which gazes are concentrated, extracts an attention area in the captured image on which gazes are concentrated based on the captured image; an object recognition unit that recognizes objects included in the attention area in the captured image; and an information provision unit that provides object information on the object identified by the identification unit or object information on the object included in the attention area. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram showing a configuration of an information providing system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of the in-vehicle terminal. [Figure 3] FIG. 3 is a block diagram showing the configuration of the information providing device. [Figure 4] FIG. 4 is a diagram illustrating the information providing method. [Figure 5] FIG. 5 is a diagram illustrating the information providing method. [Figure 6] FIG. 6 is a flowchart showing the information providing method. [Figure 7] FIG. 7 is a diagram illustrating the information providing method. [Figure 8] FIG. 8 is a block diagram showing a configuration of an information providing device according to the second embodiment. [Figure 9] FIG. 9 is a diagram illustrating the information providing method. [Figure 10] FIG. 10 is a flowchart showing the information providing method. [Figure 11] FIG. 11 is a block diagram showing a configuration of an information providing device according to the third embodiment. [Figure 12] FIG. 12 is a flowchart showing the information providing method. [Figure 13] FIG. 13 is a diagram illustrating an example of a table stored in the information providing device. [Figure 14] FIG. 14 is a diagram illustrating an example of a table stored in the information providing device. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, a mode for carrying out the present invention (hereinafter referred to as an embodiment) will be described with reference to the drawings. Note that the present invention is not limited to the embodiment described below. Furthermore, in the description of the drawings, the same parts are given the same reference numerals.
[0009] (Embodiment 1) [Outline of information provision system] Fig. 1 is a block diagram showing the configuration of an information provision system 1 according to a first embodiment. The information provision system 1 is a system that provides an occupant PA (see Fig. 7) of a vehicle VE (Fig. 1), which is a moving body, with object information (e.g., the name of the object) relating to objects such as buildings present around the vehicle VE. As shown in Fig. 1, the information provision system 1 includes an in-vehicle terminal 2 and an information provision device 3. The in-vehicle terminal 2 and the information provision device 3 communicate with each other via a network NE (Fig. 1), which is a wireless communication network.
[0010] 1 illustrates a case where there is one in-vehicle terminal 2 that communicates with the information providing device 3, but there may be multiple in-vehicle terminals mounted on multiple vehicles. Also, in order to provide object information to multiple occupants in a single vehicle, multiple in-vehicle terminals 2 may be mounted on a single vehicle.
[0011] [Configuration of in-vehicle terminal] 2 is a block diagram showing the configuration of the in-vehicle terminal 2. The in-vehicle terminal 2 is, for example, a stationary navigation device or a drive recorder installed in the vehicle VE. Note that the in-vehicle terminal 2 is not limited to a navigation device or a drive recorder, and may also be a portable terminal such as a smartphone used by an occupant PA of the vehicle VE. As shown in FIG. 2, the in-vehicle terminal 2 includes an audio input unit 21, an audio output unit 22, an imaging unit 23, an input unit 24, a terminal main body 25, a sensor unit 26, and a display unit 27.
[0012] The voice input unit 21 includes a microphone 211 (see FIG. 7) that receives voice and converts it into an electrical signal, and generates voice information by performing A / D (Analog / Digital) conversion or the like on the electrical signal. In the first embodiment, the voice information generated by the voice input unit 21 is a digital signal. Then, the voice input unit 21 outputs the voice information to the terminal main body 25.
[0013] The audio output unit 22 includes a speaker 221 (see FIG. 7), converts a digital audio signal input from the terminal main body 25 into an analog audio signal by D / A (Digital / Analog) conversion, and outputs audio corresponding to the analog audio signal from the speaker 221.
[0014] The imaging unit 23 captures an image of the surroundings of the vehicle VE and generates a captured image under the control of the terminal main body 25. Then, the imaging unit 23 outputs the generated captured image to the terminal main body 25.
[0015] The input unit 24 includes input devices such as a touch panel, keyboard, and mouse, and receives input of various data in response to operations of the occupant PA. The input unit 24 then outputs the received input of various data to the terminal main body 25. The display unit 27 is configured with a display using a liquid crystal or organic EL (Electro Luminescence) display, etc., and displays various images under the control of the terminal main body 25.
[0016] The sensor unit 26 includes sensor devices such as a GPS (Global Positioning System) sensor, a gyro sensor, an acceleration sensor, and a direction sensor, and has a function of sensing information used for processing in the terminal main body 25. The GPS sensor receives GPS signals from GPS satellites to measure information indicating the latitude, longitude, and altitude of an object. The information acquired by the GPS sensor 109 is also referred to as "location information" below.
[0017] 2, the terminal main body 25 includes a communication unit 251, a control unit 252, and a storage unit 253. The communication unit 251 transmits and receives information to and from the information providing device 3 via the network NE under the control of the control unit 252.
[0018] The control unit 252 is realized by a controller such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in the storage unit 253, and controls the overall operation of the in-vehicle terminal 2. The control unit 252 is not limited to a CPU or an MPU, and may be configured by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0019] The storage unit 253 stores various programs executed by the control unit 252, data required when the control unit 252 performs processing, and the like.
[0020] [Configuration of information providing device] 3 is a block diagram showing the configuration of the information providing device 3. The information providing device 3 is, for example, a server device. As shown in FIG. 3, the information providing device 3 includes a communication unit 31, a control unit 32, and a storage unit 33.
[0021] The communication unit 31, under the control of the control unit 32, transmits and receives information to and from the on-board terminal 2 (communication unit 251) via the network NE.
[0022] The control unit 32 is realized by a controller such as a CPU or an MPU executing various programs (including the information provision program according to this embodiment) stored in the storage unit 33, and controls the overall operation of the information provision device 3. The control unit 32 is not limited to a CPU or an MPU, and may be configured using an integrated circuit such as an ASIC or an FPGA. As shown in FIG. 3 , the control unit 32 includes a request information acquisition unit 321, a voice analysis unit 322, an image acquisition unit 323, an information acquisition unit 324, an identification unit 325, a region extraction unit 326, an object recognition unit 327, an information provision unit 328, and a reflection unit 329.
[0023] The request information acquisition unit 321 acquires request information requesting the provision of object information from the occupant PA of the vehicle VE. In the first embodiment, the request information is voice information generated by the voice input unit 21 based on words (voice) uttered by the occupant PA of the vehicle VE, which is captured by the voice input unit 21. That is, the request information acquisition unit 321 acquires the request information (voice information) from the in-vehicle terminal 2 via the communication unit 31. The voice analysis unit 322 analyzes the request information (voice information) acquired by the request information acquisition unit 321.
[0024] The image acquisition unit 323 acquires the captured image generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31. The information acquisition unit 324 acquires the location information of the moving object when a predetermined request is received from the occupant PA. In this embodiment, the information acquisition unit 324 acquires the location information of the moving object from the in-vehicle terminal 2 when the request information (voice information) contains a specific keyword as a result of the voice analysis unit 322 analyzing the request information. Here, the specific keyword is a word used by the occupant PA of the vehicle VE to request object information, and examples of the specific keyword include words such as "what," "what is it," "what is it," and "tell me." In addition to the location information of the moving object, the information acquisition unit 324 may also acquire a specific keyword for identifying the object that the occupant PA is focusing on from the content of the occupant's predetermined request. For example, when the words uttered by the occupant PA contain a specific keyword such as "building," "temple," or "shop," the information acquisition unit 324 acquires the specific keyword.
[0025] The identification unit 325 uses the position information to identify an object on which gazes are statistically concentrated. In the first embodiment, the identification unit 325 uses the position information as input data and a third learning model that statistically identifies an object on which gazes are concentrated to identify an object on which gazes are concentrated. That is, the identification unit 325 inputs the position information into the third learning model and obtains information on the object on which gazes are concentrated as an output result from the third learning model. Note that an object on which gazes are statistically concentrated here refers to an object that is predicted to be an object on which the gazes of the occupant PA are concentrated, in other words, an object that the identification unit 325 determines to be an object on which the gazes of the occupant PA are likely to be concentrated. Furthermore, the identification unit 325 may use the position information and a keyword to identify an object on which gazes are statistically concentrated according to the position information and the keyword. For example, when the occupant PA utters a phrase such as "What's that building?", "Tell me about that temple over there," or "What's that shop?", the identification unit 325 may statistically identify an object on which the gaze is focused using keywords for identifying an object on which the occupant PA is focusing, such as "building," "temple," or "shop," in addition to the position information of the vehicle VE at the time the phrase was uttered. In this case, for example, the identification unit 325 uses the position information and the keywords as input data and identifies the object on which the gaze is focused using a third learning model that statistically identifies an object on which the gaze is focused. That is, the identification unit 325 may input the keywords for identifying the object on which the occupant PA is focusing, along with the position information, into the third learning model, and obtain information on the object on which the gaze is focused as an output result from the third learning model.
[0026] The third learning model is a model obtained by, for example, using an eye tracker to determine the area where the subject's gaze is concentrated, using pre-labeled position information and keywords for the area as training data, and performing machine learning (e.g., deep learning) on the area using the training data image. Note that in this embodiment, the third learning model is updated by the reflection unit 329, which will be described later.
[0027] When the identification unit 325 cannot identify an object on which gazes are concentrated, the region extraction unit 326 extracts (predicts) an attention region in the captured image where gazes are concentrated (where gazes are likely to be concentrated) based on the captured image acquired by the image acquisition unit 323. In the first embodiment, the region extraction unit 326 extracts an attention region in the captured image by utilizing a so-called visual saliency technique. More specifically, the region extraction unit 326 extracts an attention region in the captured image by image recognition using a first learning model shown below (image recognition using AI (Artificial Intelligence)).
[0028] Here, the case where the identification unit 325 cannot identify an object on which the gaze is concentrated includes not only a case where the identification unit 325 actually cannot identify the object, but also a case where the object on which the gaze is concentrated cannot be identified with sufficient accuracy due to insufficient learning of the third learning model, for example. The first learning model is a model obtained by using an eye tracker to determine the area on which the subject's gaze is concentrated, using an image in which the area is pre-labeled as a teacher image, and using the teacher image to machine-learn the area (for example, deep learning, etc.).
[0029] The object recognition unit 327 recognizes, within the captured image, an object included in the region of interest extracted by the region extraction unit 326. In the first embodiment, the object recognition unit 327 recognizes an object included in the region of interest within the captured image by image recognition using the second learning model described below (image recognition using AI).
[0030] The second learning model is a model obtained by using photographed images of various objects such as animals, mountains, rivers, lakes, and facilities as teacher images and using machine learning (e.g., deep learning) to learn the characteristics of the objects based on the teacher images.
[0031] The information providing unit 328 provides object information related to the object identified by the identifying unit 325 or object information related to the object recognized by the object recognizing unit 327. More specifically, the information providing unit 328 reads out object information corresponding to the object identified by the identifying unit 325 or object information corresponding to the object recognized by the object recognizing unit 327 from an object information DB (Data Base) 333 in the storage unit 33. Then, the information providing unit 328 transmits the object information to the in-vehicle terminal 2 via the communication unit 31.
[0032] The reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 in the learning model. For example, as a result of the object information provided by the information providing unit 328, the reflecting unit 329 determines whether the object information provided by the information providing unit 328 is information that meets the user's wishes. As a result, if the information meets the user's wishes, the reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 as correct answer data in the learning model. If the information does not meet the user's wishes, the reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 as incorrect answer data in the learning model. Note that the method of determining whether the information meets the user's wishes may be a method of receiving a manual input from the user, or may be an automatic method based on the user's behavior after the object information is provided. Note that, if the reflecting unit 329 inputs keywords for identifying the object that the occupant PA is focusing on into the third learning model together with the position information and provides the object information, the reflecting unit 329 also reflects the keywords in the learning model.
[0033] The storage unit 33 stores various programs (information provision programs according to this embodiment) executed by the control unit 32, as well as data and the like required when the control unit 32 performs processing. As shown in FIG. 3, the storage unit 33 includes a first learning model DB 331, a second learning model DB 332, an object information DB 333, and a third learning model DB 334. The first learning model DB 331 stores the above-described first learning model. The second learning model DB 332 stores the above-described second learning model. The third learning model DB 334 stores the above-described third learning model.
[0034] The object information DB 333 stores the above-mentioned object information. Here, a plurality of pieces of object information associated with various objects are stored in the object information DB 333. The object information is information describing the object, such as the name of the object, and is configured from character data, audio data, or image data.
[0035] Here, an example of an information providing method by the information providing device 3 will be described with reference to Fig. 4 and Fig. 5. Fig. 4 and Fig. 5 are diagrams for explaining the information providing method. With reference to Fig. 4, an information providing method will be described in the case where an object on which gazes are concentrated can be statistically identified using a third learning model that takes position information into consideration.
[0036] In the example of FIG. 4, a case will be described in which an occupant PA of the vehicle VE focuses on a building and utters, "What's that?" As illustrated in FIG. 4, when the occupant PA of the vehicle VE utters the phrase "What's that?" to request the provision of object information, the in-vehicle terminal 2 of the vehicle VE notifies the information providing device 3 of the location information. Then, using a third learning model trained in consideration of the location information, the information providing device 3 identifies an object that is statistically predicted to attract the occupant's gaze at the current location of the vehicle VE. Then, the information providing device 3 transmits object information (such as the building name, detailed description, and image) corresponding to the identified object to the in-vehicle terminal 2 via the communication unit 31.
[0037] In this way, the information providing device 3 can use a third learning model that takes into account position information to statistically identify objects on which gazes are concentrated and provide object information corresponding to those objects to the occupant PA, thereby reducing the processing load compared to when visual saliency technology is used.
[0038] Here, an information provision method for identifying an object on which gazes are focused by using visual saliency technology and object recognition when the information provision device 3 is unable to statistically identify an object on which gazes are focused using the third learning model will be described with reference to Fig. 5. Note that the example in Fig. 5 will be described using an example where an occupant PA of a vehicle VE looks at a lion and utters "What's that?"
[0039] If the information providing device 3 is unable to statistically identify an object on which gazes are concentrated using the third learning model, it extracts an area of interest in the captured image through image recognition using the first learning model. Next, the information providing device 3 recognizes a lion included in the area of interest in the captured image through image recognition using the second learning model. Then, the information providing device 3 transmits object information corresponding to the recognized lion to the in-vehicle terminal 2 via the communication unit 31.
[0040] In this way, even if the information providing device 3 is unable to statistically identify an object on which gazes are focused using the third learning model, it is possible to provide object information in real time in response to the request of the occupant PA using visual saliency technology.
[0041] [Information provision method] Next, an information providing method executed by the information providing device 3 (control unit 32) will be described. FIG. 6 is a flowchart showing the information providing method. FIG. 7 is a diagram explaining the information providing method. Specifically, FIG. 7 is a diagram showing a captured image IM generated by the imaging unit 23 and acquired in step S106. Here, FIG. 7 illustrates a case where the imaging unit 23 is installed inside the vehicle VE so that an image of the front of the vehicle VE is captured from inside the vehicle VE through the windshield. FIG. 7 also illustrates a case where an occupant PA sitting in the passenger seat of the vehicle VE is included as a subject in the captured image IM. Furthermore, FIG. 7 illustrates a case where the occupant PA utters the words "What's that?"
[0042] The installation position of the imaging unit 23 is not limited to the above-mentioned installation position. For example, the imaging unit 23 may be installed inside the vehicle VE so that the left side, right side, or rear of the vehicle VE is captured from inside the vehicle VE, or the imaging unit 23 may be installed outside the vehicle VE so that the surroundings of the vehicle VE are captured. Furthermore, the vehicle occupant in this embodiment is not limited to an occupant sitting in the passenger seat of the vehicle VE, but also includes an occupant sitting in the driver's seat or a rear seat. Furthermore, the number of imaging units 23 is not limited to one, and may be multiple.
[0043] First, the request information acquisition unit 321 acquires request information (voice information) from the in-vehicle terminal 2 via the communication unit 31 (step S101). After step S101, the voice analysis unit 322 analyzes the request information (voice information) acquired in step S101 (step S102). After step S102, the voice analysis unit 322 determines, as a result of analyzing the request information (voice information) in step S102, whether or not the request information (voice information) includes a specific keyword (step S103). Here, the specific keyword is a word used by the occupant PA of the vehicle VE to request the provision of object information, and examples of the specific keyword include words such as "what," "what is it," "I wonder," and "tell me."
[0044] If it is determined that the specific keyword is not included (step S103: No), the control unit 32 returns to step S101. On the other hand, if it is determined that the specific keyword is included (step S103: Yes), the information acquisition unit 324 acquires position information of the vehicle VE when a predetermined request is received from the occupant PA (step S104). Then, the identification unit 325 uses the position information to statistically identify objects on which gazes are concentrated, and then determines whether the objects on which gazes are concentrated can be identified (step S105).
[0045] As a result, if the identification unit 325 determines that it is unable to identify the object on which the gaze is focused (step S105: No), the image acquisition unit 323 acquires the captured image IM generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31 (step S106).
[0046] 6 and 7, the image acquisition unit 323 is configured to acquire the captured image IM generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31 at the timing when the occupant PA of the vehicle VE utters the words "What's that?" (step S103: Yes), but this is not limited to this. For example, the information providing device 3 sequentially acquires the captured images generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31. Then, the image acquisition unit 323 may be configured to acquire, from the sequentially acquired captured images, the captured image acquired at the timing when the occupant PA of the vehicle VE utters the words "What's that?" (step S103: Yes) as the captured image to be used in the processing from step S104 onwards.
[0047] After step S106, the area extraction unit 326 extracts an attention area Ar1 (see Figure 7) where gazes are concentrated in the captured image IM by image recognition using the first learning model stored in the first learning model DB331 (step S107).
[0048] After step S107, the object recognition unit 327 recognizes the object OB1 contained in the attention area Ar1 extracted in step S5 in the captured image IM by image recognition using the second learning model stored in the second learning model DB332 (step S108).
[0049] After step S108, the information providing unit 328 obtains object information corresponding to the object OB1 identified by the identification unit 325 or recognized by the object recognition unit 327 from the object information DB 333 (step S109), and transmits the object information to the in-vehicle terminal 2 via the communication unit 31 (step S110). The control unit 252 then controls the operation of at least one of the audio output unit 22 and the display unit 27 to notify the occupant PA of the vehicle VE of the object information transmitted from the information providing device 3 by audio, text, or an image. For example, if the object OB1 is a "Moulin Rouge," the object information notified to the occupant PA of the vehicle VE is audio such as "That's the Moulin Rouge. They're putting on a spectacular dance show at night." Furthermore, if the object OB1 is not a building but a buffalo, the object information notified to the occupant PA of the vehicle VE is audio such as "That's a buffalo. Buffalos live in herds." After that, the reflection unit 329 reflects the result of providing the object information in the learning model (step S111).
[0050] According to the first embodiment described above, the following effects are achieved. The information providing device 3 according to the first embodiment acquires captured images of the surroundings of the vehicle VE, and acquires position information of the vehicle VE when a predetermined request is received from the occupant PA. The information providing device 3 then uses the position information to identify an object on which gazes are statistically concentrated. Furthermore, if the information providing device 3 cannot identify an object on which gazes are concentrated, it extracts an attention area in the captured image on which gazes are concentrated, based on the captured image, and recognizes an object included in the attention area in the captured image. Thereafter, the information providing device 3 provides object information on the identified object or object information on an object included in the attention area.
[0051] Therefore, an occupant PA of the vehicle VE who wishes to obtain object information about an object does not need to point at the object with his or her hand or finger as in the conventional case, thereby improving convenience.
[0052] In addition, the information providing device 3 can statistically identify objects on which gazes are concentrated using a third learning model that takes location information into account, and provide object information corresponding to the objects to the occupant PA, thereby reducing the processing load compared to when visual saliency technology is used.
[0053] Furthermore, even if the information providing device 3 is unable to statistically identify an object on which gazes are concentrated, it can use visual saliency technology to extract an attention area Ar1 on which gazes are concentrated in the captured image IM. Therefore, even if the occupant PA of the vehicle VE does not point at the object OB1 with his / her hand or finger, it is possible to accurately extract the area including the object OB1 as the attention area Ar1.
[0054] Furthermore, the information providing device 3 provides the object information in response to request information from the occupant PA of the vehicle VE, which requests the provision of the object information. Therefore, the processing load of the information providing device 3 can be reduced compared to a configuration in which the object information is always provided regardless of the request information.
[0055] (Embodiment 2) Next, a second embodiment will be described. In the following description, the same components as those in the first embodiment described above will be assigned the same reference numerals, and detailed description thereof will be omitted or simplified. FIG. 8 is a block diagram showing the configuration of an information providing device 3A according to the second embodiment. In the information providing device 3A according to the second embodiment, the functions of the information acquiring unit 324 and the identifying unit 325 are changed. For ease of explanation, the information acquiring unit according to the second embodiment will be referred to as information acquiring unit 324A, and the identifying unit according to the second embodiment will be referred to as identifying unit 325A below (see FIG. 8). Furthermore, in the information providing device 3A, a fourth learning model DB 335 (see FIG. 8) is added to the storage unit 33.
[0056] The fourth learning model is a model obtained by, for example, using an eye tracker to determine an area where the subject's gaze is concentrated, using pre-labeled position information, keywords, and attribute information of the area as training data, and performing machine learning (e.g., deep learning) on the area using the training data. In this embodiment, the fourth learning model is updated by the reflection unit 329. The fourth learning model DB 335 stores the fourth learning model.
[0057] The information acquisition unit 324A acquires attribute information of the occupant PA along with the position information of the vehicle VE. For example, the information acquisition unit 324A acquires, as attribute information, one or more of the following information from the in-vehicle terminal 2: the age of the occupant PA, the gender of the occupant PA, the nationality of the occupant PA, the appearance of the occupant PA, the occupant PA's preferences (e.g., whether the occupant PA likes coffee or travel), and the language of the occupant PA. The in-vehicle terminal 2 may acquire the attribute information by any method. For example, the in-vehicle terminal 2 may accept manual input of the attribute information by the occupant PA via the input unit 24, or may automatically acquire the occupant's attribute information by analyzing audio and images input by the audio input unit 21 or the imaging unit 23. The in-vehicle terminal 2 may store the attribute information in advance, or may acquire the attribute information from an external database.
[0058] The identification unit 325A uses the position information and the attribute information to identify an object on which gazes are statistically concentrated according to the position information and the attribute information. For example, the identification unit 325A uses the position information and the attribute information as input data and a learning model that statistically identifies an object on which gazes are concentrated to identify an object on which gazes are concentrated.
[0059] Here, an example of an information providing method by the information providing device 3A will be described with reference to Fig. 9. Fig. 9 is a diagram for explaining the information providing method. With reference to Fig. 9, an information providing method will be described in the case where an object on which gazes are concentrated can be statistically identified using a fourth learning model that takes into account position information and attribute information.
[0060] In the example of FIG. 9, a case will be described in which an occupant PA of the vehicle VE focuses on a building and utters, "What's that?" As illustrated in FIG. 9, when the occupant PA of the vehicle VE utters the phrase "What's that?" to request object information, the in-vehicle terminal 2 of the vehicle VE notifies the information providing device 3A of position information and attribute information. The information providing device 3A then uses a fourth learning model trained in consideration of the position information and attribute information to identify an object that is statistically predicted to attract gaze. To explain FIG. 9 using a specific example, assume that the attribute information of the occupant PA who uttered "What's that?" about a cafe is "female" gender and "likes coffee." In this case, the information providing device 3A inputs the attribute information, "female" gender and "likes coffee," along with the position information, to the fourth learning model trained in consideration of the attribute information as well as the position information of the vehicle VE at the time the phrase was uttered, thereby identifying "cafe XX" as an object that is statistically predicted to attract gazes of people with the same or similar attribute information at the current location. In other words, the information providing device 3A identifies "Cafe XX" as an object that, when placed at the location, is likely to attract the attention of female coffee lovers. In the example of Fig. 9, the information providing device 3A acquires the information about Cafe XX as well as the menu and reservation information for Cafe XX as object information corresponding to the information about Cafe XX, and transmits these to the in-vehicle terminal 2 of the vehicle VE.
[0061] In this way, the information providing device 3A can statistically identify an object on which gazes are concentrated using the fourth learning model that takes into account position information and attribute information, and provide object information corresponding to the object to the occupant PA, thereby reducing the processing load compared to when visual saliency technology is used. Note that if the information providing device 3A cannot statistically identify an object on which gazes are concentrated using the fourth learning model, it identifies the object on which gazes are concentrated using visual saliency technology and object recognition, as in the first embodiment.
[0062] Next, an information providing method executed by the information providing device 3A will be described. Fig. 10 is a flowchart showing the information providing method. In the information providing method according to the second embodiment, as shown in Fig. 10, the processing of step S205 is added to the information providing method described in the first embodiment (see Fig. 6). Steps S201 to S204 in Fig. 10 are the same as steps S101 to S104 in Fig. 6, and steps S207 to S212 are the same as steps S106 to S111 in Fig. 6. For this reason, only steps S205 to S206 will be mainly described below.
[0063] Step S206 is executed after step S205. In step S205, the information acquisition unit 324A acquires attribute information of the occupant PA (step S205). Then, the identification unit 325A uses the position information and the attribute information to identify an object on which gazes are statistically focused according to the position information and the attribute information, and determines whether the object on which gazes are focused can be identified (step S206).
[0064] As a result, if the identification unit 325A determines that the object on which the gazes are concentrated cannot be identified (step S206: No), the process proceeds to step S207. On the other hand, if the identification unit 325A determines that the object on which the gazes are concentrated can be identified (step S206: Yes), the process proceeds to step S210.
[0065] In addition, after acquiring the location information in step S204, the information providing device 10A may use only the location information to identify an object on which gazes are statistically concentrated according to the location information, and as a result, determine whether the object on which gazes are concentrated can be identified, and proceed to the processing of step S205 only if it determines that the object on which gazes are concentrated cannot be identified.
[0066] According to the second embodiment described above, in addition to the same effects as those of the first embodiment, the following effects are achieved. The information providing device 3A according to the second embodiment statistically identifies an object on which gazes are concentrated according to position information and attribute information, and then determines whether the object on which gazes are concentrated can be identified. Therefore, it is possible to accurately identify, as a region of interest, an area including an object about which the occupant PA of the vehicle VE wishes to obtain object information. Therefore, it is possible to provide appropriate object information to the occupant PA of the vehicle VE.
[0067] (Embodiment 3) Next, a description will be given of the present embodiment 3. In the following description, the same components as those in the above-described embodiments 1 and 2 are denoted by the same reference numerals, and detailed description thereof will be omitted or simplified.
[0068] Fig. 11 is a block diagram showing the configuration of an information providing device 3B according to embodiment 3. Moreover, the information providing device 3B according to embodiment 3 is different from the information providing device 3A (see Fig. 8) described in the above-mentioned embodiment 2 in that it does not have the image acquiring unit 323, the area extracting unit 326, and the object recognizing unit 327.
[0069] The information acquisition unit 324A acquires attribute information of the occupant PA together with the position information of the vehicle VE, as in the second embodiment. For example, the information acquisition unit 324A acquires, as the attribute information, one or more pieces of information from the in-vehicle terminal 2, including the age of the occupant PA, the sex of the occupant PA, the nationality of the occupant PA, the appearance of the occupant PA, and the language of the occupant PA.
[0070] The identification unit 325A uses the position information and the attribute information to identify an object on which gazes are statistically concentrated according to the position information and the attribute information. For example, the identification unit 325A uses the position information and the attribute information as input data and a learning model that statistically identifies an object on which gazes are concentrated to identify an object on which gazes are concentrated.
[0071] The reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 in the learning model. For example, as a result of the object information provided by the information providing unit 328, the reflecting unit 329 determines whether the object information provided by the information providing unit 328 is information that meets the user's wishes. As a result, if the information meets the user's wishes, the reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 in the learning model as correct answer data, and if the information does not meet the user's wishes, the reflecting unit 329 reflects the result of the object information provided by the information providing unit 328 in the learning model as incorrect answer data. Note that the method of determining whether the information meets the user's wishes may be a method of receiving a manual input from the user, or may be an automatic method of determining based on the user's behavior after the object information is provided.
[0072] Next, an information providing method executed by the information providing device 3B will be described. FIG. 12 is a flowchart showing the information providing method. First, the request information acquisition unit 321 acquires request information (voice information) from the in-vehicle terminal 2 via the communication unit 31 (step S301). After step S301, the voice analysis unit 322 analyzes the request information (voice information) acquired in step S301 (step S302). After step S302, the voice analysis unit 322 determines, as a result of analyzing the request information (voice information) in step S302, whether or not the request information (voice information) includes a specific keyword (step S303). Here, the specific keyword is a word used by the occupant PA of the vehicle VE to request the provision of object information, and examples of the specific keyword include words such as "what," "what is it," "I wonder," and "tell me."
[0073] If it is determined that the specific keyword is not included (step S303: No), the control unit 32 returns to step S301. On the other hand, if it is determined that the specific keyword is included (step S303: Yes), the information acquisition unit 324 acquires the position information of the vehicle VE when a predetermined request is made by the occupant PA (step S304).
[0074] Then, the information acquisition unit 324A acquires attribute information of the occupant PA (step S305). Then, the identification unit 325A uses the position information and the attribute information to identify an object on which gazes are statistically concentrated according to the position information and the attribute information, and determines whether the object on which gazes are concentrated can be identified (step S306).
[0075] As a result, if the identification unit 325A determines that the object on which the gazes are concentrated cannot be identified (step S306: No), the process ends. On the other hand, if the identification unit 325A determines that the object on which the gazes are concentrated can be identified (step S306: Yes), the process proceeds to step S307. Thereafter, the information providing unit 328 acquires object information corresponding to the object identified in step S306 from the object information DB 333 (step S307), and transmits the object information to the in-vehicle terminal 2 via the communication unit 31 (step S308).
[0076] The control unit 252 then controls the operation of at least one of the audio output unit 22 and the display unit 27, and notifies the occupant PA of the vehicle VE of the object information transmitted from the information providing device 3 by audio, text, or an image. For example, if the object OB1 is the "Moulin Rouge," the object information notified to the occupant PA of the vehicle VE is audio such as "That's the Moulin Rouge. They're putting on a spectacular dance show at night." Furthermore, if the object OB1 is not a building but a buffalo, the object information notified to the occupant PA of the vehicle VE is audio such as "That's a buffalo. Buffalos move in groups." The reflection unit 329 then reflects the result of providing the object information in the learning model (step S309).
[0077] According to the third embodiment described above, in addition to the same effects as those of the first and second embodiments described above, the following effect is achieved. When a predetermined request is received from the occupant PA, the information providing device 3B according to the third embodiment acquires position information of the vehicle VE and attribute information indicating the attributes of the occupant PA, and uses the position information and attribute information to identify an object on which the gazes are statistically concentrated, and provides object information related to the identified object. Therefore, there is no need to provide the second learning model DB 332 described in the image acquisition unit 323, the area extraction unit 326, and the object recognition unit 327 described in the first and second embodiments, and the configuration of the information providing device 3B can be simplified.
[0078] (Other embodiments) Up to this point, the embodiments for carrying out the present invention have been described, but the present invention should not be limited to only the above-mentioned embodiments 1 to 3. The information providing devices 3, 3A, and 3B according to the above-mentioned embodiments 1 to 3 have statistically identified objects on which gazes are concentrated using a third learning model that takes into account location information and keywords, or a fourth learning model that takes into account location information, keywords, and attribute information. However, it is also possible to statistically identify objects on which gazes are concentrated using a table without using a learning model.
[0079] For example, the identification unit 325A may identify an object on which gazes are statistically concentrated using a table illustrated in Fig. 13, without using the third learning model. In the table illustrated in Fig. 13, "location information," "keywords," and "objects" on which gazes are statistically concentrated are associated with each other.
[0080] Furthermore, for example, the identification unit 325A may identify an object on which gazes are statistically concentrated using a table such as that shown in Fig. 14, without using the fourth learning model. In the table shown in Fig. 14, "location information," "keywords," "attribute information," and "objects" on which gazes are statistically concentrated are associated with each other. These tables may be stored in the storage unit 33 in advance, or may be added to or changed as appropriate.
[0081] In the above-described first to third embodiments, the information providing devices 3, 3A, and 3B execute each process when they receive request information (voice information) containing a specific keyword. However, the information providing devices according to the present embodiments may be configured to execute each process at all times, even if they do not receive request information (voice information) containing a specific keyword. Furthermore, the request information according to the present embodiments is not limited to voice information, and may be operation information in response to an operation of an operation unit, such as a switch, provided on the in-vehicle terminal 2 by the occupant PA of the vehicle VE.
[0082] In the above-described first to third embodiments, all of the components of the information providing devices 3, 3A, and 3B may be provided in the in-vehicle terminal 2. In this case, the in-vehicle terminal 2 corresponds to the information providing device according to the present embodiment. Also, some of the functions of the control unit 32 and some of the storage unit 33 in the information providing devices 3, 3A, and 3B may be provided in the in-vehicle terminal 2. In this case, the entire information providing system 1 corresponds to the information providing device according to the present embodiment. [Explanation of symbols]
[0083] 3, 3A, 3B Information providing device 321 Request information acquisition unit 322 Audio Analysis Unit 323 Image Acquisition Unit 324, 324A Information acquisition section 325, 325A specific part 326 Region extraction part 327 Object recognition unit 328 Information Provision Department 329 Reflection section
Claims
1. an image acquisition unit that acquires a captured image of the surroundings of the moving object; an information acquisition unit that acquires position information of the moving body when a predetermined request is made by a passenger; an identification unit that uses the position information to identify an object on which gazes are statistically concentrated; a region extraction unit that extracts, based on the captured image, an attention region where the gazes are concentrated, when the identification unit cannot identify the object on which the gazes are concentrated; and an object recognition unit that recognizes an object included in the region of interest within the captured image; an information providing unit that provides object information regarding the object identified by the identifying unit or object information regarding the object included in the region of interest. An information providing device characterized by:
2. The information acquisition unit further acquires attribute information of the occupant along with location information of the moving body, The identification unit uses the position information and the attribute information to identify an object on which gazes are statistically concentrated according to the position information and the attribute information.
2. The information providing device according to claim 1.
3. the information acquisition unit further acquires a specific keyword from a predetermined request content of the occupant along with the location information of the moving body; The identification unit uses the position information and the keyword to identify an object on which gazes are statistically concentrated according to the position information and the keyword.
2. The information providing device according to claim 1.
4. The information providing device according to claim 2, characterized in that the information acquisition unit acquires, as the attribute information, one or more pieces of information from the age of the occupant, the gender of the occupant, the nationality of the occupant, the appearance of the occupant, the preferences of the occupant, and the language of the occupant.
5. the identification unit uses the position information as input data and a learning model that statistically identifies an object on which gazes are concentrated, and identifies the object on which gazes are concentrated; a reflection unit that reflects the object information provided by the information providing unit in the learning model; 2. The information providing device according to claim 1.
6. the identification unit uses the position information and the attribute information as input data and identifies the object on which gazes are concentrated using a learning model that statistically identifies objects on which gazes are concentrated; a reflection unit that reflects the object information provided by the information providing unit in the learning model; 3. The information providing device according to claim 2.
7. the identification unit uses the location information and the keyword as input data and identifies the object on which gazes are concentrated using a learning model that statistically identifies objects on which gazes are concentrated; The information providing device according to claim 3 , further comprising a reflection unit that reflects the result of the object information provided by the information providing unit in the learning model.
8. an information acquisition unit that acquires location information of the moving body and attribute information indicating attributes of the occupant when a predetermined request is received from the occupant; an identification unit that identifies an object on which gazes are statistically concentrated using the position information and the attribute information; an information providing unit that provides object information related to the object identified by the identifying unit.
9. The predetermined request is a voice request.
9. The information providing device according to claim 1 or 8.
10. An information providing method executed by an information providing device, an image acquisition step of acquiring a captured image of the surroundings of the moving object; an acquisition step of acquiring position information of the moving body and attribute information of the occupant when a predetermined request is received from the occupant; a step of identifying an object on which gazes are statistically concentrated using the position information; a region extraction step of extracting a region of interest in the captured image where the gazes are concentrated, if the object on which the gazes are concentrated cannot be identified by the identification step; an object recognition step of recognizing an object included in the region of interest in the captured image; and an information provision step of providing object information regarding the object identified by the identification step or object information regarding the object included in the region of interest.
1. An information providing method comprising:
11. an image acquisition step of acquiring a captured image of the surroundings of the moving object; an acquisition step of acquiring position information of the moving body and attribute information of the occupant when a predetermined request is received from the occupant; a step of identifying an object on which gazes are statistically concentrated using the position information; a region extraction step of extracting a region of interest in the captured image where the gazes are concentrated, if the object on which the gazes are concentrated cannot be identified by the identification step; an object recognition step of recognizing an object included in the region of interest in the captured image; and an information provision step of providing object information regarding the object identified by the identification step or object information regarding the object included in the region of interest. An information providing program that causes a computer to execute the above.
12. an image acquisition step of acquiring a captured image of the surroundings of the moving object; an acquisition step of acquiring position information of the moving body and attribute information of the occupant when a predetermined request is received from the occupant; a step of identifying an object on which gazes are statistically concentrated using the position information; a region extraction step of extracting a region of interest in the captured image where the gazes are concentrated, if the object on which the gazes are concentrated cannot be identified by the identification step; an information provision program for causing a computer to execute an object recognition step of recognizing an object included in the region of interest in the captured image; and an information provision step of providing object information on the object identified by the identification step or object information on the object included in the region of interest. A storage medium characterized by:
Citation Information
Patent Citations
Device and method for guiding object
JP2003329463A
Object specification device
JP2007080060A
Watching object detector and watching object detection method
JP2008082822A
System and method for informational guidance for vehicle, and computer program
JP2009031065A
Driving support device
JP2013011483A
Cited By
Combined machining process and combined machining program
DE102017011495B4