Information provision apparatus, information provision method, information provision program and storage medium
The information providing apparatus addresses the inconvenience of manual object indication by using image and posture detection to automatically provide object information based on passenger gaze and voice requests, improving convenience and accuracy.
Patent Information
- Application Number
- JP2025074280
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-01-21
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2041-01-14
AI Technical Summary
Existing object specifying devices require manual operation by a vehicle occupant to indicate an object, limiting convenience.
An information providing apparatus that includes an image acquisition unit, posture detection unit, region extraction unit, and object recognition unit to automatically identify and provide information about objects in a vehicle's surroundings based on passenger gaze and voice requests.
Enhances convenience by allowing passengers to receive object information without manual indication, reduces processing load by responsive information provision, and ensures accurate object recognition using visual saliency and learning models.
Smart Images

Figure 2025105844000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information providing apparatus, an information providing method, an information providing program, and a storage medium.
Background Art
[0002] Conventionally, an object specifying device that specifies an object existing around a vehicle and reads out information such as the name of the object by voice is known (see, for example, Patent Document 1). In the object specifying device described in Patent Document 1, a facility or the like on a map existing in the indicated direction indicated by a vehicle occupant with a hand or a finger is specified as an object.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the technique described in Patent Document 1, it is necessary to cause a vehicle occupant who desires to obtain information about an object to perform an operation of indicating the object with a hand or a finger, and there is a problem that convenience cannot be improved, for example.
[0005] The present invention has been made in view of the above, and an object thereof is to provide an information providing apparatus, an information providing method, an information providing program, and a storage medium that can improve convenience, for example.
Means for Solving the Problems
[0006] The information providing apparatus according to claim 1 includes an image acquisition unit that acquires a captured image obtained by photographing the surroundings of a moving body, a posture detection unit that detects the posture of a passenger in the moving body, a region extraction unit that extracts a plurality of attention regions where the line of sight in the captured image converges, an object recognition unit that, based on the posture, specifies one of the plurality of attention regions and recognizes an object included in the specified attention region, and an information providing unit that provides object information regarding the object included in the attention region.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments for carrying out the present invention (hereinafter referred to as embodiments) will be described with reference to the drawings. Note that the present invention is not limited by the embodiments described below. Further, in the description of the drawings, the same parts are denoted by the same reference numerals.
[0009] (Embodiment 1) 〔Schematic Configuration of Information Providing System〕 FIG. 1 is a block diagram showing the configuration of an information providing system 1 according to Embodiment 1. The information providing system 1 is a system that provides object information (for example, the name of the object, etc.) regarding an object such as a building existing around the vehicle VE (FIG. 1), which is a moving body, to the passengers PA (see FIG. 5) of the vehicle VE. As shown in FIG. 1, this information providing system 1 includes an in-vehicle terminal 2 and an information providing device 3. And these in-vehicle terminal 2 and information providing device 3 communicate with each other via a network NE (FIG. 1), which is a wireless communication network. Note that, as an example in FIG. 1, the in-vehicle terminal 2 that communicates with the information providing device 3 is shown as one unit, but it may be a plurality of units respectively mounted on a plurality of vehicles. Also, in order to provide object information to a plurality of passengers riding in one vehicle respectively, a plurality of in-vehicle terminals 2 may be mounted on one vehicle.
[0010] 〔Configuration of In-vehicle Terminal〕 FIG. 2 is a block diagram showing the configuration of the in-vehicle terminal 2. The in-vehicle terminal 2 is, for example, a stationary navigation device or a drive recorder installed in the vehicle VE. Note that the in-vehicle terminal 2 is not limited to a navigation device or a drive recorder, and a portable terminal such as a smartphone used by the occupant PA of the vehicle VE may be adopted. As shown in FIG. 2, this in-vehicle terminal 2 includes an audio input unit 21, an audio output unit 22, an imaging unit 23, a display unit 24, and a terminal main body 25.
[0011] The audio input unit 21 includes a microphone 211 (see FIG. 5) that inputs audio and converts it into an electrical signal, and generates audio information by performing A / D (Analog / Digital) conversion or the like on the electrical signal. In the first embodiment, the audio information generated by the audio input unit 21 is a digital signal. Then, the audio input unit 21 outputs the audio information to the terminal main body 25. The audio output unit 22 includes a speaker 221 (see FIG. 5), converts the digital audio signal input from the terminal main body 25 into an analog audio signal by D / A (Digital / Analog) conversion, and outputs audio corresponding to the analog audio signal from the speaker 221.
[0012] The imaging unit 23 captures the surroundings of the vehicle VE under the control of the terminal main body 25 and generates a captured image. Then, the imaging unit 23 outputs the generated captured image to the terminal main body 25. The display unit 24 is composed of a display display using liquid crystal or organic EL (Electro Luminescence) or the like, and displays various images under the control of the terminal main body 25.
[0013] As shown in FIG. 2, the terminal main body 25 includes a communication unit 251, a control unit 252, and a storage unit 253. The communication unit 251 transmits and receives information to and from the information providing device 3 via the network NE under the control of the control unit 252. The control unit 252 is realized by executing various programs stored in the storage unit 253 by a controller such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and controls the operation of the in-vehicle terminal 2 as a whole. Note that the control unit 252 may be configured by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), not limited to a CPU or an MPU. The storage unit 253 stores various programs executed by the control unit 252, data necessary when the control unit 252 performs processing, and the like.
[0014] 〔Configuration of Information Provision Device〕 FIG. 3 is a block diagram showing the configuration of the information provision device 3. The information provision device 3 is, for example, a server device. As shown in FIG. 3, this information provision device 3 includes a communication unit 31, a control unit 32, and a storage unit 33.
[0015] The communication unit 31 transmits and receives information to and from the in-vehicle terminal 2 (communication unit 251) via the network NE under the control of the control unit 32. The control unit 32 is realized by executing various programs (including the information provision program according to the present embodiment) stored in the storage unit 33 by a controller such as a CPU or an MPU, and controls the operation of the information provision device 3 as a whole. Note that the control unit 32 may be configured by an integrated circuit such as an ASIC or an FPGA, not limited to a CPU or an MPU. As shown in FIG. 3, this control unit 32 includes a request information acquisition unit 321, a voice analysis unit 322, an image acquisition unit 323, a region extraction unit 324, an object recognition unit 325, and an information provision unit 326.
[0016] The request information acquisition unit 321 acquires request information for requesting the provision of object information from the vehicle VE's occupant PA. In the first embodiment, the request information is voice information generated by the voice input unit 21 based on the voice (sound) uttered by the vehicle VE's occupant PA and captured by the voice input unit 21. That is, the request information acquisition unit 321 acquires the request information (voice information) from the in-vehicle terminal 2 via the communication unit 31. The voice analysis unit 322 analyzes the request information (voice information) acquired by the request information acquisition unit 321.
[0017] The image acquisition unit 323 acquires the captured image generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31. The region extraction unit 324 extracts (predicts) a region of interest where the line of sight converges (is likely to converge) within the captured image acquired by the image acquisition unit 323. In the first embodiment, the region extraction unit 324 extracts the region of interest within the captured image using a so-called visual saliency technique. More specifically, the region extraction unit 324 extracts the region of interest within the captured image by performing image recognition (image recognition using AI (Artificial Intelligence)) using the first learning model shown below. The first learning model is a model obtained by discriminating the region where the subject's line of sight converges using an eye tracker, using the image in which the region is pre-labeled as a teacher image, and performing machine learning (e.g., deep learning, etc.) on the region using the teacher image.
[0018] The object recognition unit 325 recognizes the objects included in the region of interest extracted by the region extraction unit 324 within the captured image. In the first embodiment, the object recognition unit 325 recognizes the objects included in the region of interest within the captured image by performing image recognition (image recognition using AI) using the second learning model shown below. The second learning model is a model obtained by using captured images of various objects such as animals, mountains, rivers, lakes, and facilities as teacher images and performing machine learning (e.g., deep learning, etc.) on the characteristics of the objects based on the teacher images.
[0019] The information providing unit 326 provides object information regarding the object recognized by the object recognition unit 325. More specifically, the information providing unit 326 reads out the object information corresponding to the object recognized by the object recognition unit 325 from the object information DB (Data Base) 333 in the storage unit 33. Then, the information providing unit 326 transmits the object information to the in-vehicle terminal 2 via the communication unit 31.
[0020] The storage unit 33 stores various programs (the information providing program according to the present embodiment) executed by the control unit 32, as well as data and the like necessary when the control unit 32 performs processing. As shown in FIG. 3, this storage unit 33 includes a first learning model DB 331, a second learning model DB 332, and an object information DB 333. The first learning model DB 331 stores the first learning model described above. The second learning model DB 332 stores the second learning model described above. The object information DB 333 stores the object information described above. Here, the object information DB 333 stores a plurality of object information associated with various objects. The object information is information for explaining the object such as the name of the object, and is composed of character data, voice data, or image data.
[0021] 〔Information providing method〕 Next, an information providing method executed by the information providing apparatus 3 (control unit 32) will be described. FIG. 4 is a flowchart showing the information providing method. FIG. 5 is a diagram for explaining the information providing method. Specifically, FIG. 5 is a diagram showing a captured image IM generated by the imaging unit 23 and acquired in step S4. Here, in FIG. 5, a case where the imaging unit 23 is installed in the vehicle VE so that the front of the vehicle VE is photographed through the windshield from inside the vehicle VE is illustrated. Further, in FIG. 5, a case where a passenger PA sitting in the passenger seat of the vehicle VE is included as a subject in the captured image IM is illustrated. Furthermore, in FIG. 5, a case where the passenger PA is saying the words "What's that?" is illustrated. Note that the installation position of the imaging unit 23 is not limited to the above-described installation position. For example, the imaging unit 23 may be installed inside the vehicle VE so that the left side, right side, or rear of the vehicle VE is photographed from inside the vehicle VE, or the imaging unit 23 may be installed outside the vehicle VE so that the surroundings of the vehicle VE are photographed. Further, the passengers of the vehicle according to the present embodiment include not only the passengers sitting in the passenger seat of the vehicle VE, but also the passengers sitting in the driver's seat, rear seats, etc. Further, the number of imaging units 23 is not limited to one, and a plurality of imaging units may be provided.
[0022] First, the request information acquisition unit 321 acquires request information (voice information) from the in-vehicle terminal 2 via the communication unit 31 (step S1). After step S1, the voice analysis unit 322 analyzes the request information (voice information) acquired in step S1 (step S2). After step S2, the voice analysis unit 322 determines whether or not a specific keyword is included in the request information (voice information) as a result of analyzing the request information (voice information) in step S2 (step S3). Here, the specific keyword is a word in which the passenger PA of the vehicle VE requests the provision of object information, and examples thereof include words such as "what", "what is it", "what could it be", "tell me".
[0023] When it is determined that the specific keyword is not included (step S3: No), the control unit 32 returns to step S1. On the other hand, when it is determined that the specific keyword is included (step S3: Yes), the image acquisition unit 323 acquires the captured image IM generated by the imaging unit 23 from the in-vehicle terminal 2 via the communication unit 31 (step S4: image acquisition step). In FIGS. 4 and 5, the image acquisition unit 323 is configured to acquire, via the communication unit 31, the captured image IM generated by the imaging unit 23 of the in-vehicle terminal 2 at the timing (step S3: Yes) when the vehicle occupant PA utters "What's that?", but the present invention is not limited to this configuration. For example, the information providing apparatus 3 sequentially acquires the captured images generated by the imaging unit 23 of the in-vehicle terminal 2 via the communication unit 31. Then, the image acquisition unit 323 may be configured to acquire, as the captured image to be used for the processing after step S4, the captured image acquired at the timing (step S3: Yes) when the vehicle occupant PA utters "What's that?" among the sequentially acquired captured images.
[0024] After step S4, the region extraction unit 324 extracts, by image recognition using the first learning model stored in the first learning model DB 331, the attention region Ar1 (FIG. 5) where the line of sight in the captured image IM converges (step S5: region extraction step). After step S5, the object recognition unit 325 recognizes, by image recognition using the second learning model stored in the second learning model DB 332, the object OB1 included in the attention region Ar1 extracted in step S5 within the captured image IM (step S6: object recognition step). After step S6, the information providing unit 326 reads out the object information corresponding to the object OB1 recognized in step S6 from the object information DB 333, and transmits the object information to the in-vehicle terminal 2 via the communication unit 31 (step S7: information providing step). Then, the control unit 252 controls the operations of at least one of the voice output unit 22 and the display unit 24, and notifies the passenger PA of the vehicle VE of the object information transmitted from the information providing device 3 by at least one of voice, characters, and images. For example, when the object OB1 is "Moulin Rouge", voices such as "That is Moulin Rouge. It is having a magnificent dance show at night." are notified to the passenger PA of the vehicle VE as the object information. Also, for example, when the object OB1 is a buffalo, which is an animal rather than a building, voices such as "That is a buffalo. Buffaloes move in groups." are notified to the passenger PA of the vehicle VE as the object information.
[0025] According to the first embodiment described above, the following effects can be obtained. The information providing device 3 according to the first embodiment acquires a captured image IM that captures the surroundings of the vehicle VE, and extracts a attention area Ar1 where the line of sight in the captured image IM converges. Then, the information providing device 3 recognizes the object OB1 included in the attention area Ar1 in the captured image IM, and transmits the object information regarding the object OB1 to the in-vehicle terminal 2. As a result, the passenger PA of the vehicle VE who desires to obtain the object information regarding the object OB1 recognizes the object information regarding the object OB1 when the object information is notified from the in-vehicle terminal 2. Therefore, it is not necessary to make the passenger PA of the vehicle VE who desires to obtain the object information regarding the object OB1 perform the operation of pointing at the object OB1 with a hand or a finger as in the prior art, and the convenience can be improved.
[0026] In particular, the information providing device 3 extracts the attention area Ar1 where the line of sight in the captured image IM converges by using a so-called visual saliency technique. Therefore, even if the passenger PA of the vehicle VE does not point at the object OB1 with a hand or a finger, the area including the object OB1 can be accurately extracted as the attention area Ar1.
[0027] In addition, the information providing device 3 provides the object information in response to the request information for requesting the provision of the object information from the occupant PA of the vehicle VE. Therefore, compared with the configuration that always provides the object information regardless of the request information, the processing load of the information providing device 3 can be reduced.
[0028] (Embodiment 2) Next, Embodiment 2 will be described. In the following description, the same components as those in Embodiment 1 described above are denoted by the same reference numerals, and the detailed description thereof is omitted or simplified. FIG. 6 is a block diagram showing the configuration of the information providing device 3A according to Embodiment 2. In the information providing device 3A according to Embodiment 2, as shown in FIG. 6, the function of the posture detection unit 327 is added to the control unit 32 with respect to the information providing device 3 (see FIG. 3) described in Embodiment 1 above. In addition, in the information providing device 3A, the function of the object recognition unit 325 is changed. Hereinafter, for convenience of explanation, the object recognition unit according to Embodiment 2 is described as the object recognition unit 325A (see FIG. 6). Further, in the information providing device 3A, a third learning model DB334 (see FIG. 6) is added to the storage unit 33.
[0029] The posture detection unit 327 detects the posture of the occupant PA of the vehicle VE. In Embodiment 2, the posture detection unit 327 detects the posture by so-called skeleton detection. More specifically, the posture detection unit 327 detects the posture of the occupant PA of the vehicle VE included as a subject in the captured image IM by image recognition using the third learning model shown below (image recognition using AI), that is, by detecting the skeleton of the occupant PA. The third learning model is a model obtained by machine learning (for example, deep learning, etc.) of the positions of the joint points of a person based on a teacher image in which the positions of the joint points of the person are pre-labeled for a captured image of the person. And the third learning model DB334 stores the third learning model.
[0030] The object recognition unit 325A has the same functions as the object recognition unit 325 described in the above-described Embodiment 1, and also has a function (hereinafter referred to as an additional function) that is executed when a plurality of regions of interest are extracted from the captured image IM by the region extraction unit 324. The additional function is as follows. That is, based on the posture of the passenger PA detected by the posture detection unit 327, the object recognition unit 325A specifies one of the plurality of regions of interest. Then, in the same manner as the object recognition unit 325 described in the above-described Embodiment 1, the object recognition unit 325A recognizes an object included in the specified one region of interest in the captured image IM by image recognition using the second learning model.
[0031] Next, an information providing method executed by the information providing apparatus 3A will be described. FIG. 7 is a flowchart showing the information providing method. FIG. 8 is a diagram for explaining the information providing method. Specifically, FIG. 8 is a diagram corresponding to FIG. 5, and shows the captured image IM generated by the imaging unit 23 and acquired in step S4. In the information providing method according to the second embodiment, as shown in FIG. 7, steps S6A1 to S6A3 are added to the information providing method (see FIG. 4) described in the above-described Embodiment 1. Therefore, hereinafter, only steps S6A1 to S6A3 will be mainly described. The steps S6A1 to S6A3 and S6 correspond to the object recognition steps according to the present embodiment.
[0032] Step S6A1 is executed after step S5. Specifically, in step S6A1, the control unit 32 determines whether or not the number of regions of interest extracted in step S5 is plural. In FIG. 8, the case where three regions of interest Ar1 to Ar3 are extracted in step S5 is illustrated as an example. When it is determined that the number of regions of interest is one (step S6A1: No), the control unit 32 proceeds to step S6 and recognizes an object (for example, object OB1) included in the one region of interest (for example, region of interest Ar1 as in the above-described Embodiment 1).
[0033] On the other hand, when the control unit 32 determines that there are a plurality of regions of interest (step S6A1: Yes), the control unit 32 proceeds to step S6A2. Then, in step S6A2, the posture detection unit 327 detects the posture of the occupant PA of the vehicle VE included as a subject in the captured image IM by performing image recognition using the third learning model stored in the third learning model DB 334, thereby detecting the skeleton of the occupant PA.
[0034] After step S6A2, the object recognition unit 325A specifies the direction DI (FIG. 8) of the face FA and fingers FI of the occupant PA from the posture of the occupant PA detected in step S6A2. Then, in the captured image IM, the object recognition unit 325A specifies one region of interest Ar2 located in the direction DI with respect to the occupant PA among the three regions of interest Ar1 to Ar3 extracted in step S5 (step S6A3). After step S6A3, the control unit 32 proceeds to step S6 and recognizes the object OB2 (FIG. 8) included in the one region of interest Ar2.
[0035] According to the second embodiment described above, in addition to the same effects as those of the first embodiment described above, the following effects are obtained. When the information providing apparatus 3A according to the second embodiment extracts a plurality of regions of interest Ar1 to Ar3 in the captured image IM, the information providing apparatus 3A detects the posture of the occupant PA of the vehicle VE, and based on the posture, specifies one region of interest Ar2 from the plurality of regions of interest Ar1 to Ar3. Then, the information providing apparatus 3 recognizes the object OB2 included in the specified region of interest Ar2. For this reason, even when a plurality of regions of interest Ar1 to Ar3 are extracted in the captured image IM, it is possible to accurately specify, as the region of interest Ar1, the region including the object OB2 that the occupant PA of the vehicle VE desires to obtain object information. Therefore, appropriate object information can be provided to the occupant PA of the vehicle VE.
[0036] In particular, the information providing device 3A detects the posture of the vehicle VE occupant PA by so-called skeleton detection. Therefore, even when a plurality of attention areas Ar1 to Ar3 are extracted in the captured image IM, appropriate object information can be provided to the vehicle VE occupant PA with high precision.
[0037] (Embodiment 3) Next, Embodiment 3 will be described. In the following description, the same components as those in Embodiment 1 described above are denoted by the same reference numerals, and their detailed descriptions are omitted or simplified. FIG. 9 is a block diagram showing the configuration of the in-vehicle terminal 2B according to Embodiment 3. In the in-vehicle terminal 2B according to Embodiment 3, as shown in FIG. 9, a sensor unit 26 is added to the in-vehicle terminal 2 (see FIG. 2) described in Embodiment 1 above. As shown in FIG. 9, the sensor unit 26 includes a lidar 261 and a GNSS (Global Navigation Satellite System) sensor 262. The lidar 261 discretely measures the distance to an object existing in the external world, recognizes the surface of the object as a three-dimensional point cloud, and generates point cloud data. Note that, as long as it is a sensor capable of measuring the distance to an object existing in the external world, other external sensors such as a millimeter-wave radar and a sonar may be employed instead of the lidar 261. The GNSS sensor 262 receives radio waves including positioning data transmitted from navigation satellites using GNSS. The positioning data is used to detect the absolute position of the vehicle VE from latitude and longitude information and the like, and corresponds to the position information according to the present embodiment. Note that the GNSS to be used may be, for example, GPS (Global Positioning System) or another system. Then, the sensor unit 26 outputs output data such as the point cloud data and the positioning data to the terminal main body 25.
[0038] FIG. 10 is a block diagram showing the configuration of the information providing device 3B according to Embodiment 3. Further, in the information providing apparatus 3B according to the third embodiment, the function of the object recognition unit 325 is changed with respect to the information providing apparatus 3 (see FIG. 3) described in the above-described first embodiment. Hereinafter, for convenience of explanation, the object recognition unit according to the third embodiment will be described as the object recognition unit 325B (see FIG. 10). In the information providing apparatus 3B, the second learning model DB 332 is omitted, and a map DB 335 (see FIG. 10) is added to the storage unit 33.
[0039] The map DB 335 stores map data. The map data includes road data represented by links corresponding to roads and nodes corresponding to connection portions (intersections) of the roads, facility information in which each facility and the position of each facility (hereinafter referred to as the facility position) are associated with each other, and the like. The object recognition unit 325B acquires output data (point cloud data generated by the lidar 261 and positioning data received by the GNSS sensor 262) of the sensor unit 26 from the in-vehicle terminal 2 via the communication unit 31. Then, the object recognition unit 325B recognizes an object included in the attention area extracted by the area extraction unit 324 in the captured image IM based on the output data, the captured image IM, and the map data stored in the map DB 335. The object recognition unit 325B described above corresponds to a position information acquisition unit and a facility information acquisition unit in addition to the object recognition unit according to the present embodiment.
[0040] Next, an information providing method executed by the information providing apparatus 3B will be described. FIG. 11 is a flowchart showing the information providing method. In the information providing method according to the third embodiment, as shown in FIG. 11, steps S6B1 to S6B5 are added instead of step S6 with respect to the information providing method (see FIG. 4) described in the above-described first embodiment. Therefore, hereinafter, only steps S6B1 to S6B5 will be mainly described. The steps S6B1 to S6B5 correspond to the object recognition steps according to the present embodiment.
[0041] Step S6B1 is executed after step S5. Specifically, in step S6B1, the object recognition unit 325B acquires, via the communication unit 31, the output data of the sensor unit 26 (point cloud data generated by the lidar 261 and positioning data generated by the GNSS sensor 262) from the in-vehicle terminal 2. In FIG. 11, the object recognition unit 325B is configured to acquire the output data of the sensor unit 26 from the in-vehicle terminal 2 via the communication unit 31 at the timing when the occupant PA of the vehicle VE utters words including a specific keyword (step S3: Yes). However, the present invention is not limited to this configuration. For example, the information providing device 3B sequentially acquires the output data of the sensor unit 26 from the in-vehicle terminal 2 via the communication unit 31. Then, the object recognition unit 325B may be configured to acquire, as the output data to be used in the processes after step S6B1, the output data acquired at the timing when the occupant PA of the vehicle VE utters words including a specific keyword (step S3: Yes) among the sequentially acquired output data.
[0042] After step S6B1, the object recognition unit 325B estimates the position of the vehicle VE based on the output data (positioning data received by the GNSS sensor 262) acquired in step S6B1 and the map data stored in the map DB 335 (step S6B2). After step S6B2, the object recognition unit 325B estimates the position of the object included in the attention area in the captured image IM extracted in step S5 (step S6B3). Here, the object recognition unit 325B estimates the position of the object by using the output data (point cloud data) acquired in step S6B1, the position of the vehicle VE estimated in step S6B2, and the position of the attention area in the captured image IM extracted in step S5.
[0043] After step S6B3, the object recognition unit 325B acquires facility information including a facility position substantially the same as the position of the object estimated in step S6B3 from the map DB 335 (step S6B4). After step S6B4, the object recognition unit 325B recognizes the facility included in the facility information acquired in step S6B4 as an object included in the attention area in the captured image IM extracted in step S5 (step S6B5). Then, after step S6B5, the control unit 32 proceeds to step S7.
[0044] According to the third embodiment described above, in addition to the same effects as those of the first embodiment described above, the following effects are obtained. The information providing apparatus 3B according to the third embodiment recognizes an object included in the attention area in the captured image IM based on the position information (positioning data received by the GNSS sensor 262) and the facility information. In other words, the information providing apparatus 3B recognizes an object included in the attention area in the captured image IM based on the information (position information and facility information) commonly used in the navigation apparatus. Therefore, it is not necessary to provide the second learning model DB332 described in the first embodiment above, and the configuration of the information providing apparatus 3B can be simplified.
[0045] (Embodiment 4) Next, the fourth embodiment will be described. In the following description, the same reference numerals are given to the same configurations as those in the first embodiment described above, and the detailed description thereof is omitted or simplified. FIG. 12 is a block diagram showing the configuration of an information providing apparatus 3C according to the fourth embodiment. In the information providing apparatus 3C according to the fourth embodiment, as shown in FIG. 12, the functions of the object recognition unit 325 and the information providing unit 326 are changed with respect to the information providing apparatus 3 (see FIG. 3) described in the first embodiment above. Hereinafter, for convenience of explanation, the object recognition unit according to the fourth embodiment is referred to as the object recognition unit 325C (see FIG. 12), and the information providing unit according to the fourth embodiment is referred to as the information providing unit 326C (see FIG. 12).
[0046] The object recognition unit 325C has, in addition to the same functions as the object recognition unit 325 described in the above-described Embodiment 1, a function (hereinafter referred to as an additional function) that is executed when a plurality of regions of interest are extracted from the captured image IM by the region extraction unit 324. The additional function is as follows. That is, the object recognition unit 325C recognizes, by image recognition using the second learning model, the objects included in each of the plurality of regions of interest in the captured image IM.
[0047] The information providing unit 326C has, in addition to the same functions as the information providing unit 326 described in the above-described Embodiment 1, a function (hereinafter referred to as an additional function) that is executed when a plurality of regions of interest are extracted from the captured image IM by the region extraction unit 324. The additional function is as follows. That is, the information providing unit 326C specifies one object from each of the objects recognized by the object recognition unit 325C based on the analysis result by the voice analysis unit 322 and the object information stored in the object information DB 333. Then, the information providing unit 326C transmits the object information corresponding to the specified one object to the in-vehicle terminal 2 via the communication unit 31.
[0048] Next, an information providing method executed by the information providing apparatus 3C will be described. FIG. 13 is a flowchart showing the information providing method. FIG. 14 is a diagram for explaining the information providing method. Specifically, FIG. 14 is a diagram corresponding to FIG. 5 and shows the captured image IM generated by the imaging unit 23 and acquired in step S4. Here, in FIG. 14, different from the example of FIG. 5, a case is exemplified in which a passenger PA sitting in the passenger seat of the vehicle VE utters the words "What is that red building?" In the information providing method according to the fourth embodiment, as shown in FIG. 13, steps S6C1, S6C2, and S7C are added to the information providing method (see FIG. 4) described in the above-described first embodiment. Therefore, hereinafter, only steps S6C1, S6C2, and S7C will be mainly described. Steps S6C1 and S6C2 and step S6 respectively correspond to the object recognition steps according to the present embodiment. Also, step S7C and step S7 respectively correspond to the information providing steps according to the present embodiment.
[0049] Step S6C1 is executed after step S5. Specifically, in step S6C1, the control unit 32 determines whether or not there are a plurality of regions of interest extracted in step S5, in the same manner as step S6A1 described in the above-described second embodiment. Note that in FIG. 14, as in FIG. 8, the case where three regions of interest Ar1 to Ar3 are extracted in step S5 is illustrated. When it is determined that there is one region of interest (step S6C1: No), the control unit 32 proceeds to step S6 and recognizes the object (for example, object OB1) included in the one region of interest (for example, region of interest Ar1 as in the above-described first embodiment).
[0050] On the other hand, when it is determined that there are a plurality of regions of interest (step S6C1: Yes), the control unit 32 proceeds to step S6C2. Then, the object recognition unit 325C recognizes the objects OB1 to OB3 respectively included in the three regions of interest Ar1 to Ar3 extracted in step S5 in the captured image IM by image recognition using the second learning model stored in the second learning model DB332 (step S6C2).
[0051] After step S6C2, the information providing unit 326C executes step S7C. Specifically, in step S7C, the information providing unit 326C identifies one object from each of the objects recognized in step S6C2. Here, the information providing unit 326C identifies the one object based on the attributes of the objects included in the request information (voice information) and the three pieces of object information corresponding to each of the objects OB1 to OB3 recognized in step S6C2 among the object information stored in the object information DB 333. Note that the attributes of the objects included in the request information (voice information) are generated by analyzing the request information (voice information) in step S2. For example, as shown in FIG. 14, when the passenger PA of the vehicle VE says "What's that red building?", the words "red" and "building" become the attributes of the objects. Specifically, the attributes of the objects are information indicating colors such as red, shapes such as squares, and types such as buildings. Then, in step S7C, the information providing unit 326C refers to the three pieces of object information corresponding to each of the objects OB1 to OB3, and identifies one object (for example, object OB3) corresponding to the object information including the character data of "red" and "building". Further, the information providing unit 326C transmits the object information corresponding to the identified one object to the in-vehicle terminal 2 via the communication unit 31.
[0052] According to the fourth embodiment described above, in addition to the same effects as those of the first embodiment described above, the following effects are obtained. When the information providing device 3C according to the fourth embodiment extracts a plurality of attention areas Ar1 to Ar3 in the captured image IM, based on the analysis result of the request information (voice information), it provides object information regarding one object among the objects OB1 to OB3 respectively included in the plurality of attention areas Ar1 to Ar3. Therefore, even when a plurality of attention areas Ar1 to Ar3 are extracted in the captured image IM, the object OB3 that the passenger PA of the vehicle VE desires to obtain object information can be accurately identified. Therefore, appropriate object information can be provided to the passenger PA of the vehicle VE.
[0053] (Other Embodiments) So far, the embodiments for carrying out the present invention have been described. However, the present invention should not be limited only by the above-described Embodiments 1 to 4. The information providing apparatuses 3, 3A to 3C according to the above-described Embodiments 1 to 4 executed each process such as an image acquisition step, a region extraction step, an object recognition step, and an information providing step triggered by acquiring request information (voice information) including a specific keyword. However, as the information providing apparatus according to the present embodiment, it may be configured to always execute each of the above processes without acquiring request information (voice information) including a specific keyword. Further, the request information according to the present embodiment may be not limited to voice information, but may be operation information corresponding to an operation on an operation unit such as a switch provided in the in-vehicle terminals 2, 2B by a passenger PA of the vehicle VE.
[0054] In the above-described Embodiments 1 to 4, all configurations of the information providing apparatuses 3, 3A to 3C may be provided in the in-vehicle terminals 2, 2B. In this case, the in-vehicle terminals 2, 2B correspond to the information providing apparatus according to the present embodiment. Further, a part of the functions of the control unit 32 and a part of the storage unit 33 in the information providing apparatuses 3, 3A to 3C may be provided in the in-vehicle terminals 2, 2B. In this case, the entire information providing system 1 corresponds to the information providing apparatus according to the present embodiment.
Explanation of Reference Numerals
[0055] 3, 3A to 3C Information providing apparatus 321 Request information acquisition unit 322 Voice analysis unit 323 Image acquisition unit 324 Region extraction unit 325, 325A to 325C Object recognition unit 326, 326C Information providing unit 327 Posture detection unit
Claims
1. An image acquisition unit that acquires a captured image of the surroundings of a moving body; A posture detection unit that detects the posture of an occupant in the moving body; A region extraction unit that extracts a plurality of attention regions where the line of sight in the captured image converges; An object recognition unit that identifies one of the plurality of attention regions based on the posture and recognizes an object included in the identified attention region; An information providing unit that provides object information regarding an object included in the attention region. An information providing apparatus characterized by the above.
2. The captured image includes an occupant in the moving body as a subject, and the posture detection unit detects the posture by detecting the skeleton of the occupant based on the captured image. The information providing apparatus according to claim 1, characterized by the above.
3. A position information acquisition unit that acquires position information regarding the position of the moving body; A facility information acquisition unit that acquires facility information regarding a facility; and further includes, The object recognition unit recognizes an object included in the attention region based on the position information and the facility information. The information providing apparatus according to claim 1 or 2, characterized by the above.
4. A request information acquisition unit that acquires request information for requesting the provision of the object information from an occupant in the moving body; and further includes, The information providing unit provides the object information according to the request information. The information providing apparatus according to any one of claims 1 to 3, characterized by the above.
5. The request information is voice information regarding voice uttered by the occupant, and further includes a voice analysis unit that analyzes the voice information, the region extraction unit extracts a plurality of the attention regions, the object recognition unit recognizes each object included in the plurality of attention regions, and the information providing unit provides the object information regarding any one of the objects included in the plurality of attention regions based on the analysis result of the voice information. The information providing apparatus according to claim 4, characterized by the above.
6. An information providing method executed by an information providing apparatus, comprising: An image acquisition step of acquiring a captured image of the surroundings of a moving body; A posture detection step of detecting the posture of an occupant in the moving body; A region extraction step of extracting a plurality of attention regions where the line of sight in the captured image converges; An object recognition step of identifying one of the plurality of attention regions based on the posture and recognizing an object included in the identified attention region; including an information providing step of providing object information regarding an object included in the target area An information providing method characterized by the above.
7. An image acquisition step of acquiring a captured image obtained by photographing the surroundings of a moving body, A posture detection step of detecting the posture of a passenger in the moving body, An area extraction step of extracting a plurality of target areas where the line of sight in the captured image converges, An object recognition step of recognizing an object included in the target area in the captured image, An information providing step of specifying one of the plurality of target areas based on the posture and providing object information regarding the object included in the specified target area An information providing program for causing a computer to execute the above.
8. An image acquisition step of acquiring a captured image obtained by photographing the surroundings of a moving body, A posture detection step of detecting the posture of a passenger in the moving body, An area extraction step of extracting a plurality of target areas where the line of sight in the captured image converges, An object recognition step of specifying one of the plurality of target areas based on the posture and recognizing the object included in the specified target area, A storage medium storing an information providing program for causing a computer to execute an information providing step of providing object information regarding the object included in the target area Characterized by the above.
Citation Information
Patent Citations
Driving assistance system
JP1994251287A
Information providing apparatus for vehicle
JP2004030212A
Device and method for inputting voice
JP2006251298A
Image processor, image processing method, image processing program and recording medium
JP2014207614A
Display image generator and display image generation method
JP2021081372A