Information processing device, information processing method, and information processing program
The information processing device aligns content with the character of nearby anthropomorphic features by outputting sound from their direction, addressing discomfort in existing anthropomorphization technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PIONEER IP
- Filing Date
- 2023-06-22
- Publication Date
- 2026-06-02
AI Technical Summary
Existing information processing devices anthropomorphize objects, but may mismatch the content with the character of the object, causing discomfort when speaking arbitrary or timely content.
An information processing device that acquires content, identifies anthropomorphic features within a user's vicinity based on location and feature-related information, and outputs sound as if it originates from those features, aligning the content with the object's character.
The device provides a natural feeling of the object speaking the content, eliminating discomfort by ensuring the content aligns with the object's character.
Smart Images

Figure 0007869336000001 
Figure 0007869336000002 
Figure 0007869336000003
Abstract
Description
Technical Field
[0001] This application relates to the technical field of information processing devices and the like that anthropomorphize ground objects and output voices corresponding to contents such as news.
Background Art
[0002] Patent Document 1 discloses an invention that sets the types of humans when an object is anthropomorphized and outputs an explanation of the object in words that are assumed to be spoken by a human of that type. Specifically, based on the product name, name, manufacturing date, production country (production place), function / usage, material, etc. of the object, for example, a gentlemanly man in his 60s is set as the type of human, and the explanation of the object is output in words that are assumed to be spoken by a gentlemanly man in his 60s. Thereby, the user who hears the explanation can feel as if he / she is hearing the voice of the object itself.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the invention of Patent Document 1, although the user can listen to the explanation of the object without discomfort because the object itself is speaking, for example, when trying to make an arbitrary timely content (such as a message like "Good morning, happy birthday" or content that is preferably conveyed to the user immediately, such as major news of the day) be spoken by a ground object that the user happens to encounter during movement, a mismatch may occur between the content of the content and the character of the ground object, and the user may feel discomfort when the ground object speaks the content.
[0005] Therefore, in view of one example of these problems, the present invention aims to provide an information processing device, etc., that can give the user the feeling that a feature, which speaks content that should be communicated to the user in a way that doesn't sound unnatural, is actually speaking that content. [Means for solving the problem]
[0006] The invention described in claim 1 is characterized by comprising: content acquisition means for acquiring content to be notified to a user from among a plurality of contents classified by content type; user information acquisition means for acquiring the user's current location information; anthropomorphic feature identification means for identifying a feature as an anthropomorphic feature that is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired by the content acquisition means, based on feature-related information which associates content type information that permits the feature to speak and the location information of the feature with each of a plurality of features; and output control means for outputting sound corresponding to the content acquired by the content acquisition means, with the sound image localization such that it is heard from the direction in which the anthropomorphic feature is located as seen from the user.
[0007] The invention described in claim 5 is an information processing method using an information processing device, characterized in that it includes: a content acquisition step of acquiring content to be notified to a user from among a plurality of contents classified by content type; a user information acquisition step of acquiring the user's current location information; an anthropomorphic feature identification step of identifying a feature as an anthropomorphic feature, based on feature-related information which associates content type information that permits the feature to speak and the location information of the feature with each of a plurality of features, the feature that is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired in the content acquisition step; and an output control step of outputting sound corresponding to the content acquired in the content acquisition step, with the sound image localization such that it is heard from the direction in which the anthropomorphic feature exists as seen from the user.
[0008] The invention described in claim 6 is characterized in that a computer included in an information processing device functions as: content acquisition means for acquiring content to be notified to the user from among a plurality of contents classified by content type; user information acquisition means for acquiring the user's current location information; anthropomorphic feature identification means for identifying a feature as an anthropomorphic feature, based on content type information that permits the feature to speak, and feature-related information that associates the location information of the feature with each of a plurality of features, the feature is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired by the content acquisition means; and output control means for outputting sound corresponding to the content acquired by the content acquisition means, with the sound image localization such that it is heard from the direction in which the anthropomorphic feature exists as seen from the user. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing an example configuration of the information processing device 100 in an embodiment. [Figure 2] This is a block diagram showing an example of the configuration of the in-vehicle device 200 in the first embodiment. [Figure 3] This figure shows an example of a geological table 230 in the first embodiment. [Figure 4] This figure shows an example of the content type table 240 in the first embodiment. [Figure 5] This figure shows an example of the content table 250 in the first embodiment. [Figure 6] This is a schematic diagram showing an example of the position of speakers 217A-217H of the vehicle RC in the first embodiment. [Figure 7] This figure shows an example of the positional relationship between the vehicle RC and the features F1-F5 in the first embodiment. [Figure 8] This flowchart shows an example of personified object speech processing (Pattern 1) by the in-vehicle device 200 in the first embodiment. [Figure 9]This flowchart shows an example of personified object speech processing (pattern 2) by the in-vehicle device 200 in the first embodiment. [Figure 10] This is a flowchart showing an example of personified object speech processing (pattern 3) by the in-vehicle device 200 in the first embodiment. [Figure 11] This is a block diagram showing an example configuration of the smartphone 400 in the second embodiment. [Modes for carrying out the invention]
[0010] A description of an embodiment for carrying out the present invention will be made using Figure 1.
[0011] As shown in Figure 1, the information processing device 100 includes a content acquisition means 101, a user information acquisition means 102, an anthropomorphic object identification means 103, and an output control means 104.
[0012] The content acquisition means 101 acquires content that should be notified to the user from among multiple pieces of content classified by content type.
[0013] The user information acquisition means 102 acquires the user's current location information.
[0014] The anthropomorphic feature identification means 103 identifies, for each of a plurality of features, as an anthropomorphic feature based on feature-related information that associates content type information that allows the feature to speak with the feature, and location information of the feature, any feature that is within a predetermined distance from the user's current location and is permitted to speak with the content type of content acquired by the content acquisition means 101.
[0015] The output control means 104 outputs the audio corresponding to the content acquired by the content acquisition means 101, after localizing the sound image so that it is heard from the direction in which the anthropomorphic object exists, as seen from the user's perspective.
[0016] According to the information processing apparatus 100, voice corresponding to content to be notified to the user, which is a content type permitted to be spoken about the feature from the direction where the feature exists as seen by the user, is output. Thereby, it is possible to give the user a feeling as if a feature without a sense of incongruity even when speaking the content to be notified to the user is speaking the content.
Embodiment
[0017] [1. First Embodiment] The first embodiment will be described with reference to FIGS. 2 to 8.
[0018] [1.1. Outline of the in-vehicle device 200] The outline of the in-vehicle device 200, which is an example of the information processing apparatus 100, will be described. The in-vehicle device 200 is mounted on the vehicle RC. The in-vehicle device 200 acquires content to be notified to the user, who is a passenger of the vehicle RC, from a plurality of contents (news, congratulatory messages, recommended information, etc.) classified by content type, identifies one feature within the range visible to the user as an anthropomorphic feature, and audio-locates the voice corresponding to the acquired content so that it can be heard from the direction where the anthropomorphic feature exists as seen by the user and outputs it from the speaker. That is, the in-vehicle device 200 anthropomorphizes the features existing here and there, giving the user a feeling as if the features are reading the text corresponding to the content.
[0019] [1.2. Configuration of the in-vehicle device 200] The configuration of the in-vehicle device 200, which is an example of the information processing apparatus 100, will be described with reference to FIG. 2. FIG. 2 is a block diagram showing a configuration example of the in-vehicle device 200.
[0020] The in-vehicle unit 200 comprises a control unit 211, a storage unit 212 consisting of an HDD (Hard Disk Drive) or SSD (Solid State Drive), a communication unit 213, a display unit 214, a touch panel 215, an in-vehicle microphone 216, speakers 217A-217H, and a GPS (Global Positioning System) receiver 218.
[0021] The control unit 211 consists of a CPU 211a that controls the entire control unit 211, a ROM 211b in which control programs for controlling the control unit 211 are pre-stored, and a RAM 211c that temporarily stores various data. The control unit 211 or CPU 211a corresponds to a "computer".
[0022] The control unit 211 performs processing related to route searching for the vehicle RC to its destination and controlling the virtual speaking of content to be communicated to the user (the passenger of the vehicle RC) by anthropomorphic features (including the anthropomorphic feature speech processing described later), in response to user input.
[0023] The memory unit 212 stores various programs such as the OS (Operating System) and application programs, as well as data and information used by these programs. The memory unit 212 also analyzes the user's voice input to the in-vehicle microphone 216, interprets the spoken content, and stores the corresponding program. These programs may be acquired, for example, from a server device via a network, or they may be read from a recording medium such as a USB memory stick.
[0024] The memory unit 212 stores map data. The map data includes information necessary for processing related to route searching for the vehicle RC to its destination by the control unit 211, and guidance at intersections along the route. For example, it includes intersection data showing the location, shape, and name of intersections, road link data (including gradient information), and feature data related to facilities, etc. The memory unit 212 also stores traffic congestion information, weather information, etc., acquired via the communication unit 213.
[0025] The memory unit 212 stores the feature table 230 shown in Figure 3, the content type table shown in Figure 4, and the content table 250 shown in Figure 5.
[0026] The feature table 230 shown in Figure 3 stores information about multiple features. A feature is a natural or artificial object with location information (for example, a building, mountain, tree, statue, traffic light, etc.), and the feature table 230 stores information about some of these features. The feature table 230 contains feature information that is registered by the manufacturer or user of the in-vehicle unit 200, including the "Feature ID," "Location," "Speech permission content type," "Environmental conditions" (including "Speech permission time" and "Speech permission weather"), "Voice type," and "Name." The "Feature ID" is an identifier used to identify a feature.
[0027] Users can also register data in the feature table 230, and they can register features that hold special memories for them (for example, the hospital where they were born, or a pedestrian bridge they crossed with their lover).
[0028] The "Location" field registers the latitude, longitude, and altitude of the geographical feature.
[0029] The "Permitted Speech Content Type" field registers a content type ID that indicates the content type for which the feature is permitted to speak (content type and content type ID will be explained later).
[0030] "Environmental conditions" (an example of "environmental conditions information") include "speaking permission time" and "speaking permission weather," and indicate the environmental conditions under which the user can see geographical features.
[0031] The "Speech Permit Time" registers the time period during which the feature is permitted to speak. The "Speech Permit Time" registers the time period during which the user can see the feature. In other words, the registered time period takes into account the surrounding environment of the feature and whether or not the feature is lit. For example, for a feature located in a place where there is no lighting and it is dark at night, the time period from sunrise to sunset may be registered. Alternatively, if there is lighting in the surrounding area, lighting directed at the feature, or the feature itself emits light, the time period during which these lights are on may be set as the "Speech Permit Time".
[0032] The "Speech-Permitting Weather" setting registers the weather conditions under which a feature is permitted to speak. The "Speech-Permitting Weather" setting registers the weather conditions under which a feature can be seen by the user. For example, high-rise towers and high-rise apartment buildings may not be clearly visible in rainy, snowy, or foggy weather, so for such features, sunny or cloudy weather may be designated as the "Speech-Permitting Weather."
[0033] The "Voice Type" field registers the type of voice that would be used if the local object were personified. The "Voice Type" field also registers information indicating the characteristics of the voice, such as gender (male / female), tone (calm / slow / cute / powerful / cheerful, etc.), and dialect (Tokyo dialect / Satsuma dialect / Kyoto dialect / Tohoku dialect, etc.). The memory unit 212 stores voice feature data corresponding to each voice type, and voice data is generated by synthesizing the text corresponding to the content based on the voice feature data. For example, if the "Voice Type" is "female / cute / Kyoto dialect," voice data will be generated in which a woman reads the text corresponding to the content in a cute Kyoto dialect voice, based on voice feature data that can generate a voice of a woman speaking in a cute Kyoto dialect.
[0034] The name of the feature is registered in the "Name" field. The "Name" field may be the official name, or, if a user registers a feature, the user may register a name of their choice.
[0035] The content type table 240 shown in Figure 4 stores content type information. Specifically, the content type table 240 registers information indicating the "content type ID" and the "content type". The "content type" is the type of content that is notified to the user, and for example, it indicates what topic the content is about. For example, it could be memorial, weather, sad topics, happy topics, political issues, etc. The "content type ID" is an identifier used to identify the content type.
[0036] The "Content Type ID" is registered in the "Speech-Permitted Content Type" in the feature table 230. In the example in Figure 3, the feature (○× Tower) with "Feature ID" "00001" has "00" and "01" registered as "Speech-Permitted Content Types," so speech is permitted for content whose "Content Type" is either memorial or weather. Note that the manufacturer or user can decide which content type to set as the "Speech-Permitted Content Type" for each feature.
[0037] The content table 250 shown in Figure 5 stores information about the content provided to the user. Specifically, the content type table 240 registers information indicating the "Content ID," "Content Type ID," and "Content." The "Content ID" is an identifier used to identify the content. Examples of "Content" that may be registered in text format include current events, congratulatory messages to the user, recommended information for the user, and weather information. Current events and weather information may be obtained and registered from an external content provider server. Congratulatory messages to the user may be generated by the in-vehicle device 200 based on the user's birthday or wedding anniversary, or obtained from an external server. Recommended information for the user may be generated by the in-vehicle device 200 based on the user's preferences or user behavior history, or obtained from an external server.
[0038] The communication unit 213 has wireless communication devices and communicates with external server equipment, smartphones, and other information processing devices equipped with communication functions to send and receive data to and from each other.
[0039] The display unit 214 is configured with a graphics controller 214a and a buffer memory 214b consisting of memory such as VRAM (Video RAM). In this configuration, the graphics controller 214a controls the display unit 214 and the touch panel 215 based on control information sent from the control unit 211. The buffer memory 214b temporarily stores image data that can be immediately displayed on each touch panel 215. Then, based on the image data output from the graphics controller 214a, an image is displayed on the touch panel 215.
[0040] The touch panel 215 displays an image based on image data received from the display unit 214. The touch panel 215 also detects user touch operations and transmits touch operation data indicating the touch location, etc., to the control unit 211. The control unit 211 determines what operation was performed based on the touch operation data and processes accordingly.
[0041] The in-vehicle microphone 216 is installed inside the vehicle and converts the voice of the user inside the vehicle into an electrical signal and outputs it to the control unit 211.
[0042] Speakers 217A-217H are installed inside the vehicle and output sound under the control of the control unit 211. Hereafter, speakers 217A-217H may be collectively referred to as speaker 217.
[0043] The arrangement of speakers 217A-217H will be explained using Figure 6. Figure 6 is a schematic diagram showing an example of the position of speakers 217A-217H in a vehicle RC equipped with an on-board unit 200. The vehicle RC is a real vehicle in which a user sits, and has a driver's seat 231, a passenger seat 232, and a rear seat 233. The vehicle RC is also equipped with eight speakers 217A-217H and an on-board unit 200 (not shown, except for speakers 217A-217H). Speakers 217H, 217A, and 217B are installed from left to right at the front of the vehicle RC's passenger compartment. Speaker 217G is installed on the left side of the passenger compartment, and speaker 217C is installed on the right side of the passenger compartment. Furthermore, speakers 217F, 217E, and 217D are installed from left to right at the rear of the passenger compartment.
[0044] In this embodiment, the CPU 211a performs processing to realize sound localization using speakers 217A-217H. In other words, the CPU 211a realizes sound localization of the voice corresponding to the content in which the anthropomorphic object virtually speaks. Furthermore, the sound localization of the voice is shifted as the vehicle RC moves, providing the user with the experience of the anthropomorphic object speaking in the same location. Note that the processing to realize sound localization using speakers 217A-217H may be performed by an audio processor or the like, separate from the CPU 211a.
[0045] The GPS receiver 218 receives navigation radio waves from GPS satellites, acquires GPS positioning data such as the current location information of the in-vehicle unit 200 (latitude, longitude, altitude data, absolute bearing data of the direction of travel, and GPS speed data), and outputs it to the control unit 211.
[0046] The external camera 219 is mounted on the vehicle RC, captures images of the area around the vehicle RC, and outputs the camera images to the control unit 211.
[0047] [1.3. Personified Object Speech Processing] Next, the processing of anthropomorphic feature speech by the control unit 211 will be described. In the processing of anthropomorphic feature speech, an anthropomorphic feature that will virtually speak voice corresponding to the content is identified from among multiple features. In this embodiment, three patterns of methods for identifying anthropomorphic features will be described. The three patterns are: "Pattern 1: Identifying anthropomorphic features by permitted content type," "Pattern 2: Identifying anthropomorphic features by permitted speech time and permitted weather conditions," and "Pattern 3: Identifying anthropomorphic features by camera image." Patterns 2 and 3 are examples of methods for identifying anthropomorphic features from among multiple features that are visible from the user's current location in the current situation.
[0048] In this scenario, assuming that the feature table 230 shown in Figure 3, the content type table 240 shown in Figure 4, and the content table 250 shown in Figure 5 are stored in the memory unit 212, it is assumed that the following are located in the direction of travel of the moving vehicle RC, as shown in Figure 7: ○× Tower F1, the traffic light of memories F2, the statue of Saigo Takamori F3, the cat monument F4, and the Jizo statue F5. Furthermore, it is assumed that the control unit 211 has acquired content with content ID "0003" ("There was a car accident at ○○.") from the content table 250 as content to be notified to the user.
[0049] [1.3.1. Pattern 1: Identification of anthropomorphic objects based on the type of content that allows speech] In Pattern 1, the control unit 211 identifies anthropomorphic features as those located within a predetermined distance from the user's current location and for which speech of the content type of the acquired content is permitted, based on a feature table 230 (an example of "feature-related information") which associates each of multiple features with a speech permission content type (an example of "content type information that allows speech by features") and a location (an example of "feature location information").
[0050] For example, if the control unit 211 finds the ○× Tower F1, the traffic light of memories F2, and the statue of Saigo Takamori F3 as features located within a predetermined distance (e.g., 150m) from the user's current location, it retrieves the speech permission content for each feature from the feature table 230 (see Figure 3). Specifically, it retrieves "00" and "01" as the speech permission content type for ○× Tower F1, "01", "02", and "04" as the speech permission content types for the traffic light of memories F2, and "01" and "03" as the speech permission content types for the statue of Saigo Takamori F3. Then, the control unit 211 retrieves the content type "03" for the content with content ID "0003" from the content table 250 (see Figure 5), and identifies the statue of Saigo Takamori F3, which has content type "03" registered as the speech permission content type, as a personified feature.
[0051] Here, using the flowchart in Figure 8, we will explain the anthropomorphic feature speech processing when the method for identifying anthropomorphic features is Pattern 1.
[0052] First, the control unit 211 retrieves content to be notified to the user from the content table 250 (step S101). At this time, the control unit 211 can retrieve any content, but it may also retrieve content according to a priority set in advance by the manufacturer or user. For example, it may set a priority for content type IDs and retrieve content with a high priority, or it may retrieve content starting with the oldest (or newest) content registration date and time. In addition, user attributes (gender, age, preferences, etc.) may be registered, and appropriate attributes for each piece of content may be registered, and the control unit 211 may retrieve content appropriate to the user's attributes.
[0053] Next, the control unit 211 acquires the current location information of the user (vehicle RC) (step S102). Specifically, the control unit 211 acquires the current location information (latitude, longitude, altitude data, absolute direction of travel data, and GPS speed data) received from the GPS receiver 218.
[0054] Next, the control unit 211 searches for features located within a predetermined distance from the user's current location (step S103). The predetermined distance can be set by the manufacturer or the user based on the feature density (number of features in the area) in the area where the vehicle RC is traveling. The features to be searched are those registered in the feature table 230, and the control unit 211 determines whether or not they are located within the predetermined distance based on the vehicle RC's current location and the location information of each feature registered in the feature table 230.
[0055] Next, the control unit 211 determines whether or not a feature was found in the process of step S103 (step S104). If the control unit 211 determines that a feature was found in the process of step S103 (step S104: YES), it proceeds to the process of step S105. On the other hand, if the control unit 211 determines that no feature was found in the process of step S103 (step S104: NO), it proceeds to the process of step S110.
[0056] If the control unit 211 determines "YES" in the process of step S104, it then obtains the speech permission content type for each feature found in the process of step S103 from the feature table 230 (step S105).
[0057] Next, the control unit 211 identifies the features found in step S103 as anthropomorphic features if the feature is permitted to speak of the content type of the content acquired in step S101 (step S106). In other words, the control unit 211 identifies the features whose content type is registered as a speech-permitted content type as anthropomorphic features. If there are multiple features identified as anthropomorphic features, one feature is identified according to a pre-set priority order. The priority order may be set by the manufacturer or user, for example, the feature closest to the current location, the feature with the smallest feature ID, etc.
[0058] Next, the control unit 211 determines whether or not it was able to identify the anthropomorphic feature in the process of step S106 (step S107). If the control unit 211 determines that it was not able to identify the anthropomorphic feature (i.e., there were no features registered as speech permission content types among the content types acquired in the process of step S101) (step S107: NO), it proceeds to the process of step S110.
[0059] On the other hand, if the control unit 211 determines that it has identified the anthropomorphic feature (step S107: YES), it then generates audio data corresponding to the content obtained in step S101 based on the audio feature data corresponding to the voice type of the anthropomorphic feature (step S108). In other words, the control unit 211 identifies the voice type of the anthropomorphic feature from the feature table 230, obtains the audio feature data corresponding to that voice type from the storage unit 212, and generates audio data by synthesizing the text corresponding to the content obtained in step S101 based on the audio feature data.
[0060] Next, the control unit 211 outputs the audio data generated in step S108 from the speaker 217, positioning the sound image so that it is heard from the direction in which the anthropomorphic object is located relative to the user, thus ending the anthropomorphic object speech processing. The direction in which the anthropomorphic object is located relative to the user is determined based on the current position of the user (vehicle RC) and the position of the anthropomorphic object. If the user (vehicle RC) is moving, the direction in which the sound is heard will also change accordingly.
[0061] If the control unit 211 determines "NO" in step S104 or step S107, it then generates audio data corresponding to the content acquired in step S101 based on the standard speech feature data (step S110). The standard speech feature data is data capable of generating audio data (gender is arbitrary) that calmly speaks the text in Tokyo dialect, and is stored in the storage unit 212.
[0062] Next, the control unit 211 outputs the audio data generated in step S110 from a predetermined speaker 217, and terminates the anthropomorphic object speech processing. The predetermined speaker 217 can be a speaker 217 that has been set in advance by the user (for example, the front speaker 217A).
[0063] [1.3.2. Pattern 2: Identification of personified objects based on speaking permission time and speaking permission weather] In Pattern 2, the control unit 211 identifies anthropomorphic features that are located within a predetermined distance from the user's current location and for which speaking is permitted at that time and weather, based on a feature table (an example of feature-related information) that associates environmental conditions ("environmental condition information"), which include at least one of the speaking permission time ("speaking permission time conditions") and speaking permission weather ("speaking permission weather conditions"), and location ("location information") for multiple features. Environmental conditions are criteria used to determine whether a feature is visible (easy to see). Here, we describe the case where a feature that satisfies both the speaking permission time and speaking permission weather conditions is identified as an anthropomorphic feature, but it is also possible to identify a feature that satisfies either the speaking permission time or speaking permission weather conditions as an anthropomorphic feature.
[0064] For example, if the control unit 211 finds that the ○× Tower F1, the traffic light of memories F2, and the Saigo Takamori statue F3 are within a predetermined distance (for example, 150m) from the user's current location, it obtains the speech permission time and speech permission weather for each feature from the feature table 230 (see Figure 3). Specifically, it obtains "05:00-00:00 (5am to midnight)" as the speech permission time for ○× Tower F1 and "sunny" and "cloudy" as the speech permission weather. It also obtains "00:00-00:00 (24 hours)" as the speech permission time for the traffic light of memories F2 and "sunny," "cloudy," "rainy," and "snowy" as the speech permission weather. Furthermore, it obtains "05:00-20:00 (5am to 8pm)" as the speech permission time for the Saigo Takamori statue F3 and "sunny," "cloudy," and "rainy" as the speech permission weather. Then, the control unit 211 identifies the memorable traffic light F2 as a personified object, for example, if the current time is "15:00" and the weather is "snow," and the registered speech permission time is "00:00-00:00 (24 hours)" and the speech permission weather is "sunny," "cloudy," "rainy," or "snowy."
[0065] Here, using the flowchart in Figure 9, we will explain the processing of anthropomorphic object speech when the method for identifying anthropomorphic objects is Pattern 2. Note that steps S201-S204 in Figure 9 are the same as steps S101-S104 in Figure 8, and steps S209-S212 in Figure 9 are the same as steps S108-S111 in Figure 8 (however, step S201 is read as step S101), so the explanation will be omitted.
[0066] If the control unit 211 determines "YES" in step S204, it obtains date, time, and weather information for the user's current location (step S205). For example, the control unit 211 obtains the date and time from the system time of the in-vehicle device 200 and the weather from an external server.
[0067] Next, the control unit 211 obtains the speech permission time and speech permission weather for each feature found in the processing of step S203 from the feature table 230 (step S206).
[0068] Next, the control unit 211 identifies the features found in step S203 as anthropomorphic features for which speech is permitted under the current date, time, and weather conditions (step S207). In other words, the control unit 211 identifies features as anthropomorphic features for which the date, time, and weather conditions obtained in step S205 satisfy the speech permission time and speech permission weather conditions, respectively. If there are multiple features identified as anthropomorphic features, one feature is selected according to a pre-set priority order. The priority order may be set by the manufacturer or user, for example, by selecting the feature closest to the current location or the feature with the smallest feature ID.
[0069] Next, the control unit 211 determines whether or not it was able to identify the anthropomorphic feature in the process of step S207 (step S208). If the control unit 211 determines that it was not able to identify the anthropomorphic feature (i.e., there were no features that met the conditions for speech permission time and speech permission weather, respectively, based on the date and time and weather conditions obtained in the process of step S205) (step S208: NO), it proceeds to the process of step S211. On the other hand, if the control unit 211 determines that it was able to identify the anthropomorphic feature (step S208: YES), it proceeds to the process of step S209.
[0070] [1.3.3. Pattern 3: Identification of anthropomorphic objects using camera images] In Pattern 3, recognizable features from the camera image captured by the external camera 219 from the user's current location are identified as anthropomorphic features.
[0071] Here, using the flowchart in Figure 10, we will explain the processing of anthropomorphic object speech when the method for identifying anthropomorphic objects is pattern 3. Note that steps S301-S304 in Figure 10 are the same as steps S101-S104 in Figure 8, and steps S307-S310 in Figure 10 are the same as steps S108-S111 in Figure 8 (however, step S301 is read as step S101), so the explanation will be omitted.
[0072] If the control unit 211 determines "YES" in step S304, it identifies the features visible in the camera image taken from the current location as anthropomorphic features from among the features found in step S303 (step S305). For example, the control unit 211 identifies the direction in which the feature exists from the location information of the feature found in step S303, and determines whether the feature is visible by performing image analysis on the camera image taken in that direction by the external camera 219.
[0073] Next, the control unit 211 determines whether or not it was able to identify the anthropomorphic object in the process of step S305 (step S306). If the control unit 211 determines that it was not able to identify the anthropomorphic object (i.e., there were no objects visible in the camera image taken from the current location) (step S306: NO), it proceeds to the process of step S309. On the other hand, if the control unit 211 determines that it was able to identify the anthropomorphic object (step S306: YES), it proceeds to the process of step S307.
[0074] As described above, in this embodiment, the in-vehicle device 200 has a control unit 211 (an example of "content acquisition means," "user information acquisition means," "anthropomorphic feature identification means," and "output control means") which acquires content to be notified to the user from among a plurality of contents classified by content type (step S101 in Figure 8), acquires the user's current location information (step S102 in Figure 8), and identifies anthropomorphic features that are located within a predetermined distance from the user's current location and for which speech of the content type of content acquired in step S101 is permitted, based on a feature table 230 (an example of "feature-related information") which associates the content type for which speech is permitted by the feature (an example of "content type information for which speech is permitted") and the location of the feature (an example of "location information") with each of the plurality of features as an anthropomorphic feature (step S106 in Figure 8), and outputs the sound corresponding to the content acquired in step S101, with the sound image localization so that it is heard from the direction in which the anthropomorphic feature exists from the user's perspective (step S109 in Figure 8).
[0075] Therefore, according to the in-vehicle device 200, voice corresponding to the content to be communicated to the user, which is a content type that is permitted to speak about the feature, is output from the direction in which the feature exists from the user's perspective. This makes it possible to give the user the feeling that the feature, which would sound natural speaking the content to be communicated to the user, is actually speaking that content.
[0076] Furthermore, in this embodiment, the in-vehicle device 200 has a control unit 211 that acquires content to be notified to the user from among multiple contents (step S201 in Figure 9, step S301 in Figure 10), acquires the user's current location information (step S202 in Figure 9, step S302 in Figure 10), identifies an anthropomorphic feature from among multiple features that is visible from the user's current location in the current situation (step S207 in Figure 9, step S305 in Figure 10), and outputs the sound corresponding to the content acquired in step S201 (step S301) with the sound image localized so that it is heard from the direction in which the anthropomorphic feature exists from the user's perspective (step S210 in Figure 9, step S308 in Figure 10).
[0077] Therefore, according to the in-vehicle device 200, audio corresponding to the content to be communicated to the user is output from the direction of any visible features to the user at that time. This gives the user the feeling that the visible features are speaking the content that should be communicated to the user.
[0078] Furthermore, in this embodiment, the in-vehicle unit 200, based on a feature table 230 (an example of "feature-related information") which associates environmental conditions (an example of "environmental condition information"), including a speech permission time (an example of "speech permission time conditions") and a speech permission weather (an example of "speech permission weather conditions"), and location (an example of "location information") for multiple features, identifies features that are within a predetermined distance from the user's current location and for which speech is permitted at that time and weather as anthropomorphic features (step S207 in Figure 9). As a result, audio corresponding to the content to be notified to the user is output from the direction of a feature that satisfies the time and weather conditions visible to the user at that time. This gives the user the feeling that a feature visible to the user is speaking the content that should be notified to the user.
[0079] Furthermore, in this embodiment, the in-vehicle device 200 has a control unit 211 that identifies recognizable features in a camera image (an example of an "image") taken from the user's current location as anthropomorphic features (step S305 in Figure 10). As a result, based on the camera image, audio corresponding to the content to be communicated to the user is output from the direction where the features that are determined to be visible to the user at that time are located. This gives the user the feeling that the features visible to the user are speaking the content that should be communicated to the user.
[0080] Furthermore, in this embodiment, the in-vehicle unit 200 has a control unit 211 (an example of a "voice data generation means") that generates voice data corresponding to the voice based on the voice feature data associated with each of the multiple features (step S108 in Figure 8, step S209 in Figure 9, and step S307 in Figure 10). As a result, the voice corresponding to the output content has features corresponding to the anthropomorphic feature, allowing the user to enjoy the feeling that the anthropomorphic feature is speaking the content without any sense of incongruity.
[0081] Furthermore, in this embodiment, if the control unit 211 is unable to identify the anthropomorphic feature (step S107 in Figure 8: NO, step S208 in Figure 9: NO, step S306 in Figure 10: NO), the in-vehicle device 200 generates voice data corresponding to the voice based on standard voice feature data (an example of "predetermined voice feature data") (step S110 in Figure 8, step S211 in Figure 9, step S309 in Figure 10). This makes it possible to provide the user with content to be notified even if the anthropomorphic feature cannot be identified.
[0082] [1.4. Variant Example] Next, a modified example of the first embodiment will be described. Note that the modified examples described below can be combined as appropriate.
[0083] [1.4.1. Variation 1] In the first embodiment, the environmental conditions "speaking permission time" and "speaking permission weather" were registered in the feature table 230, and features that satisfy both the current time and weather conditions of "speaking permission time" and "speaking permission weather" were identified as anthropomorphic features. Alternatively, only "speaking permission time" could be registered as an environmental condition in the feature table 230, and features that satisfy the current time of "speaking permission time" could be identified as anthropomorphic features, or only "speaking permission weather" could be registered as an environmental condition in the feature table 230, and features that satisfy the current weather conditions of "speaking permission weather" could be identified as anthropomorphic features.
[0084] [1.4.2. Variation 2] In the first embodiment, a feature that is located within a predetermined distance from the user's current location and is permitted to speak the content type of the content acquired in step S101 of Figure 8 is identified as a personified feature. However, this condition may be further modified to include the requirement that at least one of the current time and weather conditions satisfies at least one of the environmental conditions of "speaking permission time" and "speaking permission weather," or that the feature is visible from a camera image taken from the user's current location. The added condition here is a condition for identifying a feature visible to the user as a personified feature. This allows for the identification of a feature that is permitted to speak the content type of the content acquired in step S101 of Figure 8 and is visible to the user as a personified feature. Therefore, it is possible to prevent features that are not visible to the user from speaking, and to make features visible to the user speak the voice corresponding to the content.
[0085] [1.4.3. Variation 3] In the first embodiment, in the processing of step S103 in Figure 8, step S203 in Figure 9, and step S303 in Figure 10, the control unit 211 searches for features that are within a predetermined distance from the user's current location. However, it may also search for features that are within a predetermined distance from the user's current location and are in the direction the user is facing (or the direction the user is traveling). To this end, the control unit 211 may, for example, obtain the direction the user is facing from the orientation of the user's face using an in-car camera (not shown), or obtain the direction the user is traveling from absolute bearing data received from the GPS receiver 218. This makes it possible to identify features that are easily visible to the user and are located in the direction the user is facing or traveling as anthropomorphic features.
[0086] [1.4.4. Variation 4] The feature table 230 may include an "effective range" for each feature. The "effective range" registers the distance from the user's current location at which the feature can be selected as a personified feature. For example, if 2km is registered as the "effective range," the feature will not be selected as a personified feature unless it is within 2km of the user's current location. The "effective range" can be set based on the size of the feature, etc. This allows the system to identify features that are visible to the user as personified features, taking into account the size of the feature, etc.
[0087] [1.4.5. Variation 5] In the first embodiment, the control unit 211 searches for features located within a predetermined distance from the user's current location in step S303 of Figure 10, and if it determines "YES" in step S304 of Figure 10, it performs the process in step S305 of Figure 10. However, steps S303 and S304 may be skipped, and in step S305, features visible from the camera image taken from the current location may be identified as anthropomorphic features. In other words, even if a feature does not exist within a predetermined distance from the user's current location, if it is visible from the camera image, that feature may be identified as an anthropomorphic feature.
[0088] [1.4.6. Variation 6] In the first embodiment, the control unit 211 performs the anthropomorphic object speech processing shown in Figures 8-10. However, at least a portion of these processes may be performed in cooperation with an external server device or the like to distribute the processing load.
[0089] [2. Second Example] Next, a second embodiment will be described using Figure 11. In the first embodiment, the case where the information processing device is an in-vehicle device 200 was described, but in the second embodiment, the case where the information processing device is a portable information processing device, a smartphone 400, will be described, focusing on the differences from the first embodiment. In the second embodiment, the case where the voice of the anthropomorphic object is output from an earphone, which is an audio output device, will be described. [2.1. Smartphone 400 Configuration] Figure 11 is a block diagram showing an example configuration of a user's smartphone 400.
[0090] The smartphone 400 is comprised of a control unit 411, a storage unit 412 consisting of an SSD or the like, a communication unit 413, a display unit 414, a touch panel 415, a microphone 416, a speaker 417, a GPS receiver 418, and a camera 419.
[0091] The control unit 411 corresponds to the control unit 211 in the first embodiment and consists of a CPU 411a that controls the entire control unit 411, a ROM 411b in which control programs for controlling the control unit 411 are pre-stored, and a RAM 411c that temporarily stores various data. The control unit 411 or CPU 411a corresponds to a "computer".
[0092] The storage unit 412 corresponds to the storage unit 212 in the first embodiment and stores various programs such as the OS and application programs, as well as data and information used by these programs. The storage unit 412 also has dedicated applications for a browser and various services installed. These programs may be obtained, for example, from a server device via a network, or they may be read from a recording medium such as a USB memory stick.
[0093] The communication unit 413 corresponds to the communication unit 213 in the first embodiment and has a wireless communication device, etc., and communicates with other information processing equipment equipped with communication functions to send and receive data to and from each other. The communication unit 413 also communicates wirelessly with the earphone 420.
[0094] The earphone 420 corresponds to the speaker 217 in the first embodiment and is worn in the user's ear. The earphone 420 is equipped with a geomagnetic sensor and a 3G acceleration sensor to determine the direction the user is facing, and instead of the speaker 217 in the first embodiment, it realizes sound image localization of voices from anthropomorphic objects. The CPU 411a executes processing to realize sound image localization by the earphone 420. In other words, the CPU 411a realizes sound image localization of voices from anthropomorphic objects. Furthermore, the sound image localization of voices is shifted in accordance with the relative position change between the user and the anthropomorphic object due to the user's movement. As a result, the voices of the anthropomorphic objects are always heard coming from the direction of the anthropomorphic objects. Note that the processing to realize sound image localization by the earphone 420 may be executed by an audio processor or the like, separate from the CPU 411a.
[0095] The display unit 414 corresponds to the display unit 214 of the first embodiment and is configured with a graphics controller 414a and a buffer memory 414b consisting of memory such as VRAM (Video RAM). In this configuration, the graphics controller 414a controls the display unit 414 and the touch panel 415 based on control information sent from the control unit 411. The buffer memory 414b temporarily stores image data that can be immediately displayed on each touch panel 415. Then, an image is displayed on the touch panel 415 based on the image data output from the graphics controller 414a.
[0096] The touch panel 415 corresponds to the touch panel 215 of the first embodiment and displays an image based on image data received from the display unit 414. The touch panel 415 also detects the operator's touch operation and transmits touch operation data indicating the touched location, etc., to the control unit 411. The control unit 411 determines what operation was performed based on the touch operation data and processes according to the content of the operation.
[0097] Microphone 416 corresponds to the in-vehicle microphone 216 in the first embodiment, and converts the user's voice into an electrical signal and outputs it to the control unit 411.
[0098] The speaker 417 outputs sound under the control of the control unit 411, for example, when the earphones 420 are not connected.
[0099] The GPS receiver 418 corresponds to the GPS receiver 218 in the first embodiment and receives navigation radio waves from GPS satellites. It acquires GPS positioning data such as the current location information of the smartphone 400, including latitude, longitude, altitude data, absolute direction of travel data, and GPS speed data.
[0100] The camera 419 outputs the captured camera image to the control unit 211.
[0101] [2.2. Personified Object Speech Processing] The processing of anthropomorphic object speech in Example 2 is almost the same as in Example 1. However, in Pattern 3 of the three methods for identifying anthropomorphic objects, objects visible to the user from the camera image are identified as anthropomorphic objects, but in Example 2, it is not practical for the user to constantly photograph their surroundings with the camera 419. Therefore, instead of the camera 419, for example, smart glasses with a camera, which are a glasses-type wearable device, may be connected to the smartphone 400 and worn by the user. The camera of the smart glasses photographs the direction the user is facing and transmits the camera image to the smartphone 400, so the control unit 411 of the smartphone 400 identifies objects visible to the user from the received camera image. Furthermore, smart glasses equipped with a geomagnetic sensor and a 3G acceleration sensor can also determine the direction the user is facing.
[0102] As described above, in the second embodiment, the smartphone 400 has a control unit 411 (an example of "content acquisition means," "user information acquisition means," "anthropomorphic feature identification means," and "output control means") which acquires content to be notified to the user from among a plurality of contents classified by content type (step S101 in Figure 8), acquires the user's current location information (step S102 in Figure 8), and identifies anthropomorphic features that are located within a predetermined distance from the user's current location and for which speech of the content type of content acquired in step S101 is permitted, based on a feature table 230 (an example of "feature-related information") which associates the content type for which speech is permitted by the feature (an example of "content type information for which speech is permitted") and the location of the feature (an example of "location information") with each of the plurality of features as an anthropomorphic feature (step S106 in Figure 8), and outputs the sound corresponding to the content acquired in step S101, with the sound image localization so that it is heard from the direction in which the anthropomorphic feature exists from the user's perspective (step S109 in Figure 8).
[0103] Therefore, according to the smartphone 400, from the direction in which the feature exists from the user's perspective, audio corresponding to the content that should be communicated to the user, which is a content type that is permitted to speak about the feature, is output. This makes it possible to give the user the feeling that the feature, which would sound natural speaking the content that should be communicated to the user, is actually speaking that content.
[0104] Furthermore, in this embodiment, the smartphone 400 has a control unit 411 that acquires content to be notified to the user from among multiple contents (step S201 in Figure 9, step S301 in Figure 10), acquires the user's current location information (step S202 in Figure 9, step S302 in Figure 10), identifies an anthropomorphic object from among multiple features that is visible from the user's current location in the current situation (step S207 in Figure 9, step S305 in Figure 10), and outputs the sound corresponding to the content acquired in step S201 (step S301) with the sound image localized so that it is heard from the direction in which the anthropomorphic object exists from the user's perspective (step S210 in Figure 9, step S308 in Figure 10).
[0105] Therefore, according to the smartphone 400, audio corresponding to the content to be communicated to the user is output from the direction of any visible features to the user at that moment. This gives the user the feeling that the visible features are speaking the content that should be communicated to the user.
[0106] [2.3. Variant] Modifications of the first embodiment can be appropriately applied to the second embodiment. Furthermore, each modification can be combined as appropriate. [Explanation of symbols]
[0107] 100: Information Processing Device 101: Content Acquisition Methods 102: User Information Acquisition Method 103:Anthropomorphic feature identification means 104: Output control means 200: Onboard equipment 211: Control Unit 211a: CPU 211b :ROM 211c: RAM 212: Storage section 213: Communications Department 214: Display Unit 214a: Graphics controller 214b: Buffer memory 215: Touch panel 216: In-car microphone 217A: Speaker 217B: Speaker 217C: Speaker 217D: Speaker 217E: Speaker 217F: Speaker 217G: Speaker 217H: Speaker 218: GPS receiver 219: Exterior car camera 230: Local Table 231: Driver's seat 232:Passenger seat 233: Rear seat 240: Content Type Table 250: Content Table 400: Smartphone 411: Control Unit 411a: CPU 411b :ROM 411c: RAM 412: Storage section 413: Communications Department 414: Display Unit 414a: Graphics Controller 414b: Buffer memory 415: Touch panel 416: Mike 417: Speaker 418: GPS receiver 419: Camera 420: Earphones F1: ○× Tower F2: Traffic lights of memories F3: Statue of Saigo Takamori F4: Cat Monument F5: Jizo RC: Vehicle
Claims
1. A content acquisition method that acquires content to be notified to the user from among multiple content classified by content type, A user information acquisition means for acquiring the user's current location information, An anthropomorphic feature identification means identifies, for each of multiple features, a feature that is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired by the content acquisition means, based on feature-related information which associates content type information that allows the feature to speak and location information of the feature, as an anthropomorphic feature. Output control means that outputs audio corresponding to the content acquired by the content acquisition means, with the sound image localization such that it is heard from the direction in which the personified object exists as seen by the user. An information processing device characterized by comprising:
2. An information processing apparatus according to claim 1, The information processing device is characterized in that the personified feature identification means identifies a feature as the personified feature that is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired by the content acquisition means, and is located in the direction the user is facing or the direction the user is moving.
3. An information processing apparatus according to claim 1 or 2, An information processing apparatus further comprising a voice data generation means for generating voice data corresponding to the voice based on voice feature data associated with each of the plurality of features.
4. An information processing apparatus according to claim 3, The audio data generation means is characterized in that, when the anthropomorphic feature identification means fails to identify the anthropomorphic feature, it generates audio data corresponding to the audio based on predetermined audio feature data.
5. An information processing method using an information processing device, A content acquisition process that acquires content that should be notified to the user from among multiple content classified by content type, A user information acquisition step, which involves acquiring the user's current location information, An anthropomorphic feature identification step, which identifies a feature as an anthropomorphic feature based on feature-related information that associates each of multiple features with content type information that allows the feature to speak and location information of the feature, and which features that are located within a predetermined distance from the user's current location and are permitted to speak of the content type of content acquired in the content acquisition step. An output control step that outputs the audio corresponding to the content acquired in the content acquisition step, with the sound image localization adjusted so that it is heard from the direction in which the personified object exists, as seen from the user. An information processing method characterized by including
6. The computer included in the information processing device A content acquisition method that acquires content to be notified to the user from among multiple content classified by content type. User information acquisition means for acquiring the user's current location information, An anthropomorphic feature identification means identifies, for each of multiple features, a feature that is located within a predetermined distance from the user's current location and is permitted to speak of the content type of content acquired by the content acquisition means, based on feature-related information which associates content type information that allows the feature to speak and location information of the feature, as an anthropomorphic feature. Output control means that outputs the audio corresponding to the content acquired by the content acquisition means, with the sound image localization such that it is heard from the direction in which the anthropomorphic object exists, as seen from the user. An information processing program characterized by functioning as such.