Sound output control device, sound output control method, and sound output control program
The sound output control device addresses user confusion by adding background sound to voice messages based on the type of agent information, allowing users to easily recognize the category and source of the voice content.
Patent Information
- Application Number
- JP2021044051
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-03-17
AI Technical Summary
Users face difficulties in recognizing whether the currently output voice content is of the desired type, especially when multiple apps provide voice content of various categories.
A sound output control device that acquires agent information from an agent device capable of outputting multiple types of information and notifies the user with a voice message that includes background sound corresponding to the type of agent information being output.
Enables users to easily identify the type of currently output voice content, improving user experience by clarifying the source and category of the voice content.
Smart Images

Figure 0007696216000001 
Figure 0007696216000002 
Figure 0007696216000003
Abstract
Description
Technical Field
[0001] The present invention relates to a sound output control device, a data structure, a sound output control method, and a sound output control program.
Background Art
[0002] Conventionally, there has been known a technique in which agents with different timbres and intonations respond according to the type of task. There is also known a technique for changing the voice color of an agent for each of a plurality of request processing devices.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, as an example of a problem in the above prior art, there are cases where it is difficult for a user to appropriately recognize whether the currently output voice content is the desired type of voice content.
[0005] The present invention has been made in view of the above, and an object thereof is to provide a sound output control device, a data structure, a sound output control method, and a sound output control program that can support a user to appropriately grasp what type of voice content the currently output voice content is.
Means for Solving the Problems
[0006] The sound output control device according to claim 1 includes an information acquisition unit that acquires agent information to be output from an agent device capable of outputting a plurality of types of agent information distinguished by the content of the information or the source of the information as information provided to the user, and a notification control unit that causes a notification unit to output a voice message corresponding to the agent information to be output. The notification control unit is characterized in that background sound corresponding to the type to which the agent information to be output belongs is added to the voice message and notified from the notification unit.
[0007] Further, the data structure according to claim 8 is a data structure of agent information that is output from an agent device capable of outputting a plurality of types of agent information distinguished by the content of the information or the source of the information as information provided to the user, and is used when a sound output control device performs a notification process to the user. The data structure has message information indicating the content to be notified to the user and identification information indicating to which of the types the agent information belongs. The identification information can be used in a process of setting background sound to be added to the voice message when the sound output control device outputs the voice message corresponding to the message information.
[0008] Further, the sound output control method according to claim 9 is a sound output control method executed by a sound output control device. The method includes an information acquisition step of acquiring agent information to be output from an agent device capable of outputting a plurality of types of agent information distinguished by the content of the information or the source of the information as information provided to the user, and a notification control step of causing a notification unit to output a voice message corresponding to the agent information to be output. The notification control step is characterized in that background sound corresponding to the type to which the agent information to be output belongs is added to the voice message and notified from the notification unit.
[0009] The sound output control program according to claim 10 is a sound output control program executed by a sound output control device including a computer, which causes the computer to function as information acquisition means for acquiring output target agent information from an agent device capable of outputting a plurality of types of agent information distinguished by the content of the information or the source of the information as information provided to the user, and notification control means for causing a notification unit to output a voice message corresponding to the output target agent information. The notification control means is characterized in that background sound corresponding to the type to which the output target agent information belongs is added to the voice message and notified from the notification unit.
Brief Description of Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments for carrying out the present invention (hereinafter, embodiments) will be described with reference to the drawings. Note that the present invention is not limited by the embodiments described below. Further, in the description of the drawings, the same parts are denoted by the same reference numerals.
[0012] (Overview of the Embodiment) [1. Introduction] There are known applications (hereinafter abbreviated as "apps") that provide various contents via a terminal device (navigation terminal) installed in a vehicle or a terminal device such as a smartphone owned by a user (for example, a passenger in the vehicle). For example, by having an agent function that assists the user, it is possible to assist driving according to the driving state of the vehicle and the situation of the user driving the vehicle, or to assist route guidance according to various inputs (for example, character input or voice input). There are also apps that assist in more comfortable driving by providing various contents such as sightseeing guidance, store guidance, or other useful information according to the driving of the vehicle.
[0013] In addition, many of such apps attempt to assist the user using voice content by a voice agent function from the viewpoint of safety, considering that the output destination user is a passenger in the vehicle. In such a case, for example, since a plurality of apps are associated with a user in the vehicle and voice contents of various categories are provided, the following problems may occur.
[0014] For example, as a first problem, when various categories of voice content are output and the user is waiting for the voice content of a desired category to be output, it may be difficult for the user to determine whether the currently output voice content is the voice content of the desired category.
[0015] Also, as a second problem, when voice content of various app types is output, for example, when the user is waiting for the voice content provided by a specific app among multiple apps to be output, it may be difficult for the user to determine whether the currently output voice content is the voice content of the desired type.
[0016] Also, as a third problem, when using multiple apps, it may be difficult for the user to identify an app according to their preferences because it is difficult to distinguish which app is the provider of the voice content.
[0017] Therefore, an object of the present invention is to provide a sound output control device, a data structure, a sound output control method, and a sound output control program that can solve the above problems. Hereinafter, as information processing realized by the sound output control device, data structure, sound output control method, and sound output control program corresponding to the present invention, three information processes (first information process, second information process, third information process) will be described in detail. Specifically, the first information process will be described as the information process according to the first embodiment, and the second information process will be described as the information process according to the second embodiment. Also, the third information process will be described as the information process according to the third embodiment.
[0018] [2. Overall image of the information processing according to the embodiment] Prior to the description of each embodiment, the overall picture of the information processing according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing the overall picture of the information processing according to the embodiment. The information processing system 1 (an example of the information processing system according to the embodiment) shown in FIG. 1 is common to the first embodiment, the second embodiment, and the third embodiment. Also, in the following description, when there is no need to distinguish between the first embodiment, the second embodiment, and the third embodiment, it may simply be referred to as the "embodiment".
[0019] According to the example of FIG. 1, the information processing system 1 according to the embodiment includes a terminal device 10, a situation grasping device 30, an agent device 60-x, and a sound output control device 100. Also, these devices included in the information processing system 1 are communicably connected by wire or wirelessly via a network.
[0020] (Regarding the terminal device) The terminal device 10 is an information processing terminal used by a user. The terminal device 10 may be, for example, a stationary navigation device installed in a vehicle, or a portable terminal device owned by the user (e.g., a smartphone, a tablet terminal, a notebook PC, a desktop PC, a PDA, etc.). In this embodiment, it is assumed that the terminal device 10 is a navigation device installed in a vehicle.
[0021] Also, in the example of FIG. 1, it is assumed that a plurality of apps are associated with the terminal device 10. For example, a plurality of apps may be associated by being installed by the user in the terminal device 10, or the association may be performed by including in the information processing system 1 an app capable of push-notifying various contents regardless of the installation status.
[0022] In addition, the terminal device 10 has a notification unit (output unit), and voice content provided by each application is output from this notification unit. From this, the notification unit may be, for example, a speaker. Also, the user mentioned here may be a passenger (e.g., a driver) of the vehicle in which the terminal device 10 is installed. That is, in the example of FIG. 1, the terminal device 10 is installed in the vehicle VE1, and an example is shown in which the user of the terminal device 10 is the user U1 who is driving the vehicle VE1.
[0023] (Regarding the agent device) The agent device 60-x exists for each application associated with the terminal device 10 and may be an information processing device that realizes the functions and roles of the application. In FIG. 1, an example is shown in which the agent device 60-x is a server device, but it may be realized by, for example, a cloud system.
[0024] Also, the application associated with the terminal device 10 may be an application that assists the user using voice content, and the agent device 60-x has a voice agent function corresponding to this assist function. Also, from this, the application associated with the terminal device 10 can be said to be a so-called voice agent application.
[0025] Also, in the following embodiments, when distinguishing and representing the agent device 60-x corresponding to the application APx and the processing units (e.g., the application control function 631-x, the agent information generation unit 632-x) that the agent device 60-x has, an arbitrary value will be used for "x".
[0026] For example, in FIG. 1, an agent device 60-1 is shown as an example of the agent device 60-x. The agent device 60-1 is an agent device corresponding to the application AP1 (hazard detection application), and provides the user with voice content that detects hazards in driving and indicates warnings. Also, according to the example of FIG. 1, the agent device 60-x has an application control function 631-1 and an agent information generation unit 632-1.
[0027] In addition, FIG. 1 shows another example of the agent device 60-x, which is the agent device 60-2. The agent device 60-2 is an agent device corresponding to the app AP2 (navigation app) and provides voice content related to route guidance to the user. Although not shown in FIG. 1, the agent device 60-2 may have an app control function 631-2 and an agent information generation unit 632-2.
[0028] Note that the app associated with the terminal device 10 is not limited to the above example. For example, other examples include apps that provide voice content related to tourist guidance, voice content related to store guidance, or various useful information. Also, the expression "provided by the app" shall include the concept of "provided by the agent device 60-x corresponding to this app".
[0029] Subsequently, the functions of the agent device 60-x will be described. According to the example in FIG. 1, the agent device 60-x has an app control function 631-x and an agent information generation unit 632-x.
[0030] The app control function 631-x executes various controls related to the app APx. For example, based on the usage history of the user, the app control function 631-x personalizes the content provided for each user. Also, the app control function 631-x performs a process of determining what kind of voice message content should be used as a response based on the utterance content indicated by the voice input by the user so as to realize interaction with the user. Further, the app control function 631-x can also determine the content of the content provided to the user and the content of the voice message to respond to the user based on the user's situation.
[0031] The agent information generation unit 632-x performs a generation process for generating voice content (an example of agent information). For example, based on the data received from the situation awareness device 30 described later, the agent information generation unit 632-x determines what category of voice content should be output, and generates voice content of the content belonging to the determined category. For example, the agent information generation unit 632-x generates message information of content corresponding to the driving state of the vehicle grasped by the situation awareness device 30 and the situation of the user driving the vehicle. Note that the message information is, for example, text data that serves as the basis for the voice content finally notified to the user U1, and defines the content of the voice obtained by being later converted into voice data. That is, the agent information generation unit 632-x is not limited to generating voice data as voice content, and may generate other forms of data that serve as the basis for voice messages as voice content. Further, the agent information generation unit 632-x determines category identification information (category ID) for identifying the category to which the message information belongs based on the content indicated by the generated message information. Note that the category to which the message information belongs can also be said to be the category to which the voice content including the message information belongs.
[0032] In addition, the agent information generation unit 632-x assigns application identification information (app ID) for identifying the application APx corresponding to the agent device 60-x to the voice content.
[0033] Taking the agent device 60-1 as an example, the agent information generation unit 632-1 assigns an app ID (for example, "AP1") that identifies the app AP1 corresponding to the agent device 60-1 to the voice content. Further, when the generated voice content belongs to the category "Entertainment", the agent information generation unit 632-1 assigns a category ID indicating the category "Entertainment" to the voice content. That is, the agent information generation unit 632-1 generates agent information as voice content by adding an app ID for identifying the generation source of the message information and a category ID for identifying the content of the message information to the generated message information. In other words, the agent information as voice content generated by the agent information generation unit 632-1 is provided with app identification information for identifying the application that provided the agent information and category identification information for identifying the category to which the agent information belongs.
[0034] In addition, the agent information generation unit 632-x is not limited to the above example. For example, when voice input by speech is performed by the user via the terminal device 10, the agent information generation unit 632-x may generate message information of content that responds to the input voice. As a result, the agent device 60-x can generate voice content capable of realizing a dialogue with the user.
[0035] Furthermore, the agent information generation unit 632-x can also specify the timing for outputting the voice content. For example, the agent information generation unit 632-x can generate allowable range information that specifies a temporal range or a geographical range for allowing the output of the voice content, using, for example, a time range, a vehicle travel distance range, a vehicle passage area, the speed of the vehicle, etc. Also, in such a case, the agent information generation unit 632-x requests (reserves) the voice output control device SV to output the voice content to the terminal device 10 of the vehicle that meets the conditions indicated by the allowable range information, by transmitting the voice content and the allowable range information to the voice output control device SV. Regarding the timing specification and the request, it may be performed by a processing unit other than the agent information generation unit 632-x.
[0036] (Regarding the situation awareness device) The situation awareness device 30 performs analysis processing for grasping the driving state of the vehicle and the situation of the user driving the vehicle. In FIG. 1, an example where the situation awareness device 30 is a server device is shown, but it may be realized by, for example, a cloud system. Also, according to the example of FIG. 1, such analysis processing is performed by the situation awareness engine E30 installed in the situation awareness device 30. For example, the situation awareness engine E30 senses the driving state and the user situation based on the sensor information obtained from various sensors. Here, the sensors mentioned may be, for example, sensors provided in the vehicle or sensors that the terminal device 10 has, and examples include an acceleration sensor, a gyro sensor, a magnetic sensor, GPS, a camera, a microphone, etc.
[0037] For example, the situation awareness engine E30 can perform the following series of analysis processes. For example, based on the sensor information obtained from the above sensors, the situation awareness engine E30 performs sensing and uses the sensing result as a core element to perform basic analysis. In the basic analysis, the situation awareness engine E30 extracts the necessary data using the core element as an information source, and performs conversion and processing of the extracted data. Subsequently, the situation awareness engine E30 performs higher-order analysis using the data after conversion and processing. In the higher-order analysis, the situation awareness engine E30 analyzes specific situations based on the data after conversion and processing. For example, the situation awareness engine E30 performs various situation awareness such as the impact situation on the vehicle, the vehicle lighting situation, the change in the driving state, and the user's own situation from the data after conversion and processing. Also, as situation awareness, the situation awareness engine E30 can perform user behavior prediction (for example, prediction of a stopover location).
[0038] (Regarding the sound output control device) The sound output control device SV performs the information processing according to the embodiment. Specifically, as the information processing according to the embodiment, the sound output control device SV performs the information processing according to the first embodiment (first information processing), the information processing according to the second embodiment (second information processing), and the information processing according to the third embodiment (third information processing), which will be described later. Also, the information processing according to the embodiment is a process related to notification control for outputting a voice message from the notification unit of the terminal device 10. In FIG. 1, an example where the sound output control device SV is a server device is shown, but it may be realized by, for example, a cloud system.
[0039] Also, as shown in FIG. 1, each information processing according to the embodiment is performed by the information integration engine ESV mounted on the sound output control device SV. As shown in FIG. 1, the information integration engine ESV includes functions such as a request manager function ESV1 and a response manager function ESV2.
[0040] The request manager function ESV1 receives requests from the agent device 60-x and performs queuing according to the received requests. Here, the request may be an output request that requests to output the generated voice content to the user, and is, for example, transmitted in a state including the voice content. Further, the request manager function ESV1 queues the received voice content in the content buffer 122 (FIG. 4).
[0041] The response manager function ESV2 determines the priority of how to actually output the voice content reserved for output based on the data related to the situation grasped by the situation grasping device 30 (for example, data indicating the result of the analysis process) and the tolerance range information included in the request. Then, the response manager function ESV2 performs output control on the terminal device 10 so as to output each voice content in the determined priority order. Note that the output control for the terminal device 10 includes the concept of output control for the notification unit included in the terminal device 10.
[0042] [3. Flow of information processing according to the embodiment] So far, each device included in the information processing system 1 has been described in focus. Next, the overall flow of the information processing according to the embodiment performed in the information processing system 1 will be described. Here, a scenario is assumed in which voice content is output to the user U1 who is driving the vehicle VE1 via the terminal device 10 installed in the vehicle VE1.
[0043] In such a scenario, the terminal device 10 transmits the sensor information detected by the sensors included in the device to the situation grasping device 30 at any time (step S11).
[0044] When the situation awareness engine E30 of the situation awareness device 30 acquires the sensor information transmitted from the terminal device 10, it performs analysis processing for grasping various situations including the driving state of the vehicle VE1 and the state of the user U1 driving the vehicle VE1 (step S12). For example, the situation awareness engine E30 performs a series of analysis processing such as sensing using the sensor information, base analysis using the sensing result as a core element, and higher-order analysis using the data obtained from the result of the base analysis, thereby performing detailed situation awareness.
[0045] Also, when the analysis processing is completed, the situation awareness device 30 transmits data regarding the situation grasped by the situation awareness engine E30 (for example, data indicating the result of the analysis processing) to the agent device 60-x (step S13). In the example of FIG. 1, the situation awareness device 30 transmits data regarding the situation to each agent device 60-x such as the agent device 60-1 and the agent device 60-2.
[0046] When the agent information generation unit 632-x of the agent device 60-x acquires data regarding the situation from the situation awareness device 30, it performs generation processing for generating voice content to be output based on such data (step S14). For example, the agent information generation unit 632-x determines, based on the acquired data, which category of voice content belonging to the categories that the own device can handle should be output, and generates voice content of the content belonging to the determined category. For example, the agent information generation unit 632-x generates message information (text data) of content corresponding to the situation indicated by the acquired data.
[0047] Also, the agent information generation unit 632-x transmits the generated voice content to the voice output control device SV in a state where a category ID for identifying the category to which the voice content belongs (the category to which the message information belongs) and an app ID for identifying the app APx corresponding to the own device are attached (step S15).
[0048] In the example of FIG. 1, since the generation process in step S14 is performed by each agent device 60-x such as agent device 60-1 and agent device 60-2, an example is shown in step S15 where each agent device 60-x transmits its own voice content to the voice output control device SV.
[0049] Subsequently, when the voice content to be output is acquired, the information integration engine ESV of the voice output control device SV performs notification control processing on the voice content to be output (step S16). For example, when converting the message information included in the voice content to be output into voice data (voice message), the information integration engine ESV converts while changing the voice mode according to the category to which the voice content belongs, and performs notification control so that the converted voice data is notified. Also, for example, the information integration engine ESV performs notification control so that a sound effect (for example, background sound) corresponding to the application type indicating to which application the voice content to be output belongs is added to the converted voice data (voice message) and then notified. Such notification control processing will be described in detail in the first embodiment and the second embodiment described later.
[0050] Finally, the voice output control device SV performs voice output control on the terminal device 10 according to the notification control by the information integration engine ESV (step S17). Specifically, the voice output control device SV controls the terminal device 10 so that the voice data notified by the information integration engine ESV is output by the notification unit of the terminal device 10.
[0051] (First Embodiment) [1. Outline of the First Embodiment] From here, the first embodiment will be described. The information processing according to the first embodiment (i.e., the first information processing) is performed for the purpose of solving the above-described first problem. Specifically, the first information processing is performed by a sound output control device 100 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 100 performs the first information processing according to the sound output control program according to the first embodiment. Further, the sound output control device 100 has a structure including a category classification database 121 (FIG. 3) and a content buffer 122 (FIG. 4).
[0052] [2. Configuration of the Sound Output Control Device According to the First Embodiment] Next, with reference to FIG. 2, the sound output control device 100 according to the first embodiment will be described. FIG. 2 is a diagram showing a configuration example of the sound output control device 100 according to the first embodiment. As shown in FIG. 2, the sound output control device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0053] (Regarding the Communication Unit 110) The communication unit 110 is realized by, for example, a NIC or the like. And the communication unit 110 is connected to the network by wire or wirelessly, and performs information transmission and reception with, for example, the terminal device 10, the situation grasping device 30, and the agent device 60-x.
[0054] (Regarding the Storage Unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 includes a category classification database 121 and a content buffer 122.
[0055] (Regarding the Category Classification Database 121) The category classification database 121 stores information regarding the category to which the voice content (agent information) provided by the application APx belongs. Here, FIG. 3 shows an example of the category classification database 121 according to the first embodiment. In the example of FIG. 3, the category classification database 121 has items such as "category ID", "category", and "voice color feature".
[0056] The "category ID" indicates identification information for identifying the candidate "category" used to specify to which category the voice content to be output provided from the application APx side belongs.
[0057] The "category" is the candidate "category" used to specify to which category the voice content to be output provided from the application side belongs. In the example of FIG. 3, the candidate "categories" include "attention", "warning", "entertainment", "advertisement", "guidance", "news", etc. Note that even when the applications that provided the voice content to be output are different, there are situations where the voice content to be output itself belongs to the same category. For example, even when different applications such as application AP1 and application AP5 provide the voice content to be output, this voice content to be output may all belong to the category "entertainment".
[0058] The "voice color feature" indicates the candidate voice color parameters used in the notification control process of changing the voice color when outputting the voice corresponding to this voice content from the notification unit of the terminal device 10 according to the category to which the voice content to be output belongs.
[0059] In the example of FIG. 3, for the category "Attention" identified by the category ID "CT1", the timbre feature "male voice + slowly" is associated. Such an example shows a case where when the voice content to be output belongs to the category "Attention", it is defined that the voice message output from the notification unit of the terminal device 10 is changed to the timbre with the feature of "male voice + slowly". Therefore, the timbre parameters in such an example correspond to the parameters indicating "male voice + slowly".
[0060] Also, in the example of FIG. 3, for the category "Entertainment" identified by the category ID "CT3", the timbre feature "female voice + fast speech" is associated. Such an example shows a case where when the voice content to be output belongs to the category "Entertainment", it is defined that the voice message output from the notification unit of the terminal device 10 is changed to the timbre with the feature of "female voice + fast speech". Therefore, the timbre parameters in such an example correspond to the parameters indicating "female voice + fast speech".
[0061] (Content buffer 122) The content buffer 122 functions as a storage area for queuing information regarding the voice content transmitted from the agent device 60-x. Here, FIG. 4 shows an example of the data stored in the content buffer 122 according to the first embodiment. In the example of FIG. 4, the content buffer 122 has items such as "destination user ID", "app ID", "category ID", and "voice content".
[0062] The "destination user ID" indicates the identification information for identifying the user (or the terminal device 10 of the user) to which the "voice content" is output (notified). The "app ID" indicates the identification information for identifying the application that provided the "voice content" to be output (or the agent device 60-x corresponding to the application). Note that the application that provided the content can be rephrased as the application that generated the "voice content" to be output.
[0063] The "Category ID" indicates identification information that identifies the category to which the "audio content" that is the output target provided by the application identified by the "App ID" belongs. The "Category ID" is given to the "audio content" that is the output target by the agent device 60-x corresponding to the application identified by the "App ID".
[0064] The "audio content" is information regarding the "audio content" that is the output target provided by the application identified by the "App ID". The "audio content" includes, for example, text data as message information.
[0065] That is, in the example of FIG. 4, it shows an example where the content of the message information #11 provided by the application (app AP1) identified by the app ID "AP1" belongs to the category identified by the category ID "CT3" and is to be output to the user (user U1) identified by the user ID "U1".
[0066] (Regarding the control unit 130) Returning to FIG. 2, the control unit 130 is realized by various programs (for example, an audio output control program) stored in the storage device inside the audio output control device 100 being executed with the RAM as a work area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), etc. Further, the control unit 130 is realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0067] As shown in FIG. 2, the control unit 130 is equipped with an information integration engine E100. The information integration engine E100 corresponds to the information integration engine ESV described in FIG. 1. The information integration engine E100 includes a request manager function E101 (corresponding to the request manager function ESV1 described in FIG. 1) and a response manager function E102 (corresponding to the response manager function ESV2 described in FIG. 1).
[0068] Also, as shown in FIG. 2, the request manager function E101 has a request reception unit 131 and a queuing unit 132. Further, the response manager function E102 has an information acquisition unit 133, a determination unit 134, and a notification control unit 135.
[0069] (Regarding the request reception unit 131) The request reception unit 131 receives requests from the agent device 60-x. Specifically, the request reception unit 131 receives from the agent device 60-x a request to request the user to output the voice content to be output. Such a request may include the voice content to be output, a user ID for identifying the destination user, allowable range information for conditioning the period or timing for allowing the output of the voice content, and the like.
[0070] (Regarding the queuing unit 132) The queuing unit 132 queues the voice content to be output in response to the request received by the request reception unit 131. For example, the queuing unit 132 queues the voice content associated with the request in the content buffer 122.
[0071] (Regarding the information acquisition unit 133) The information acquisition unit 133 acquires the voice content (agent information) to be output, which belongs to each of a plurality of different categories, from the agent device 60-x capable of outputting the voice content. Specifically, the information acquisition unit 133 is an agent device 60-x capable of outputting agent information provided from each of a plurality of applications, and acquires voice content belonging to a plurality of different categories from each of the agent devices 60-x corresponding to the agent functions of the applications, and delivers it to the queuing unit 132.
[0072] (Regarding the determination unit 134) Based on the allowable range information included in the request received by the request reception unit 131, the determination unit 134 determines the priority order of how to actually output the voice content reserved for output, and reads out the voice content that has reached the timing to be output from the content buffer 122.
[0073] (Regarding the notification control unit 135) The notification control unit 135 converts the message information included in the voice content (agent information) transmitted from the agent device 60-x into voice data. For example, for the voice content determined to be output by the determination unit 134, the notification control unit 135 uses the Text to Speech (TTS) technology to synthesize voice based on the text data, thereby converting the message information into voice data. Then, the voice data (voice message) obtained by converting the message information is output from the notification unit of the terminal device 10.
[0074] In addition, the notification control unit 135 changes the mode of the voice message according to the category to which the voice content to be output belongs, and causes the notification unit to perform notification. For example, regardless of which application among a plurality of applications the voice content to be output is provided from, the notification control unit 135 changes the voice synthesis tone color parameter according to the category to which the voice content to be output belongs, thereby changing the mode of the voice data to be converted.
[0075] Also, as described above, category identification information (category ID) for identifying the category to which each voice content acquired from the agent device 60-x belongs is attached to each voice content. Therefore, the notification control unit 135 changes the mode of the voice message according to the category indicated by the category identification information attached to the voice content to be output.
[0076] For example, the notification control unit 135 changes the tone color of the voice message according to the category to which the voice content to be output belongs among a plurality of different categories.
[0077] In addition, the notification control unit 135 may cause the notification unit to notify the voice data with a sound effect added according to the category to which the voice content to be output belongs among a plurality of different categories. Here, the sound effect refers to a sound effect added to the beginning or end of the voice message, or background sound such as background music (BGM) superimposed on the voice message.
[0078] On the other hand, different sound effects may be set among a plurality of applications. In such a case, the notification control unit 135 may cause the voice message to be notified from the notification unit with the sound effect corresponding to the application that provided the voice content to be output added among the different sound effects among the plurality of applications. Specifically, application identification information (app ID) for identifying the application that provided each voice content is attached to each voice content acquired from the agent device 60-x. Therefore, the notification control unit 135 causes the voice message to be notified from the notification unit with the sound effect corresponding to the application indicated by the app ID attached to the voice content to be output added among the plurality of applications. This point will be described in detail in the second embodiment.
[0079] [3. Specific Example of Sound Output Control Method] Subsequently, with reference to FIG. 5, a specific example of the sound output control method performed in the first information processing will be described. FIG. 5 is a diagram showing an example of the sound output control method according to the first embodiment.
[0080] FIG. 5 shows five applications as multiple apps, namely, app AP1, app AP2, app AP3, app AP4, and app AP5 (apps AP1 to AP5). Further, FIG. 5 shows agent devices 60-1, 60-2, 60-3, 60-4, and 60-5 as agent devices 60-x that can output voice contents provided from each of apps AP1 to AP5 to a notification unit. Specifically, according to the example of FIG. 5, agent device 60-1 is an agent device that can output voice content corresponding to app AP1 so as to be provided to the user. Agent device 60-2 is an agent device that can output voice content corresponding to app AP2 so as to be provided to the user. Agent device 60-3 is an agent device that can output voice content corresponding to app AP3 so as to be provided to the user. Agent device 60-4 is an agent device that can output voice content corresponding to app AP4 so as to be provided to the user. Agent device 60-5 is an agent device that can output voice content corresponding to app AP5 so as to be provided to the user.
[0081] Also, according to the example of FIG. 5, the agent device 60-1 is a device capable of outputting voice content belonging to the category "Entertainment" and voice content belonging to the category "Advertisement". Also, the agent device 60-2 is a device capable of outputting voice content belonging to the category "Caution" and voice content belonging to the category "Warning". Also, the agent device 60-3 is a device capable of outputting voice content belonging to the category "Guidance". Also, the agent device 60-4 is a device capable of outputting voice content belonging to the category "News" and voice content belonging to the category "Advertisement". Also, the agent device 60-5 is a device capable of outputting voice content belonging to the category "Caution" and voice content belonging to the category "Entertainment".
[0082] Here, for example, it is assumed that the agent information generation unit 632-1 of the agent device 60-1 generates voice content A-1 corresponding to message information of content belonging to the category "Entertainment" based on the data regarding the situation acquired from the situation grasping device 30. In such a case, as shown in FIG. 5, the agent device 60-1 transmits to the sound output control device 100 in a state where the category ID "CT3" for identifying the category "Entertainment" is attached to the voice content A-1 so that the voice content A-1 is output to the user U1. Also, at this time, the agent device 60-1 may further attach the app ID "AP1" for identifying the app AP1 which is the source application providing the voice content A-1.
[0083] The information acquisition unit 133 of the sound output control device 100 acquires the voice content A-1 with the category ID "CT3" from the agent device 60-1 as the voice content to be output. Subsequently, when the voice content A-1 is determined as the voice content to be output from the notification unit of the terminal device 10 through the processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined voice content A-1 to be output using the voice color parameters corresponding to the category to which the voice content A-1 belongs.
[0084] According to the example of FIG. 5, the notification control unit 135 can identify that the category to which the voice content A-1 belongs is "entertainment" by comparing the category ID "CT3" assigned to the voice content A-1 with the category classification database 121. Also, the notification control unit 135 refers to the category classification database 121 and recognizes that the voice message to be output from the notification unit of the terminal device 10 is defined to be changed to a voice color with the characteristics of "female voice + fast speech". Then, the notification control unit 135 changes the voice color by performing voice synthesis on the voice data included in the voice content A-1 using the parameters indicating "female voice + fast speech".
[0085] Subsequently, the notification control unit 135 performs sound output control so that the voice message after the voice synthesis of the voice content A-1 is notified from the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls to output the voice message corresponding to the voice content A-1 by transmitting the voice message after the voice synthesis of the voice content A-1 to the terminal device 10 of the user U1. The terminal device 10 notifies the voice message from the notification unit according to the sound output control from the notification control unit 135. Thereby, the user U1 can easily grasp that the currently output voice content belongs to the category "entertainment".
[0086] Note that the notification control unit 135 not only changes the tone color of the voice message corresponding to the voice content A-1 in response to the voice content A-1 belonging to the category "Entertainment", but may also add a sound effect (for example, background sounds such as sound effects and BGM) corresponding to the voice content A-1 belonging to the category "Entertainment" to the voice message. In such a case, for example, in the category classification database 121 shown in FIG. 3, data on sound effects corresponding to the category indicated by the "category ID" may be associated for each "category ID" (not shown).
[0087] Next, regarding the point that the mode of the voice data is changed according to the category to which the voice content to be output belongs, regardless of which application among a plurality of applications the voice content to be output is provided from, examples of the voice contents A-2 and D-3 will be used for explanation.
[0088] For example, assume that the agent information generation unit 632-1 of the agent device 60-1 generates voice content A-2 corresponding to message information of content belonging to the category "Advertisement" based on data regarding the situation acquired from the situation grasping device 30. In such a case, as shown in FIG. 5, the agent device 60-1 transmits the voice content A-2 to the voice output control device 100 in a state where the category ID "CT4" for identifying the category "Advertisement" is attached so that the voice content A-2 is output to the user U1. At this time, the agent device 60-1 may further attach an app ID "AP1" for identifying the app AP1 which is the source application that provides the voice content A-2.
[0089] The information acquisition unit 133 acquires the voice content A-2 with the category ID "CT4" from the agent device 60-1 as the voice content to be output. Subsequently, when the voice content A-2 is determined as the voice content to be output from the notification unit of the terminal device 10 through the processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined voice content A-2 to be output using the voice color parameters corresponding to the category to which the voice content A-2 belongs.
[0090] According to the example of FIG. 5, the notification control unit 135 can identify that the category to which the voice content A-2 belongs is "advertisement" by comparing the category ID "CT4" assigned to the voice content A-2 with the category classification database 121. Also, the notification control unit 135 refers to the category classification database 121 and recognizes that the voice message to be output from the notification unit of the terminal device 10 is defined to be changed to a voice color with the characteristic of "robot voice + slowly". Then, the notification control unit 135 changes the voice color by performing voice synthesis on the voice data included in the voice content A-2 using the parameters indicating "robot voice + slowly".
[0091] Subsequently, the notification control unit 135 performs sound output control so that the voice message after the voice synthesis of the voice content A-2 is notified from the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls to output the voice message corresponding to the voice content A-2 by transmitting the voice message after the voice synthesis of the voice content A-2 to the terminal device 10 of the user U1. The terminal device 10 notifies the voice message from the notification unit according to the sound output control from the notification control unit 135. As a result, the user U1 can easily grasp that the currently output voice content belongs to the category "advertisement".
[0092] Also, for example, assume that the agent information generation unit 631-4 of the agent device 60-4 generates voice content D-3 corresponding to message information belonging to the category "advertisement" based on data related to the situation acquired from the situation awareness device 30. In such a case, as shown in FIG. 5, the agent device 60-4 transmits the voice content D-3 to the sound output control device 100 with the category ID "CT4" that identifies the category "advertisement" attached thereto so that the voice content D-3 is output to the user U1. At this time, the agent device 60-1 may further attach an app ID "AP4" that identifies the app AP4, which is the application that provides the voice content D-3.
[0093] The information acquisition unit 133 acquires the voice content D-3 with the category ID "CT4" attached thereto from the agent device 60-4 as the voice content to be output. Subsequently, when the voice content D-3 is determined as the voice content to be output from the notification unit of the terminal device 10 by the processing of the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined output target voice content D-3 using the tone color parameters corresponding to the category to which the voice content D-3 belongs.
[0094] According to the example of FIG. 5, the notification control unit 135 can identify that the category to which the voice content D-3 belongs is "advertisement" by comparing the category ID "CT4" attached to the voice content D-3 with the category classification database 121. Also, the notification control unit 135 refers to the category classification database 121 and recognizes that the voice message to be output from the notification unit of the terminal device 10 is defined to be changed to a tone color with the feature of "robot voice + slowly". Then, the notification control unit 135 changes the tone color of the voice by performing voice synthesis on the voice data included in the voice content D-3 using the parameter indicating "robot voice + slowly".
[0095] Subsequently, the notification control unit 135 performs sound output control so that the voice message after voice synthesis of the voice content D-3 is notified from the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls to output the voice message corresponding to the voice content D-3 by transmitting the voice message after voice synthesis of the voice content D-3 to the terminal device 10 of the user U1. The terminal device 10 notifies the voice message from the notification unit according to the sound output control from the notification control unit 135. As a result, the user U1 can easily grasp that the currently output voice content belongs to the category "advertisement".
[0096] Here, according to the above two examples, the apps that are the providers of the voice content are different, such as app AP1 and app AP4. However, since each voice content provided by both belongs to the same category (advertisement), it is output in the same mode (robot voice + slowly) regardless of the type of the app.
[0097] Also, regardless of which application among a plurality of applications the voice content to be output is provided from, sound effects (for example, background sounds such as sound effects and BGM) corresponding to the category to which the voice content to be output belongs may be added to the voice content.
[0098] So far, a specific example of the sound output control method performed in the first information processing has been described by taking some of the voice contents shown in FIG. 5 as examples. Since the other voice contents shown in FIG. 5 can also be described by following the examples of some voice contents, detailed descriptions are omitted.
[0099] [4. Processing procedure] Next, with reference to FIG. 6, the procedure of the information processing according to the first embodiment will be described. FIG. 6 is a flowchart showing the information processing procedure according to the first embodiment. Note that the flow shown in the flowchart of FIG. 6 is repeatedly executed, for example, while the user U1 is driving the vehicle VE1.
[0100] First, the control unit 130 of the sound output control device 100 determines whether it has acquired agent information from the agent device 60-x (step S101). If the control unit 130 determines that new agent information has been acquired (step S101; Yes), it performs queuing processing on the acquired agent information (step S102). In step S102, the newly acquired agent information is queued together with the already acquired agent information, and the priority for output as a voice message is determined, and then it proceeds to step S103. On the other hand, if it is determined in step S101 that new agent information cannot be acquired from the agent device 60-x (step S101; No), it directly proceeds to step S103.
[0101] Next, the control unit 130 determines whether there is agent information (voice content) acquired from the agent device 60-x for which the output timing has arrived (step S103). If the control unit 130 determines that there is no agent information for which the output timing has arrived (step S103; No), it temporarily ends the flow and repeats the flow from the beginning.
[0102] On the other hand, if the control unit 130 determines that there is agent information for which the output timing has arrived (step S103; Yes), it identifies the category to which this agent information belongs based on the category ID assigned to the agent information to be output (step S104). For example, the control unit 130 identifies the category to which the agent information to be output belongs by comparing the category ID assigned to the agent information to be output with the category classification database 121.
[0103] Also, the control unit 130 identifies the timbre feature (timbre parameter) corresponding to the identified category among the timbre features (timbre parameters) set to be different among categories, as in the example of the category classification database 121 shown in FIG. 3 (step S105).
[0104] Then, when converting the message information included in the agent information to be output into voice data, the control unit 130 changes the voice synthesis parameters to the specified voice color parameters (voice color parameters corresponding to the category to which the agent information to be output belongs) and performs voice conversion (step S106).
[0105] Finally, the control unit 130 performs sound output control so that the voice data corresponding to the agent information to be output is notified from the notification unit of the terminal device 10 of the user designated as the destination of the agent information (step S107). After that, the control unit 130 repeats the flow from the beginning.
[0106] In the flowchart of FIG. 6, after specifying the agent information at the timing to be output, the procedure of converting the message information included in the specified agent information into voice data has been described. However, the timing of converting the message information into voice data is not limited to this. For example, as soon as the information acquisition unit 133 acquires new agent information, the conversion process to voice data corresponding to steps S104 to S106 is immediately performed, and for the voice content including the converted voice data, the determination of the output priority order and the determination process of the timing to be output, corresponding to steps S102 and S103, may be performed.
[0107] 〔5. Summary〕 The sound output control device 100 according to the first embodiment acquires output target agent information from an agent device capable of outputting agent information belonging to each of a plurality of different categories as information provided to the user. Then, the sound output control device 100 causes the notification unit to output a voice message corresponding to the output target agent information. Specifically, the sound output control device 100 causes the notification unit to notify by changing the mode of the voice message according to the category to which the output target agent information belongs. According to such a sound output control device 100, the user can easily grasp whether the currently output voice content is the voice content of the desired category.
[0108] (Second Embodiment) [1. Outline of the Second Embodiment] Hereinafter, the second embodiment will be described. The information processing according to the second embodiment (that is, the second information processing) is performed for the purpose of solving the above-described second problem. Specifically, the second information processing is performed by a sound output control device 200 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 200 performs the second information processing according to the sound output control program according to the second embodiment. Further, the sound output control device 200 has a structure including an application classification database 233 in addition to the category classification database 121 (FIG. 3) and the content buffer 122 (FIG. 4).
[0109] [2. Configuration of the Sound Output Control Device According to the Second Embodiment] Next, the sound output control device 200 according to the second embodiment will be described with reference to FIG. 7. FIG. 7 is a diagram showing a configuration example of the sound output control device 200 according to the second embodiment. As shown in FIG. 7, the sound output control device 200 includes a communication unit 110, a storage unit 220, and a control unit 130. In the following description, the description of the processing unit assigned the same reference numeral as the sound output control device 100 may be omitted or simplified.
[0110] (Regarding the storage unit 220) The storage unit 220 is implemented by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 220 includes a category classification database 121, a content buffer 122, and an app classification database 233.
[0111] (Regarding the app classification database 233) The app classification database 233 stores information related to sound effects. Here, FIG. 8 shows an example of the app classification database 233 according to the second embodiment. In the example of FIG. 8, the app classification database 233 has items such as "app ID", "app type", and "sound effect".
[0112] The "app ID" indicates identification information for identifying the application that provided the "voice content" to be output (or the agent device 60-x corresponding to the application). Note that the application that provided the "voice content" to be output can be rephrased as the application that generated the "voice content" to be output. The "app type" is information regarding the type of the application identified by the "app ID", and may be, for example, the name of the application. Also, the "app type" corresponds to the type to which the voice content (agent information) to be output provided by the application identified by the "app ID" belongs.
[0113] The "sound effect" is a candidate for background sound to be superimposed on the voice content to be output according to the application that provided the voice content to be output, and the background sound may be, for example, a sound effect or music.
[0114] In the example of FIG. 8, the sound effect "sound effect #1" is associated with the app ID "AP1". Such an example shows a case where, when the application that provided the voice content to be output is app AP1, it is specified that the voice data (voice message) included in the voice content to be output is notified in a state where the sound effect #1 is superimposed as background sound.
[0115] (Regarding the information acquisition unit 133) The information acquisition unit 133 acquires the voice content to be output from an agent device capable of outputting a plurality of types of voice contents distinguishable by the content of the information or the source of the information as information provided to the user.
[0116] For example, the information acquisition unit 133 is an agent device 60-x capable of outputting voice content provided from each of a plurality of applications, and acquires the voice content to be output from the agent device 60-x corresponding to the agent function of the application.
[0117] (Regarding the notification control unit 135) The notification control unit 135 causes the notification unit to output voice data corresponding to the voice content to be output.
[0118] In addition, the notification control unit 135 adds background sound corresponding to the type to which the voice content to be output belongs to the voice data and causes the notification unit to perform notification.
[0119] For example, the notification control unit 135 adds background sound corresponding to the application of the source that provided the voice content to be output as the type to which the voice content to be output belongs to the voice data and causes the notification unit to perform notification. In such a case, application identification information for identifying the application of the source that provided the voice content is attached to each voice content acquired from the agent device 60-x. Therefore, the notification control unit 135 adds background sound corresponding to the application indicated by the application identification information attached to the voice content to be output among the plurality of applications to the voice data and causes the notification unit to perform notification.
[0120] In addition, the plurality of types of voice contents may include voice contents belonging to a plurality of different categories distinguished based on the content of the voice contents. In such a case, the information acquisition unit 133 acquires the voice content to be output from the agent device 60-x among the voice contents belonging to the plurality of different categories. Then, the notification control unit 135 adds background sound corresponding to the category to which the voice content to be output belongs among the background sounds different among the plurality of different categories to the voice message and causes the notification unit to perform notification. As a specific example, the information acquisition unit 133 is an agent device 60-x capable of outputting voice contents provided from each of a plurality of applications, and acquires the voice content to be output from the agent device 60-x corresponding to the agent function of the application. Then, regardless of which application among the plurality of applications the voice content to be output is agent information provided from, the notification control unit 135 adds background sound corresponding to the category to which the voice content to be output belongs to the voice data and causes the notification unit to perform notification.
[0121] Also, as described in the first embodiment, the notification control unit 135 may control the notification unit so that the voice message is notified in a tone color corresponding to the category to which the voice content to be output belongs among the plurality of different categories. In such a case, category identification information for identifying the category to which the voice content belongs is given to each voice content acquired from the agent device 60-x. Therefore, the notification control unit 135 controls the notification unit so that the voice data is notified in a tone color corresponding to the category indicated by the category identification information given to the voice content to be output.
[0122] [3. Specific Example of Sound Output Control Method] Subsequently, with reference to FIG. 9, a specific example of the sound output control method performed in the second information process will be described. FIG. 9 is a diagram showing an example of the sound output control method according to the second embodiment.
[0123] Many of FIG. 9 correspond to the example of FIG. 5. Specifically, FIG. 9 shows five applications as a plurality of applications, such as application AP1, application AP2, application AP3, application AP4, and application AP5 (applications AP1 to AP5). Further, FIG. 5 shows agent devices 60-1, agent device 60-2, agent device 60-3, agent device 60-4, and agent device 60-5 as agent devices 60-x capable of outputting voice content provided from each of applications AP1 to AP5 to the notification unit. The description of each agent device 60-x is omitted.
[0124] Here, for example, it is assumed that the agent information generation unit 632-1 of the agent device 60-1 generates voice content A-1 using voice data corresponding to message information of content belonging to the category "entertainment" based on data regarding the situation acquired from the situation grasping device 30. In such a case, as shown in FIG. 9, the agent device 60-1 assigns a category ID "CT3" for identifying the category "entertainment" to the voice content A-1 so that the voice content A-1 is output to the user U1. Further, the agent device 60-1 further assigns an application ID "AP1" for identifying the application AP1, which is the application that provides the voice content A-1, to the voice content A-1. Then, the agent device 60-1 transmits the voice content A-1 with the category ID and the application ID assigned thereto to the sound output control device 200.
[0125] The information acquisition unit 133 of the sound output control device 200 acquires the voice content A-1 with the application ID "AP1" and the category ID "CT3" assigned thereto from the agent device 60-1 as the voice content to be output. Subsequently, when the voice content A-1 is determined as the voice content to be output from the notification unit of the terminal device 10 by the processing of the queuing unit 132 and the determination unit 134, the notification control unit 135 performs voice synthesis on the message information included in the determined voice content A-1 to be output using voice color parameters corresponding to the category to which the voice content A-1 belongs.
[0126] According to the example of FIG. 9, similar to the first embodiment, the notification control unit 135 can identify that the category to which the voice content A-1 belongs is "entertainment" by comparing the category ID "CT3" assigned to the voice content A-1 with the category classification database 121. Further, the notification control unit 135 refers to the category classification database 121 and recognizes that the voice message output from the notification unit of the terminal device 10 is defined to be changed to a voice color with the characteristics of "female voice + fast speech". Then, the notification control unit 135 changes the voice color by performing voice synthesis on the voice data included in the voice content A-1 using the parameter indicating "female voice + fast speech".
[0127] In the second embodiment, in addition to this, the sound output control device 200 superimposes and outputs the background sound corresponding to the application that provides the voice content A-1 on the voice message.
[0128] For example, the notification control unit 135 further identifies that the application type to which the voice content A-1 belongs is "application AP1" by comparing the application ID "AP1" assigned to the voice content A-1 with the application classification database 223.
[0129] Also, the notification control unit 135 extracts the sound effect #1 from the application classification database 223 according to the fact that the application type to which the voice content A-1 belongs is "application AP1". Then, the notification control unit 135 adds the extracted sound effect #1 as the background sound to the voice message after voice synthesis.
[0130] Next, the notification control unit 135 performs sound output control so that the voice content A-1 after the conversion process such as voice synthesis and addition of background sound as described above is notified from the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output the voice content A-1 after the conversion process by transmitting the voice content A-1 after the conversion process. The terminal device 10 notifies the voice content A-1 after the conversion process from the notification unit in response to the sound output control from the notification control unit 135. As a result, for the user U1, a voice message of "female voice + fast speech" is output with an effect sound such as "pipipipip..." (an example of effect sound #1) as the background sound. That is, the user U1 can simultaneously hear a voice message with a tone color corresponding to the category of the voice content and a background sound corresponding to the providing application. As a result, the user U1 can easily grasp that the currently output voice content is related to "entertainment" provided by the application AP1.
[0131] Subsequently, another example shown in FIG. 9 will be described. For example, it is assumed that the agent information generation unit 631-5 of the agent device 60-5 generates voice content E-1 using voice data corresponding to message information belonging to the category "Attention" based on the data regarding the situation acquired from the situation grasping device 30. In such a case, as shown in FIG. 9, the agent device 60-5 assigns a category ID "CT1" for identifying the category "Attention" to the voice content E-1 so that the voice content E-1 is output to the user U1. Further, the agent device 60-5 further assigns an app ID "AP5" for identifying the application AP5 which is the application that provides the voice content E-1 to the voice content E-1. Then, the agent device 60-5 transmits the voice content A-1 to which the category ID and the app ID are assigned to the sound output control device 200.
[0132] The information acquisition unit 133 acquires, from the agent device 60-5, the voice content E-1 with the app ID "AP5" and the category ID "CT1" as the voice content to be output. Subsequently, when the voice content E-1 is determined as the voice content to be output from the notification unit of the terminal device 10 through the processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs voice synthesis on the message information included in the determined voice content E-1 to be output, using the voice color parameters corresponding to the category to which the voice content E-1 belongs.
[0133] According to the example of FIG. 9, the notification control unit 135 can identify that the category to which the voice content E-1 belongs is "Attention" by comparing the category ID "CT1" assigned to the voice content E-1 with the category classification database 121. Also, the notification control unit 135 refers to the category classification database 121 and recognizes that the voice message to be output from the notification unit of the terminal device 10 is defined to be changed to a voice color with the characteristics of "male voice + slowly". Then, the notification control unit 135 changes the voice color by performing voice synthesis on the voice data included in the voice content E-1 using the parameters indicating "male voice + slowly".
[0134] Furthermore, the notification control unit 135 can identify that the app type to which the voice content E-1 belongs is "App AP5" by comparing the app ID "AP5" assigned to the voice content E-1 with the app classification database 223.
[0135] Also, the notification control unit 135 extracts Music♯5 from the app classification database 223 according to the fact that the app type to which the voice content E-1 belongs is "App AP5". Then, the notification control unit 135 adds the extracted Music♯5 as background music to the voice message after voice synthesis.
[0136] Next, the notification control unit 135 performs sound output control so that the voice content E-1 after the conversion process such as voice synthesis and addition of background sound as described above is notified from the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output the voice content E-1 after the conversion process by transmitting the voice content E-1 after the conversion process. The terminal device 10 notifies the voice content E-1 after the conversion process from the notification unit in response to the sound output control from the notification control unit 135. As a result, a voice message of "male voice + slowly" is output to the user U1 with Music #5 as the background sound. That is, the user U1 can simultaneously listen to the voice message with the timbre corresponding to the category of the voice content and the background sound corresponding to the providing application. As a result, the user U1 can easily grasp that the currently output voice content is related to "Attention" provided from the application AP5.
[0137] So far, taking some of the voice contents shown in FIG. 9 as an example, a specific example of the sound output control method performed in the second information process has been described. Since the other voice contents shown in FIG. 9 can also be described by following the example of some voice contents, detailed description thereof will be omitted.
[0138] [4. Processing Procedure] Next, with reference to FIG. 10, the procedure of the information process according to the second embodiment will be described. FIG. 10 is a flowchart showing the information process procedure according to the second embodiment. Steps S101 to S106 shown in FIG. 10 are common to the example of FIG. 6, so the description thereof will be omitted, and steps S207 to S210 newly added in the information process according to the second embodiment will be described.
[0139] Based on the app ID assigned to the agent information to be output, the control unit 130 identifies the app type to which the agent information to be output belongs (step S207). For example, the control unit 130 identifies the app type to which the agent information to be output belongs by comparing the app ID with the app classification database 223.
[0140] Also, as in the example of the app classification database 223 shown in FIG. 8, the control unit 130 extracts the background sound corresponding to the identified app type from among the background sounds set to be different between apps (step S208).
[0141] Also, the control unit 130 adds the extracted background sound to the agent information after voice conversion (step S209).
[0142] Finally, the notification control unit 135 performs sound output control so that the agent information after adding the background sound is notified from the notification unit of the terminal device 10 of the user U1 (step S210).
[0143] In the flowchart of FIG. 10, after identifying the agent information at the timing to be output, the message information included in the identified agent information is converted into voice data, and then the background sound is added. However, the timing of converting the message information into voice data and the timing of adding the background sound are not limited to this. For example, as soon as the information acquisition unit 133 acquires new agent information, the conversion process into voice data corresponding to steps S104 to S106 and the addition process of the background sound corresponding to steps S207 to S209 are performed, and for the voice content including the voice data after adding the background sound, the determination process of the output priority order and the determination process of the timing to be output corresponding to steps S102 and S103 may be performed.
[0144] Note that, as the second embodiment so far, it has been described that the sound output control device 200 superimposes and outputs background sound corresponding to the source application on a voice message with a timbre corresponding to the category of the voice content. However, as another example, the sound output control device 200 may superimpose and output background sound corresponding to the category to which the voice content belongs on a voice message with a timbre corresponding to the source application. For example, when the application that is the source of the voice content is application AP1 and the message information included in the voice content belongs to the category "Entertainment", the notification control unit 135 may perform voice synthesis using the timbre parameters corresponding to application A1 and add background sound corresponding to the category "Entertainment" to the voice message. In such a case, for example, in the category classification database 121 shown in FIG. 3, data of background sound corresponding to the category indicated by the category ID is associated with each "category ID", and in the application classification database 223 shown in FIG. 8, voice characteristics (voice parameters) corresponding to the application type indicated by the application ID are associated with each "application ID".
[0145] Also in this case, user U1 can listen to a voice message with a timbre corresponding to the source application and background sound corresponding to the category of the voice content at the same time. As a result, user U1 can easily grasp the application that is the source of the currently output voice content and the category of the voice content.
[0146] Alternatively, as yet another example, the sound output control device 200 may always fix the tone color of the voice of the voice message to a standard tone color and change only the background sound according to the category of the voice content. That is, regardless of which application among a plurality of applications the voice content to be output is provided from, background sound corresponding to the category to which the voice content to be output belongs may be added. Also in this case, the user U1 can hear the background sound corresponding to the category of the voice content at the same time as the voice message. As a result, the user U1 can easily grasp the category to which the currently output voice content belongs.
[0147] [5. Summary] The sound output control device 200 according to the second embodiment acquires the agent information to be output from an agent device capable of outputting a plurality of types of agent information distinguished by the content of the information or the source of the information as information provided to the user. Then, the sound output control device 200 causes the notification unit to output a voice message corresponding to the agent information to be output. Specifically, the sound output control device 200 adds background sound corresponding to the type to which the agent information to be output belongs to the voice message and causes the notification unit to notify it. According to such a sound output control device 200, the user can easily grasp whether the currently output voice content is the desired type of voice content.
[0148] (Third Embodiment) [1. Outline of the Third Embodiment] Hereinafter, the third embodiment will be described. The information processing according to the third embodiment (that is, the third information processing) is performed for the purpose of solving the above-described third problem. Specifically, the third information processing is performed by a sound output control device 300 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 300 performs the third information processing according to the sound output control program according to the third embodiment.
[0149] 〔2. Configuration of Sound Output Control Device According to the Third Embodiment〕 Next, with reference to FIG. 11, the sound output control device 300 according to the third embodiment will be described. FIG. 11 is a diagram showing a configuration example of the sound output control device 300 according to the third embodiment. As shown in FIG. 11, the sound output control device 300 includes a communication unit 110, a storage unit 220, and a control unit 330. In the following description, the description of the processing units with the same reference numerals as those of the sound output control devices 100 and 200 may be omitted or simplified.
[0150] (Regarding the control unit 330) The control unit 330 is realized by various programs (for example, a sound output control program) stored in a storage device inside the sound output control device 300 being executed with the RAM as a work area by a CPU, an MPU, or the like. Further, the control unit 330 is realized by an integrated circuit such as an ASIC or an FPGA, for example.
[0151] As shown in FIG. 11, the control unit 330 further includes a presentation control unit 336, a sound effect setting unit 337, and a usage stop reception unit 338, and realizes or executes the functions and operations of information processing described below. Note that the internal configuration of the control unit 330 is not limited to the configuration shown in FIG. 11, and other configurations may be used as long as they perform the information processing described later. Also, the connection relationship of each processing unit included in the control unit 330 is not limited to the connection relationship shown in FIG. 11, and other connection relationships may be used.
[0152] (Regarding the presentation control unit 336) As described in the second embodiment, the notification control unit 135 causes the notification unit to notify voice data with different sound effects among a plurality of applications according to the application that is the source of the voice content acquired by the information acquisition unit 133 among the plurality of applications. Therefore, the presentation control unit 336 causes the user to be presented with an application list indicating the sound effects corresponding to each of the plurality of applications.
[0153] For example, the presentation control unit 336 controls so that image information indicating an application list is presented to the user via the display unit (the display screen of the terminal device 10).
[0154] In addition, the presentation control unit 336 causes the notification unit to notify a voice message indicating the name of the application included in the application list with a sound effect corresponding to the application added thereto.
[0155] (Regarding the sound effect setting unit 337) The sound effect setting unit 337 receives a user operation for setting a sound effect for each application.
[0156] (Regarding the usage stop reception unit 338) The usage stop reception unit 338 receives a user operation for stopping the usage of an arbitrary application among a plurality of applications being used by the user. For example, the usage stop reception unit 338 stops the usage of the application selected by the user among the applications included in the application list.
[0157] [3. Specific Example of the Third Information Processing] Subsequently, with reference to FIG. 12, a specific example of the third information processing performed among the presentation control unit 336, the sound effect setting unit 337, and the usage stop reception unit 338 will be described. FIG. 12 is a diagram showing an example of the third information processing.
[0158] FIG. 12 shows an example in which a setting screen C1 capable of performing various settings is displayed on the terminal device 10 for each of a plurality of applications (applications being used by the user U1) associated with the terminal device 10 of the user U1. The setting screen C1 may be provided by the presentation control unit 336 in response to a request from the user U1, for example. Also, the mode (screen configuration) of the setting screen C1 is not limited to the example of FIG. 12.
[0159] For example, assume that the applications associated with the terminal device 10 of user U1 are application AP1, application AP2, application AP3, application AP4, and application AP5. In such a case, as shown in FIG. 12, on the setting screen C1, the application names indicating the names of each of applications AP1 to AP5 are displayed as "List of Applications in Use". Also, such a list of applications corresponds to the application list.
[0160] Further, on the setting screen C1, for each application associated with the terminal device 10 of user U1, it is possible to set the background sound corresponding to the application. Regarding this point, FIG. 12 shows an example in which a pull-down button PD1 for displaying a list of candidates for the background sound corresponding to application AP1 in a pull-down format is associated with the application name indicating application AP1. Thereby, user U1 can set the selected background sound by selecting an arbitrary background sound from the candidates for the background sound displayed in a pull-down manner using the pull-down button PD1.
[0161] For example, as shown in FIG. 12, when BGM "MUSIC♯3" is selected, the sound effect setting unit 337 accepts the setting of BGM "MUSIC♯3" for application AP1 in response to such a selection operation. Also, in response to the setting of BGM "MUSIC♯3" being accepted, the presentation control unit 336 controls, for example, the notification control unit 135 so that voice data (voice message) indicating the application name of application AP1 is output from the notification unit with BGM "MUSIC♯3" added. In such a case, the notification control unit 135 extracts the data of BGM "MUSIC♯3" from the storage unit and adds it to the voice data indicating the application name of application AP1. Then, the notification control unit 135 performs sound output control so that the voice data after adding BGM "MUSIC♯3" is notified from the notification unit of the terminal device 10 of user U1.
[0162] As a result, user U1 can, for example, listen to a voice message (e.g., "This is the entertainment information providing app of Company A") that reads out the app name of app AP1 while the BGM "MUSIC#3" is playing, and can imagine the atmosphere of the BGM "MUSIC#3" and how the voice message sounds within that atmosphere. Also, as a result, user U1 can easily distinguish which app is the provider of the voice content when using multiple apps, as in the example of FIG. 12.
[0163] So far, a specific example of the third information processing has been described using the example of app AP1 shown in FIG. 12, but other apps will also be described.
[0164] For example, FIG. 12 shows an example in which a pull-down button PD3 for displaying a list of background sound candidates corresponding to app AP3 in a pull-down format is associated with the app name of app AP3 shown. Thereby, user U1 can set an arbitrary background sound by selecting it from the background sound candidates displayed in a pull-down manner using pull-down button PD3.
[0165] For example, as shown in FIG. 12, when BGM "MUSIC#1" is selected, sound effect setting unit 337 accepts the setting of BGM "MUSIC#1" for app AP3 in response to such a selection operation. Also, presentation control unit 336 controls, for example, notification control unit 135 so that voice data indicating the app name of app AP3 is output from the notification unit with BGM "MUSIC#1" added in response to the acceptance of the setting of BGM "MUSIC#1". In such a case, notification control unit 135 extracts the data of BGM "MUSIC#1" from the storage unit and adds it to the voice data indicating the app name of app AP3. Then, notification control unit 135 performs sound output control so that the voice data after adding BGM "MUSIC#1" is notified from the notification unit of user U1's terminal device 10.
[0166] As a result, user U1 can, for example, listen to a voice message (e.g., "This is the vacation facility information providing app of Company C") that reads out the app name of app AP3 while the BGM "MUSIC#1" is playing, and can imagine the atmosphere of the BGM "MUSIC#1" and how the voice message sounds in that atmosphere. Also, as a result, user U1 can easily distinguish which app is the provider of the voice content when using a plurality of apps, as in the example of FIG. 12.
[0167] Hereinafter, the deletion of the suspension of use (unnecessary app) of the app will be described using the example of FIG. 12. In the setting screen C1 shown in FIG. 12, among the apps included in the application list, in addition to the function of deleting the selected app, there is a function of putting the selected app into a state of suspension of use.
[0168] For example, user U1 does not need the provision of voice content from app AP1 among the apps AP1 to AP5 that are in use, and wants to put app AP1 into a state of suspension of use. In such a case, user U1 presses the delete execution button BT while app AP1 is selected from the app names included in the application list.
[0169] Then, the suspension reception unit 338 receives a user operation to suspend the use of app AP1. And the suspension reception unit 338 suspends the use of the app AP1 selected by user U1 among the apps included in the application list. For example, the suspension reception unit 338 suspends the use of app AP1 by deleting app AP1 from the application list. As a result, user U1 can, for example, set an environment in which only the voice content he / she needs is output.
[0170] [4. Summary] The sound output control device 300 according to the third embodiment acquires agent information provided from each of a plurality of applications having a voice agent function. Then, the sound output control device 300 causes the notification unit to notify a voice message with different sound effects among the plurality of applications according to the application that is the source of the acquired agent information. In addition, the sound output control device 300 causes a user to present an application list indicating sound effects corresponding to each of the plurality of applications. According to such a sound output control device 300, when the user uses a plurality of applications, the user can easily distinguish which application is the source of the voice content. As a result, the user can identify an application according to his or her preference.
[0171] (Others) [1. Hardware Configuration] In addition, the sound output control device 100 in the first embodiment described above and the sound output control device 200 according to the second embodiment are realized by, for example, a computer 1000 having a configuration as shown in FIG. 13. Hereinafter, the sound output control device 100 will be described as an example. FIG. 13 is a hardware configuration diagram showing an example of a computer that realizes the functions of the sound output control device 100. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0172] The CPU 1100 operates based on a program stored in the ROM 1300 or the HDD 1400 and controls each part. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 is started up, a program depending on the hardware of the computer 1000, and the like.
[0173] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, and the like. The communication interface 1500 receives data from other devices via a predetermined communication network and sends it to the CPU 1100, and sends data generated by the CPU 1100 to other devices via the predetermined communication network.
[0174] The CPU 1100 controls output devices such as displays and printers, and input devices such as keyboards and mice, via the input / output interface 1600. The CPU 1100 acquires data from the input device via the input / output interface 1600. Further, the CPU 1100 outputs the generated data to the output device via the input / output interface 1600.
[0175] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1200. The CPU 1100 loads such a program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc), a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0176] For example, when the computer 1000 functions as the sound output control device 100 in the first embodiment, the CPU 1100 of the computer 1000 realizes the functions of the control unit 130 by executing the program loaded on the RAM 1200. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800. As another example, these programs may be acquired from other devices via a predetermined communication network.
[0177] Further, for example, when the computer 1000 functions as the sound output control device 300 in the third embodiment, the CPU 1100 of the computer 1000 realizes the functions of the control unit 330 by executing the program loaded on the RAM 1200.
[0178] 〔2. Others〕 Also, among the processes described in each of the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.
[0179] Also, each component of each device shown in the drawings is a functional concept, and it is not necessarily physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown in the drawings, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.
[0180] Also, the above embodiments can be appropriately combined within a range that does not conflict with the processing content.
[0181] As described above, some of the embodiments of the present application have been described in detail with reference to the drawings. However, these are examples, and the present invention can be implemented in other forms with various modifications and improvements based on the knowledge of those skilled in the art, starting from the aspects described in the column of the disclosure of the invention.
[0182] Also, the above-mentioned "section, module, unit" can be read as "means", "circuit", etc. For example, the information acquisition section can be read as an information acquisition means or an information acquisition circuit.
Explanation of Reference Numerals
[0183] 1 Information processing system 100 Sound output control device 120 Memory unit 121 Category classification database 122 Content buffer 130 Control unit 133 Information acquisition unit 135 Notification control unit 200 Sound output control device 220 Memory unit 223 Application classification database 300 Sound output control device 336 Presentation control unit 337 Sound effect setting unit 338 Use stop reception unit
Claims
1. An information acquisition unit that acquires agent information to be output from an agent device including a plurality of applications each capable of outputting agent information, as information provided to a user; A notification control unit that causes a notification unit to output a voice message corresponding to the agent information to be output; comprising: The notification control unit adds background sound corresponding to the application that provided the agent information to be output to the voice message and causes the notification unit to notify. A sound output control device characterized by the above.
2. Application identification information for identifying the application that is the source of the agent information is attached to the agent information; The notification control unit adds the background sound corresponding to the application identification information attached to the agent information to be output among the plurality of applications to the voice message and causes the notification unit to notify. The sound output control device according to claim 1, characterized by the above.
3. The notification control unit controls the notification unit so that the voice message is notified in a voice of a tone color corresponding to the category to which the content of the agent information to be output belongs, with the background sound added to the voice message. The sound output control device according to claim 1 or 2, characterized by the above.
4. Category identification information for identifying the category to which the content of the agent information belongs is attached to the agent information; The notification control unit controls the notification unit so that the voice message is notified in a voice of a tone color corresponding to the category identification information attached to the agent to be output. The sound output control device according to claim 3, characterized by the above.
5. A sound output control method executed by a sound output control device, An information acquisition step of acquiring agent information to be output from an agent device including a plurality of applications each capable of outputting agent information, as information to be provided to a user; An announcement control step of causing a notification unit to output a voice message corresponding to the agent information to be output; comprising: In the announcement control step, background sound corresponding to the application that provided the agent information to be output is added to the voice message and the notification unit is caused to make an announcement. A sound output control method characterized by the above.
6. A sound output control program executed by a sound output control device including a computer, causing the computer to function as an information acquisition means for acquiring agent information to be output from an agent device including a plurality of applications each capable of outputting agent information, as information to be provided to a user; an announcement control means for causing a notification unit to output a voice message corresponding to the agent information to be output, and in the announcement control means, background sound corresponding to the application that provided the agent information to be output is added to the voice message and the notification unit is caused to make an announcement. A sound output control program characterized by the above.
Citation Information
Patent Citations
Information processor and control method therefor
JP1996339288A
Agent device
JP2000020888A
Information providing device for vehicle
JP2005238962A
Network system, information processing method and server
JP2019152969A
Control device, agent apparatus, and program
JP2020067785A