Sound output control device, sound output control method, and sound output control program

The sound output control device addresses the challenge of identifying audio content sources and categories by applying distinct sound effects and presenting a list of sound effects, improving user comprehension and satisfaction.

JP2025157585APending Publication Date: 2025-10-15PIONEER IP
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2025128333
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

When multiple applications with agent functions are used simultaneously, it becomes difficult for users to recognize the source and category of audio content being provided, leading to confusion and difficulty in identifying preferred content.

Method used

A sound output control device that includes an information acquisition unit to identify the source application, an alarm control unit to apply distinct sound effects based on the application, and a presentation control unit to provide a list of sound effects corresponding to each application, ensuring clear identification of the audio content source and category.

Benefits of technology

The solution enables users to easily distinguish between audio content sources and categories, enhancing user understanding and satisfaction by providing clear audio content differentiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157585000001_ABST
    Figure 2025157585000001_ABST
Patent Text Reader

Abstract

To support a user in appropriately grasping what kind of voice content is currently being output.SOLUTION: A sound output control device includes: an information acquisition section that acquires agent information provided from each of a plurality of applications having a voice agent function; a reporting control section that causes a reporting section to report a voice message to which different sound effects are imparted between the plurality of applications depending on an application, among the plurality of applications, that is a providing source of the agent information acquired by the information acquisition section; and a presentation control section that presents, to a user, an application list indicating the sound effects corresponding to each of the plurality of applications.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound output control device, a sound output control method, and a sound output control program. [Background technology]

[0002] Conventionally, there is known a technique in which an agent responds with a different tone of voice or a different tone of voice for each type of task, and a technique in which the agent's tone of voice is changed for each of a plurality of request processing devices. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 8-339288 [Patent Document 2] Japanese Patent Publication No. 2020-67785 Summary of the Invention [Problem to be solved by the invention]

[0004] However, with the above-mentioned conventional technology, for example, when multiple applications with agent functions are used simultaneously and audio content is provided from each agent, there is a problem in that it may be difficult for the user to properly recognize the source of the audio content currently being provided and the type of audio content currently being provided.

[0005] The present invention has been made in consideration of the above, and aims to provide, for example, a sound output control device, a sound output control method, and a sound output control program that can support a user in properly understanding what kind of sound content is currently being output. [Means for solving the problem]

[0006] The sound output control device described in claim 1 is characterized by comprising: an information acquisition unit that acquires agent information provided from each of a plurality of applications having a voice agent function; an alarm control unit that causes an alarm unit to issue a voice message with different sound effects applied to the plurality of applications depending on which of the plurality of applications is the source of the agent information acquired by the information acquisition unit; and a presentation control unit that causes a user to be presented with an application list showing the sound effects corresponding to each of the plurality of applications.

[0007] Furthermore, a sound output control method according to claim 8 is a sound output control method executed by a sound output control device, and is characterized by including: an information acquisition step of acquiring agent information provided from each of a plurality of applications having a voice agent function; an alarm control step of causing an alarm unit to announce a voice message to which different sound effects are applied among the plurality of applications depending on which application among the plurality of applications is the source of the agent information acquired in the information acquisition step; and a presentation control step of presenting to a user an application list indicating the sound effects corresponding to each of the plurality of applications.

[0008] The sound output control program according to claim 10 is a sound output control program executed by a sound output control device having a computer, and is characterized in that it causes the computer to function as: information acquisition means for acquiring agent information provided from each of a plurality of applications having a voice agent function; notification control means for causing a notification unit to issue a voice message to which different sound effects are applied among the plurality of applications depending on which of the plurality of applications is the source of the agent information acquired by the information acquisition means; and presentation control means for presenting to a user an application list showing the sound effects corresponding to each of the plurality of applications. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing an overview of information processing according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the sound output control device according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a category classification database according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a content buffer according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a sound output control method according to the first embodiment. [Figure 6] FIG. 6 is a flowchart showing an information processing procedure according to the first embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of the configuration of a sound output control device according to the second embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of an application classification database according to the second embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of a sound output control method according to the second embodiment. [Figure 10] FIG. 10 is a flowchart showing an information processing procedure according to the second embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of a sound output control device according to the third embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the third information processing. [Figure 13] FIG. 13 is a hardware configuration diagram showing an example of a computer that realizes the functions of the sound output control device. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, a mode for carrying out the present invention (hereinafter referred to as an embodiment) will be described with reference to the drawings. Note that the present invention is not limited to the embodiment described below. Furthermore, in the description of the drawings, the same parts are given the same reference numerals.

[0011] (Outline of the embodiment) 1. Introduction There are known applications (hereinafter abbreviated as "apps") that provide various types of content via a terminal device (navigation terminal) installed in a vehicle or a terminal device such as a smartphone owned by a user (e.g., a vehicle passenger). For example, there are apps that have an agent function to assist the user, and thereby assist driving according to the vehicle's driving state and the situation of the user driving the vehicle, or assist with route guidance according to various inputs (e.g., text input or voice input). There are also apps that assist in more comfortable driving by providing various types of content such as tourist guides, store guides, and other useful information as the vehicle travels.

[0012] Furthermore, many of these applications consider that the user to whom the output is sent is a passenger in the vehicle, and for safety reasons, they attempt to assist the user by using voice content provided by a voice agent function. In such cases, for example, multiple applications may be linked to the user while in the vehicle, and voice content of various categories may be provided, which may result in the following problems:

[0013] For example, a first problem is that when audio content of various categories is output, if a user is waiting for audio content of a desired category to be output, it becomes difficult for the user to determine whether the audio content currently being output is audio content of the desired category.

[0014] A second problem is that when audio content of various application types is output, for example, when a user is waiting for audio content provided by a specific application among multiple applications to be output, it becomes difficult for the user to determine whether the audio content currently being output is the type of audio content that the user desires.

[0015] A third issue is that when using multiple apps, it is difficult for users to distinguish which app is the source of audio content, making it difficult for users to identify apps that suit their preferences.

[0016] Therefore, an object of the present invention is to provide a sound output control device, a data structure, a sound output control method, and a sound output control program that can solve the above-mentioned problems. Below, three types of information processing (first information processing, second information processing, and third information processing) will be described in detail as information processing realized by the sound output control device, data structure, sound output control method, and sound output control program corresponding to the present invention. Specifically, the first information processing will be described as information processing according to the first embodiment, and the second information processing will be described as information processing according to the second embodiment. Furthermore, the third information processing will be described as information processing according to the third embodiment.

[0017] 2. Overview of Information Processing According to the Embodiment Before describing each embodiment, an overview of information processing according to the embodiment will be described using FIG. 1. FIG. 1 is a diagram showing an overview of information processing according to the embodiment. The information processing system 1 (an example of an information processing system according to the embodiment) shown in FIG. 1 is common to the first embodiment, the second embodiment, and the third embodiment. In the following description, when there is no need to distinguish between the first embodiment, the second embodiment, and the third embodiment, they may be simply referred to as "embodiments."

[0018] 1, the information processing system 1 according to the embodiment includes a terminal device 10, a situation grasping device 30, an agent device 60-x, and a sound output control device 100. These devices included in the information processing system 1 are connected to each other via a network so as to be able to communicate with each other via wired or wireless communication.

[0019] (About terminal devices) The terminal device 10 is an information processing terminal used by a user. The terminal device 10 may be, for example, a stationary navigation device installed in a vehicle, or may be a portable terminal device (for example, a smartphone, a tablet terminal, a notebook PC, a desktop PC, a PDA, etc.) owned by a user. In this embodiment, the terminal device 10 is assumed to be a navigation device installed in a vehicle.

[0020] 1, it is assumed that multiple apps are linked to the terminal device 10. For example, multiple apps may be linked to the terminal device 10 because the user has installed them, or the information processing system 1 may include an app that can send push notifications of various contents regardless of whether the app is installed.

[0021] The terminal device 10 also has a notification unit (output unit), and audio content provided by each app is output from this notification unit. For this reason, the notification unit may be, for example, a speaker. The user here may be a passenger (for example, a driver) of a vehicle in which the terminal device 10 is installed. That is, in the example of FIG. 1, the terminal device 10 is installed in a vehicle VE1, and the user of the terminal device 10 is a user U1 who is driving the vehicle VE1.

[0022] (About agent devices) The agent device 60-x exists for each application linked to the terminal device 10, and may be an information processing device that realizes the functions and roles of the application. In Fig. 1, an example is shown in which the agent device 60-x is a server device, but it may also be realized by, for example, a cloud system.

[0023] The application associated with the terminal device 10 may be an application that assists a user using voice content, and the agent device 60-x has a voice agent function corresponding to this assist function. Therefore, the application associated with the terminal device 10 can be called a voice agent application.

[0024] In the following embodiments, when distinguishing between the agent device 60-x corresponding to the application APx and the processing units (e.g., the application control function 631-x and the agent information generation unit 632-x) possessed by the agent device 60-x, any value will be used for "x."

[0025] For example, Fig. 1 shows agent device 60-1 as an example of agent device 60-x. Agent device 60-1 is an agent device corresponding to application AP1 (hazard detection application), and provides audio content that detects hazards while driving and provides cautions and warnings to the user. Also, according to the example of Fig. 1, agent device 60-x has application control function 631-1 and agent information generation unit 632-1.

[0026] 1 also shows agent device 60-2 as another example of agent device 60-x. Agent device 60-2 is an agent device corresponding to application AP2 (navigation application) and provides audio content related to route guidance to the user. Although not shown in FIG. 1, agent device 60-2 may have application control function 631-2 and agent information generation unit 632-2.

[0027] The app linked to the terminal device 10 is not limited to the above example, and may be, for example, audio content related to tourist information, audio content related to store information, or other apps that provide various types of useful information. The expression "provided by the app" also includes the concept of "provided by the agent device 60-x" corresponding to this app.

[0028] Next, the functions of the agent device 60-x will be described. According to the example of Fig. 1, the agent device 60-x has an application control function 631-x and an agent information generation unit 632-x.

[0029] The application control function 631-x executes various controls related to the application APx. For example, the application control function 631-x personalizes the content provided to each user based on the user's usage history. The application control function 631-x also performs processing to determine the content of a voice message to respond to based on the content of the utterance indicated by the voice input by the user, so as to realize a dialogue with the user. The application control function 631-x can also determine the content to be provided to the user and the content of the voice message to respond to the user based on the user's situation.

[0030] The agent information generation unit 632-x performs a generation process to generate audio content (an example of agent information). For example, the agent information generation unit 632-x determines to which category audio content should be output based on data received from the situation awareness device 30 (described later), and generates audio content with content belonging to the determined category. For example, the agent information generation unit 632-x generates message information with content corresponding to the vehicle's running state and the situation of the user driving the vehicle as understood by the situation awareness device 30. Note that the message information is, for example, text data that serves as the basis for audio content to be ultimately notified to the user U1, and defines the content of the audio that is later converted into audio data. In other words, the agent information generation unit 632-x is not limited to generating audio data as audio content, and may generate data in other formats that serve as the basis for audio messages as audio content. Furthermore, the agent information generation unit 632-x determines category identification information (category ID) that identifies the category to which the message information belongs based on the content indicated by the generated message information. Note that the category to which the message information belongs can also be said to be the category to which the audio content including the message information belongs.

[0031] Furthermore, the agent information generating unit 632-x assigns application identification information (application ID) for identifying the application APx corresponding to the agent device 60-x to the audio content.

[0032] Taking agent device 60-1 as an example, agent information generation unit 632-1 assigns to the audio content an application ID (e.g., "AP1") that identifies application AP1 corresponding to agent device 60-1. Furthermore, if the generated audio content belongs to the category "Entertainment," agent information generation unit 632-1 assigns to the audio content a category ID that indicates the category "Entertainment." That is, agent information generation unit 632-1 generates agent information as audio content by adding to the generated message information an application ID that identifies the generator of the message information and a category ID that identifies the content of the message information. In other words, the agent information as audio content generated by agent information generation unit 632-1 is assigned application identification information that identifies the application that provided the agent information and category identification information that identifies the category to which the agent information belongs.

[0033] Furthermore, the agent information generating unit 632-x is not limited to the above example, and may generate message information that responds to the input voice when, for example, the user inputs voice by speaking via the terminal device 10. This enables the agent device 60-x to generate voice content that can realize a dialogue with the user.

[0034] Furthermore, the agent information generation unit 632-x can also specify the timing for outputting audio content. For example, the agent information generation unit 632-x can generate tolerance range information that specifies the time range or geographic range in which output of audio content is permitted, using a time range, a vehicle travel distance range, a vehicle passing area, a vehicle speed, etc. In such a case, the agent information generation unit 632-x transmits the audio content and the tolerance range information to the sound output control device SV, thereby requesting (reserving) the sound output control device SV to output the audio content to the terminal device 10 of a vehicle that meets the conditions indicated in the tolerance range information. The timing specification and request may be performed by a processing unit other than the agent information generation unit 632-x.

[0035] (Regarding situational awareness devices) The situation awareness device 30 performs an analysis process to understand the vehicle's driving state and the situation of the user driving the vehicle. While FIG. 1 illustrates an example in which the situation awareness device 30 is a server device, it may also be realized by, for example, a cloud system. According to the example in FIG. 1, the analysis process is performed by a situation awareness engine E30 mounted on the situation awareness device 30. For example, the situation awareness engine E30 senses the driving state and the user's situation based on sensor information obtained from various sensors. The sensors referred to here may be, for example, sensors provided in the vehicle or sensors included in the terminal device 10, such as an acceleration sensor, a gyro sensor, a magnetic sensor, a GPS, a camera, a microphone, etc.

[0036] For example, the situation awareness engine E30 can perform the following series of analysis processes. For example, the situation awareness engine E30 performs sensing based on sensor information acquired from the above sensors and performs base analysis by using the sensing results as core elements. In the base analysis, the situation awareness engine E30 extracts necessary data using the core elements as information sources and converts and processes the extracted data. Next, the situation awareness engine E30 performs high-level analysis using the converted and processed data. In the high-level analysis, the situation awareness engine E30 analyzes specific situations based on the converted and processed data. For example, the situation awareness engine E30 uses the converted and processed data to understand various situations, such as the status of impacts on the vehicle, the status of vehicle lighting, changes in driving conditions, and the user's own situation. In addition, the situation awareness engine E30 can predict the user's behavior (e.g., predicting stopovers) as part of the situation awareness.

[0037] (Sound output control device) The sound output control device SV performs information processing according to the embodiments. Specifically, the sound output control device SV performs information processing according to a first embodiment (first information processing), information processing according to a second embodiment (second information processing), and information processing according to a third embodiment (third information processing), which will be described later, as information processing according to the embodiments. Furthermore, the information processing according to the embodiments is processing related to notification control that causes a notification unit included in the terminal device 10 to output a voice message. While FIG. 1 shows an example in which the sound output control device SV is a server device, it may also be realized by, for example, a cloud system.

[0038] 1, each information processing according to the embodiment is performed by an information matching engine ESV mounted on the sound output control device SV. As shown in FIG. 1, the information matching engine ESV includes functions such as a request manager function ESV1 and a response manager function ESV2.

[0039] The request manager function ESV1 receives requests from the agent device 60-x and performs queuing according to the received requests. Note that the request here may be an output request requesting that the generated audio content be output to the user, and is transmitted, for example, including the audio content. The request manager function ESV1 also queues the received audio content in the content buffer 122 (FIG. 4).

[0040] The response manager function ESV2 determines the order of priority for outputting the reserved audio contents based on data relating to the situation grasped by the situation grasping device 30 (for example, data indicating the results of the analysis process) and the allowable range information included in the request. Then, the response manager function ESV2 controls the output of the terminal device 10 so that each audio content is output in the determined order of priority. Note that the output control of the terminal device 10 includes the concept of output control of the notification unit of the terminal device 10.

[0041] 3. Information Processing Flow According to the Embodiment Up to this point, the description has focused on each device included in the information processing system 1. Next, a description will be given of the overall flow of information processing according to the embodiment that is performed within the information processing system 1. Here, a situation is assumed in which audio content is output to a user U1 who is driving a vehicle VE1 via a terminal device 10 installed in the vehicle VE1.

[0042] In such a situation, the terminal device 10 transmits sensor information detected by a sensor included in the terminal device 10 to the situation grasping device 30 as needed (step S11).

[0043] When the situation assessment engine E30 of the situation assessment device 30 acquires the sensor information transmitted from the terminal device 10, it performs an analysis process to assess various situations, including the traveling state of the vehicle VE1 and the state of the user U1 who is driving the vehicle VE1 (step S12). For example, the situation assessment engine E30 assesses the situation in detail by performing a series of analysis processes, such as sensing using the sensor information, base analysis using the sensing results as core elements, and higher-level analysis using data obtained as a result of the base analysis.

[0044] Furthermore, when the analysis process is completed, the situation assessment device 30 transmits data relating to the situation assessed by the situation assessment engine E30 (for example, data indicating the results of the analysis process) to the agent device 60-x (step S13). In the example of Fig. 1, the situation assessment device 30 transmits data relating to the situation to each of the agent devices 60-x, such as the agent device 60-1, the agent device 60-2, etc.

[0045] When the agent information generation unit 632-x of the agent device 60-x acquires data relating to the situation from the situation grasping device 30, it performs a generation process to generate audio content to be output based on the acquired data (step S14). For example, the agent information generation unit 632-x determines, based on the acquired data, to which category of audio content that the agent device can support should be output, and generates audio content whose content belongs to the determined category. For example, the agent information generation unit 632-x generates message information (text data) whose content corresponds to the situation indicated by the acquired data.

[0046] In addition, the agent information generation unit 632-x sends the generated audio content to the sound output control device SV with a category ID that identifies the category to which the audio content belongs (the category to which the message information belongs) and an application ID that identifies the application APx corresponding to the device (step S15).

[0047] In the example of FIG. 1, the generation process of step S14 is performed by each agent device 60-x, such as agent device 60-1, agent device 60-2, etc., and in step S15, each agent device 60-x transmits its own audio content to the sound output control device SV.

[0048] Next, when the information matching engine ESV of the sound output control device SV acquires the audio content to be output, it performs notification control processing on the audio content to be output (step S16). For example, when converting message information included in the audio content to be output into audio data (audio message), the information matching engine ESV converts the message information while changing the audio manner according to the category to which the audio content belongs, and controls notification so that the converted audio data is notified. Furthermore, for example, the information matching engine ESV controls notification so that sound effects (e.g., background sounds) according to the application type to which the audio content to be output belongs are added to the converted audio data (audio message) and notified. Such notification control processing will be described in detail in the first and second embodiments described later.

[0049] Finally, the sound output control device SV performs sound output control on the terminal device 10 in accordance with the notification control by the information matching engine ESV (step S17). Specifically, the sound output control device SV controls the terminal device 10 so that the audio data notification-controlled by the information matching engine ESV is output from the notification unit of the terminal device 10.

[0050] (First embodiment) 1. Overview of the First Embodiment From here, a first embodiment will be described. Information processing according to the first embodiment (i.e., first information processing) is performed with the aim of solving the first problem described above. Specifically, the first information processing is performed by a sound output control device 100 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 100 performs the first information processing in accordance with a sound output control program according to the first embodiment. In addition, the sound output control device 100 has a structure consisting of a category classification database 121 (FIG. 3) and a content buffer 122 (FIG. 4).

[0051] 2. Configuration of the Sound Output Control Device According to the First Embodiment Next, the sound output control device 100 according to the first embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the sound output control device 100 according to the first embodiment. As shown in Fig. 2, the sound output control device 100 has a communication unit 110, a storage unit 120, and a control unit 130.

[0052] (Regarding the communication unit 110) The communication unit 110 is realized by, for example, a NIC etc. The communication unit 110 is connected to a network by wire or wirelessly, and transmits and receives information to and from, for example, the terminal device 10, the situation grasping device 30, and the agent device 60-x.

[0053] (Regarding the storage unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk, an optical disk, etc. The storage unit 120 has a category classification database 121 and a content buffer 122.

[0054] (About Category Classification Database 121) The category classification database 121 stores information about the category to which the audio content (agent information) provided by the application APx belongs. An example of the category classification database 121 according to the first embodiment is shown in Fig. 3. In the example of Fig. 3, the category classification database 121 has items such as "category ID", "category", and "timbre feature".

[0055] The "category ID" indicates identification information for identifying a candidate "category" used to identify to which category the audio content to be output provided from the application APx side belongs.

[0056] "Category" is a candidate "category" used to identify to which category the audio content to be output provided by the application belongs. In the example of Figure 3, the candidate "categories" include "Caution," "Warning," "Entertainment," "Advertisement," "Guide," and "News." Note that even if the audio content to be output is provided by different applications, there are situations in which the audio content to be output itself belongs to the same category. For example, even if different applications, such as application AP1 and application AP5, provide audio content to be output, both of the audio content to be output may belong to the category "Entertainment."

[0057] "Tone characteristics" indicates candidate tone parameters used in the notification control process that changes the tone of the audio when outputting audio corresponding to the audio content from the notification unit of the terminal device 10, depending on the category to which the audio content to be output belongs.

[0058] In the example of Fig. 3, the timbre feature "male voice + slow" is associated with the category "Caution" identified by the category ID "CT1". This example shows an example in which it is specified that when the audio content to be output belongs to the category "Caution", the audio message to be output from the notification unit of the terminal device 10 should be changed to a timbre feature of "male voice + slow". Therefore, the timbre parameter in this example corresponds to the parameter indicating "male voice + slow".

[0059] 3, the timbre feature "female voice + fast speaking" is associated with the category "entertainment" identified by the category ID "CT3." This example shows an example in which it is specified that when the audio content to be output belongs to the category "entertainment," the audio message to be output from the notification unit of the terminal device 10 should be changed to a timbre feature of "female voice + fast speaking." Therefore, the timbre parameter in this example corresponds to a parameter indicating "female voice + fast speaking."

[0060] (Content Buffer 122) The content buffer 122 functions as a storage area for queuing information related to audio content transmitted from the agent device 60-x. An example of data stored in the content buffer 122 according to the first embodiment is shown in Fig. 4. In the example of Fig. 4, the content buffer 122 has items such as "destination user ID," "application ID," "category ID," and "audio content."

[0061] The "destination user ID" indicates identification information for identifying the destination user (or the terminal device 10 of the user) to whom the "audio content" is to be output (notified). The "application ID" indicates identification information for identifying the application that provided the "audio content" to be output (or the agent device 60-x corresponding to the application). The "providing application" can be rephrased as the application that generated the "audio content" to be output.

[0062] The "category ID" indicates identification information for identifying the category to which the "audio content" to be output, provided by the application identified by the "application ID", belongs. The "category ID" is assigned to the "audio content" to be output by the agent device 60-x corresponding to the application identified by the "application ID".

[0063] “Audio content” is information about “audio content” to be output, provided by an application identified by an “application ID.” “Audio content” includes, for example, text data as message information.

[0064] That is, in the example of Figure 4, the content of message information #11 provided by an application (application AP1) identified by application ID "AP1" belongs to a category identified by category ID "CT3" and is to be output to a user (user U1) identified by user ID "U1".

[0065] (Regarding the control unit 130) 2, the control unit 130 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs (e.g., a sound output control program) stored in a storage device inside the sound output control device 100 using RAM as a work area. The control unit 130 is also realized by an integrated circuit, such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0066] As shown in Fig. 2, the control unit 130 is equipped with an information matching engine E100. The information matching engine E100 corresponds to the information matching engine ESV described in Fig. 1. The information matching engine E100 includes a request manager function E101 (corresponding to the request manager function ESV1 described in Fig. 1) and a response manager function E102 (corresponding to the response manager function ESV2 described in Fig. 1).

[0067] 2, the request manager function E101 has a request receiving unit 131 and a queuing unit 132. The response manager function E102 has an information acquisition unit 133, a determination unit 134, and a notification control unit 135.

[0068] (Regarding the request receiving unit 131) The request receiving unit 131 receives a request from the agent device 60-x. Specifically, the request receiving unit 131 receives a request from the agent device 60-x to output audio content to a user. The request may include the audio content to be output, a user ID that identifies the user to whom the audio content is to be provided, and tolerance information that conditions the period and timing during which the audio content is allowed to be output.

[0069] (Regarding the queuing unit 132) The queuing unit 132 queues audio content to be output in response to a request accepted by the request accepting unit 131. For example, the queuing unit 132 queues the audio content associated with the request in the content buffer 122.

[0070] (Regarding the information acquisition unit 133) The information acquiring unit 133 acquires audio content (agent information) to be output from the agent device 60-x capable of outputting audio content (agent information) belonging to each of a plurality of different categories as information to be provided to the user. Specifically, the information acquiring unit 133 acquires audio content belonging to a plurality of different categories from each of the agent devices 60-x capable of outputting agent information provided by each of a plurality of applications and corresponding to the agent functions of the applications, and passes the audio content to the queuing unit 132.

[0071] (Regarding the decision unit 134) A determination part 134 determines the priority order of the audio contents reserved for output based on the allowable range information included in the request accepted by the request acceptance part 131, as to the order in which the audio contents are actually output, and reads out the audio contents from the content buffer 122 when it is time to output them.

[0072] (Regarding the notification control unit 135) The notification control unit 135 converts message information included in the audio content (agent information) transmitted from the agent device 60-x into audio data. For example, the notification control unit 135 converts the message information into audio data by using Text to Speech (TTS) technology to synthesize audio based on text data for the audio content determined to be output by the determination unit 134. Then, the notification control unit 135 causes the audio data (audio message) obtained by converting the message information to be output from the notification unit of the terminal device 10.

[0073] Furthermore, the notification control unit 135 changes the mode of the voice message depending on the category to which the voice content to be output belongs, and causes the notification unit to notify the voice message. For example, the notification control unit 135 changes the mode of the voice data to be converted by changing the tone parameter of voice synthesis depending on the category to which the voice content to be output belongs, regardless of which application out of multiple applications the voice content to be output is provided by.

[0074] As explained above, each sound content acquired from the agent device 60-x is assigned category identification information (category ID) that identifies the category to which the sound content belongs. Therefore, the notification control unit 135 changes the mode of the sound message depending on the category indicated by the category identification information assigned to the sound content to be output.

[0075] For example, the notification control unit 135 changes the tone of the voice message depending on the category to which the voice content to be output belongs, among a plurality of different categories.

[0076] The notification control unit 135 may also cause the notification unit to notify the audio data with sound effects added according to the category to which the audio content to be output belongs, among a plurality of different categories. Note that the sound effects referred to here refer to sound effects added to the beginning or end of an audio message, and background sounds such as background music (BGM) superimposed on an audio message.

[0077] On the other hand, different sound effects may be set among the multiple applications. In such a case, the notification control unit 135 may cause the notification unit to notify a voice message with a sound effect added that corresponds to the application that provided the audio content to be output, among the different sound effects among the multiple applications. Specifically, each audio content acquired from the agent device 60-x is assigned application identification information (application ID) that identifies the application that provided the audio content. Therefore, the notification control unit 135 causes the notification unit to notify a voice message with a sound effect added that corresponds to the application, among the multiple applications, that is identified by the application identification information assigned to the audio content to be output. This point will be described in detail in the second embodiment.

[0078] [3. Specific examples of sound output control methods] Next, a specific example of the sound output control method performed in the first information processing will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the sound output control method according to the first embodiment.

[0079] FIG. 5 illustrates five applications, namely, application AP1, application AP2, application AP3, application AP4, and application AP5 (applications AP1 to AP5), as the plurality of applications. Also, FIG. 5 illustrates agent device 60-x capable of causing a notification unit to output audio content provided by each of applications AP1 to AP5: agent device 60-1, agent device 60-2, agent device 60-3, agent device 60-4, and agent device 60-5. Specifically, according to the example of FIG. 5, agent device 60-1 is an agent device capable of outputting audio content corresponding to application AP1 to be provided to the user. Furthermore, agent device 60-2 is an agent device capable of outputting audio content corresponding to application AP2 to be provided to the user. Furthermore, agent device 60-3 is an agent device capable of outputting audio content corresponding to application AP3 to be provided to the user. Furthermore, agent device 60-4 is an agent device capable of outputting audio content corresponding to application AP4 to be provided to the user. Furthermore, the agent device 60-5 is an agent device that can output audio content corresponding to the application AP5 so that the audio content is provided to the user.

[0080] Also, according to the example of FIG. 5, the agent device 60-1 is a device capable of outputting audio content belonging to the category "Entertainment" and audio content belonging to the category "Advertisement". The agent device 60-2 is a device capable of outputting audio content belonging to the category "Caution" and audio content belonging to the category "Warning". The agent device 60-3 is a device capable of outputting audio content belonging to the category "Guidance". The agent device 60-4 is a device capable of outputting audio content belonging to the category "News" and audio content belonging to the category "Advertisement". The agent device 60-5 is a device capable of outputting audio content belonging to the category "Caution" and audio content belonging to the category "Entertainment".

[0081] For example, it is assumed here that the agent information generation unit 632-1 of the agent device 60-1 generates audio content A-1 corresponding to message information of which contents belong to the category "Entertainment" based on data relating to the situation acquired from the situation grasping device 30. In this case, the agent device 60-1 assigns a category ID "CT3" that identifies the category "Entertainment" to the audio content A-1 and transmits the same to the sound output control device 100, as shown in Fig. 5, so that the audio content A-1 is output to the user U1. At this time, the agent device 60-1 may further assign an application ID "AP1" that identifies application AP1, which is a source application that provides the audio content A-1.

[0082] The information acquisition unit 133 of the sound output control device 100 acquires the audio content A-1 assigned the category ID "CT3" from the agent device 60-1 as the audio content to be output. Subsequently, when the audio content A-1 is determined as the audio content to be output from the notification unit of the terminal device 10 through processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined audio content A-1 to be output using tone parameters according to the category to which the audio content A-1 belongs.

[0083] 5, the notification control unit 135 can identify that the category to which the audio content A-1 belongs is "Entertainment" by comparing the category ID "CT3" assigned to the audio content A-1 with the category classification database 121. Furthermore, the notification control unit 135 refers to the category classification database 121 and recognizes that it is specified that the audio message to be output from the notification unit of the terminal device 10 should be changed to a timbre characteristic of "female voice + fast speaking." Then, the notification control unit 135 performs voice synthesis on the audio data included in the audio content A-1 using parameters indicating "female voice + fast speaking," thereby changing the timbre of the voice.

[0084] Next, the notification control unit 135 performs sound output control so that a voice message after voice synthesis of the audio content A-1 is announced by the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output a voice message corresponding to the audio content A-1 by transmitting the voice message after voice synthesis of the audio content A-1 to the terminal device 10. The terminal device 10 announces the voice message from the notification unit in accordance with the sound output control from the notification control unit 135. This allows the user U1 to easily understand that the audio content currently being output belongs to the category "Entertainment."

[0085] The notification control unit 135 may not only change the tone of the voice message corresponding to the voice content A-1 in accordance with the fact that the voice content A-1 belongs to the category "Entertainment", but may also add a sound effect (for example, background sound such as sound effects or BGM) to the voice message in accordance with the fact that the voice content A-1 belongs to the category "Entertainment". In this case, for example, in the category classification database 121 shown in Fig. 3, data of sound effects corresponding to the category indicated by each "category ID" may be associated with the category ID (not shown).

[0086] Next, using the examples of audio content A-2 and D-3, we will explain how the format of the audio data can be changed depending on the category to which the audio content to be output belongs, regardless of which of multiple applications the audio content to be output is provided by.

[0087] For example, it is assumed that the agent information generating unit 632-1 of the agent device 60-1 generates audio content A-2 corresponding to message information of which contents belong to the category "advertisement" based on data relating to the situation acquired from the situation grasping device 30. In this case, the agent device 60-1 transmits the audio content A-2 to the sound output control device 100 with the category ID "CT4" that identifies the category "advertisement" assigned to the audio content A-2, as shown in Fig. 5, so that the audio content A-2 is output to the user U1. At this time, the agent device 60-1 may further assign an application ID "AP1" that identifies the application AP1 that is the application that provides the audio content A-2.

[0088] The information acquisition unit 133 acquires the audio content A-2 assigned the category ID "CT4" from the agent device 60-1 as the audio content to be output. Subsequently, when the audio content A-2 is determined as the audio content to be output from the notification unit of the terminal device 10 through processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined audio content A-2 to be output using tone parameters according to the category to which the audio content A-2 belongs.

[0089] 5, the notification control unit 135 can identify that the category to which the audio content A-2 belongs is "advertising" by comparing the category ID "CT4" assigned to the audio content A-2 with the category classification database 121. Furthermore, the notification control unit 135 refers to the category classification database 121 and recognizes that it is specified that the audio message to be output from the notification unit of the terminal device 10 should be changed to a tone characteristic of "robot voice + slow." Then, the notification control unit 135 performs voice synthesis on the audio data included in the audio content A-2 using a parameter indicating "robot voice + slow," thereby changing the tone of the voice.

[0090] Next, the notification control unit 135 performs sound output control so that the voice message after voice synthesis of the audio content A-2 is announced by the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output the voice message corresponding to the audio content A-2 by transmitting the voice message after voice synthesis of the audio content A-2. The terminal device 10 announces the voice message from the notification unit in accordance with the sound output control from the notification control unit 135. This allows the user U1 to easily understand that the audio content currently being output belongs to the category "advertisement."

[0091] Also, for example, it is assumed that the agent information generation unit 631-4 of the agent device 60-4 generates audio content D-3 corresponding to message information of which contents belong to the category "advertisement" based on data relating to the situation acquired from the situation grasping device 30. In this case, the agent device 60-4 assigns a category ID "CT4" that identifies the category "advertisement" to the audio content D-3 and transmits the same to the sound output control device 100, as shown in Fig. 5, so that the audio content D-3 is output to the user U1. At this time, the agent device 60-4 may further assign an application ID "AP4" that identifies application AP4, which is a provider application that provides the audio content D-3.

[0092] The information acquisition unit 133 acquires the audio content D-3 assigned the category ID "CT4" from the agent device 60-4 as the audio content to be output. Subsequently, when the audio content D-3 is determined as the audio content to be output from the notification unit of the terminal device 10 through processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs notification control processing such as performing voice synthesis on the message information included in the determined audio content D-3 to be output using tone parameters according to the category to which the audio content D-3 belongs.

[0093] 5, the notification control unit 135 can identify that the category to which the audio content D-3 belongs is "advertising" by comparing the category ID "CT4" assigned to the audio content D-3 with the category classification database 121. Furthermore, the notification control unit 135 refers to the category classification database 121 and recognizes that it is specified that the audio message to be output from the notification unit of the terminal device 10 should be changed to a tone characteristic of "robot voice + slow." Then, the notification control unit 135 performs voice synthesis on the audio data included in the audio content D-3 using a parameter indicating "robot voice + slow," thereby changing the tone of the voice.

[0094] Next, the notification control unit 135 performs sound output control so that a voice message after voice synthesis of the audio content D-3 is announced by the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output a voice message corresponding to the audio content D-3 by transmitting the voice message after voice synthesis of the audio content D-3 to the terminal device 10. The terminal device 10 announces the voice message from the notification unit in response to the sound output control from the notification control unit 135. This allows the user U1 to easily understand that the audio content currently being output belongs to the category "advertisement."

[0095] In the above two examples, the apps that provided the audio content are different, such as app AP1 and app AP4. However, since the audio content provided by both apps belongs to the same category (advertising), it is output in the same style (robot voice + slow voice) regardless of the type of app.

[0096] Furthermore, regardless of which of multiple applications the audio content to be output is provided by, the audio content may be given sound effects (e.g., sound effects, background sounds such as background music, etc.) according to the category to which the audio content to be output belongs.

[0097] So far, a specific example of the sound output control method performed in the first information processing has been described using some of the audio content shown in Fig. 5 as an example. The other audio content shown in Fig. 5 can also be explained following the example of some of the audio content, so detailed explanations will be omitted.

[0098] [4. Processing Procedure] Next, the procedure of information processing according to the first embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the procedure of information processing according to the first embodiment. Note that the flow shown in the flowchart of Fig. 6 is repeatedly executed, for example, while the user U1 is driving the vehicle VE1.

[0099] First, the control unit 130 of the sound output control device 100 determines whether or not agent information has been acquired from the agent device 60-x (step S101). If it is determined that new agent information has been acquired (step S101; Yes), the control unit 130 performs a queuing process for the acquired agent information (step S102). In step S102, the newly acquired agent information is queued together with already acquired agent information, and a priority order for outputting the agent information as a voice message is determined, and the process proceeds to step S103. On the other hand, if it is determined in step S101 that new agent information has not been acquired from the agent device 60-x (step S101; No), the process proceeds directly to step S103.

[0100] Next, the control unit 130 determines whether or not there is any agent information that should be output among the agent information (audio content) acquired from the agent device 60-x (step S103). If the control unit 130 determines that there is no agent information that should be output (step S103; No), it ends the flow once and then repeats the flow from the beginning.

[0101] On the other hand, when the control unit 130 determines that there is agent information that should be output (step S103; Yes), it identifies the category to which this agent information belongs based on the category ID assigned to the agent information to be output (step S104). For example, the control unit 130 collates the category ID assigned to the agent information to be output with the category classification database 121 to identify the category to which the agent information to be output belongs.

[0102] Furthermore, the control unit 130 identifies the timbre features (timbre parameters) that correspond to the identified category, among the timbre features that are set to differ between categories, as in the example of the category classification database 121 shown in FIG. 3 (step S105).

[0103] Then, the control unit 130 changes the voice synthesis parameters used when converting the message information included in the agent information to be output into voice data to the identified tone parameters (tone parameters corresponding to the category to which the agent information to be output belongs) and performs voice conversion (step S106).

[0104] Finally, the control unit 130 controls sound output so that the voice data corresponding to the agent information to be output is notified by the notifying unit of the terminal device 10 of the user designated as the recipient of the agent information (step S107). Thereafter, the control unit 130 repeats the flow from the beginning.

[0105] 6, the procedure is described in which, after identifying agent information whose output timing has come, the message information included in the identified agent information is converted into voice data, but the timing for converting message information into voice data is not limited to this. For example, as soon as the information acquisition unit 133 acquires new agent information, the conversion process into voice data corresponding to steps S104 to S106 may be performed, and then the process of determining the output priority and the output timing corresponding to steps S102 and S103 may be performed on the voice content including the converted voice data.

[0106] [5. Summary] The sound output control device 100 according to the first embodiment acquires agent information to be output from an agent device capable of outputting agent information belonging to a plurality of different categories as information to be provided to a user. The sound output control device 100 then causes a notification unit to output a voice message corresponding to the agent information to be output. Specifically, the sound output control device 100 changes the manner of the voice message depending on the category to which the agent information to be output belongs, and causes the notification unit to notify the voice message. Such a sound output control device 100 allows a user to easily know whether the currently output audio content belongs to a desired category.

[0107] (Second embodiment) 1. Overview of the Second Embodiment Next, a second embodiment will be described. Information processing according to the second embodiment (i.e., second information processing) is performed with the aim of solving the second problem described above. Specifically, the second information processing is performed by a sound output control device 200 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 200 performs the second information processing in accordance with a sound output control program according to the second embodiment. Furthermore, the sound output control device 200 has a structure including an application classification database 233 in addition to a category classification database 121 (FIG. 3) and a content buffer 122 (FIG. 4).

[0108] 2. Configuration of the Sound Output Control Device According to the Second Embodiment Next, a sound output control device 200 according to a second embodiment will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the configuration of the sound output control device 200 according to the second embodiment. As shown in Fig. 7, the sound output control device 200 has a communication unit 110, a storage unit 220, and a control unit 130. In the following description, descriptions of processing units that are assigned the same reference numerals as those in the sound output control device 100 may be omitted or simplified.

[0109] (Regarding the storage unit 220) The storage unit 220 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 220 has a category classification database 121, a content buffer 122, and an application classification database 233.

[0110] (About App Classification Database 233) The application classification database 233 stores information related to sound effects. An example of the application classification database 233 according to the second embodiment is shown in Fig. 8. In the example of Fig. 8, the application classification database 233 has items such as "application ID," "application type," and "sound effect."

[0111] The "application ID" indicates identification information for identifying the application that provided the "audio content" to be output (or the agent device 60-x corresponding to the application). The application that provided the "audio content" to be output can be rephrased as the application that generated the "audio content" to be output. The "application type" is information about the type of application identified by the "application ID", and may be, for example, the name of the application. The "application type" corresponds to the type to which the audio content (agent information) to be output, provided by the application identified by the "application ID", belongs.

[0112] "Sound effects" are candidates for background sounds to be superimposed on the audio content to be output depending on the application of the provider that provided the audio content to be output. For example, the background sounds may be sound effects or music.

[0113] 8, a sound effect "Sound Effect #1" is associated with the application ID "AP1." This example shows a case where, when the application that provided the audio content to be output is the application AP1, the sound effect #1 is superimposed as background sound on the audio data (audio message) included in the audio content to be output.

[0114] (Regarding the information acquisition unit 133) The information acquisition unit 133 acquires, as information to be provided to the user, audio content to be output from an agent device capable of outputting multiple types of audio content that are differentiated by the content of the information or the source of the information.

[0115] For example, the information acquisition unit 133 acquires the audio content to be output from an agent device 60-x that is capable of outputting audio content provided by each of a plurality of applications and corresponds to the agent function possessed by the application.

[0116] (Regarding the notification control unit 135) The notification control unit 135 causes the notification unit to output audio data corresponding to the audio content to be output.

[0117] Furthermore, the notification control unit 135 adds background sound corresponding to the type to which the audio content to be output belongs to the audio data and causes the notification unit to notify the audio data.

[0118] For example, the notification control unit 135 adds background sound corresponding to the application of the provider that provided the audio content to be output, as the type to which the audio content to be output belongs, to the audio data, and causes the notification unit to notify the audio data. In such a case, each audio content acquired from the agent device 60-x is assigned application identification information that identifies the application of the provider that provided the audio content. Therefore, the notification control unit 135 adds background sound corresponding to the application identified by the application identification information assigned to the audio content to be output, among multiple applications, to the audio data, and causes the notification unit to notify the audio data.

[0119] Furthermore, the multiple types of audio content may include audio content belonging to multiple different categories distinguished based on the content of the audio content. In this case, the information acquisition unit 133 acquires audio content to be output from the agent device 60-x among the audio content belonging to multiple different categories. Then, the notification control unit 135 adds background sounds corresponding to the category to which the audio content to be output belongs, among background sounds that differ among the multiple different categories, to a voice message and causes the notification unit to notify the audio message. As a specific example, the information acquisition unit 133 acquires the audio content to be output from an agent device 60-x that is capable of outputting audio content provided by each of multiple applications and corresponds to the agent function of the application. Then, the notification control unit 135 adds background sounds corresponding to the category to which the audio content to be output belongs to the audio data and causes the notification unit to notify the audio message, regardless of which application provides the agent information for the audio content to be output.

[0120] Furthermore, as described in the first embodiment, the notification control unit 135 may control the notification unit to notify the voice message in a tone corresponding to the category to which the voice content to be output belongs, among a plurality of different categories. In this case, each voice content acquired from the agent device 60-x is assigned category identification information that identifies the category to which the voice content belongs. Therefore, the notification control unit 135 controls the notification unit to notify the voice data in a tone corresponding to the category indicated by the category identification information assigned to the voice content to be output.

[0121] [3. Specific examples of sound output control methods] Next, a specific example of the sound output control method performed in the second information processing will be described with reference to Fig. 9. Fig. 9 is a diagram showing an example of the sound output control method according to the second embodiment.

[0122] Many parts of Fig. 9 correspond to the example of Fig. 5. Specifically, Fig. 9 shows five applications, namely, application AP1, application AP2, application AP3, application AP4, and application AP5 (applications AP1 to AP5), as multiple applications. Fig. 5 also shows agent device 60-1, agent device 60-2, agent device 60-3, agent device 60-4, and agent device 60-5 as agent devices 60-x that can cause a notification unit to output audio content provided from each of applications AP1 to AP5. Explanation of each agent device 60-x will be omitted.

[0123] For example, assume that the agent information generation unit 632-1 of the agent device 60-1 generates audio content A-1 using audio data corresponding to message information of a content belonging to the category "Entertainment" based on data related to the situation acquired from the situation grasping device 30. In this case, the agent device 60-1 assigns a category ID "CT3" that identifies the category "Entertainment" to the audio content A-1, as shown in FIG. 9 , so that the audio content A-1 is output to the user U1. The agent device 60-1 also assigns an application ID "AP1" that identifies an application AP1, which is a provider application that provides the audio content A-1, to the audio content A-1. Then, the agent device 60-1 transmits the audio content A-1, to which the category ID and application ID have been assigned, to the sound output control device 200.

[0124] The information acquisition unit 133 of the sound output control device 200 acquires the audio content A-1, which has been assigned the application ID "AP1" and the category ID "CT3," from the agent device 60-1 as the audio content to be output. Subsequently, when the audio content A-1 is determined as the audio content to be output from the notification unit of the terminal device 10 through processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs voice synthesis on the message information included in the determined audio content A-1 to be output using tone parameters according to the category to which the audio content A-1 belongs.

[0125] 9, similarly to the first embodiment, the notification control unit 135 can identify that the category to which the audio content A-1 belongs is "Entertainment" by comparing the category ID "CT3" assigned to the audio content A-1 with the category classification database 121. Furthermore, the notification control unit 135 refers to the category classification database 121 and recognizes that it is specified that the audio message to be output from the notification unit of the terminal device 10 should be changed to a timbre characteristic of "female voice + fast speaking." Then, the notification control unit 135 performs voice synthesis on the audio data included in the audio content A-1 using parameters indicating "female voice + fast speaking," thereby changing the timbre of the voice.

[0126] In addition, in the second embodiment, the sound output control device 200 outputs background sound corresponding to the application of the provider of the audio content A-1, superimposed on the audio message.

[0127] For example, the notification control unit 135 further compares the application ID "AP1" assigned to the audio content A-1 with the application classification database 223 to determine that the application type to which the audio content A-1 belongs is "application AP1."

[0128] Furthermore, in response to the fact that the application type to which the sound content A-1 belongs is “application AP1,” the notification control unit 135 extracts sound effect #1 from the application classification database 223. Then, the notification control unit 135 adds the extracted sound effect #1 as background sound to the voice message after voice synthesis.

[0129] Next, the notification control unit 135 controls the sound output so that the audio content A-1 after conversion processing, such as voice synthesis and background sound addition, is notified by the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output the converted audio content A-1 by transmitting the converted audio content A-1 to the terminal device 10 of the user U1. The terminal device 10 notifies the converted audio content A-1 from the notification unit in response to the sound output control from the notification control unit 135. As a result, a voice message of "female voice + fast speech" is output to the user U1, with a sound effect such as "beep beep beep beep..." (an example of sound effect #1) as background sound. In other words, the user U1 can simultaneously hear the voice message with a tone corresponding to the category of the audio content and the background sound corresponding to the application providing it. As a result, the user U1 can easily understand that the currently output audio content is related to "entertainment" provided by the application AP1.

[0130] Next, another example shown in Fig. 9 will be described. For example, it is assumed that the agent information generation unit 631-5 of the agent device 60-5 generates audio content E-1 using audio data corresponding to message information of a content belonging to the category "Caution" based on data related to the situation acquired from the situation grasping device 30. In this case, the agent device 60-5 assigns a category ID "CT1" that identifies the category "Caution" to the audio content E-1, as shown in Fig. 9, so that the audio content E-1 is output to the user U1. The agent device 60-5 also assigns an application ID "AP5" that identifies application AP5, which is a provider application that provides the audio content E-1, to the audio content E-1. Then, the agent device 60-5 transmits the audio content A-1 to which the category ID and application ID have been assigned to the sound output control device 200.

[0131] The information acquisition unit 133 acquires the audio content E-1, which has been assigned the application ID "AP5" and the category ID "CT1," from the agent device 60-5 as the audio content to be output. Subsequently, when the audio content E-1 is determined as the audio content to be output from the notification unit of the terminal device 10 through processing by the queuing unit 132 and the determination unit 134, the notification control unit 135 performs voice synthesis on the message information included in the determined audio content E-1 to be output using tone parameters according to the category to which the audio content E-1 belongs.

[0132] 9, the notification control unit 135 can identify that the category to which the audio content E-1 belongs is "Caution" by comparing the category ID "CT1" assigned to the audio content E-1 with the category classification database 121. Furthermore, the notification control unit 135 references the category classification database 121 and recognizes that it is specified that the audio message to be output from the notification unit of the terminal device 10 should be changed to a timbre characteristic of "male voice + slow voice." Then, the notification control unit 135 performs voice synthesis on the audio data included in the audio content E-1 using a parameter indicating "male voice + slow voice," thereby changing the timbre of the voice.

[0133] Furthermore, the notification control unit 135 can identify that the application type to which the audio content E-1 belongs is "application AP5" by comparing the application ID "AP5" assigned to the audio content E-1 with the application classification database 223.

[0134] Furthermore, in response to the fact that the application type to which the audio content E-1 belongs is “application AP5,” the notification control unit 135 extracts music #5 from the application classification database 223. Then, the notification control unit 135 adds the extracted music #5 as background sound to the voice message after voice synthesis.

[0135] Next, the notification control unit 135 controls the sound output so that the audio content E-1 after conversion processing, such as voice synthesis and background sound addition, is notified by the notification unit of the terminal device 10 of the user U1. For example, the notification control unit 135 controls the terminal device 10 of the user U1 to output the converted audio content E-1 by transmitting the converted audio content E-1 to the terminal device 10 of the user U1. The terminal device 10 notifies the converted audio content E-1 from the notification unit in response to the sound output control from the notification control unit 135. As a result, a voice message of "Male voice + slow voice" is output to the user U1 with music #5 as background sound. In other words, the user U1 can simultaneously hear the voice message with a tone corresponding to the category of the audio content and the background sound corresponding to the application providing the audio content. As a result, the user U1 can easily understand that the currently output audio content is related to a "warning" provided by the application AP5.

[0136] So far, a specific example of the sound output control method performed in the second information processing has been described using some of the audio content shown in Fig. 9 as an example. The other audio content shown in Fig. 9 can also be explained following the example of some of the audio content, so detailed explanation will be omitted.

[0137] [4. Processing Procedure] Next, the procedure of information processing according to the second embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the procedure of information processing according to the second embodiment. Steps S101 to S106 shown in Fig. 10 are the same as those in the example of Fig. 6, so their description will be omitted, and steps S207 to S210 that are newly added in the information processing according to the second embodiment will be described.

[0138] The control unit 130 identifies the application type to which the agent information to be output belongs, based on the application ID assigned to the agent information to be output (step S207). For example, the control unit 130 identifies the application type to which the agent information to be output belongs, by checking the application ID against the application classification database 223.

[0139] Furthermore, the control unit 130 extracts background sounds corresponding to the identified application type from among background sounds that are set to differ between applications, as in the example of the application classification database 223 shown in FIG. 8 (step S208).

[0140] Furthermore, the control unit 130 adds the extracted background sound to the agent information after the voice conversion (step S209).

[0141] Finally, the notification control unit 135 controls the sound output so that the agent information to which the background sound has been added is notified by the notification unit of the terminal device 10 of the user U1 (step S210).

[0142] 10, the procedure has been described in which, after identifying agent information whose output timing has come, message information included in the identified agent information is converted into audio data, and then background sound is added, but the timing for converting message information into audio data and the timing for adding background sound are not limited to this. For example, as soon as the information acquisition unit 133 acquires new agent information, the conversion process into audio data corresponding to steps S104 to S106 and the process for adding background sound corresponding to steps S207 to S209 may be performed, and the audio content including the audio data after the addition of background sound may be subjected to the processes for determining the output priority and the output timing corresponding to steps S102 and S103.

[0143] In the second embodiment, the sound output control device 200 has been described as superimposing background sound corresponding to a provider application on a voice message in a tone corresponding to the category of the audio content, and outputting the superimposed background sound corresponding to the provider application. However, as another example, the sound output control device 200 may output a voice message in a tone corresponding to the provider application, and superimposing background sound corresponding to the category to which the audio content belongs. For example, when the application that provides the audio content is app AP1 and the message information included in the audio content belongs to the category "Entertainment," the notification control unit 135 may perform voice synthesis using tone parameters corresponding to app A1 and add background sound corresponding to the category "Entertainment" to the voice message. In this case, for example, in the category classification database 121 shown in FIG. 3, background sound data corresponding to the category indicated by the category ID is associated with each "category ID," and in the application classification database 223 shown in FIG. 8, voice features (voice parameters) corresponding to the application type indicated by the app ID are associated with each "app ID."

[0144] In this case, user U1 can simultaneously hear the voice message with the tone corresponding to the application that provides the voice message and the background sound corresponding to the category of the voice content. As a result, user U1 can easily understand the application that provides the voice content currently being output and the category of the voice content.

[0145] Alternatively, as yet another example, the sound output control device 200 may always fix the voice tone of the voice message to a standard tone and change only the background sound according to the category of the voice content. In other words, regardless of which of multiple applications the voice content to be output is provided by, the sound output control device 200 may add background sound according to the category to which the voice content to be output belongs. In this case, too, the user U1 can hear the background sound corresponding to the category of the voice content simultaneously with the voice message. As a result, the user U1 can easily understand the category to which the currently output voice content belongs.

[0146] [5. Summary] The sound output control device 200 according to the second embodiment acquires agent information to be output from an agent device capable of outputting multiple types of agent information, each type being distinguished by the content of the information or the source of the information, as information to be provided to the user. The sound output control device 200 then causes a notification unit to output a voice message corresponding to the agent information to be output. Specifically, the sound output control device 200 adds a background sound corresponding to the type to which the agent information to be output belongs to the voice message and causes the notification unit to notify the user. Such a sound output control device 200 allows the user to easily determine whether the currently output audio content is the desired type of audio content.

[0147] (Third embodiment) 1. Overview of the Third Embodiment From here, a third embodiment will be described. Information processing according to the third embodiment (i.e., the third information processing) is performed with the aim of solving the third problem described above. Specifically, the third information processing is performed by a sound output control device 300 corresponding to the sound output control device SV shown in FIG. 1. The sound output control device 300 performs the third information processing in accordance with a sound output control program according to the third embodiment.

[0148] 2. Configuration of the Sound Output Control Device According to the Third Embodiment Next, a sound output control device 300 according to a third embodiment will be described with reference to Fig. 11. Fig. 11 is a diagram showing an example of the configuration of the sound output control device 300 according to the third embodiment. As shown in Fig. 11, the sound output control device 300 has a communication unit 110, a storage unit 220, and a control unit 330. In the following description, descriptions of processing units that are assigned the same reference numerals as those in the sound output control devices 100 and 200 may be omitted or simplified.

[0149] (Regarding the control unit 330) The control unit 330 is realized by a CPU, an MPU, or the like executing various programs (e.g., a sound output control program) stored in a storage device inside the sound output control device 300 using RAM as a work area. The control unit 330 is also realized by an integrated circuit such as an ASIC or FPGA.

[0150] 11, the control unit 330 further has a presentation control unit 336, a sound effect setting unit 337, and a usage suspension acceptance unit 338, and realizes or executes the functions and actions of the information processing described below. Note that the internal configuration of the control unit 330 is not limited to the configuration shown in FIG. 11, and may have other configurations as long as they perform the information processing described below. Also, the connection relationship between the processing units included in the control unit 330 is not limited to the connection relationship shown in FIG. 11, and may be other connection relationships.

[0151] (Regarding the presentation control unit 336) As described in the second embodiment, the notification control unit 135 causes the notification unit to notify audio data to which different sound effects are applied among the plurality of applications, depending on the application that is the provider of the audio content acquired by the information acquisition unit 133. Therefore, the presentation control unit 336 causes the user to be presented with an application list showing the sound effects corresponding to each of the plurality of applications.

[0152] For example, the presentation control unit 336 controls so that image information showing the application list is presented to the user via the display unit (the display screen of the terminal device 10).

[0153] Furthermore, the presentation control unit 336 causes the notification unit to notify a voice message indicating the name of an application included in the application list, with sound effects corresponding to the application added.

[0154] (About sound effect setting section 337) The sound effect setting unit 337 accepts a user operation for setting a sound effect for each application.

[0155] (Regarding the suspension of use acceptance unit 338) The use suspension acceptance unit 338 accepts a user operation to suspend the use of an arbitrary application among a plurality of applications being used by the user. For example, the use suspension acceptance unit 338 suspends the use of an application selected by the user from among the applications included in the application list.

[0156] [3. Specific example of third information processing] Next, a specific example of the third information processing performed among the presentation control unit 336, the sound effect setting unit 337, and the use suspension receiving unit 338 will be described with reference to Fig. 12. Fig. 12 is a diagram showing an example of the third information processing.

[0157] 12 shows an example in which a setting screen C1 on which various settings can be made for each of a plurality of applications (apps currently being used by the user U1) linked to the terminal device 10 of the user U1 is displayed on the terminal device 10. The setting screen C1 may be provided by the presentation control unit 336 in response to a request from the user U1, for example. The aspect (screen configuration) of the setting screen C1 is not limited to the example in FIG. 12.

[0158] For example, suppose that the applications associated with the terminal device 10 of the user U1 are applications AP1, AP2, AP3, AP4, and AP5. In this case, as shown in Fig. 12, the application names indicating the names of the applications AP1 to AP5 are displayed as a "list of applications currently in use" on the setting screen C1. This list of applications corresponds to an application list.

[0159] Furthermore, the setting screen C1 allows the user U1 to set background sounds corresponding to each application linked to the terminal device 10. In this regard, Fig. 12 shows an example in which a pull-down button PD1 for displaying a list of background sound candidates corresponding to the application AP1 in a pull-down format is associated next to the application name indicating the application AP1. This allows the user U1 to select any background sound from the background sound candidates displayed in a pull-down menu using the pull-down button PD1, and set the selected background sound.

[0160] For example, as shown in FIG. 12 , when the BGM “MUSIC#3” is selected, the sound effect setting unit 337 accepts the setting of the BGM “MUSIC#3” for the app AP1 in response to the selection operation. Furthermore, in response to the acceptance of the setting of the BGM “MUSIC#3,” the presentation control unit 336 controls, for example, the notification control unit 135 so that audio data (audio message) indicating the app name of the app AP1 is output from the notification unit with the BGM “MUSIC#3” added. In this case, the notification control unit 135 extracts the data of the BGM “MUSIC#3” from the storage unit and adds it to the audio data indicating the app name of the app AP1. Then, the notification control unit 135 performs sound output control so that the audio data with the BGM “MUSIC#3” added is announced by the notification unit of the terminal device 10 of the user U1.

[0161] This allows user U1 to hear a voice message reading out the name of app AP1 (e.g., "Entertainment information app from Company A") while background music "MUSIC♯3" plays, and allows user U1 to imagine the atmosphere of background music "MUSIC♯3" and how the voice message will sound in that atmosphere. As a result, when user U1 uses multiple apps, as in the example of Figure 12, user U1 can easily distinguish which app is the source of the audio content.

[0162] Up to this point, a specific example of the third information processing has been described using the example of the application AP1 shown in FIG. 12, but other applications will also be described.

[0163] 12 shows an example in which a pull-down button PD3 for displaying a list of background sound candidates corresponding to the application AP3 in a pull-down format is associated next to the application name indicating the application AP3. This allows the user U1 to set the selected background sound by selecting any background sound from the background sound candidates displayed in a pull-down menu using the pull-down button PD3.

[0164] For example, as shown in FIG. 12 , when the BGM “MUSIC#1” is selected, the sound effect setting unit 337 accepts the setting of the BGM “MUSIC#1” for the app AP3 in response to the selection operation. Furthermore, in response to the acceptance of the setting of the BGM “MUSIC#1,” the presentation control unit 336 controls, for example, the notification control unit 135 so that audio data indicating the app name of the app AP3 is output from the notification unit with the BGM “MUSIC#1” added. In this case, the notification control unit 135 extracts the data of the BGM “MUSIC#1” from the storage unit and adds it to the audio data indicating the app name of the app AP3. Then, the notification control unit 135 performs sound output control so that the audio data with the BGM “MUSIC#1” added is announced by the notification unit of the terminal device 10 of the user U1.

[0165] This allows user U1 to hear a voice message reading out the name of app AP3 (e.g., "Company C's vacation facility information app") while background music "MUSIC#1" plays, allowing the user U1 to imagine the atmosphere of the background music "MUSIC#1" and how the voice message will sound in that atmosphere. As a result, when using multiple apps, as in the example of Figure 12, user U1 can easily distinguish which app is the source of the audio content.

[0166] From here, we will explain how to stop the use of apps (delete unnecessary apps) using the example of Fig. 12. The setting screen C1 shown in Fig. 12 further has a function to stop the use of selected apps by deleting the selected apps from the apps included in the application list.

[0167] For example, suppose that user U1 does not need to provide audio content from application AP1 among applications AP1 to AP5 that are currently in use, and wishes to disable application AP1. In this case, user U1 selects application AP1 from the application names included in the application list, and presses the delete button BT.

[0168] Then, the use suspension acceptance unit 338 accepts a user operation to suspend the use of the application AP1. Then, the use suspension acceptance unit 338 suspends the use of the application AP1 selected by the user U1 from among the applications included in the application list. For example, the use suspension acceptance unit 338 suspends the use of the application AP1 by deleting the application AP1 from the application list. This allows the user U1 to set an environment in which, for example, only the audio content that the user U1 requires is output.

[0169] [4. Summary] A sound output control device 300 according to the third embodiment acquires agent information provided by each of a plurality of applications having a voice agent function. The sound output control device 300 then causes a notification unit to notify a user of a voice message with different sound effects applied to the plurality of applications depending on which application provided the previously acquired agent information. The sound output control device 300 also presents a user with an application list showing the sound effects corresponding to each of the plurality of applications. With this sound output control device 300, when a user uses multiple applications, the user can easily distinguish which application is the source of the audio content, thereby identifying an application that suits the user's preferences.

[0170] (others) [1. Hardware Configuration] The sound output control device 100 according to the first embodiment and the sound output control device 200 according to the second embodiment described above are realized, for example, by a computer 1000 configured as shown in Fig. 13. The sound output control device 100 will be described below as an example. Fig. 13 is a hardware configuration diagram showing an example of a computer that realizes the functions of the sound output control device 100. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0171] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.

[0172] The HDD 1400 stores programs executed by the CPU 1100, data used by these programs, etc. The communication interface 1500 receives data from other devices via a predetermined communication network and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the predetermined communication network.

[0173] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.

[0174] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0175] For example, when the computer 1000 functions as the sound output control device 100 in the first embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to realize the functions of the control unit 130. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via a predetermined communication network.

[0176] Furthermore, for example, when the computer 1000 functions as the sound output control device 300 in the third embodiment, the CPU 1100 of the computer 1000 executes a program loaded onto the RAM 1200 to realize the functions of the control unit 330.

[0177] [2. Other] Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0178] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0179] Furthermore, the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0180] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be implemented in other forms that include the aspects described in the Disclosure of the Invention section, as well as in various modifications and improvements based on the knowledge of those skilled in the art.

[0181] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an information acquisition unit can be read as information acquisition means or information acquisition circuit. [Explanation of symbols]

[0182] 1. Information Processing Systems 100 Sound output control device 120 Storage section 121 Category Classification Database 122 Content Buffer 130 Control Unit 133 Information Acquisition Department 135 Notification control unit 200 Sound output control device 220 Storage section 223 App Classification Database 300 Sound output control device 336 Presentation control unit 337 Sound Effect Settings 338 Suspension of Use Reception Department

Claims

1. an information acquisition unit that acquires agent information to be output from each of a plurality of applications that can output agent information belonging to each of a plurality of different categories as information to be provided to a user; an annunciation control unit that controls the annunciation unit so that the tone of the voice of the voice message corresponding to the agent information to be output is changed according to the category to which the agent information to be output belongs, and the voice message is annunciated with a sound effect added according to the application that provided the agent information to be output; A sound output control device comprising:

2. Each piece of agent information acquired from the plurality of applications is assigned category identification information that identifies a category to which the agent information belongs, and application identification information that identifies an application that has provided the agent information, The notification control unit changes the tone of the voice of the voice message in accordance with the category indicated by the category identification information assigned to the agent information to be output, and controls the notification unit so that the voice message is notified with a sound effect added thereto in accordance with the application indicated by the application identification information assigned to the agent information to be output.

2. The sound output control device according to claim 1.

3. an information acquisition unit that acquires agent information to be output from each of a plurality of applications that can output agent information belonging to each of a plurality of different categories as information to be provided to a user; an annunciation control unit that changes the tone of the voice of the voice message corresponding to the agent information to be output in accordance with the application that provided the agent information to be output, and controls the annunciation unit so that the voice message is annunciated with a sound effect added in accordance with the category to which the agent information to be output belongs; A sound output control device comprising:

4. Each piece of agent information acquired from the plurality of applications is assigned category identification information that identifies a category to which the agent information belongs, and application identification information that identifies an application that has provided the agent information, The notification control unit changes the tone of the voice of the voice message in accordance with the application indicated by the application identification information assigned to the agent information to be output, and controls the notification unit so that the voice message is notified with a sound effect added in accordance with the category indicated by the category identification information assigned to the agent information to be output.

4. The sound output control device according to claim 3.

5. The sound effect is a background sound or a sound effect.

5. The sound output control device according to claim 1, wherein the sound output control device is a sound output control device.

6. A sound output control method executed by a sound output control device, an information acquisition step of acquiring agent information to be output from each of a plurality of applications capable of outputting agent information belonging to each of a plurality of different categories as information to be provided to a user; a notification control step of controlling the notification unit so that the tone of the voice of the voice message corresponding to the agent information to be output is changed according to the category to which the agent information to be output belongs, and the voice message is notified with a sound effect added according to the application that provided the agent information to be output; A sound output control method comprising:

7. A sound output control method executed by a sound output control device, an information acquisition step of acquiring agent information to be output from each of a plurality of applications capable of outputting agent information belonging to each of a plurality of different categories as information to be provided to a user; a notification control step of changing the tone of the voice of the voice message corresponding to the agent information to be output in accordance with the application that provided the agent information to be output, and controlling the notification unit so that the voice message is notified with a sound effect added in accordance with the category to which the agent information to be output belongs; A sound output control method comprising:

8. A sound output control program executed by a sound output control device having a computer, The computer an information acquisition means for acquiring agent information to be output from each of a plurality of applications capable of outputting agent information belonging to each of a plurality of different categories as information to be provided to a user; an annunciation control means for controlling the annunciation unit so that the tone of the voice of the voice message corresponding to the agent information to be output is changed in accordance with the category to which the agent information to be output belongs, and the voice message is annunciated with a sound effect added in accordance with the application that provided the agent information to be output; A sound output control program characterized by causing the program to function as a sound output control program.

9. A sound output control program executed by a sound output control device having a computer, The computer an information acquisition means for acquiring agent information to be output from each of a plurality of applications capable of outputting agent information belonging to each of a plurality of different categories as information to be provided to a user; an annunciation control means for controlling the annunciation unit so that the tone of the voice of the voice message corresponding to the agent information to be output is changed in accordance with the application that provided the agent information to be output, and the voice message is annunciated with a sound effect added in accordance with the category to which the agent information to be output belongs; A sound output control program characterized by causing the program to function as a sound output control program.

Citation Information

Patent Citations

  • Voice guidance control system

    JP1992316100A

  • Data processing method and system discriminating audio response of plurality of processing

    JP1993324262A

  • Voice synthesizing system and method therefor

    JP1997081174A

  • Voice synthesizer

    JP2000305584A

  • Guiding device

    JP2001041763A