server

JP7900440B2Active Publication Date: 2026-08-04NOMURA RESEARCH INSTITUTE
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NOMURA RESEARCH INSTITUTE
Filing Date
2024-06-11
Publication Date
2026-08-04

AI Technical Summary

Benefits of technology

【0008】 本発明によれば、スマートスピーカを効果的な広告媒体として用いることができる技術を提供できる、またはスマートスピーカシステムをさらに改善することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007900440000001
    Figure 0007900440000001
  • Figure 0007900440000002
    Figure 0007900440000002
  • Figure 0007900440000003
    Figure 0007900440000003
Patent Text Reader

Abstract

To provide a technology capable of using a smart speaker as an effective advertising medium, or further improving a smart speaker system.SOLUTION: A server includes: receiving means for receiving a distribution request from a microphone and a speaker having a communication function via a network; acquisition means for acquiring an audio content without an image in response to the received distribution request; selection means for selecting a voice advertisement without an image from voice advertisement holding means; and transmitting means for transmitting the acquired audio content and the selected audio advertisement together to the speaker via the network.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a server that communicates with a speaker having a microphone.

Background Art

[0002] Smart speakers equipped with a microphone and a communication function, which enable voice operation and information search, have begun to spread (see, for example, Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Current systems including smart speakers can receive requests from users by voice and process those requests. In such a situation, it is desired to create a more beneficial smart speaker system.

[0005] The present invention has been made in view of such problems, and an object thereof is to provide a technology that can use a smart speaker as an effective advertising medium or to further improve a smart speaker system.

Means for Solving the Problems

[0006] Note: In the translation of , the original "平成30年5月9日" is translated as "May 9, 2018" considering that 平成30年 corresponds to 2018 in the Gregorian calendar.One aspect of the present invention relates to a server. This server includes: a receiving means for receiving audio information acquired via a network from a speaker having a microphone and communication function via the speaker's microphone; a speech recognition means for recognizing the user's speech content based on the received audio information; a speech information holding means for holding the recognized speech content; an acquisition means for acquiring audio content without images in response to a distribution request included in the recognized speech content; a selection means for selecting audio advertisements without images from an audio advertisement holding means; and a transmission means for transmitting the acquired audio content and the selected audio advertisements together to the speaker via the network. A control unit that controls the timing of the output of audio advertisements from the speaker, The selection means selects an audio advertisement based on the spoken content held in the audio information holding means.

[0007] Furthermore, any combination of the above components, or any substitution of components or expressions of the present invention between devices, methods, systems, computer programs, recording media storing computer programs, etc., are also valid embodiments of the present invention. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a technology that allows smart speakers to be used as an effective advertising medium, or to further improve smart speaker systems. [Brief explanation of the drawing]

[0009] [Figure 1] This is a schematic diagram showing the configuration of an audio advertising distribution system according to the first embodiment. [Figure 2] Figure 1 is a block diagram showing the functions and configuration of the smart speaker. [Figure 3] Figure 1 is a hardware configuration diagram of the management server. [Figure 4] Figure 1 is a block diagram showing the functions and configuration of the management server. [Figure 5] Figure 4 is a data structure diagram showing an example of the audio content storage unit. [Figure 6] It is a data structure diagram showing an example of the voice advertisement holding unit in FIG. 4. [Figure 7] It is a data structure diagram showing an example of the voice information holding unit in FIG. 4. [Figure 8] It is a data structure diagram showing an example of the user information holding unit in FIG. 4. [Figure 9] It is a data structure diagram showing an example of the session information holding unit in FIG. 4. [Figure 10] It is a flowchart showing the flow of a series of processes in the management server in FIG. 1. [Figure 11] It is a schematic top view of the user's room. [Figure 12] It is a schematic diagram showing the configuration of the voice operation system according to the third embodiment. [Figure 13] It is a schematic diagram showing the configuration of the voice operation system according to the fourth embodiment. [Figure 14] It is a block diagram showing the functions and configuration of the management server in FIG. 13. [Figure 15] It is a data structure diagram showing an example of the user information holding unit in FIG. 14.

Embodiments for Carrying Out the Invention

[0010] Hereinafter, the same or equivalent components, members, and processes shown in each drawing are denoted by the same reference numerals, and repeated explanations are omitted as appropriate. Also, some members that are not important for explanation in each drawing are omitted from the display.

[0011] (First Embodiment) In the voice advertisement distribution system according to the first embodiment, the user can perform the following operations using a smart speaker, for example. · Simple research · Checking the weather forecast · Listening to the news · Setting an alarm · Checking the schedule · Calculating · Playing music · Control of smart home appliances. The voice advertisement distribution system acquires the user's speech through the microphone of the smart speaker, and by recognizing the speech, understands that the user is requesting the distribution of voice content (such as search results, weather forecasts, news, schedules, calculation results, music, etc.). The system prepares the requested voice content and distributes it to the smart speaker. At this time, a voice advertisement is inserted into the voice content to be distributed so that the voice advertisement is played before the voice content is played on the smart speaker.

[0012] This voice advertisement may be, for example, a voice advertisement tailored to the voice content to be distributed, a voice advertisement based on the content of previous conversations, or a voice advertisement based on the sounds around the smart speaker collected by the smart speaker immediately before the distribution of the voice content.

[0013] The length of the voice advertisement may be adjusted according to the situation of the conversation with the user and the situation around the smart speaker. As a mode of adjustment, for example, the length of the voice advertisement may be adjusted according to the content of the voice content to be distributed, or the important part of the voice advertisement may be extracted.

[0014] The timing of audio ad playback may be determined considering factors such as the fact that ads are more effective when users are near a smart speaker, and that users may find it annoying if audio ads are played while they are talking to the smart speaker or other users. For example, audio ads may only be played when it is determined that a user is near a smart speaker. Also, audio ads may not be played when a user is talking to other users or speaking to the smart speaker. In the latter case, the audio ad may start or resume outputting when the user stops speaking. Audio ads may also be output in conjunction with other electronic devices, such as televisions (hereinafter referred to as TVs). For example, after an ad is broadcast on TV, follow-up information may be output as audio from the smart speaker. In this case, the next TV ad may be muted. Alternatively, a related ad may be broadcast on TV after an audio ad has been played on the smart speaker.

[0015] The audio advertising distribution system has a function to acquire a voiceprint from the user's speech obtained via a smart speaker and perform user authentication using voiceprint authentication. Furthermore, the audio advertising distribution system can be linked with other services such as web services and social networking services, allowing it to associate authenticated users with user accounts in other services. In this case, the audio advertising distribution system may select audio advertisements for authenticated users that are associated with the authenticated user's account. For example, the audio advertising distribution system may select audio advertisements based on information collected by the smart speaker and account attributes. The audio advertising distribution system may also update account attributes with information collected by the smart speaker.

[0016] Figure 1 is a schematic diagram showing the configuration of the audio advertising distribution system 2 according to the first embodiment. The audio advertising distribution system 2 comprises a management server 4, a smart speaker 10, and a TV 12. The management server 4, the smart speaker 10, and the TV 12 are connected to each other via a network 6 such as the Internet. Both the smart speaker 10 and the TV 12 are installed in the user 8's room 14. The smart speaker 10 is a speaker with a microphone and communication functions, and as described above, it is connected to the network 6 and is also configured to enable P2P (Peer to Peer) communication 16 with the TV 12. Figure 1 shows an example of communication between the smart speaker 10 and the management server 4, but there is no limit to the number of smart speakers 10, nor is there a limit to the number of users 8.

[0017] User 8 speaks sentences expressing requests for audio content, such as "I want to hear a song by Taro Komura," "Tell me today's news," "What's the weather like tonight?", or "Tell me about Izumo Taisha Shrine," into the smart speaker 10. The microphone in the smart speaker 10 converts the voice spoken by User 8 into an electrical signal, and the smart speaker 10 transmits the resulting electrical signal as an audio signal to the management server 4 via the network 6. The management server 4 performs speech recognition processing on the received audio signal to understand what kind of audio content the user is requesting. The management server 4 generates distribution information with an audio advertisement attached to the requested audio content and transmits it to the smart speaker 10 via the network 6. Upon receiving the distribution information, the smart speaker 10 first outputs the audio advertisement, and then outputs the audio content.

[0018] The smart speaker 10 may or may not have a display, but the content delivered from the management server 4 is not content that combines images and audio, such as still images or videos, but rather audio content without images (or content consisting only of audio). Similarly, audio advertisements are also audio advertisements without images.

[0019] Figure 2 is a block diagram showing the functions and configuration of the smart speaker 10 in Figure 1. Each block shown here can be realized in hardware terms by elements and mechanical devices such as a computer CPU, and in software terms by computer programs, etc., but here we are depicting functional blocks that are realized through the cooperation of these. Therefore, it will be understood by those skilled in the art who have read this specification that these functional blocks can be realized in various forms by combinations of hardware and software.

[0020] The smart speaker 10 comprises a speaker 102, a microphone 104, a communication unit 106, an input unit 108, and a processing unit 110. The communication unit 106 functions as an interface for communication with the network 6 and also functions as an interface for P2P communication 16. The input unit 108 includes physical input mechanisms such as a power button and volume control buttons. The processing unit 110 controls the speaker 102, microphone 104, communication unit 106, and input unit 108 to realize the various functions of the smart speaker 10.

[0021] In this embodiment, it is assumed that the microphone 104 converts the user's speech into an audio signal, the communication unit 106 transmits the audio signal to the management server 4, and the management server 4 performs speech recognition processing on the audio signal. However, the technical concept of this embodiment can also be applied when at least some of the speech recognition processing is performed in the smart speaker, when the audio content acquisition processing and audio advertisement selection processing described later are performed in the smart speaker, or when the smart speaker is standalone. It should be noted that sending the results of speech recognition performed in the smart speaker to the management server, and sending the audio signal directly from the smart speaker to the management server, can both be said to be sending audio information corresponding to the user's speech to the management server.

[0022] Figure 3 is a hardware configuration diagram of the management server 4 shown in Figure 1. The management server 4 includes memory 130, a processor 132, a communication interface 134, a display 136, and an input interface 138. Each of these elements is connected to a bus 140 and communicates with each other via the bus 140.

[0023] Memory 130 is a storage area for storing data and programs. Data and programs may be permanently stored in memory 130 or temporarily stored. The processor 132 implements various functions in the management server 4 by executing programs stored in memory 130. The communication interface 134 is an interface for sending and receiving data to and from the outside of the management server 4. For example, the communication interface 134 includes an interface for accessing the network 6. The display 136 is a device for displaying various information, such as a liquid crystal display or an organic EL (electroluminescence) display. The input interface 138 is a device for receiving input from the user. The input interface 138 includes, for example, a mouse, a keyboard, or a touch panel provided on the display 138.

[0024] Figure 4 is a block diagram showing the functions and configuration of the management server 4 in Figure 1. Each block shown here can be implemented in hardware terms by components and mechanical devices such as the CPU of a computer, and in software terms by computer programs, etc., but here, the functional blocks are depicted that are realized through the cooperation of these components. Therefore, it will be understood by those skilled in the art who have read this specification that these functional blocks can be implemented in various ways by combinations of hardware and software.

[0025] The management server 4 includes an audio content storage unit 402, an audio advertisement storage unit 404, an audio information storage unit 406, a user information storage unit 408, a session information storage unit 410, an audio signal reception unit 412, an audio recognition unit 414, a user authentication unit 416, a session management unit 418, a content acquisition unit 420, an advertisement selection unit 422, an advertisement adjustment unit 424, a transmission information generation unit 426, a transmission unit 428, a timing control unit 430, and an attribute update unit 432.

[0026] Figure 5 is a data structure diagram showing an example of the audio content storage unit 402 in Figure 4. The audio content storage unit 402 stores a content ID that identifies the audio content, keywords that characterize the audio content, and the data of the audio content in association with each other. In addition to or instead of keywords, other metadata such as tags may be used.

[0027] The data stored in the audio content storage unit 402 may be data that has been generated and registered by the management server 4 in advance or upon request. Known speech synthesis technology may be used when creating the audio content data. Alternatively, the data stored in the audio content storage unit 402 may be data that the management server 4 has obtained in advance or upon request from servers of other services.

[0028] Figure 6 is a data structure diagram showing an example of the audio advertisement storage unit 404 in Figure 4. The audio advertisement storage unit 404 stores an advertisement ID that identifies the audio advertisement, keywords that characterize the audio advertisement, attributes of the audio advertisement, and data of the audio advertisement in association with each other. In addition to or instead of keywords, other metadata such as tags may be used. The data stored in the audio advertisement storage unit 404 may be data received from advertisers by the entity operating the management server 4.

[0029] Figure 7 is a data structure diagram showing an example of the voice information storage unit 406 in Figure 4. The voice information storage unit 406 stores voice information acquired via the microphone 104 of the smart speaker 10. The voice information includes the user's utterances obtained by speech recognition of the voice signal by the speech recognition unit 414, which will be described later. The voice information storage unit 406 stores a user ID that identifies the user, a session ID of the dialogue session between the smart speaker 10 and the user, and the utterances of the user or system in the dialogue session, in association with each other. The voice information storage unit 406 also stores a system-specific ID as the user ID corresponding to the system's utterances.

[0030] Figure 8 is a data structure diagram showing an example of the user information storage unit 408 in Figure 4. The user information storage unit 408 stores the following in association: user ID, user voiceprint data, account ID which identifies the user's account in other services, account attributes, and NG ad ID which identifies ads that the user found offensive. The voiceprint data may be obtained when the user first registers with the voice ad delivery system 2.

[0031] Figure 9 is a data structure diagram showing an example of the session information storage unit 410 in Figure 4. The session information storage unit 410 maintains the state of the current dialogue session between the smart speaker 10 and the user. The session information storage unit 410 maintains the user ID of the user involved in the currently existing or maintained dialogue session, the session ID of the dialogue session, and the state of the dialogue session in association with each other. The state of the dialogue session is selected from three options: "Speaking," indicating that the user is speaking; "Conversing," indicating that the user is conversing with another user; and "Waiting for Speech," indicating that the user is not speaking and is waiting for the next utterance from the user or the system. When it is determined that the dialogue session has ended, the entry related to that dialogue session is deleted from the session information storage unit 410. In other words, if the session ID of a dialogue session is registered in the session information storage unit 410, it is determined that the dialogue session is continuing and the user involved in that dialogue session is in the vicinity of the smart speaker 10.

[0032] Returning to Figure 4, the voice signal receiving unit 412 receives voice signals representing the user's speech content from the smart speaker 10 via the network 6. As described above, the voice signal is an electrical signal obtained by converting the user's speech voice with the microphone 104, and is in particular an electrical signal representing the waveform of the speech. The speech content includes questions and responses to the smart speaker 10 (or the voice advertising distribution system 2), monologues, and conversations with other users.

[0033] The speech recognition unit 414 performs predetermined speech recognition processing on the speech signal received by the speech signal receiving unit 412. The speech recognition unit 414 derives the user's utterance from the speech signal through speech recognition. The speech recognition processing in the speech recognition unit 414 may be implemented using known speech recognition techniques such as n-grams or hidden Markov models.

[0034] The user authentication unit 416 extracts or acquires a voiceprint from the audio signal received by the audio signal receiving unit 412. The user authentication unit 416 performs user authentication (i.e., voiceprint authentication) based on the extracted voiceprint. The user authentication unit 416 refers to the user information holding unit 408 and determines whether there is a voiceprint in the voiceprints held in the user information holding unit 408 that matches the extracted voiceprint. If there is a matching voiceprint, the user authentication unit 416 identifies the user ID corresponding to that voiceprint and associates the identified user ID with the utterance derived by the speech recognition unit 414. In this case, the user who made the utterance corresponding to the audio signal received by the audio signal receiving unit 412 is considered to have been voiceprint authenticated by the management server 4. If there is no matching voiceprint, the user authentication unit 416 generates an output indicating no match or user unknown. The management server 4 may start new user registration according to this output.

[0035] The session management unit 418 manages the dialogue session between the smart speaker 10 and the user. The session management unit 418 manages the voice information storage unit 406 and the session information storage unit 410. The session management unit 418 associates the user ID and spoken content, which have been associated by the user authentication unit 416, with a session ID that identifies the dialogue session between the smart speaker 10 and its user, and registers this information in the voice information storage unit 406.

[0036] The session management unit 418 determines the state of the current dialogue session between the smart speaker 10 and its user based on the user ID and utterance content associated with the user authentication unit 416. The session management unit 418 updates the session information storage unit 410 with the determined state. For example, if the analysis result of the utterance content indicates that the utterance is still in progress, the session management unit 418 determines the state of the current dialogue session to "Speaking". If the analysis result of the utterance content indicates that the utterance has ended, the session management unit 418 determines the state of the current dialogue session to "Waiting for utterance". If the analysis result of the utterance content indicates the end of the dialogue session (for example, if the utterance is a word indicating the end of the dialogue session, such as "See you later" or "Bye-bye"), the session management unit 418 deletes all entries with the session ID of that dialogue session from the session information storage unit 410. The session management unit 418 may also delete dialogue sessions from the session information storage unit 410 if they remain in the "Waiting for utterance" state for a predetermined period of time. Here, the state of the dialogue session was determined based on the content of the utterance, but the state of the dialogue session may also be determined based on imaging information from a separately provided camera, instead of, or in addition to, the content of the utterance. The camera here may be installed in the smart speaker 10 itself, a separate camera with communication capabilities may be used, or a television or computer with camera capabilities may be used.

[0037] If the utterance derived by the speech recognition unit 414 includes a request for the distribution of audio content, the content acquisition unit 420 acquires the requested audio content from the audio content storage unit 402. For example, if the utterance is a request for the distribution of music content such as "I want to listen to a song by Taro Komura," the content acquisition unit 420 acquires the requested music content from the audio content storage unit 402. Alternatively, the content acquisition unit 420 may access a music distribution service server and acquire the requested music content along with metadata from that server. In this case, the content acquisition unit 420 may register the acquired music content and metadata in the audio content storage unit 402.

[0038] If the spoken content is a request for information content such as "Tell me today's news" or "What's the weather like tonight?", the content acquisition unit 420 acquires the requested information content from the audio content storage unit 402. Alternatively, the content acquisition unit 420 may access the information distribution service server and acquire the requested information content in text format along with metadata from that server. In this case, the content acquisition unit 420 may convert the acquired text-format information content into audio data using a predetermined speech synthesis process. The content acquisition unit 420 may register the information content and metadata, now in audio data format, with the audio content storage unit 402. The speech synthesis process may be implemented using known speech synthesis technology.

[0039] If the spoken content is a request for search results such as "Tell me about Izumo Taisha," the content acquisition unit 420 acquires the requested search results from the audio content storage unit 402. Alternatively, the content acquisition unit 420 may access the search service server and acquire the requested search results in text format along with metadata from that server. In this case, the content acquisition unit 420 may convert the acquired text format search results into audio data using a predetermined speech synthesis process. The content acquisition unit 420 may register the audio data of the search results and metadata in the audio content storage unit 402.

[0040] The ad selection unit 422 selects from the audio ad storage unit 404 an audio ad to be attached to the audio content acquired by the content acquisition unit 420. The criteria for selecting an audio ad in the ad selection unit 422 are any one or any combination thereof of (1) relevance to the content of the audio content acquired by the content acquisition unit 420, (2) relevance to the user's utterances in the current conversation session between the smart speaker 10 and the user, which are stored in the audio information storage unit 406, or (3) relevance to the attributes of the authenticated user's account.

[0041] For example, regarding (1), in response to a search result delivery request for "Tell me about Izumo Taisha," the content acquisition unit 420 acquires audio data of audio content that says "Izumo Taisha was, in ancient times..." The ad selection unit 422 acquires the keywords "Izumo Taisha, god, matchmaking" which correspond to "Izumo Taisha was, in ancient times..." acquired by the content acquisition unit 420, from the audio content holding unit 402. The ad selection unit 422 refers to the audio ad holding unit 404 and selects the audio data of the audio ad "Want to travel to Izumo Taisha? Then consult ABC Travelers" which has the keywords "Izumo Taisha, matchmaking" which correspond to the acquired keywords "Izumo Taisha, god, matchmaking." In this way, by comparing the keywords of the audio content with the keywords of the audio ad, the ad selection unit 422 selects an audio ad that corresponds to the content of the audio content acquired by the content acquisition unit 420.

[0042] For example, regarding (2), between the smart speaker 10 and the user (User) "Will I make it to the station in time by taxi?" (Smart Speaker 10) "We'll make it." (User) "What's the weather like tonight?" Let's assume the following conversation is taking place. In response to a request for information content delivery, "What's the weather like tonight?", the content acquisition unit 420 acquires audio data for audio content, "The weather in area C tonight is showers, the temperature is..." The ad selection unit 422 refers to the audio information storage unit 406 and identifies "Will I make it to the station by taxi?" as the user's utterance in the current conversation session between the smart speaker 10 and the user. The ad selection unit 422 extracts the keywords "station" and "taxi" from the identified utterance "Will I make it to the station by taxi?". The ad selection unit 422 refers to the audio ad storage unit 404 and selects the audio data for an audio advertisement called "ZZZ Taxi Dispatch Service," which has the keywords "taxi" and "dispatch" corresponding to the extracted keywords "station" and "taxi". In this way, by referring to the audio information storage unit 406, the ad selection unit 422 can select an audio advertisement based on the user's utterance in the current conversation session between the smart speaker 10 and the user.

[0043] In the example above, if criterion (1) is used instead of (2), the ad selection unit 422 retrieves the keywords "C region, rain, low temperature" from the audio content holding unit 402, which correspond to "Tonight's weather in C region is showery, the temperature is..." retrieved by the content acquisition unit 420. The ad selection unit 422 then refers to the audio ad holding unit 404 and selects the audio data for the audio ad "CB Company's Hyper Umbrella won't break for 10 years!" which has the keywords "umbrella, rain" corresponding to the retrieved keywords "C region, rain, low temperature". Thus, even if the content of the conversation between the smart speaker 10 and the user is the same, the audio ad selected may differ depending on the criteria used.

[0044] For example, regarding (3), in response to a request for information content delivery, "Tell me today's news," the content acquisition unit 420 acquires audio data of audio content, "Around 6 a.m. this morning, there was a fire in B city, A prefecture..." Simultaneously, the user who spoke "Tell me today's news" is authenticated by voiceprint authentication in the user authentication unit 416, and the user ID "B102" of that user is identified. The ad selection unit 422 acquires the account attributes "child, male, single" corresponding to the identified user ID "B102" from the user information retention unit 408. The ad selection unit 422 refers to the audio ad retention unit 404 and selects the audio data of an audio advertisement, "If you come to F city, you can ride the SL," which has the attributes "child, male" corresponding to the acquired attributes "child, male, single." Furthermore, if the account attributes corresponding to the identified user ID "B102" are "adult, female, single," the ad selection unit 422 refers to the audio ad holding unit 404 and selects the audio data of the audio ad "Want to travel to Izumo Taisha Shrine? Then consult ABC Travel Agency," which has the attributes "single, adult" that correspond to those attributes. In this way, by comparing the attributes of the authenticated user's account with the attributes of the audio ad, the ad selection unit 422 selects an audio ad that corresponds to the attributes of the authenticated user's account.

[0045] Alternatively, if the attributes of the account corresponding to the identified user ID "B102" are "adult, male, married," the ad selection unit 422 first selects four audio advertisements as candidates that correspond to those attributes: "For Christmas presents, we recommend XX precious metal rings," "For fire insurance, leave it to XYZ Fire & Marine Insurance," "ZZZ taxi dispatch service that comes quickly," and "Want to travel to Izumo Taisha Shrine? Then consult ABC Travel Agency." Furthermore, the ad selection unit 422 obtains the keywords "A prefecture, B city, fire" corresponding to "Around 6 a.m. this morning in B city, A prefecture,..." obtained by the content acquisition unit 420 from the audio content holding unit 402. Of the four candidates selected, the ad selection unit 422 selects the audio data of the audio advertisement "For fire insurance, leave it to XYZ Fire & Marine Insurance" which has the keywords "fire, fire, insurance" corresponding to the acquired keywords "A prefecture, B city, fire." Thus, it is also possible to combine criteria (1) and criteria (3) in the form of selecting candidates based on criterion (3) and then narrowing them down based on criterion (1).

[0046] For example, between the smart speaker 10 and the user (User) "I want to listen to a song by Taro Komura." (Smart Speaker 10) "How about some Christmas songs?" (User) "Okay, then." Let's assume the following dialogue is taking place. In response to a request for music content distribution, "I want to listen to a song by Taro Komura," the content acquisition unit 420 acquires audio data of Taro Komura's Christmas song. The management server 4 asks the user via the smart speaker 10 if a Christmas song is acceptable. When the management server 4 receives an affirmative response from the user, "Okay, that's fine," it attaches an audio advertisement to the acquired audio data of Taro Komura's Christmas song and sends it to the smart speaker 10. At this point, the advertisement selection unit 422 acquires the keywords corresponding to Taro Komura's Christmas song acquired by the content acquisition unit 420, "Taro Komura (lyrics and composition), Anime B (theme song), Movie C (insert song), Christmas song, ring," from the audio content holding unit 402. The ad selection unit 422 refers to the audio ad holding unit 404 and selects two audio ads as candidates: "Anime Otsu, airing Fridays from 6pm!" (keyword: "Anime Otsu, Friday, 6pm") and "If you're looking for a Christmas present, we recommend a ring from XX precious metals" (keyword: "Christmas, present, ring"), which have keywords corresponding to the acquired keywords "Taro Komura (lyrics and composition), Anime Otsu (theme song), Movie C (insert song), Christmas song, ring". Furthermore, the ad selection unit 422 obtains the account attributes "adult, male, married" corresponding to the user ID "A101" of the user authenticated by voiceprint authentication from the user information holding unit 408. Of the two selected candidates, the ad selection unit 422 selects the audio data of the audio ad "If you're looking for a Christmas present, we recommend a ring from XX precious metals" which has the attribute "adult" corresponding to the acquired attribute "adult, male, married". Furthermore, if the user ID of the user authenticated by voiceprint authentication is "A105", the ad selection unit 422 selects the audio data of the audio advertisement "Anime, airing Fridays from 6pm!" which has the attributes "female, child" corresponding to the acquired attributes "child, female, single" from among the two selected candidates. In this way, it is also possible to combine criteria (1) and (3) in a form where candidates are selected by criterion (1) and then narrowed down by criterion (3).

[0047] Furthermore, combinations of criteria (1) and (2), combinations of criteria (2) and (3), and combinations of all three criteria (1), (2), and (3) are also possible. Alternatively, in addition to the criteria in (1), (2), and (3), audio advertisements may also be selected based on ambient noises around the smart speaker 10 or conversations between users that the smart speaker 10 has picked up. For example, if the smart speaker 10 picks up a conversation between a user and another user such as "We're out of tissues," and "Yeah, let's order some from an e-commerce site," the ad selection unit 422 may extract the keyword "tissues" from the conversation and select an audio advertisement promoting the extracted "tissues." Also, for example, if the smart speaker 10 picks up the sounds of a dog or cat barking, the ad selection unit 422 may identify the keywords "dog" and "cat" from the sounds and select an audio advertisement promoting dog food or cat food related to the identified "dog" and "cat."

[0048] The ad adjustment unit 424 determines whether or not to adjust the length of the audio ad selected by the ad selection unit 422. If it determines that adjustment is necessary, the ad adjustment unit 424 extracts a portion (for example, a relatively important portion) from the selected audio ad by applying a predetermined extraction algorithm to the audio ad. The ad adjustment unit 424 may also determine whether or not to adjust the length of the audio ad based on the content of the audio content acquired by the content acquisition unit 420 and / or the state of the current conversation session between the smart speaker 10 and the user.

[0049] For example, the ad adjustment unit 424 may refer to the session information holding unit 410 and decide to adjust the ad if the session status corresponding to the selected audio ad is "in conversation," and decide not to adjust if it is "waiting for utterance." Alternatively, the ad adjustment unit 424 may determine the length of the audio ad based on the content of the audio content acquired by the content acquisition unit 420. For example, the ad adjustment unit 424 may determine the length of the audio ad to match the playback time of the audio content. For relatively long audio content, the audio ad may be played multiple times. Also, for example, if the audio content is informational content such as news or weather forecasts, the ad adjustment unit 424 may decide to adjust the ad because there is a high probability that the user wants to get the desired information as quickly as possible.

[0050] Techniques are known for extracting important parts of audio advertisements, such as sections where a person is speaking, sections where a person is speaking loudly, and sections where the background sound gradually increases. These techniques are used in hard disk recorders and the like. A predetermined extraction algorithm may be constructed using this known technique. The advertisement adjustment unit 424 may adjust the volume of the audio advertisement in addition to, or instead of, the length of the audio advertisement.

[0051] The transmission information generation unit 426 combines the audio content acquired by the content acquisition unit 420 and the audio advertisement selected by the advertisement selection unit 422 to generate a single transmission information. If the length of the audio advertisement has been adjusted by the advertisement adjustment unit 424, the audio advertisement whose length has been adjusted by the advertisement adjustment unit 424 is used instead of the audio advertisement selected by the advertisement selection unit 422. The transmission information generation unit 426 configures the transmission information such that when the transmission information is received and played back by the smart speaker 10, the playback of the audio advertisement precedes the playback of the audio content. For example, if the transmission information includes a header, audio content, and audio advertisement, the transmission information generation unit 426 may generate the transmission information so that the header, audio advertisement, and audio content are arranged in that order.

[0052] The transmitting unit 428 transmits the transmission information generated by the transmission information generation unit 426 to the smart speaker 10 via the network 6. When the smart speaker 10 receives the transmission information via the network 6, it first plays the audio advertisement included in the transmission information, and then plays the audio content included in the transmission information. Alternatively, the timing control unit 430, described later, may control the audio output from the smart speaker 10 via the network 6. In this case, the timing control unit 430 causes the smart speaker 10 to first output the audio advertisement included in the transmission information, and then output the audio content included in the transmission information. In any case, the audio advertisement and the audio content are played consecutively. That is, there is no other audio between the audio advertisement and the audio content. In particular, the audio advertisement is played immediately before the audio content.

[0053] Alternatively, audio advertisements may be embedded within audio content, or they may be played after the audio content has been output.

[0054] The timing control unit 430 communicates with the smart speaker 10 via the network 6 and controls the timing of the audio output from the smart speaker 10. The timing control unit 430 refers to the session information storage unit 410 and determines whether or not a user is present around the smart speaker 10. If the session ID of the dialogue session between the smart speaker 10 and the user is stored in the session information storage unit 410, the timing control unit 430 determines that a user is present around the smart speaker 10, or detects the presence of a user. If no such session ID is stored in the session information storage unit 410, the timing control unit 430 determines that no user is present around the smart speaker 10.

[0055] The timing control unit 430 restricts the output of audio advertisements from the smart speaker 10 if the presence of a user is not detected in the vicinity of the smart speaker 10, or if the state of the dialogue session between the smart speaker 10 and the user, as held in the session information holding unit 410, is "speaking" or "conversing". When the presence of a user is detected, the timing control unit 430 allows the output of audio advertisements from the smart speaker 10. When the state of the dialogue session changes to "waiting for utterance", the timing control unit 430 allows the output of audio advertisements from the smart speaker 10.

[0056] The timing control unit 430 controls the timing of the output of audio advertisements so that the output of the TV 12 associated with the smart speaker 10 and the audio advertisements output from the smart speaker 10 are coordinated. For example, the timing control unit 430 controls the smart speaker 10 so that when the advertisement "Continue listening on the speaker!" finishes playing on the TV 12, the smart speaker 10 starts outputting an audio advertisement saying "This product featured on TV is..." In this case, the timing control unit 430 obtains the channel number currently being broadcast from the TV 12 via the network 6. The timing control unit 430 has previously obtained the broadcast schedule from the server of another service. From the obtained channel number and broadcast schedule, the timing control unit 430 can identify the content of the advertisement to be broadcast on the TV 12, its start timing, and its end timing. The timing control unit 430 selects an audio advertisement related to the identified content from the audio advertisement holding unit 404 and sends it to the smart speaker 10 (with or without accompanying audio content). The timing control unit 430 controls the smart speaker 10 to start outputting the transmitted audio advertisement at the same time as the advertisement played on TV 12 ends.

[0057] The attribute update unit 432 updates the user's account attributes using voice information collected by the smart speaker 10. For example, if the account ID shown in the user information storage unit 408 in Figure 8 belongs to a search service account, this account ID is assigned to users who visit the search service site. The user's account attributes are derived from information about what the user is searching for and registered as attributes in the user information storage unit 408 in Figure 8.

[0058] Meanwhile, the management server 4 can understand the user's preferences by analyzing the content of the conversation between the smart speaker 10 and the user, as well as the content of conversations between the user and other users. The attribute update unit 432 updates the attributes of the user information storage unit 408 with the preferences thus understood via the smart speaker 10. The attribute update unit 432 may also provide the updated content to the search service server in association with the account ID. This allows the search service to obtain user preferences that it could not previously obtain, and to optimize the browser's ad output using these preferences. The user preferences and attributes obtained via the smart speaker 10 and the preferences and attributes understood by the search service may be stored in a way that allows them to be identified separately. This makes it possible to select audio ads using only one of the preferences and attributes.

[0059] Furthermore, when an audio advertisement is output from the smart speaker 10, the attribute update unit 432 determines whether or not it has received a direct or indirect expression of disinterest from the user via the smart speaker 10. If such an expression has been received, the attribute update unit 432 registers the ad ID of the corresponding outputted audio advertisement as an NG ad ID in the user information storage unit 408, associating it with the user's user ID. When selecting an audio advertisement, the ad selection unit 422 refers to the user information storage unit 408 and excludes audio advertisements identified by the NG ad ID stored in accordance with the authenticated user's user ID from the selection. The ad selection unit 422 may also exclude successor audio advertisements to the audio advertisement identified by the NG ad ID from the selection. The ad selection unit 422 may also exclude audio advertisements that correspond to or have the same attributes and keywords as the audio advertisement identified by the NG ad ID from the selection.

[0060] The operation of management server 4 with the above configuration will now be explained. Figure 10 is a flowchart showing the sequence of processes in the management server 4 shown in Figure 1. The management server 4 receives an audio signal from the smart speaker 10 via the network 6, representing a request for audio content delivery (S302). The management server 4 identifies the requested audio content by performing speech recognition processing on the received audio signal (S304). The management server 4 obtains the requested audio content from the audio content holding unit 402 or from an external source (S306). The management server 4 selects an audio advertisement from the audio advertisement holding unit 404 (S308). The management server 4 determines whether or not it is necessary to adjust the length of the selected audio advertisement (S310). If it is determined that adjustment is necessary (YES in S310), the management server 4 adjusts the length of the audio advertisement (S312). The management server 4 generates transmission information based on the audio content acquired in step S306 and the audio advertisement selected in step S308 (if NO in step S310) or the audio advertisement whose length was adjusted in step S312 (if YES in step S310) (S314). The management server 4 transmits the generated transmission information to the smart speaker 10 via the network 6 (S316). The management server 4 determines whether it is a suitable time to output the audio advertisement (S318). If it is a suitable time (YES in S318), the management server 4 first causes the smart speaker 10 to output the audio advertisement (S320), and then causes it to output the audio content (S322).

[0061] In the embodiments described above, examples of the holding portion include a hard disk and a semiconductor memory. Furthermore, it will be understood by those skilled in the art who have read this specification that each portion can be realized based on the description herein by a CPU (not shown), a module of an installed application program, a module of a system program, or a semiconductor memory that temporarily stores the contents of data read from the hard disk.

[0062] According to the management server 4 of this embodiment, audio advertisements are played in conjunction with the playback of audio content on the smart speaker 10. This makes it possible to provide audio advertisements that are timed to the delivery of audio content. In addition, in this embodiment, audio advertisements are played before the playback of audio content. In this case, the probability that users will hear the audio advertisement is higher than when the audio advertisement is played after the playback of audio content. Therefore, it becomes possible to provide more effective advertisements.

[0063] Furthermore, in the management server 4 according to this embodiment, audio advertisements are selected based on the content of the audio content, audio information obtained via the smart speaker 10, and the attributes of the authenticated user's account. Audio advertisements selected in this way are highly likely to align with the user's preferences and desires. Therefore, it is possible to provide audio advertisements that are more appealing to the user.

[0064] Furthermore, in the management server 4 according to this embodiment, the timing of the output of audio advertisements from the smart speaker 10 is appropriately controlled. Therefore, it becomes possible to output audio advertisements that do not interfere with the user's conversation or speech. Alternatively, it becomes possible to provide audio advertisements in cooperation with other electronic devices such as the TV 12.

[0065] In this embodiment, in relation to the fact that the smart speaker 10 can acquire ambient sounds and conversations between users, the management server 4 may determine whether or not child abuse is occurring by understanding the audio context. The management server 4 may edit the audio data related to the abuse and provide it to a designated investigative agency. The investigative agency can listen to the entire audio data.

[0066] In this embodiment, the user ID is identified by voiceprint authentication by the user authentication unit 416, and the attribute corresponding to this user ID is identified by referring to the user information storage unit 408. In this case, the management server 4 may change the manner of outputting audio content or audio advertisements according to the user's attributes. For example, if the user's attribute is a child, the management server 4 may use the simplest language possible in the audio content or audio advertisement, and delete or rephrase any offensive language. Alternatively, if the user's attribute is an elderly person, the management server 4 may increase the volume or make the pronunciation clearer in the audio content or audio advertisement.

[0067] In this embodiment, the case in which the user is authenticated by the user authentication unit 416 has been described, but it is not limited to this, and user authentication may not be required. In this case, the user's attributes will not be reflected in the selection of audio advertisements. Here, in this embodiment, the operation of outputting an audio advertisement in S320 has been described, but in addition to outputting this audio advertisement, the management server 4 can also store this output status. Examples of output statuses include "the target advertisement was played to the end", "the target advertisement was stopped midway", and "in addition to playing the target advertisement, additional information about the product targeted by the advertisement was output". "The target advertisement was played to the end" can be achieved by the smart speaker 10 reporting to the management server 4 when it has output the target audio data to the end. Furthermore, both stopping midway and outputting additional information can be controlled by the management server 4 and therefore can be managed.

[0068] The technical concept relating to this embodiment may be expressed by the following items. (Item 1) It has the ability to receive distribution requests via the network from a speaker equipped with a microphone and communication capabilities, A function to acquire audio content without images in response to received distribution requests, A feature to select audio ads without images from the audio ad retention methods, A computer program for a server to implement a function that combines acquired audio content and selected audio advertisements and transmits them to the speaker via the network. (Item 2) It accepts distribution requests via the network from a speaker equipped with a microphone and communication capabilities, In response to the received distribution request, the system will acquire audio content without accompanying images, Selecting audio ads without images from the audio ad retention methods, A method comprising transmitting the acquired audio content and selected audio advertisements together to the speaker via the network.

[0069] (Second Embodiment) In the second embodiment, multiple smart speakers are placed at different locations within a real space, and each of them is connected via a network to a management server similar to the management server 4 in the first embodiment.

[0070] Figure 11 is a schematic top view of user 204's room 202. This room 202 contains a fixed first smart speaker 208, a fixed second smart speaker 210, a fixed third smart speaker 212, a fixed fourth smart speaker 214, a movable fifth smart speaker 216, and a TV 206. Each smart speaker communicates with the management server via the network. Although five smart speakers are shown in Figure 11, there is no limit to the number of smart speakers. Each smart speaker may be installed on the walls, floor, or ceiling of room 202.

[0071] (1) Automatic determination of the smart speaker's position The management server records and manages the location of each smart speaker in room 202. This location may be registered by user 204 to the management server 4. Alternatively, the management server may automatically determine the location of each smart speaker using the microphones and speakers of the five smart speakers.

[0072] The management server determines the relative positions between smart speakers by having other smart speakers 10 detect the sound output by one smart speaker. For example, if the positions of the second smart speaker 210, the third smart speaker 212, and the fourth smart speaker 214 are known, and the position of the first smart speaker 208 is to be determined, the management server causes the speaker of the first smart speaker 208 to output a sound pulse of a predetermined wavelength. The management server obtains the time at which it received the sound pulse of the predetermined wavelength from each of the second smart speaker 210, the third smart speaker 212, and the fourth smart speaker 214. The management server calculates the propagation time of the pulse from the obtained time, and calculates the distance from the calculated propagation time and the speed of sound. The management server calculates the position of the first smart speaker 208 from each of the calculated distances and the known positions of the second, third, and fourth smart speakers 210, 212, and 214.

[0073] The fifth smart speaker 216 is, for example, a smart speaker attached to a robot and is capable of moving on its own. If the positions of the first, second, third, and fourth smart speakers 208, 210, 212, and 214 are known, the management server can track the position of the fifth smart speaker 216 by the position calculation process described above. In addition, if the fifth smart speaker 216 determines the positions of other smart speakers based on its own position, it may move to a position where it can easily receive the sound emitted by those smart speakers.

[0074] The system components shown in Figure 11 preferably include smart speakers or smartphones equipped with both a microphone and a speaker. However, televisions, radios, and other electrical devices that generally only have speakers can also provide audio playback support because they have a speaker. Some electrical devices also have communication capabilities. This allows for speaker output from multiple locations. The location of the electrical devices can be notified by the user by setting it on the management server, or it can be automatically determined by the management server through the location calculation process described above, or by radio waves from wireless communication.

[0075] (2) Applications of mobile smart speakers In audio output, the way the sound is heard by the target user may vary depending on the speaker's position. Therefore, the management server may control the fifth smart speaker 216 to move to a position where the sound is heard more appropriately.

[0076] Furthermore, devices like TV206 generally do not have microphone functionality and therefore cannot participate in the position calculation process described above. However, by having the fifth smart speaker 216 move to the position of TV206 and take over the microphone functionality of TV206, TV206 can also participate in the position calculation process.

[0077] (3) Audio output according to the user's location If the location of each smart speaker is known, then knowing the location of user 204 allows the management server to identify the smart speaker closest to user 204. When user 204 utters a voice output request such as "Turn on the TV," the five smart speakers convert the utterance into an audio signal and send it to the management server. The management server performs speech recognition processing on the audio signal to understand user 204's voice output request. The management server identifies the smart speaker corresponding to user 204's location in room 202. In particular, the management server identifies the second smart speaker 210, which is closest to user 204's location. At this time, the management server compares the volume of each smart speaker's microphone when it receives the user's utterance and identifies the second smart speaker 210, which has the highest volume, as the smart speaker closest to user 204's location. Alternatively, if user 204 is using a smartphone, the management server can determine user 204's location by obtaining the smartphone's current location.

[0078] The management server transmits the audio accompanying the video being played on TV 206 to the second smart speaker 210 identified as described above. In this way, user 204 can receive the audio output from TV 206 from the second smart speaker 210 closest to them.

[0079] Alternatively, if the smart speakers have directional speakers, the management server may control the directionality. For example, if the speakers of the first, second, third, and fourth smart speakers 208, 210, 212, and 214 have directional output, the management server controls each smart speaker so that the audio output of each speaker is directed towards the location of user 204. The management server transmits audio accompanying the video being broadcast on TV 206 to each smart speaker. In this case, each smart speaker outputs audio towards user 204.

[0080] Furthermore, a directional smart speaker may indicate the current direction of the audio output to the user in a visually identifiable manner. For example, a directional arrow may be displayed on the top surface of the smart speaker using an LED or the like.

[0081] According to this example, if, for example, a user who has instructed a smart speaker to play content and other users are in room 202, the audio can be output targeted to the user who gave the instruction. By configuring a server process to be assigned to each user, the first smart speaker can be controlled to face the location of the first user, and the second smart speaker can be controlled to face the location of the second user, so that the first and second smart speakers output audio simultaneously. Alternatively, by equipping a single smart speaker with multiple audio output devices and configuring the server to assign a server process to each user, a single smart speaker can be controlled to output audio directed to the location of the first user and audio directed to the location of the second user simultaneously.

[0082] The technical concept relating to this embodiment may be expressed by the following items. (Item 3) A system comprising multiple speakers, each having a microphone and communication function, A system configured to determine the relative position between speakers by having other speakers detect the sound output from one speaker. (Item 4) A server that communicates via a network with multiple speakers, each having a microphone and communication capabilities, wherein the multiple speakers are located at different positions within the same physical space. The aforementioned server, Means for receiving voice output requests from a user in the real space via one of the plurality of speakers, Means for identifying a speaker corresponding to the user's location, A server comprising means for transmitting audio content to a specified speaker. (Item 5) A server that communicates via a network with multiple speakers, each having a microphone and communication capabilities, wherein the multiple speakers are located at different positions within the same physical space. The aforementioned server, Means for receiving voice output requests from a user in the real space via one of the plurality of speakers, A server comprising means for controlling the directivity of at least one of the plurality of speakers so that sound is output towards the location of the user.

[0083] (Third embodiment) Figure 12 is a schematic diagram showing the configuration of a voice-operated system 232 according to a third embodiment. The voice-operated system 232 comprises a management server 234, a smart speaker 240, a TV 242, and a smartphone 248. The management server 234, smart speaker 240, TV 242, and smartphone 248 are connected to each other via a network 236 such as the Internet. Both the smart speaker 240 and TV 242 are installed in the user's room 244. The smart speaker 240 is configured to enable P2P communication 246 with the smartphone 248.

[0084] When operating TV242 by voice, user 238 speaks a command such as "Turn on the TV" to smart speaker 240. The microphone of smart speaker 240 converts the voice spoken by user 238 into an electrical signal, and smart speaker 240 transmits the resulting electrical signal as an audio signal to management server 234 via network 236. Management server 234 performs voice recognition processing on the received audio signal to understand that user 238 is requesting that TV242 be powered on. Management server 234 generates an instruction signal to perform the requested operation, i.e., to power on TV242, and transmits it to TV242 via network 236. When TV242 receives the instruction signal via network 236, it transitions from the power-off state to the power-on state.

[0085] Thus, control and operation via the smart speaker 10 are basically performed by voice. However, if there are other users in room 244 besides user 238, user 238 may be reluctant to use voice control. Also, even if user 238 is the only one in room 244, there may be cases where voice control is to be avoided depending on the control content. In such cases, the voice control system 232 makes it possible to control the smart speaker 240 system side (management server 234) via text through the smartphone 248.

[0086] To enable operation using the smartphone 248, the smartphone 248 downloads and installs an application specifically for the management server 234. When the application is launched on the smartphone 248, it obtains the URL of the management server 234 from the smart speaker 240 via P2P communication 246 or the local network. The application uses the obtained URL to establish a connection with the management server 234. Once the connection between the smartphone 248 and the management server 234 is established, the operation content entered into the smartphone 248 is transmitted to the management server 234 through that connection. The management server 234 generates and transmits instruction signals to implement the received operation content.

[0087] For example, when the management server 234 receives the text string "Turn on the TV" from the smartphone 248, it parses the received text string and understands that user 238 is requesting that the TV 242 be powered on. The management server 234 generates an instruction signal to perform the requested operation, i.e., to power on the TV 242, and sends it to the TV 242 via the network 236. When the TV 242 receives the instruction signal via the network 236, it transitions from the power-off state to the power-on state.

[0088] According to the voice control system 232 of this embodiment, the user 238 can switch between voice control and control via the smartphone 248 depending on the situation.

[0089] In this embodiment, the case in which the management server 234 receives operation instructions from the user 238 via the smart speaker 240 or smartphone 248 has been described, but it is not limited to this, and for example, as in the first embodiment, the management server 234 may receive requests for the distribution of audio content from the user 238 via the smart speaker 240 or smartphone 248.

[0090] The technical concept relating to this embodiment may be expressed by the following items. (Item 6) A server that communicates via a network with a speaker having a microphone and communication capabilities, A means for processing requests received from the user via the microphone of the speaker, A server comprising means for processing requests received from the user via other electronic devices that communicate with the speaker.

[0091] (Fourth embodiment) The fourth embodiment relates to layering in the management server. In this embodiment, the server process (or server) is changed depending on the speaker. The server can recognize who is speaking through voiceprint authentication. The server process (or server) that the target user normally uses performs the processing.

[0092] Figure 13 is a schematic diagram showing the configuration of a voice-operated system 252 according to the fourth embodiment. The voice-operated system 252 comprises a management server 254, a smart speaker 260, and a TV 262. The management server 254, the smart speaker 260, and the TV 262 are connected to each other via a network 256 such as the Internet. Both the smart speaker 260 and the TV 262 are installed in room 264, and there are three users (first user 266, second user 268, and third user 270) in room 264.

[0093] When operating TV262 by voice, the first user 266 speaks a command such as "Turn on the TV" towards the smart speaker 260. The microphone of the smart speaker 260 converts the voice spoken by the first user 266 into an electrical signal, and the smart speaker 260 transmits the resulting electrical signal as an audio signal to the management server 254 via the network 256. The management server 254 performs voiceprint authentication on the received audio signal to identify the first user 266. The management server 254 selects a server process corresponding to the identified first user 266, and the selected server process processes subsequent requests. When TV262 receives an instruction signal from the selected server process of the management server 254 via the network 236, it transitions from the power-off state to the power-on state. In the management server 254, a different server process is assigned to the voices of a second user 268 or a third user 270, who are different from the first user 266.

[0094] Figure 14 is a block diagram showing the functions and configuration of the management server 254 in Figure 13. Each block shown here can be implemented in hardware terms by components and mechanical devices such as the CPU of a computer, and in software terms by computer programs, etc., but here it depicts a functional block that is realized through the cooperation of these. Therefore, it will be understood by those skilled in the art who have read this specification that these functional blocks can be implemented in various forms by combinations of hardware and software.

[0095] The management server 254 comprises a user information holding unit 272, an audio signal receiving unit 274, a user authentication unit 276, and a group of server processes 278. The group of server processes 278 includes multiple server processes SP1, SP2, SP3, ... each assigned to a specific user. The following describes the case where each user has a different server process, but in other embodiments, each user may have a different server. If the servers are different, the server processes will naturally be different as well.

[0096] Figure 15 is a data structure diagram showing an example of the user information storage unit 272 in Figure 14. The user information storage unit 272 stores the user ID, the user's voiceprint data, and the ID of the server process assigned to the user in association with each other.

[0097] Returning to Figure 14, the voice signal receiving unit 274 receives a voice signal from the smart speaker 260 via the network 256 that represents the content of a speech by one of the three users 266, 268, or 270.

[0098] The user authentication unit 276 extracts or acquires a voiceprint from the voice signal received by the voice signal receiving unit 274. The user authentication unit 276 performs voiceprint authentication based on the extracted voiceprint. The user authentication unit 276 refers to the user information holding unit 272 and determines whether there is a voiceprint in the voiceprints held in the user information holding unit 272 that matches the extracted voiceprint. If there is a matching voiceprint, the user authentication unit 276 identifies the user ID and server process ID corresponding to that voiceprint. If there is no matching voiceprint, the user authentication unit 276 generates output indicating no match or unknown user.

[0099] When the user authentication unit 276 identifies a server process ID, the server process with the identified server process ID among the server processes included in the server process group 278 is started. The started server process then performs subsequent processing on the audio signal received by the audio signal receiving unit 274.

[0100] Each server process included in the server process group 278 implements an electronic device operation function as described in the third embodiment. In other embodiments, the server process may implement, for example, an audio content distribution function as described in the first embodiment.

[0101] For example, if User 1 266 is a resident of Room 264, User 1 266's server process is granted the authority to control TV 262. If User 2 268 and User 3 270 are visitors to User 1 266's Room 264, their server processes are not granted the authority to control TV 262. Therefore, if User 2 268 or User 3 270 wants to control TV 262 by voice, their server process sends a request to User 1 266's server process to control TV 262. The server process that receives the request asks User 1 266 if it is okay to perform the operation, and if it receives consent from User 1 266, it executes the operation.

[0102] Alternatively, the first user 266, who owns the smart speaker 260, may set permissions for the second user 268 and the third user 270, who are guests (visitors). For example, they may be allowed to control the lights via the smart speaker 260, but not be allowed to make purchases on e-commerce sites via the smart speaker 260.

[0103] According to the voice control system 252 of this embodiment, by assigning a server process to each user, it is possible to differentiate the operations that can be performed, the information that can be accessed, and the permissions for each user.

[0104] In this embodiment, we have described a case where the management server 254 has multiple server processes, and the management server 254 receives the voice signal, performs voiceprint authentication, and identifies the server process to be used, but we are not limited to this. For example, if there are multiple servers, the voice signal from the smart speaker 260 may be sent to all servers, and voiceprint authentication may be performed at each server. Alternatively, a server or server process for any one user, or any one server or server process, may receive the voice signal, extract only the voice signal of the target user, and forward it to the target server or server process.

[0105] In this embodiment, we have described a case where the first user 266 issuing the operation command and the TV 262 to be operated are in the same room 264, but we are not limited to this, and remote control of the electronic device to be operated may also be possible. For example, suppose a second user 268 is visiting the first user 266's room 264 and the second user 268 wants to turn on the air conditioner in their own room (different from room 264). The second user 268 speaks the command "Turn on the air conditioner in my room" to the smart speaker 260. The management server 254 understands the second user 268's request through voiceprint authentication and speech recognition. The management server 254 forwards the second user 268's request to another management server connected to the smart speaker in the second user 268's room. At this time, an audio signal is attached as authentication data.

[0106] The technical concept relating to this embodiment may be expressed by the following items. (Item 7) A server that communicates via a network with a speaker having a microphone and communication capabilities, A means for identifying the speaker by analyzing the audio signal acquired through the microphone of the speaker, A server comprising means for processing the audio signal using a server process assigned to a specific speaker.

[0107] In addition to the smart speaker, the above embodiments also describe operation in cooperation with a network-connected television or computer. After outputting an audio advertisement on the smart speaker, additional information can be displayed on the television or computer in response to the user's voice control. This is achieved when the management server 4 sends a URL to be displayed on the television or computer. The parameters of this URL can also include information indicating that an additional information request has been made by the smart speaker or the management server 4. This allows the receiving system to recognize that the request originated from the management server 4 or the smart speaker. Here, URL parameters were used as an example of a transmission method, but notification may be done by other methods.

[0108] The configuration and operation of the system according to the embodiments have been described above. These embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of each component and each process, and that such modifications also fall within the scope of the present invention. Combinations of different embodiments are also possible. [Explanation of symbols]

[0109] 2 audio advertising distribution system, 4 management server, 6 network, 8 users, 10 smart speaker.

Claims

1. A receiving means for receiving voice information acquired via a network from a speaker having a microphone and communication function, A speech recognition means that recognizes the content of the user's speech based on the received voice information, A voice information holding means that holds the recognized speech content, An acquisition means for acquiring audio content without images in response to a distribution request contained in the recognized speech content, A selection method for selecting audio advertisements without images from the audio advertisement retention methods, A transmission means that transmits the acquired audio content and the selected audio advertisement together to the speaker via the network, The system includes a timing control unit that controls the timing of the output of the audio advertisement from the speaker, The selection means is a server that selects an audio advertisement based on the spoken content held in the audio information holding means.

2. The server according to claim 1, wherein the selection means selects an audio advertisement based on the user's past utterances in the current dialogue session between the speaker and the user, which are held in the audio information holding means.

3. The server according to claim 1 or 2, wherein the selection means selects an audio advertisement based on the least recent utterance among a plurality of utterances made by the user included in the current dialogue session between the speaker and the user.

4. The server according to any one of claims 1 to 3, further comprising determination means for determining the state of the current dialogue session between the speaker and the user based on the recognized utterance.

5. The server according to claim 4, wherein the determination means determines whether the recognized utterance indicates the end of the dialogue session, and if the recognized utterance indicates the end of the dialogue session, it terminates the current dialogue session.

6. The server according to claim 1, wherein the speaker, upon receiving the acquired audio content and the selected audio advertisement via the network, plays the audio advertisement and then plays the audio content.

7. The server according to claim 1, wherein the timing control unit determines that there is no user around the speaker and restricts the output of the audio advertisement from the speaker.

8. The server according to claim 7, wherein the timing control unit permits the output of the audio advertisement from the speaker when it determines that a user is present around the speaker.

9. The server according to claim 1, wherein the timing control unit controls the timing of the audio advertisement according to the state of the dialogue session.

10. The server according to claim 9, wherein the timing control unit restricts the output of the audio advertisement when it is not in a dialogue session.

11. The server according to claim 9, wherein the timing control unit restricts the output of the voice advertisement from the speaker when the state of the interaction session with the user is "speaking" or "conversing".

12. The server according to claim 9, wherein the timing control unit permits the output of the voice advertisement from the speaker when the state of the dialogue session is "waiting for utterance".