Information processing apparatus and control method

The information processing device uses AI to analyze user speech and provide audio assistance, addressing the challenge of enhancing two-party communication in interactive agent systems by balancing participation and introducing relevant topics, thereby facilitating smoother conversations.

JP2025170112APending Publication Date: 2025-11-14SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025151970
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2025-09-12
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Conventional interactive agent systems primarily facilitate one-to-one user interactions, lacking the ability to enhance communication between two parties effectively.

Method used

An information processing device and method that analyzes the speech of two users, using AI to detect conversation patterns and intervene with appropriate audio assistance to balance participation and introduce relevant topics, leveraging network services and environmental factors to facilitate smoother dialogue.

Benefits of technology

Enables smooth communication between two parties by balancing speaking times, introducing relevant topics, and providing contextually appropriate audio assistance, enhancing user engagement and conversation flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170112000001_ABST
    Figure 2025170112000001_ABST
Patent Text Reader

Abstract

To facilitate smooth communication between two parties.SOLUTION: An information processing apparatus of this technique analyzes utterances of each of two users who are conversing, detected by one or more agent apparatuses used by the two users, and outputs audio having a content about a network service from the one or more agent apparatuses as conversation assistance audio which is audio that assists the conversation, when a word related to the network service used by the one or more of the two users is included in the utterances of the one or more of the two users. This technique can be applied, for example, to a server that controls an operation of an interactive robot used by two people who have the conversation remotely.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] In particular, the present technology relates to an information processing device and a control method that enable smooth communication between two parties. [Background technology]

[0002] In recent years, interactive agents that support users by responding to their spoken questions have become popular. By talking to devices equipped with such interactive agent functionality, users can get weather forecasts, play music, check their schedules, and more.

[0003] Patent Document 1 describes an interactive agent system that collects personal information in a conversational format and proposes appropriate products and the like to individual users based on the collected personal information.

[0004] Non-Patent Document 1 discloses a matching service in which video calls are made via a third party called a matchmaker. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-52449 [Non-patent literature]

[0006] [Non-Patent Document 1] "Yi Dui",Internet,<URL https: / / www.520yidui.com / > ,Retrieved March 16, 2020 Summary of the Invention [Problem to be solved by the invention]

[0007] In conventional interactive agent systems, the relationship between the user and the system is generally one-to-one, and the system responds to questions from the user.

[0008] This technology was developed in light of these circumstances, and enables smooth communication between two parties. [Means for solving the problem]

[0009] An information processing device according to one aspect of the present technology includes an analysis unit that analyzes the speech of each of two users who are having a conversation, detected by one or more agent devices used by the two users, and a control unit that, when a word related to a network service used by one or more of the two users is included in the speech of one or more of the two users, outputs audio related to the network service from one or more of the agent devices as conversation assistance audio, which is audio that assists the conversation.

[0010] In one aspect of the present technology, the speech of each of two users who are having a conversation, detected by one or more agent devices used by the two users, is analyzed, and if the speech of one or more of the two users contains words related to a network service used by the one or more users, a process is performed in which audio of content related to the network service is output from one or more of the agent devices as conversation assistance audio, which is audio that assists the conversation. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating a configuration example of a voice communication system according to an embodiment of the present technology. [Figure 2] FIG. 10 is a diagram illustrating an example of an output of an assist utterance. [Figure 3] FIG. 1 is a diagram illustrating an example of AI that realizes a conversation assistance function. [Figure 4] FIG. 10 is a diagram showing a conversation. [Figure 5] FIG. 2 is an enlarged perspective view showing the external appearance of the interactive agent device. [Figure 6] 10A and 10B are diagrams showing an example of bottle attachment. [Figure 7] FIG. 10 is a diagram showing an example of a display of a drinking record. [Figure 8] FIG. 10 is a diagram showing an example of a display of a conversation record. [Figure 9] FIG. 10 is a diagram showing a specific example of a conversation between user A and user B. [Figure 10] FIG. 10 is a diagram showing a specific example of a conversation following FIG. 9. [Figure 11] FIG. 11 is a diagram showing a specific example of a conversation following FIG. 10. [Figure 12] FIG. 10 is a diagram showing a specific example of a conversation between user C and user D. [Figure 13] FIG. 13 is a diagram showing a specific example of a conversation following FIG. 12. [Figure 14] FIG. 10 is a diagram showing a specific example of a conversation between user A and user B. [Figure 15] FIG. 15 is a diagram showing a specific example of a conversation following FIG. 14. [Figure 16] FIG. 10 is a diagram showing a specific example of a conversation between user A and user B. [Figure 17] FIG. 10 is a diagram illustrating an example of matching. [Figure 18] FIG. 2 is a block diagram showing an example of the configuration of an interactive agent device. [Figure 19] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a communication management server. [Figure 20] FIG. 2 is a block diagram illustrating an example of a functional configuration of a communication management server. [Figure 21] 10 is a flowchart illustrating processing of a communication management server. [Figure 22] 10 is a flowchart illustrating processing of the interactive agent device. [Figure 23] FIG. 1 is a diagram illustrating an example of use of an interactive agent device. DETAILED DESCRIPTION OF THE INVENTION

[0012] <Overview of this technology> The server that manages the voice communication system of this technology is an information processing device that realizes smooth conversation between two people using a conversation assistance function using AI (Artificial Intelligence).The conversation assistance function outputs utterances from the system and prompts the user in the conversation to speak.

[0013] For example, the speaking time of each user during a conversation between two people is measured. If there is a difference in the speaking time of each user, the system will prompt the user who has spoken less to speak. The phrases spoken by the system are selected from preset phrases. For example, a phrase including the user's account name, such as "What do you think, Mr. A?", is output as the system's utterance.

[0014] The system also measures the amount of silence between the two users. If a certain period of silence, such as 10 seconds, occurs, the system will provide a new topic. For example, the system will extract the latest article with a title that matches a topic of mutual interest from a news site on the web, and provide content related to that article as a new topic.

[0015] In other words, the voice communication system of this technology has a user:AI = 2:1 ratio, and AI plays a role in assisting communication between users. Dedicated hardware for voice input and output is provided near each user. In addition, detailed settings and functions for checking conversation archives are provided by dedicated applications installed on each user's mobile device, such as a smartphone.

[0016] Hereinafter, embodiments of the present technology will be described in the following order. 1. Voice communication system configuration 2. External configuration of the interactive agent device 3. About the dedicated application 4. Specific examples of conversations including assisted speech 5. Configuration examples of each device 6. Operation of each device 7.Other

[0017] <Configuration of voice communication system> FIG. 1 is a diagram showing an example of the configuration of a voice communication system according to an embodiment of the present technology.

[0018] 1 is configured by two interactive agent devices 1, ie, interactive agent devices 1A and 1B, being connected via a network 21. A communication management server 11 is also connected to the network 21, which may be the Internet or the like.

[0019] The interactive agent device 1A is a device used by user A and is installed in the home of user A. Similarly, the interactive agent device 1B is a device used by user B and is installed in the home of user B. Although two interactive agent devices 1 are shown in FIG. 1, in reality, many more interactive agent devices 1 are connected to the network 21.

[0020] 1, users A and B have mobile terminals 2A and 2B, such as smartphones, respectively. The mobile terminals 2A and 2B are also connected to the network 21.

[0021] The interactive agent device 1 is a device with an interactive agent function that allows voice communication with a user. The interactive agent device 1 is provided with a microphone for detecting the user's voice, a speaker for outputting the voices of other users, etc. The agent function of the interactive agent device 1 is realized by the interactive agent device 1 and a communication management server 11 cooperating as appropriate. Various types of information are sent and received between the interactive agent device 1 and the communication management server 11.

[0022] For example, a conversation between two matched users is realized by the agent function of the interactive agent device 1. User A and user B shown in FIG. 1 are users matched by the communication management server 11.

[0023] The voice of user A is collected by the interactive agent device 1A and transmitted to the interactive agent device 1B via the communication management server 11. The interactive agent device 1B outputs the voice of user A transmitted via the communication management server 11.

[0024] Similarly, the voice of user B is collected by the interactive agent device 1B and transmitted to the interactive agent device 1A via the communication management server 11. The interactive agent device 1A outputs the voice of user B transmitted via the communication management server 11. This allows user A and user B to have a remote conversation from their respective homes.

[0025] During a conversation between user A and user B, utterances to assist (support) the conversation between the two are transmitted as system-side utterances from the communication management server 11 to the interactive agent device 1A and the interactive agent device 1B, and are output by the interactive agent device 1A and the interactive agent device 1B, respectively. User A and user B will listen to the system-side utterances and react to them.

[0026] That is, the communication management server 11 has a conversation assist function that not only matches two people to have a conversation, but also analyzes the conversation situation between the two people and makes utterances to assist the conversation between the two people according to the conversation situation between the two people. Hereinafter, utterances that the communication management server 11 causes the interactive agent device 1 to output by the conversation assist function will be referred to as assist utterances as appropriate. Assist utterances are conversation assist sounds that assist the conversation.

[0027] FIG. 2 is a diagram illustrating an example of an output of an assisted utterance.

[0028] The upper part of Fig. 2 shows the state of user A and user B engaged in an active conversation. Although the interactive agent device 1 and other devices are not shown in Fig. 2, the utterances of each user are transmitted from the interactive agent device 1 used by that user to the interactive agent device 1 used by the other user and output therefrom.

[0029] When the conversation between user A and user B is interrupted as shown in the middle part of Fig. 2, an assist utterance is output from the interactive agent device 1A and the interactive agent device 1B as shown in the bottom part of Fig. 2. In the example of Fig. 2, an assist utterance is output to encourage the two users to start a conversation by talking about "baseball," a common hobby of the two users. User A and user B will resume their conversation by talking about "baseball."

[0030] In this way, the communication management server 11 analyzes the conversation situation, such as whether the conversation has been interrupted, and outputs an assist utterance based on the analysis result. The conversation assistance function is realized by an AI provided in the communication management server 11. The communication management server 11 is managed by, for example, the manufacturer of the interactive agent device 1.

[0031] FIG. 3 is a diagram showing an example of AI that realizes a conversation assistance function.

[0032] As shown in the upper part of Fig. 3, the communication management server 11 is provided with a conversation assistance AI, which is an AI that realizes a conversation assistance function. The conversation assistance AI is an inference model configured with a neural network or the like that receives as input the conversation situation and personal information of each of user A and user B, such as hobbies and interests, and outputs the content to be provided as a topic of conversation. The conversation situation includes the speaking times of each of user A and user B, periods of silence (periods when the conversation was interrupted), etc.

[0033] The inference models that make up the conversation assistance AI are generated through machine learning using information representing various conversation situations, personal information of various users, and information on news articles obtained from news sites.

[0034] As shown by dashed lines #1 and #2, the interactive agent device 1A and the interactive agent device 1B are each connected to a conversation assistance AI. The conversation assistance AI analyzes the situation of the conversation between the two people based on information transmitted from the interactive agent device 1A and the interactive agent device 1B, and provides topics of conversation using the conversation assistance function as appropriate.

[0035] As shown in the lower part of Fig. 3, User A and User B each input profile information such as topics of interest (events, topics) in advance using their own mobile terminal 2 on which a dedicated application is installed. When User A and User B launch the dedicated application and log in by entering account information, the profile information of User A and User B that has been linked to the account information and managed in the communication management server 11 is identified.

[0036] A conversation between two people using such a conversation assistance function takes place, for example, when two users are drinking alcoholic beverages prepared by the interactive agent device 1 at their respective homes. That is, the interactive agent device 1 is provided with a function to provide alcoholic beverages in response to a user's request. Assist utterances are output according to the situation of the conversation after alcoholic beverages have been provided to the two users by the interactive agent devices 1.

[0037] User A and User B will each have a one-on-one conversation at home while drinking alcohol prepared by the interactive agent device 1. Since assist utterances, which are utterances by a third party, are inserted into the one-on-one conversation between User A and User B as appropriate depending on the situation of the conversation, the situation in which User A and User B are having a conversation becomes similar to the situation in which they are facing a bartender who joins the conversation at the appropriate time, as shown in Figure 4.

[0038] User A and User B can have a conversation while drinking alcohol with the support of assisted speech, enabling smooth communication.

[0039] 4 shows a situation in which user A and user B are sitting next to each other, but in reality, user A and user B are in their respective homes and talking to the interactive agent device 1. The interactive agent device 1, which plays the role of a bartender by joining in on one-on-one conversation at appropriate times to create the feeling of being in a bar, can also be called a bartender robot.

[0040] <External view of the interactive agent device> FIG. 5 is an enlarged perspective view showing the external appearance of the interactive agent device 1. As shown in FIG.

[0041] As shown in Fig. 5, the interactive agent device 1 has a vertically long, approximately rectangular parallelepiped housing 51 with a gently sloping upper surface. A recessed portion 51A is formed in the upper surface of the housing 51. A bottle 61 containing alcoholic beverages such as whiskey is attached to the recessed portion 51A, as shown by the arrow in Fig. 6.

[0042] Furthermore, a rectangular opening 51B is formed at the bottom front of the housing 51. Opening 51B is used as an outlet for a glass 62. A glass 62 is placed in opening 51B, and alcohol from a bottle 61 is poured into the glass 62 in response to a request for alcohol from a user. A server mechanism that automatically pours alcohol is also provided inside the housing 51.

[0043] When the bottle 61 becomes empty, the user can attach a new bottle 61 that has been delivered to the recessed portion 51A, thereby continuing to use the interactive agent device 1. For example, as a service for users of the interactive agent device 1, a subscription service for alcoholic beverages is provided in which bottles 61 are delivered periodically.

[0044] The side of the housing 51 is provided with openings for ice, water to be used as a mixer, carbonated water, etc. The user can try various drinking styles such as straight, on the rocks, or highball by making a voice request. The interactive agent device 1 is provided with recipe data that controls the server mechanism to reproduce the bartender's pouring method.

[0045] <About the dedicated application> As described above, a dedicated application for the voice communication system is installed on each mobile terminal 2. The dedicated application is prepared by the manufacturer of the interactive agent device 1, for example.

[0046] The user operates a dedicated application to register profile information such as age, address, hobbies, etc. The registered profile information is sent to the communication management server 11 and managed in association with the user's account information.

[0047] 7 and 8 are diagrams showing examples of screens of the dedicated application.

[0048] The dedicated application screen has a drinking record tab T1 and a conversation record tab T2. When the drinking record tab T1 is tapped, the drinking record is displayed as shown in Figure 7. In the example of Figure 7, information such as the date and time of drinking, the amount of alcohol consumed, and how it was consumed is displayed as the drinking record.

[0049] On the other hand, when the conversation record tab T2 is tapped, the conversation record is displayed as shown in Fig. 8. In the example of Fig. 8, information such as the name of the other person, the date and time of the conversation, and tags indicating the content of the conversation are displayed as the conversation record.

[0050] The function of displaying such drinking records and conversation records is realized based on information managed by the communication management server 11. The dedicated application communicates with the communication management server 11 and displays various screens based on the information sent from the communication management server 11.

[0051] <Examples of conversations including assisted speech> Here, a specific example of a conversation between two people in a voice communication system will be described.

[0052] 1. Assisted speech based on the conversation situation (1) Assisted speech according to speech duration For example, if the speaking time of user B is longer than that of user A, the following assist utterance is output, which uses a fixed phrase to address user A: "What do you think, A?" (a request for B's opinion on what he said) "What do you like, A?" (a question to A) "What has A been doing lately?" (topic change utterance)

[0053] Such an assist utterance is output when there is a large difference between the speaking time of user A and the speaking time of user B, such as when user B's speaking time exceeds 80% of the total. In the specific examples of assist utterances, "Mr. A" represents user A, and "Mr. B" represents user B.

[0054] (2) Assisted speech in response to silence periods If neither user speaks for a certain period of time, such as 10 seconds, the following assist utterance is output to provide a topic of conversation: "Do you know about (news title)?" (Utterances encouraging continuation and in-depth discussion) "It's (news title)." (Informative speech)

[0055] Such assisted speech is generated by searching the web for news articles related to the most frequently occurring words in the last 10 minutes of conversation, and includes, for example, the titles of the latest news articles that are attracting attention on a news site.

[0056] 9 to 11 are diagrams showing specific examples of conversations between users A and B. FIG.

[0057] 9 to 11, the utterances shown in the left column represent utterances made by user A, and the utterances shown in the right column represent utterances made by user B. The utterances shown in the center are utterances from the system (system utterances) output from the interactive agent device 1 under the control of the communication management server 11. The system utterances also include the above-mentioned assist utterances. The same applies to the figures described later that show other specific examples of conversations.

[0058] A conversation between user A and user B starts when a system utterance S1 such as "Mr. A, Mr. B is calling you" is output from the interactive agent device 1A, and user A, hearing the system utterance S1, agrees to start a conversation with user B.

[0059] The system utterance S1 is an utterance that informs user A that user B wishes to start a conversation with user A. For example, the system utterance S1 is output when user A is selected by user B from among the conversation partner candidates matched by the communication management server 11.

[0060] The matching by the communication management server 11 is performed based on topics of interest, such as "economy" or "entertainment," that are registered in advance by each user. However, matching may also be performed based on text data entered when selecting a conversation partner, rather than on pre-registered topics. This allows each user to select a user who shares a common topic of interest as a conversation partner.

[0061] In the example of FIG. 9, from time t1 to time t2, user A utters "Yes, please," and from time t2 to time t3, user B utters "Nice to meet you, nice to meet you. I see you like baseball, too, A." User A's voice data is transmitted from interactive agent device 1A to interactive agent device 1B via the communication management server 11 and is output as user A's utterance in interactive agent device 1B. On the other hand, user B's voice data is transmitted from interactive agent device 1B to interactive agent device 1A via the communication management server 11 and is output as user B's utterance in interactive agent device 1A.

[0062] The communication management server 11 measures the speech time of user A and the speech time of user B as the speech status of user A and user B. In the central band shown in Fig. 9, the hatched section represents the speech time of user A, and the lightly shaded section represents the speech time of user B. The same applies to the other figures.

[0063] Furthermore, the communication management server 11 extracts keywords from the speech of user A and the speech of user B as the context of the speech of user A and user B. The words enclosed in boxes in Fig. 9 are words extracted by the communication management server 11 as keywords.

[0064] After time t3, user A and user B alternately speak, and the conversation continues between user A and user B. In the examples of FIGS. 9 and 10, user B speaks for a longer period of time than user A.

[0065] When the difference between the speaking time of user A and the speaking time of user B becomes larger than a threshold, a system utterance S2 such as "What do you like, A?" is output at time t12 in FIG. 10. The system utterance S2 is an assist utterance that uses a fixed phrase to ask a question to user A. For example, when the speaking time of user B exceeds 80% of the total conversation time between the two people, such an assist utterance is output.

[0066] The voice data of the system utterance S2 is transmitted from the communication management server 11 to both the interactive agent device 1A and the interactive agent device 1B, and is output as an assist utterance in each of the interactive agent device 1A and the interactive agent device 1B. User A, having heard the system utterance S2, responds to the conversation by uttering something like, "Um, I like Tokyo Skurna Hayabusazu" between time t13 and time t14.

[0067] The communication management server 11 provides an opportunity to speak to user A, who has a short speaking time, and balances the speaking time of user A with the speaking time of user B, thereby enabling smooth communication to be realized.

[0068] Between time t14 and time t17, the assist utterance is triggered, and user A and user B alternately speak.

[0069] As shown in the top part of Figure 11, when both User A and User B are silent and the conversation is interrupted for a certain period of time, such as 10 seconds, a system utterance S3 such as "What do you think about 'Tohoku being the 2019 Central League champion'?" is output. System utterance S3 is an assisted utterance that provides a topic of conversation for the two users after the silence has continued.

[0070] In this way, the communication management server 11 also measures the duration of silence between users A and B as part of the speech situation between the two users.

[0071] Of User A and User B, who received the topic provided by system utterance S3, User B will utter something like, "We were completely defeated this year. But next year, of course, it'll be Keihan!" between time t21 and time t22.

[0072] The communication management server 11 encourages two silent people to speak and starts a conversation, thereby enabling smooth communication to be realized.

[0073] Between time t22 and time t24, the user A and the user B alternately speak, triggered by the assist utterance.

[0074] For example, when a predetermined time such as one hour has passed, the system outputs a system utterance S4 such as "It's time to end the conversation. Thank you very much," as shown in the lower part of Fig. 11. After hearing the system utterance S4, User A and User B will exchange greetings and end the conversation.

[0075] In this way, during a conversation between user A and user B, the conversation situation between the two is analyzed in the communication management server 11. Assist utterances according to the conversation situation are output as appropriate, thereby realizing smooth communication between user A and user B.

[0076] 2. Assisted speech linked to web services If the words extracted from the conversation between users include words related to the linked web service, an assistive utterance including information such as the user's usage status of the linked web service is provided to the user as a new topic.

[0077] (1) Integration with music streaming services Based on the information about the song the user is listening to, an assisted speech is output that provides information related to the content of the conversation as a topic. The information about the song the user is listening to is obtained, for example, by a dedicated application from a server that provides a music streaming service, or from an application that the user has installed on the mobile terminal 2 to use the music streaming service.

[0078] (2) Collaboration with shopping services Based on the information of the user's shopping history, an assist utterance is output that provides information related to the content of the conversation as a topic. The information of the user's shopping history is obtained, for example, by a dedicated application from a server that manages a shopping site, or from an application that the user has installed on the mobile terminal 2 for shopping.

[0079] (3) Linking with event information obtained from the web Based on information obtained from the web, an assisted speech is output that provides information about events related to the content of the conversation as a topic.

[0080] 12 and 13 are diagrams showing a specific example of a conversation between user C and user D. FIG.

[0081] As shown in Fig. 12, the conversation between user C and user D starts in the same manner as the conversation between user A and user B described with reference to Fig. 9. User C and user D are users who are matched as conversation partners based on, for example, a common interest in "foreign TV dramas."

[0082] During the period from time t1 to time t7, users C and D alternately speak. User C's voice data is transmitted from interactive agent device 1C, which is the interactive agent device 1 used by user C, to interactive agent device 1D via the communication management server 11, and is output as user C's utterance in interactive agent device 1D. Interactive agent device 1D is the interactive agent device 1 used by user D. On the other hand, user D's voice data is transmitted from interactive agent device 1D to interactive agent device 1C via the communication management server 11, and is output as user D's utterance in interactive agent device 1C.

[0083] For example, from time t6 to time t7, user D talks about a scene from the movie and says, "I know! I also like the third season the best. The final scene, "story," was the best." Also, from time t7 to time t8, user C says, "That scene was great! I'm so hooked, I've been listening to the Stranger XXXX soundtrack all day lately."

[0084] The communication management server 11 analyzes the content of the conversation and detects, for example, words from the movie soundtrack that user C likes to listen to. Here, it is assumed that user C is listening to the movie soundtrack using a music streaming service that can be linked with the communication management server 11.

[0085] After detecting words in the soundtrack that user C is listening to, at time t8, a system utterance S12 such as "It seems that user C has listened to 'XX story' more than 10 times in the past week" is output. System utterance S12 is an assisted utterance that provides information related to the content of the conversation as a topic based on information about the song that user C is listening to.

[0086] The voice data of the system utterance S12 is transmitted from the communication management server 11 to both the interactive agent device 1C and the interactive agent device 1D, and is output as an assist utterance in each of the interactive agent device 1C and the interactive agent device 1D. Having heard the system utterance S12, user D, in response to the topic being provided, will utter something like, "I listen to the soundtrack too! I find myself playing that song on repeat over and over again" during the period from time t9 to time t10.

[0087] The communication management server 11 provides user D with information about user C that can serve as a catalyst for speech, thereby encouraging user D to speak and enabling smooth communication.

[0088] After time t10, as shown in FIG. 13, the assist utterance is triggered, and then user C and user D alternately speak.

[0089] For example, after analyzing the utterances made by user C between time t10 and time t11 to detect words related to products purchased by user C, at time t12, a system utterance S13 such as "I hear you bought a mug a week ago, C. Other popular items include T-shirts." is output. System utterance S13 is an assisted utterance that provides information related to the content of the conversation as a topic based on information about user C's shopping history.

[0090] Furthermore, after the preferences of User C and User D are identified by analyzing the content of the utterances, at time t14, a system utterance S14 such as "If you two like 'Stranger XXXX', I recommend the event being held in Shibuya." System utterance S14 is an assisted utterance that provides information about events related to the content of the conversation as a topic, based on information obtained from the web.

[0091] After the conversation triggered by such an assisted utterance continues, as shown in the lower part of FIG. 13, user C and user D each exchange greetings and end the conversation.

[0092] In this way, during a conversation between user C and user D, the communication management server 11 analyzes the content of the conversation between the two people and acquires information related to the content of the conversation based on the usage status of the web service. In addition, an assist utterance is output that provides the information acquired based on the usage status of the web service as a topic of conversation. This allows smooth communication between user C and user D.

[0093] 3. Assisted speech based on remaining alcohol level The following assist utterances are output depending on the amount of alcohol remaining in the user's drink. (1) Assisted speech to end the conversation (when both people have run out of alcohol) (2) An assistant voice recommending a second drink (when one user has run out of drink and the other user has more than half of their drink remaining)

[0094] For example, a sensor for detecting the remaining amount of alcohol is provided in the glass 62 used by the user. Information on the remaining amount of alcohol detected by the sensor is acquired by the interactive agent device 1 and transmitted to the communication management server 11.

[0095] The remaining amount of alcohol may be detected by analyzing an image captured by a camera provided in the interactive agent device 1. The analysis of the image to detect the remaining amount of alcohol may be performed in the interactive agent device 1 or in the communication management server 11.

[0096] 14 and 15 are diagrams showing a specific example of a conversation between user A and user B. FIG.

[0097] The conversation shown in Fig. 14 is the same as the conversation between user A and user B described with reference to Fig. 9. The left side of Fig. 14 shows a time series of the remaining amount of alcohol consumed by user A. The right side of Fig. 14 shows a time series of the remaining amount of alcohol consumed by user B. The remaining amount of alcohol is determined by the communication management server 11 based on information transmitted from the interactive agent devices 1 used by each user.

[0098] In the example of FIG. 14, at time t10 when user A finishes speaking, user A has 80% of his alcohol remaining, and user B has 50% of his alcohol remaining.

[0099] After time t10, user A and user B alternately speak, as shown in Fig. 15. In the example of Fig. 15, after a predetermined period of silence, such as 10 seconds, a system utterance S22, which is the same as the assist utterance described with reference to Fig. 11, is output.

[0100] At time t24 when user B makes an utterance, the remaining amount of alcohol in user B is 0%, as shown on the right side of Fig. 15. In this case, at time t24, a system utterance S23 such as "Mr. B, would you like a second drink?" is output. System utterance S23 is an assist utterance recommending a second drink.

[0101] The voice data of the system utterance S23 is transmitted from the communication management server 11 to both the interactive agent device 1A and the interactive agent device 1B, and is output as an assist utterance in each of the interactive agent device 1A and the interactive agent device 1B. Having heard the system utterance S23, user B can request a second drink and have the interactive agent device 1B prepare the drink. At time t24, as shown on the left side of FIG. 15, user A's remaining drink is 60%, meaning more than half is left.

[0102] The communication management server 11 can offer a second drink when one of the users has run out of alcohol, and can adjust the progress of both users' drinks, thereby enabling smooth communication between user A and user B. A user who has run out of alcohol usually becomes concerned about this and is unable to concentrate on the conversation, but this can be prevented.

[0103] The conversation between user A and user B shown in FIG. 15 ends in response to an assist utterance that is output when, for example, both users have run out of alcohol.

[0104] 4. Example of using emotion analysis results The user's emotions are analyzed based on the utterance, and the following processing is performed according to the emotion analysis results. The communication management server 11 is equipped with an emotion analysis function (emotion analysis engine). The user's emotions are analyzed based on the amount of time the user is speaking, the amount of time the user is listening, keywords included in the utterance, etc.

[0105] (1) For a user who has negative emotions, an assist utterance is output that provides a topic that is likely to give a positive emotion. For example, a topic related to the content that a user who has negative emotions likes is provided by the assist utterance.

[0106] (2) The system matches users with the most suitable users based on the personality and preferences of the users identified based on the results of emotion analysis. In this case, for example, the user's personality and preferences are analyzed based on the utterances made just before the moment when the user's emotion changes from negative to positive. The user's personality and preferences are analyzed based on the change in emotion during a conversation, and when matching for the next conversation, the system matches users with whom both parties are likely to have positive emotions.

[0107] (3) Based on the emotion analysis results, IoT (Internet of Things) devices are controlled. In the space where the user is, the interactive agent device 1 is installed together with IoT devices that can be controlled by the interactive agent device 1. For example, LED lighting with adjustable brightness and color temperature is installed as an IoT device.

[0108] The communication management server 11 controls the operation of the IoT device via the interactive agent device 1 by transmitting a control command to the interactive agent device 1. The communication management server 11 may control the operation of the IoT device via the mobile terminal 2 by transmitting a control command to the mobile terminal 2.

[0109] FIG. 16 is a diagram showing a specific example of a conversation between user A and user B.

[0110] The conversation shown in FIG. 16 is basically the same as the conversation between user A and user B described with reference to FIG. 9. The waveform shown to the right of user A's utterance represents user A's emotion while speaking, and the waveform shown to the left of user B's utterance represents user B's emotion while speaking. Of the waveforms representing emotions, waveforms shown with hatching represent negative emotions, and waveforms shown with lighter colors represent positive emotions. The amplitude of the waveform represents an emotion value, which is the degree of emotion.

[0111] 16, user B makes utterances from time t1 to time t2, from time t3 to time t4, and from time t5 to time t6. The emotion of user B during each utterance is positive.

[0112] On the other hand, user A speaks during short periods from time t2 to time t3, from time t4 to time t5, and from time t6 to time t7. User A's emotions during the utterances from time t2 to time t3 and from time t4 to time t5 are negative. User A's emotions during the utterances from time t6 to time t7 are positive.

[0113] The communication management server 11 analyzes the user's emotions, personality, preferences, etc., along with the conversation situation based on each utterance. For example, for user B, characteristics such as long speaking time, short listening time, and always positive emotions are inferred. Also, characteristics such as the user's likes to talk and interest in topics such as "baseball" are inferred.

[0114] On the other hand, it is inferred that User A has characteristics such as a short speaking time and a long listening time. In addition, since User A felt positive emotions in response to listening to User B's speech from time t5 to time t6, it is inferred that User A is interested in the name of the baseball player "Takamori," which is included as a keyword in the speech.

[0115] In this case, at time t7, a system utterance S31 such as "I have looked up the latest news about Takamori" is output. The system utterance S31 is an assisted utterance that provides a topic that is likely to evoke positive emotions. After outputting the system utterance S31, a system utterance that conveys the content of the latest news article that has been found is output.

[0116] As a result, the communication management server 11 can change user A's emotions to positive ones, and from then on, it becomes possible to realize smooth communication between user A and user B.

[0117] FIG. 17 is a diagram showing an example of matching.

[0118] In this example, the communication management server 11 infers from the conversation history with various users that user A is generally not good at listening to others, but has the characteristic of actively participating in conversations that interest him / her.

[0119] Furthermore, based on the content of the speech at the time of the change in emotion as described above, it is assumed that the person is interested in specific matters related to professional baseball, such as "Rookie of the Year," "Draft," and "Koshien."

[0120] In this case, as shown in Fig. 17, since each utterance is relatively short in order to summarize the main points, matching is performed with User C, who is interested in training professional baseball players. The matching between User A and User C was performed according to the personalities and preferences of the users, which were inferred based on their respective emotions during the conversation.

[0121] A conversation between user A and user C begins when a system utterance S41 such as "Mr. A, Mr. C is calling you" is output from the interactive agent device 1A, and user A, hearing the system utterance S41, agrees to start a conversation with user C.

[0122] This allows the communication management server 11 to match users with the most suitable users based on their personalities and preferences. The communication management server 11 has information on user combinations that are considered to be optimal.

[0123] The LED lighting is controlled based on the emotion analysis results so that if the conversation content is cheerful, the light is adjusted to a brighter color. On the other hand, if the conversation content is dark, the LED lighting is controlled so that the light is adjusted to a more subdued, dimmer color. For example, conversations about hobbies, family, and romance tend to be cheerful, while conversations about problems, worries, funerals, etc. tend to be dark.

[0124] This allows the communication management server 11 to adjust the environment around the user according to the content of the conversation.

[0125] <Configuration example of each device> Here, the configuration of each device in the voice communication system of FIG. 1 will be described.

[0126] ·Configuration of interactive agent device 1 FIG. 18 is a block diagram showing an example of the configuration of the interactive agent device 1. As shown in FIG.

[0127] The interactive agent device 1 is configured by connecting a speaker 52, a microphone 102, a communication unit 103, and a drink serving unit 104 to a control unit 101.

[0128] The control unit 101 is configured with a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory). The control unit 101 executes a predetermined program using the CPU, and controls the overall operation of the interactive agent device 1.

[0129] In the control unit 101, an agent function unit 111, a conversation control unit 112, a device control unit 113, and a sensor data acquisition unit 114 are realized by executing a predetermined program.

[0130] The agent function unit 111 realizes the agent function of the interactive agent device 1. For example, the agent function unit 111 executes various tasks requested by the user by voice and presents the results of the task execution to the user by synthesized voice. For example, the agent function unit 111 executes various tasks such as checking the weather or preparing alcohol. The agent function is realized by appropriately communicating with an external server such as the communication management server 11.

[0131] The conversation control unit 112 controls the conversation with the user selected as the conversation partner. For example, the conversation control unit 112 controls the communication unit 103 to transmit the user's voice data supplied from the microphone 102 to the communication management server 11. The voice data transmitted to the communication management server 11 is then transmitted to the interactive agent device 1 used by the user who is the conversation partner.

[0132] In addition, when the communication unit 103 receives voice data of the other user transmitted from the communication management server 11, the conversation control unit 112 outputs the speech of the other user from the speaker 52 based on the voice data supplied from the communication unit 103.

[0133] When the communication unit 103 receives the voice data of the system utterance transmitted from the communication management server 11, the conversation control unit 112 outputs the system utterance from the speaker 52 based on the voice data supplied from the communication unit 103.

[0134] The device control unit 113 controls the communication unit 103 to transmit control commands to the external device to be controlled and control the operation of the device. The device control unit 113 controls the IoT device and the like according to the user's emotions as described above, based on the information transmitted from the communication management server 11.

[0135] The sensor data acquisition unit 114 controls the communication unit 103 to receive sensor data transmitted from the sensor provided in the glass 62. For example, sensor data indicating the remaining amount of alcohol is transmitted from the sensor provided in the glass 62. The sensor data acquisition unit 114 transmits information indicating the remaining amount of alcohol to the communication management server 11. The sensor data acquisition unit 114 functions as a detection unit that detects the remaining amount of alcohol of the user based on the sensor data transmitted from the sensor provided in the glass 62.

[0136] The microphone 102 detects the user's speech and outputs the voice data to the control unit 101 .

[0137] The communication unit 103 is configured by a network interface for communicating with devices on the network 21, and a wireless communication interface for short-range wireless communication such as wireless LAN or Bluetooth (registered trademark). The communication unit 103 transmits and receives various data such as voice data to and from the communication management server 11. The communication unit 103 also transmits and receives various data to and from external devices provided in the same space as the interactive agent device 1, such as devices to be controlled and sensors provided in the glasses 62.

[0138] The alcohol serving unit 104 pours alcohol from the bottle 61 into the glass 62 under the control of the agent function unit 111. The alcohol server mechanism described above is realized by the alcohol serving unit 104. The alcohol serving unit 104 prepares alcohol in accordance with recipe data. The recipe data stored in the control unit 101 describes information on how to prepare alcohol according to drinking habits.

[0139] ·Communication Management Server 11 configuration FIG. 19 is a block diagram showing an example of the hardware configuration of the communication management server 11. As shown in FIG.

[0140] The CPU 201 , ROM 202 , and RAM 203 are connected to one another by a bus 204 .

[0141] An input / output interface 205 is further connected to the bus 204. An input unit 206 including a keyboard, a mouse, etc., and an output unit 207 including a display, a speaker, etc. are connected to the input / output interface 205.

[0142] Furthermore, the input / output interface 205 is connected to a storage unit 208 including a hard disk or nonvolatile memory, a communication unit 209 including a network interface, and a drive 210 for driving a removable medium 211 .

[0143] The communication management server 11 is configured by a computer having such a configuration. The communication management server 11 may be configured by a plurality of computers instead of one computer.

[0144] FIG. 20 is a block diagram showing an example of the functional configuration of the communication management server 11. As shown in FIG.

[0145] As shown in Fig. 20, a control unit 221 is implemented in the communication management server 11. The control unit 221 is configured with a profile management unit 231, a matching unit 232, a Web service analysis unit 233, a robot control unit 234, a conversation analysis unit 235, an emotion analysis unit 236, a drinking progress analysis unit 237, and a system utterance generation unit 238. At least a part of the configuration shown in Fig. 20 is implemented by the CPU 201 in Fig. 19 executing a predetermined program.

[0146] The profile management unit 231 manages the profile information of each user who uses the voice communication system. In addition to information registered using a dedicated application, the profile management unit 231 also manages information such as emotions during conversations and user characteristics identified based on the content of conversations as profile information.

[0147] The matching unit 232 matches users to be conversation partners based on the profile information managed by the profile management unit 231. Information on users matched by the matching unit 232 is supplied to the Web service analysis unit 233 and the robot control unit 234.

[0148] The Web service analysis unit 233 analyzes the usage status of the Web service by each user who is having a conversation. For example, the Web service analysis unit 233 obtains and analyzes information about the usage status of the Web service from a dedicated application installed on the mobile terminal 2.

[0149] The analysis by the Web service analysis unit 233 identifies information such as the song the user is listening to using a music streaming service, or the product the user has purchased using a shopping site. The analysis result by the Web service analysis unit 233 is supplied to the system utterance generation unit 238. Based on the analysis result by the Web service analysis unit 233, an assist utterance linked to the Web service is generated as described with reference to Figs. 12 and 13 .

[0150] The robot control unit 234 controls the interactive agent device 1, which is a bartender robot used by the user who is having a conversation. For example, the robot control unit 234 controls the communication unit 209 to transmit voice data transmitted from the interactive agent device 1 of one user to the interactive agent device 1 of the other user. The voice data of the user's utterance received by the robot control unit 234 is supplied to a conversation analysis unit 235 and an emotion analysis unit 236.

[0151] Furthermore, the robot control unit 234 transmits the voice data of the system utterance generated by the system utterance generating unit 238 to the interactive agent devices 1 of both users engaged in conversation, causing them to output the system utterance.

[0152] Furthermore, when information indicating the remaining amount of alcohol is transmitted from the interactive agent device 1, the robot control unit 234 outputs the information indicating the remaining amount of alcohol received by the communication unit 209 to the alcohol progress analysis unit 237. The robot control unit 234 communicates with the interactive agent device 1 and performs various processes such as controlling IoT devices via the interactive agent device 1.

[0153] The conversation analysis unit 235 analyzes the speech situation of each user who is having a conversation, such as the speaking time and silence time, based on the voice data supplied from the robot control unit 234. The conversation analysis unit 235 also analyzes the content of the conversation to analyze keywords included in the utterances. The analysis result by the conversation analysis unit 235 is supplied to the system utterance generation unit 238. Based on the analysis result by the conversation analysis unit 235, an assist utterance is generated according to the conversation situation, as described with reference to FIGS. 9 to 11.

[0154] The emotion analysis unit 236 analyzes the emotion of each user who is having a conversation, based on the voice data supplied from the robot control unit 234. The analysis result by the emotion analysis unit 236 is supplied to the system utterance generation unit 238. Based on the analysis result by the emotion analysis unit 236, an assist utterance according to the emotion is generated, as described with reference to FIG. 16 .

[0155] The drinking progress analysis unit 237 analyzes the drinking progress of each user who is having a conversation, based on the information supplied from the robot control unit 234. As described above, the information indicating the remaining amount of alcohol transmitted from the interactive agent device 1 is sensor data transmitted from the sensor provided in the glass 62. The drinking progress analysis unit 237 analyzes the drinking progress of each user based on the sensor data transmitted from the sensor provided in the glass 62.

[0156] The analysis result by the drinking progress analysis unit 237 is supplied to the system utterance generation unit 238. Based on the analysis result by the drinking progress analysis unit 237, an assist utterance according to the remaining amount of alcohol is generated, as described with reference to Figs. 14 and 15.

[0157] The system utterance generation unit 238 generates an assist utterance based on the analysis results of the Web service analysis unit 233, the conversation analysis unit 235, the emotion analysis unit 236, and the drinking progress analysis unit 237, and supplies the voice data of the generated assist utterance to the robot control unit 234. The system utterance generation unit 238 also generates system utterances other than the assist utterances as appropriate, and supplies the voice data of the generated system utterances to the robot control unit 234.

[0158] <Operation of each device> Here, the basic operations of the communication management server 11 and the interactive agent device 1 having the above-described configuration will be described.

[0159] · Operation of Communication Management Server 11 First, the processing of the communication management server 11 will be described with reference to the flowchart of FIG.

[0160] In step S1, the matching unit 232 refers to the profile information managed by the profile management unit 231 to match users who will be conversation partners, and starts a conversation.

[0161] In step S2, the robot control unit 234 transmits and receives voice data of the user's utterance to and from the interactive agent device 1 used by the user with whom the conversation is taking place.

[0162] In step S3, the conversation analysis unit 235 analyzes the conversation situation between the two users based on the speech audio data.

[0163] In step S4, the system utterance generation unit 238 determines whether or not an assist utterance is necessary based on the result of analyzing the conversation situation.

[0164] If it is determined in step S4 that an assisting utterance is necessary, in step S5, the system utterance generation unit 238 generates an assisting utterance and causes the robot control unit 234 to transmit voice data of the assisting utterance to the interactive agent device 1 of each user.

[0165] In step S6, the robot control unit 234 determines whether the conversation has ended.

[0166] If it is determined in step S6 that the conversation has not ended, the process returns to step S2, and the above-described processing is repeated. Similarly, if it is determined in step S4 that an assist utterance is not necessary, the processing from step S2 onwards is repeated.

[0167] If it is determined in step S6 that the conversation has ended, the process ends.

[0168] · Operation of the interactive agent device 1 Next, the processing of the interactive agent device 1 will be described with reference to the flowchart of FIG.

[0169] In step S11, the microphone 102 detects the user's speech.

[0170] In step S12, the conversation control unit 112 transmits the voice data of the user's speech supplied from the microphone 102 to the communication management server 11.

[0171] In step S13, the conversation control unit 112 determines whether or not voice data of the user's speech or the system's speech has been transmitted from the communication management server 11.

[0172] If it is determined in step S13 that voice data has been transmitted, the speaker 52 outputs the speech of the user who is the conversation partner or the system speech under the control of the conversation control unit 112 in step S14.

[0173] If it is determined in step S15 that the conversation has ended, the process ends.

[0174] Through the above process, the user of the interactive agent device 1 can casually enjoy conversations with other users at home, for example, by using the alcohol prepared by the interactive agent device 1 for an evening drink. For example, even if the conversation comes to a halt, the communication management server 11 provides assistance, allowing the user to communicate smoothly with the person they are talking to.

[0175] In particular, elderly people who live alone find it difficult to go out, etc. By using the interactive agent device 1 as a communication tool and having a conversation with a person in a remote location as shown in Fig. 23, elderly people who live alone can eliminate their sense of loneliness.

[0176] To be able to talk freely about one's anxieties and worries with others, an environment is required that satisfies certain conditions, such as the other person being a good listener, guaranteeing that the conversation will not take place in person, protecting personal information, being a person who is trusted by those around them, and having a third party acting as an intermediary. The interactive agent device 1 allows users to easily introduce such an environment into their homes.

[0177] Furthermore, users can use a dedicated application to manage their alcohol intake and review conversation records.

[0178] <Other> Although all of the components shown in FIG. 20 are provided in the communication management server 11, at least a part of the components shown in FIG. 20 may be provided in the interactive agent device 1.

[0179] Although the interactive agent device 1 has been described as providing alcoholic beverages, other beverages such as coffee, tea, and juice may also be provided. Food may also be provided. By providing food, each user can enjoy conversation with other users while eating the food.

[0180] About the program The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware or a general-purpose personal computer.

[0181] The program to be installed is provided by being recorded on removable media such as an optical disc (CD-ROM (Compact Disc-Read Only Memory), DVD (Digital Versatile Disc), etc.) or semiconductor memory. It may also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital broadcasting. The program can be pre-installed in a ROM or memory unit.

[0182] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0183] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.

[0184] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0185] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.

[0186] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0187] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0188] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0189] <Configuration combination example> The present technology can also be configured as follows.

[0190] (1) an analysis unit that analyzes utterances of two users who are having a conversation via a network, the utterances being detected by the interactive robots used by the two users; a control unit that causes each of the interactive robots to output a conversation assistance voice, which is a voice that assists the conversation, according to a conversation situation between the two users; An information processing device comprising: (2) The control unit outputs the conversation assistance voice according to a conversation situation after alcohol is served to the two users by each of the interactive robots. The information processing device according to (1) above. (3) The system further includes a matching unit that matches two users to have a conversation with each other based on the profile information of each user. The information processing device according to (1) or (2). (4) The control unit outputs the conversation assistance voice to encourage the user with a shorter speaking time to speak based on the speaking times of the two users. The information processing device according to any one of (1) to (3). (5) The control unit outputs the conversation assistance voice to prompt the two users to speak in response to a pause in speech between the two users for a certain period of time. The information processing device according to any one of (1) to (4). (6) The control unit outputs the conversation assistance voice with content related to information that is attracting attention on a news site on the network. The information processing device according to any one of (1) to (5). (7) When a word related to a web service used by the user is included in an utterance of the two users, the control unit outputs the conversation assistance voice based on a usage status of the web service. The information processing device according to any one of (1) to (6). (8) The control unit outputs the conversation assistance voice based on the emotions of the two users analyzed based on their utterances. The information processing device according to any one of (1) to (7). (9) The control unit outputs the conversation assistance voice relating to content preferred by the user having negative emotions, the content being specified based on preference information of each of the two users. The information processing device according to (8). (10) The control unit controls a device installed together with the interactive robot in a space where each of the users is present, based on the emotions of the two users analyzed based on their speech. The information processing device according to any one of (1) to (9). (11) The control unit transmits a control command for controlling the device to the interactive robot and controls the device via the interactive robot, or transmits the control command to a mobile terminal held by the user and controls the device via the mobile terminal. The information processing device according to (10) above. (12) The control unit outputs the conversation assistance voice according to the degree of drinking of each of the two users analyzed based on the sensor data. The information processing device according to (2) above. (13) The information processing device analyzing utterances of two users who are having a conversation via a network, detected by respective interactive robots used by the two users; According to the situation of the conversation between the two users, a conversation assistance voice, which is a voice that assists the conversation, is output from each of the interactive robots. Control method. (14) a serving department that serves alcoholic beverages to users; a conversation control unit that detects the user's speech after the alcoholic beverage is served, transmits audio data of the detected speech to an information processing device that analyzes the user's speech and the speech of another user who is the conversation partner, and outputs a conversation assistance voice that is a voice that assists the conversation and that is transmitted from the information processing device according to the state of the conversation between the two people; An interactive robot comprising: (15) a detection unit that detects the remaining amount of alcohol of the user and transmits information indicating the detected remaining amount of alcohol to the information processing device; The conversation control unit outputs the conversation assistance voice transmitted from the information processing device according to the degree to which each of the two people has consumed alcohol. The interactive robot according to (14) above. (16) The interactive robot Providing alcohol to users, detecting an utterance of the user after the alcoholic beverage is served, and transmitting audio data of the detected utterance to an information processing device that analyzes the utterance of the user and the utterance of another user who is the conversation partner; A conversation assistance voice, which is a voice that assists the conversation and is transmitted from the information processing device according to the situation of the conversation between the two people, is output. Control method. [Explanation of symbols]

[0191] 1A, 1B Interactive agent device, 2A, 2B Mobile terminal, 11 Communication management server, 21 Network, 51 Housing, 61 Bottle, 62 Glass, 101 Control unit, 102 Microphone, 103 Communication unit, 104 Drink serving unit, 111 Agent function unit, 112 Conversation control unit, 113 Device control unit, 114 Sensor data acquisition unit, 221 Control unit, 231 Profile management unit, 232 Matching unit, 233 Web service analysis unit, 234 Robot control unit, 235 Conversation analysis unit, 236 Emotion analysis unit, 237 Drink progress analysis unit, 238 System utterance generation unit

Claims

1. an analysis unit that analyzes utterances of two users who are having a conversation, detected by one or more agent devices used by the two users; a control unit that, when a word related to a network service used by one or more of the two users is included in an utterance of one or more of the two users, outputs a voice of content related to the network service from one or more of the agent devices as a conversation assistance voice, which is a voice that assists the conversation; An information processing device comprising:

2. The conversation assistance voice is a voice with content related to the usage status of the network service. The information processing device according to claim 1 .

3. The network service is a web service. The information processing device according to claim 1 .

4. The control unit outputs the conversation assistance voice according to a conversation situation between the two users. The information processing device according to claim 1 .

5. The control unit outputs the conversation assistance voice according to a conversation situation after alcohol is served to the two users by the agent device. The information processing device according to claim 1 .

6. The system further includes a matching unit that matches two users to have a conversation with each other based on profile information of each user. The information processing device according to claim 1 .

7. The control unit outputs the conversation assistance voice to encourage the user who has a shorter speaking time to speak, based on the speaking times of the two users. The information processing device according to claim 1 .

8. The control unit outputs the conversation assistance voice to prompt the two users to speak in response to a pause in speech between the two users for a certain period of time. The information processing device according to claim 1 .

9. The control unit outputs the conversation assistance voice based on emotions of the two users analyzed based on their utterances. The information processing device according to claim 1 .

10. The control unit outputs the conversation assistance voice relating to content preferred by the user having negative emotions, the content being specified based on preference information of each of the two users. The information processing device according to claim 9 .

11. The control unit controls devices installed together with the agent devices in the spaces where the two users are present, based on the emotions of the two users analyzed based on their speech. The information processing device according to claim 1 .

12. The control unit transmits a control command for controlling the device to the agent device and controls the device via the agent device, or transmits the control command to a mobile terminal held by the user and controls the device via the mobile terminal. The information processing device according to claim 11.

13. The control unit outputs the conversation assistance voice in accordance with the degree of drinking of each of the two users analyzed based on the sensor data. The information processing device according to claim 5 .

14. The information processing device analyzing utterances of two users engaged in a conversation detected by one or more agent devices used by said users; When a word related to a network service used by one or more of the two users is included in an utterance of one or more of the two users, a voice of a content related to the network service is output from one or more of the agent devices as a conversation assistance voice, which is a voice that assists the conversation. A control method comprising:

Citation Information

Patent Citations

  • Interactive agent system and method

    JP2008052449A