Information processor, method, and program

The information processing device enhances chatbot accuracy by using emotional information and keyword extraction to improve user input processing, leading to more accurate and relevant responses.

JP2025086223AActive Publication Date: 2025-06-06CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023200138
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-06
Estimated Expiration
2043-11-27

AI Technical Summary

Technical Problem

Existing chatbot systems struggle to accurately process user input, particularly when voice is used, as they fail to distinguish between important keywords and unnecessary words, leading to decreased input quality for generative AI.

Method used

An information processing device that acquires emotional information from users through captured images or voice, generates prompts based on this information and extracted keywords, and outputs answers generated by generative AI.

Benefits of technology

Improves the accuracy of answers provided to users by effectively filtering out unnecessary information and emphasizing important keywords based on user emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025086223000001_ABST
    Figure 2025086223000001_ABST
Patent Text Reader

Abstract

To improve the accuracy of a response created in response to a user's inquiry.SOLUTION: An information processor comprises: acquisition means for acquiring emotional information of a user based on at least one of a captured image and the voice of the user; prompt generation means for generating a prompt based on the emotional information and one or more keywords extracted from the voice; and output means for outputting an answer obtained by using the prompt generated by the prompt generation means.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device, a method, and a program. [Background technology]

[0002] Chatbots are being used to reduce labor costs and provide 24-hour support when responding to inquiries to companies and creating product plans. A chatbot is a program and / or device that automatically converses with a user via text and / or voice. Patent Document 1 discloses a method of conversing with a user based on the user's emotional information and browsing information estimated based on keywords extracted from the user's voice information. Meanwhile, the spread of generative AI has made it possible to create more advanced suggestions than conventional chatbots. For example, if you input "Please tell me famous tourist spots in Tokyo" into Chat-GPT, it will output information that associates the names of several representative tourist spots with simple descriptions of the tourist spots. Here, it is known that the quality of the input information (also called prompts) is important in generative AI. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2021-103411 A [Non-patent literature]

[0004] [Non-Patent Document 1] M. Sarma, "Emotion identificationfrom raw speech signals using DNNs," Proc. Interspeech 2018 [Non-Patent Document 2] S. Li, "Deep Facial Expression Recognition: A Survey", IEEE transactions on affective computing, 2020 [Non-Patent Document 3] Sharma, P., & Li, Y., "Self-Supervised Contextual Keyword and Keyphrase Retrieval with Self-Labeling", 2019 Summary of the Invention [Problem to be solved by the invention]

[0005] However, if unnecessary words and / or modifiers are included in the input information for the generative AI, the expected output may not be obtained. In particular, when voice is used as input, important keywords included in the user's utterance cannot be distinguished from unnecessary keywords, and the quality of the input information for the generative AI decreases. In addition, Non-Patent Document 1 does not consider determining the importance of keywords included in the user's utterance.

[0006] Therefore, an object of the present invention is to improve the accuracy of answers created in response to questions from users. [Means for solving the problem]

[0007] In order to achieve the object of the present invention, an information processing device according to an embodiment of the present invention has the following configuration: an acquisition means for acquiring emotional information of a user based on at least one of a captured image and a voice of the user, a prompt generation means for generating a prompt based on the emotional information and a keyword extracted from the voice, and an output means for outputting an answer obtained by using the prompt generated by the prompt generation means. Effect of the Invention

[0008] According to the present invention, it is possible to improve the accuracy of answers prepared in response to questions from users. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a hardware configuration of an information processing apparatus according to a first embodiment. [Diagram 2] FIG. 1 is a block diagram showing the functional configuration of an information processing apparatus according to a first embodiment. [Diagram 3] FIG. 4 is a diagram showing emotion information according to the first embodiment. [Figure 4] FIG. 4 is a view showing keyword information according to the first embodiment. [Diagram 5] FIG. 4 is a diagram showing a prompt according to the first embodiment. [Figure 6] A figure showing an answer created by the generative AI in the first embodiment. [Figure 7] 5 is a flowchart for explaining a response creation process for a conversation according to the first embodiment. [Figure 8A] FIG. 4 is a view showing a UI displayed on a display unit according to the first embodiment. [Figure 8B] FIG. 8B is an enlarged view of a portion of the UI in FIG. 8A. [Figure 9] FIG. 4 is a view for explaining template information according to the first embodiment. [Figure 10] FIG. 11 is a diagram showing a prompt according to the second embodiment. [Figure 11] A figure showing customer service support information created by the generative AI in the second embodiment. [Figure 12] 10 is a flowchart illustrating a process for creating customer service assistance information according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0011] First Embodiment The information processing device 100 recognizes the emotion of the subject from a video in which the subject appears and / or the voice of the subject. The information processing device 100 also creates information (also called a prompt) to be input to the generative AI based on the emotion and remarks of the subject. The information processing device 100 then inputs the prompt to the generative AI to create a response to a conversation with the subject (e.g., a customer who visits a store). In this embodiment, a system in which customer service is provided in a real store in the form of a chatbot is described as an example. This embodiment can also be applied to a situation in which customer service is provided to customers in a virtual store.

[0012] FIG. 1 is a block diagram showing the hardware configuration of an information processing device according to the first embodiment.

[0013] The information processing device 100 includes an input unit 101, a display unit 102, a network I / F unit 103, a CPU 104, a RAM 105, a ROM 106, a HDD 107, and a data bus 108.

[0014] The input unit 101 includes at least one of a keyboard, a mouse, a touch panel, and a microphone, for example, and receives input from a user.

[0015] The display unit 102 is, for example, a liquid crystal display, and displays information such as the results of various processes, etc. The input unit 101 and the display unit 102 are connected to other functional units via a data bus 108 so as to be able to communicate with each other.

[0016] The network I / F unit 103 connects to an external device (not shown) via the Internet so as to be able to transmit and receive various types of information.

[0017] The CPU 104 reads out a control computer program stored in the ROM 106, loads it into the RAM 105, and executes various control processes. The CPU 104 executes an image processing program stored in the ROM 106 or the HDD 107, thereby implementing image processing on image data.

[0018] The RAM 105 is used as a temporary storage area for storing programs executed by the CPU 104, a work memory, and the like.

[0019] The HDD 107 stores various information such as image data, setting parameters, various programs, etc. The HDD 107 can also receive data input from an external device (not shown) via the network I / F unit 103.

[0020] Image data and the like received from an external device (not shown) via the network I / F unit 103 is transmitted to and received from the CPU 104, RAM 105, and ROM 106 via a data bus 108.

[0021] FIG. 2 is a block diagram showing the functional configuration of the information processing device according to the first embodiment.

[0022] The information processing device 100 includes an audio acquisition unit 201 , a video acquisition unit 202 , an emotion determination unit 203 , an extraction unit 204 , a generation unit 205 , and a correction unit 206 .

[0023] The voice acquisition unit 201 acquires the user's voice via the input unit 101 (for example, a microphone). The voice acquisition unit 201 may acquire the user's voice stored in the HDD 107, or may acquire the user's voice stored in an external device (for example, a server).

[0024] The video acquisition unit 202 acquires a video of the user captured by an external camera via the network I / F unit 103. The video acquisition unit 202 may acquire a video of the user stored in the HDD 107, or may acquire a video of the user stored in an external device (e.g., a server).

[0025] The emotion determination unit 203 generates emotion information of a person (that is, a user) with whom the information processing device 100 is talking, based on the voice acquired by the voice acquisition unit 201 and / or the video acquired by the video acquisition unit 202.

[0026] Here, FIG. 3 is a diagram showing emotion information according to the first embodiment.

[0027] The emotion information 300 includes emotion labels 310 that represent the user's emotion, such as "happiness," "surprise," "fear," "sadness," "anger," "contempt," "disgust," and "neutral," and emotion scores 320 for each emotion label 310. The emotion scores 320 are expressed as scalar values ​​ranging from 0 to 1. The closer the black bar of the emotion label 310 is to 1, the higher the likelihood that the user has the emotion of the emotion label 310. Now, we will return to the explanation of FIG. 2.

[0028] The extraction unit 204 determines the importance of keywords included in the previous answer and / or the user's voice, based on the previous answer generated by the generation unit 205, the user's voice acquired from the voice acquisition unit 201, and emotion information 300 acquired from the emotion determination unit 203. The extraction unit 204 stores keyword information, in which keywords and emotion information 300 are linked for each time series of the conversation with the user, in a database server (not shown) and / or HDD 107.

[0029] Here, FIG. 4 is a diagram showing keyword information according to the first embodiment.

[0030] The keyword information 400 includes a date and time 410, a keyword 420, an emotion label 430, and an emotion score 440. The date and time 410 indicates the time when the extraction unit 204 extracted the keyword 420 from the most recent answer and / or the user's voice. The keyword 420 is linked to the user's emotion label 430 and emotion score 440. For example, the keyword 420 "record of the athletic meet" is linked to the emotion label 430 "happiness" and the emotion score 440 "0.9". Note that a symbol representing an emotion is displayed to the right of the emotion in the emotion label 430, but it does not have to be displayed as necessary. Returning to the description of FIG. 2, the following is now given.

[0031] The extraction unit 204 creates a prompt based on the voice acquired from the voice acquisition unit 201, the keyword information of the database server, the conversation state information, and the correction information (described later) acquired from the correction unit 206. Here, the conversation state information is information that indicates the purpose of the current conversation with the user. For example, assuming customer service in a store, the conversation state information includes information that indicates a series of customer service flows such as "initial state," "confirming the purpose of visiting the store," "suggesting a product," and "explaining the details of the product." At this time, the extraction unit 204 internally holds information on the product category and / or product name according to the conversation state information. The product category means the type of product sold in the store, such as, for example, a camera, a vacuum cleaner, and a refrigerator. The product name means a specific product name, such as the model name of a camera.

[0032] Here, FIG. 5 is a diagram showing a prompt according to the first embodiment.

[0033] Prompt 500 is input information to the generative AI of generation unit 205, which will be described later. Prompt 500 in FIG. 5 shows input information when "Explain the product in detail" is selected as conversation state information. For example, prompt 500 includes a message "Please write an introduction to ABC in 200 characters or less, keeping in mind the following points," as an instruction to the generative AI when creating an answer, and keywords 510. Keywords 510 are keywords used when creating an answer, and include, for example, "Excellent autofocus, lightweight body, good image quality." Now, we return to the explanation of FIG. 2.

[0034] The generation unit 205 generates an answer by inputting the prompt 500 in FIG. 5 generated by the extraction unit 204 into a generative AI (for example, Chat-GPT).

[0035] FIG. 6 is a diagram showing an answer created by the generative AI according to the first embodiment.

[0036] Answer 600 is an answer created by the generative AI based on prompt 500 in Fig. 5. Answer 600 includes a product description of ABC. In this case, answer 600 includes only a text description of the product, but may also include diagrams, tables, images, etc. according to the instructions of prompt 500. Now, we return to the explanation of Fig. 2.

[0037] The display unit 102 presents the answer 600 (see FIG. 6) that the generation unit 205 created using the generative AI to the user.

[0038] Correction unit 206 corrects the emotion labels and / or emotion scores of keyword information 400 and / or corrects conversation state information based on keyword information 400 in a database server (not shown) and / or HDD 107 and transitions in emotion information.

[0039] 7 is a flowchart illustrating a conversation answer creation process according to the first embodiment. The process in FIG. 7 is started, for example, when the information processing device 100 receives an answer creation instruction from a user via the input unit 101.

[0040] In S701, the display unit 102 displays the answer (shown in FIG. 6) created by the generation unit 205 using a generative AI. When the conversation state is the initial state, the display unit 102 displays "What are you looking for?" as a general question to a user (for example, a customer looking for a desired product). In this embodiment, it is assumed that the display of the answer portion of the display unit 102 is updated based on the answer created in the processes of S702 to S709. At this time, the information processing device 100 may display on the display unit 102 data obtained by converting the voice obtained in the conversation with the user into text and / or the extracted keyword information 400.

[0041] Here, Fig. 8A is a diagram showing a UI displayed on the display unit according to the first embodiment, and Fig. 8B is a diagram showing an enlarged view of a part of the UI in Fig. 8A.

[0042] In an area 810 on the left side of a screen 800 of the display unit 102, the contents of past conversations (i.e., conversation history) with a user (here, a customer looking for a desired camera) are displayed. In an area 820 on the upper right of the screen 800, a product category ("mirrorless camera") and a product name ("ABC") are displayed. In an area 830 located below the area 820, keywords extracted from the conversation in the area 810 are displayed. In a position below the area 830, a store clerk avatar 840 that serves the user is displayed. For example, the contents of the conversation in the area 810 that correspond to the keywords in the area 830 ("good image quality", "light", "autofocus") may be highlighted by bold text and / or coloring. Also, a product description may be displayed with the hand 850 of the store clerk avatar 840 positioned at a characteristic position of the product on the product introduction page (see FIG. 8B). Here, we return to the explanation of FIG. 7.

[0043] In S702, the voice acquisition unit 201 acquires the voice of the user who is having a conversation. The video acquisition unit 202 acquires a video in which the user who is having a conversation appears.

[0044] In S703, the emotion determination unit 203 analyzes the emotion of the user during the conversation. The emotion determination unit 203 performs an emotion analysis of the user based on the user's voice and / or a video showing the user, using a trained model that has undergone deep learning. Here, the emotion analysis process by the emotion determination unit 203 can be performed by a known technique for analyzing the user's emotion from the voice and / or video (image). The emotion determination unit 203 can acquire the user's emotion information from the voice by using the method of Non-Patent Document 1. Furthermore, the emotion determination unit 203 can acquire the user's emotion information from the video (specifically, the image that constitutes the video) by using the method of Non-Patent Document 2.

[0045] In S704, the extraction unit 204 extracts keywords based on the previous answer and the current user's voice. That is, the extraction unit 204 extracts one or more keywords from the contents of answers previously presented to the user and the user's voice. These keywords are used to generate prompts. The extraction unit 204 extracts keywords from the user's voice using a known AI service. For example, keyword extraction can be performed based on a known technology for converting voice data to text and a known technology for automatically extracting characteristic phrases and / or words from the text. The extraction unit 204 converts the user's voice into text by using an API for converting voice data to text, such as Google Cloud Speech-to-Text API. In addition, the extraction unit 204 extracts keywords from the text by using an API for automatically extracting characteristic phrases and / or words from the text, similar to the method of Non-Patent Document 3.

[0046] In S705, the extraction unit 204 stores keyword information 400 in which the extracted keywords are linked to the emotion information analyzed by the emotion determination unit 203 in a database server (not shown) and / or the HDD 107.

[0047] In S706, the extraction unit 204 creates a prompt by inserting a keyword from the keyword information 400 into a template, which will be described later. The extraction unit 204 acquires template information for creating a prompt based on the state of the conversation and the correction information. Furthermore, the extraction unit 204 corrects the content of the prompt based on the keyword information 400 and the correction information.

[0048] The extraction unit 204 acquires "template information" that is stored in advance in a database server (not shown) or the like, based on the state of the conversation.

[0049] Here, FIG. 9 is a diagram for explaining template information according to the first embodiment.

[0050] The template information 900 is information that associates a conversation state 910 with a template 920 for creating a prompt. The conversation state 910 includes a conversation state that represents a series of customer service flows, such as "initial state (state 1)," "confirming the purpose of the visit (state 2)," "proposing a product (state 3)," and "explaining the details of the product (state 4)." Note that the conversation state 910 may include less than four or five or more conversation states. The template 920 represents a standard form of a prompt for providing an optimal answer to the user according to the conversation state 910. For example, when the conversation state 910 is the initial state (state 1), the template 920 corresponding to state 1 is "Are you looking for something?"

[0051] Next, the extraction unit 204 selects keywords to be applied to the template 920 based on the keyword information 400. The selection of keywords may be performed based on emotion information (emotion label 430 and emotion score 440) included in the keyword information 400 and predefined keyword selection conditions. For example, the keyword selection conditions include "keywords in which the emotion is happiness and the emotion score is 0.8 or more" and "keywords in which the emotion is surprise and the emotion score is 0.8 or more".

[0052] The extraction unit 204 can select "record of athletic meet", "photo", "good image quality", "mirrorless camera", and "autofocus" from the keyword information 400 in FIG. 4 based on the preset keyword selection condition. The keyword selection condition may also include "an antonym of a keyword whose emotion is anger and whose emotion score is 0.9 or more". For example, the extraction unit 204 can select "light", which is an antonym of "heavy", from the keyword information 400 in FIG. 4 based on "an antonym of a keyword whose emotion is anger and whose emotion score is 0.9 or more". That is, the extraction unit 204 can create a prompt based on one or more keywords associated with a specific type of emotion among a plurality of keywords extracted from the voice. The extraction unit 204 can also generate a prompt based on one or more keywords associated with a specific type of emotion among a plurality of keywords and whose likelihood of being the emotion is equal to or greater than a threshold.

[0053] Next, the extraction unit 204 creates a prompt by inserting the product category, product name, and selected keyword, which are managed together with the conversation state, into the corresponding locations of the template 920. Specifically, the extraction unit 204 creates a prompt by inserting each keyword selected (extracted) from the keyword information 400 into the "product category," "product name," and "keyword" of the template 920 in Fig. 9. Now, we return to the explanation of Fig. 7.

[0054] In S707, the generation unit 205 generates an answer to be presented to the user based on the prompt generated by the extraction unit 204. The answer generated in this manner is displayed by the display unit 102 in S701. Note that while FIG. 2 shows an example in which the information processing device 100 has the display unit 102, the information processing device 100 and the display unit 102 may be connected by a wired or wireless connection medium. In either case, the CPU 104 of the information processing device 100 executes display control for displaying the answer obtained based on the prompt on the display screen. Moreover, the information processing device 100 of this embodiment can output the answer as a voice instead of or in addition to displaying the answer.

[0055] In S708, the generation unit 205 determines whether the conversation with the user has ended. If the generation unit 205 determines that the conversation with the user has ended (Yes in S708), the process proceeds to S710. On the other hand, if the generation unit 205 determines that the conversation with the user has not ended (No in S708), the process proceeds to S709.

[0056] In S709, the correction unit 206 corrects the emotion information and conversation state based on the following conditions 1 to 4. Here, the corrected information refers to information obtained by correcting the emotion information (emotion label 430, emotion score 440) of the keyword information 400, or the conversation state 910 of the template information 900, by the correction unit 206.

[0057] (Method of correcting emotion information based on condition 1) A method in which correction section 206 corrects current emotion information based on past emotion information when past emotion information and current emotion information corresponding to the same keyword in keyword information 400 differ will be described with reference to FIG. 4.

[0058] Based on the keyword 420 ("single-lens reflex camera") of the current (first time) keyword information 460, the correction unit 206 acquires past (second time) keyword information 450 including "single-lens reflex camera" from the database server (not shown) and / or HDD 107. Here, the keyword information 450 and the keyword information 460 include the keyword 420 ("single-lens reflex camera"). However, the emotion label 430 ("anger") of the keyword information 460 is different from the emotion label 430 ("happiness") of the keyword information 450. In this case, the correction unit 206 compares the keyword information 450 and the keyword information 460, and determines the emotion label 430 with the higher emotion score 440 value from among the two pieces of keyword information. The emotion score 440 ("0.6") of the keyword information 450 is higher than the emotion score 440 ("0.4") of the keyword information 460. Therefore, correction unit 206 corrects emotion label 430 of keyword information 460 from "anger" to "happiness", and corrects emotion score 440 from "0.4" to "0.6." Note that in the correction method described here, the oldest emotion information in the time series (emotion label 430 of keyword information 450) may be adopted as the current emotion information (emotion label 430 of keyword information 460, etc.), emotion information at the latest time may be adopted, or emotion information based on the average value of all past emotion scores may be adopted.

[0059] (How to correct the conversation state based on condition 2) A method for changing the conversation state to the next conversation state when the correction unit 206 obtains information required for the conversation state transition (such as product category and product name) will be described below. The correction unit 206 transitions the current conversation state to the next conversation state based on the predefined conversation state transition conditions.

[0060] For example, it is defined that the conversation state transitions in the order of "initial state (hereinafter, state 1)", "confirm purpose of visit (hereinafter, state 2)", "suggest product (hereinafter, state 3)", and "explain details of product (hereinafter, state 4)". In states 1 to 4, the transition conditions from one conversation state to the other conversation state will be explained. First, the transition condition from state 1 to state 2 includes that the conversation has started. The transition condition from state 2 to state 3 includes that the product category has been obtained. The transition condition from state 3 to state 4 includes that the product name has been obtained. Then, when the current keyword information 400 includes a predetermined product category and / or product name, the correction unit 206 transitions the current conversation state (e.g., state 2) to the next conversation state (e.g., state 3).

[0061] (How to correct the conversation state based on condition 3) When a specific phrase is obtained during a conversation with a user, the correction unit 206 performs correction to return the current conversation state (e.g., state 2) to the previous conversation state (e.g., state 1). For example, phrases that return the conversation state, such as "I want to see other products" and "I want to find products other than cameras," are set in advance. Then, when the correction unit 206 obtains the above phrase during a conversation with a user, the correction unit 206 returns the conversation state to the previous conversation state. In addition, when the user operates the product category and / or product name on the screen 800 of the display unit 102 to clear (delete) or change the product category and / or product name, the current conversation state may be returned to the previous conversation state. That is, the correction unit 206 can change the conversation state in response to receiving a user operation to delete one or more keywords extracted from the voice from the display screen.

[0062] (How to correct the conversation state based on condition 4) When the emotion information obtained for the most recent answer is a negative emotion, the correction unit 206 performs correction to return the conversation state to the previous state. Here, the negative emotion refers to "fear," "sadness," "anger," "contempt," and "disgust" in the emotion label 310 of FIG. 3. First, the correction unit 206 acquires keyword information from the keyword information 400 in the database server (not shown) and / or the HDD 107 for a preset period, for example, from one minute ago to the current time. When the number of negative emotions in the emotion information (emotion label 430) included in the acquired keyword information exceeds a preset ratio, the correction unit 206 returns the current conversation state to the previous conversation state. For example, when the correction unit 206 decides to return the current conversation state (state 2) to the previous conversation state (state 1), it causes the generation unit 205 to create answers such as "Would you like to search for another product category?" and "Shall I suggest another product?", and displays one of the created answers on the display unit 102. Alternatively, the correction unit 206 may select an answer that is appropriate for the state of the conversation after correction from among the answers created by the generation unit 205 from the past to the present, and display the selected answer on the display unit 102.

[0063] As described above, according to the first embodiment, it is possible to present an appropriate answer to the user in response to changes in the user's emotions. This makes it possible to prevent the user from feeling stressed by the answers presented by the information processing device 100 (specifically, the chatbot), and to increase the user's purchasing motivation.

[0064] <Second embodiment> The information processing device 100 according to this embodiment recognizes the emotion of the subject from the video in which the subject appears and / or the voice of the subject, and evaluates the importance of keywords in the conversation between the subject (customer) and the store clerk. Next, the information processing device 100 creates information (also called a prompt) to be input to the generative AI based on the content of the remark and the importance of the keyword. Then, the information processing device 100 creates customer service support information for the store clerk by inputting the prompt to the generative AI. In this embodiment, a system for supporting store clerks who serve customers in a store is described, but it may also be applied to a system for supporting doctors in questioning patients in a medical field, for example. In the second embodiment, differences from the first embodiment are described.

[0065] The extraction unit 204 determines the importance of keywords in the voice, based on the voice of the conversation between the store clerk and the user (customer) acquired from the voice acquisition unit 201 and the emotion information acquired from the emotion determination unit 203. The extraction unit 204 also stores keyword information in which the keywords are linked to their importance in a database server (not shown) and / or HDD 107. Then, the extraction unit 204 creates a prompt to be input to the generative AI, using the same procedure as in the first embodiment.

[0066] Here, FIG. 10 is a diagram showing a prompt according to the second embodiment.

[0067] Prompt 1000 is input information for the generative AI of generating unit 205. Prompt 1000 includes keywords 1010, such as "Please keep the following points in mind when creating an introductory material for ABC" as an instruction to the generative AI when creating materials. Keywords 1010 are keywords used when creating materials, and include, for example, "Excellent autofocus, lightweight body, and good image quality." Here, compared to prompt 500, prompt 1000 does not limit the number of characters in the output (delivery) by the generative AI, and allows expressions other than text responses (for example, figures, tables, and images). Therefore, the deliverable (i.e., product introduction materials) generated by the generative AI based on prompt 1000 is different from the deliverable generated by the generative AI based on prompt 500.

[0068] Next, the generation unit 205 generates customer service support information by inputting the prompt 1000 in FIG. 10 generated by the extraction unit 204 into a generative AI (for example, Chat-GPT).

[0069] Here, FIG. 11 is a diagram showing customer service support information created by the generative AI according to the second embodiment.

[0070] Customer service support information 1100 includes areas 1110, 1120, and 1130. Area 1110 includes information for explaining the features of a product to a user in an easy-to-understand manner. Area 1110 includes, for example, a photo of the product and a catchphrase (e.g., "creative," "full size," "ABC"). Areas 1120 and 1130 include product descriptions corresponding to keyword 1010 in FIG. 10. For example, area 1120 includes a product description corresponding to keyword 1010, "lightweight body, good image quality." Area 1130 includes a product description corresponding to keyword 1010, "excellent autofocus."

[0071] Fig. 12 is a flowchart for explaining the process of creating customer service support information according to the second embodiment. The process of Fig. 12 is started, for example, when the information processing device 100 receives an instruction to start creating materials from a store clerk via the input unit 101.

[0072] S1201 to S1204 and S1206 are similar to S702 to S705 and S709 in the first embodiment, and therefore the description thereof will be omitted.

[0073] In S1205, input unit 101 determines whether or not there is an instruction to create customer service support information. If input unit 101 determines that there is an instruction to create customer service support information (Yes in S1205), it proceeds to S1207. On the other hand, if input unit 101 determines that there is no instruction to create customer service support information (No in S1205), it proceeds to S1206.

[0074] In S1207, the extraction unit 204 creates the prompt 1000 by inserting the keyword of the keyword information 400 into the template 920. Specifically, the extraction unit 204 acquires a specific template 920 from the template information 900 for creating the prompt 1000, based on the state of the conversation and the correction information. Furthermore, the extraction unit 204 corrects the content of the prompt 1000, based on the keyword information 400 and the correction information.

[0075] The extraction unit 204 acquires a specific template 920 from template information 900 stored in advance in a database server (not shown) or the like based on the state of the conversation. The template 920 includes, for example, "Please create a catalogue material for 'product name' based on 'keywords'." Next, the extraction unit 204 creates a prompt 1000 by inserting into the template 920 the product category, product name, and selected keyword that are managed together with the state of the conversation.

[0076] In S1208, the generation unit 205 generates customer service assistance information 1100 based on the prompt 1000 generated by the extraction unit 204. Then, the generation unit 205 displays the generated customer service assistance information 1100 on the display unit 102, and ends the process.

[0077] In the present embodiment, the creation of a product catalog has been described as an example of customer service support information 1100. However, as other customer service support information, the creation of product introduction videos and / or the creation of customer service messages may be implemented.

[0078] As described above, according to the second embodiment, customer service support information (i.e., a product catalog to be presented to the user) that supports the salesperson in serving customers can be efficiently created based on the conversation between the salesperson and the user (customer). This customer service support information can reduce the salesperson's burden in serving customers and further increase the user's desire to purchase.

[0079] <Other embodiments> Although the embodiment has been described above in detail, the present invention can be embodied, for example, as a system, an apparatus, a method, a program, or a recording medium (storage medium), etc. Specifically, the present invention may be applied to a system composed of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus composed of a single device.

[0080] Needless to say, the object of the present invention can be achieved by the following: That is, a recording medium (or storage medium) on which is recorded a program code (computer program) of software that realizes the functions of the above-mentioned embodiments is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. Then, a computer (or a CPU or MPU) of the system or device reads and executes the program code stored in the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-mentioned embodiments, and the recording medium on which the program code is recorded constitutes the present invention.

[0081] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0082] The disclosure of this specification includes the following information processing device, method, and program. (Item 1) an acquisition means for acquiring emotion information of a user based on at least one of a captured image and a voice of the user; a prompt generating means for generating a prompt based on the emotion information and keywords extracted from the speech; and an output unit that outputs an answer obtained by using the prompt generated by the prompt generating unit. (Configuration 2) 2. The information processing device according to item 1, wherein the prompt generating means generates a prompt based on one or more keywords identified based on the emotion information from a plurality of keywords extracted from the voice. (Item 3) 3. The information processing device according to item 1 or 2, wherein the prompt generating means generates a prompt based on one or more keywords associated with a specific type of emotion from among a plurality of keywords extracted from the voice. (Item 4) The information processing device described in any one of items 1 to 3, wherein the prompt generation means generates a prompt based on one or more keywords out of a plurality of keywords extracted from the voice, the keywords being associated with the specific emotion type and having a likelihood of being the emotion equal to or greater than a threshold value. (Item 5) A correction means for correcting a state of conversation based on keywords extracted from the voice is further provided, 5. The information processing device according to any one of items 1 to 4, wherein the prompt generating means generates a prompt according to a state of a conversation with a user. (Item 6) a correction means for correcting, when emotion information associated with a keyword at a first time point differs from emotion information associated with an identical keyword at a second time point that is earlier than the first time point, the emotion information of the keyword at the first time point based on the emotion information of the keyword at the second time point; The prompt generating means generates a prompt based on the one or more keywords identified based on the emotion information corrected by the correcting means. 6. An information processing device according to any one of items 1 to 5. (Item 7) 6. The information processing device according to item 5, wherein the correction means changes the state of the conversation in response to receiving a user operation to erase one or more keywords extracted from the voice from a display screen. (Item 8) 6. The information processing device according to item 5, wherein the correction means changes the conversation state to another conversation state different from the conversation state based on a proportion of the emotion information having a specific emotion. (Item 9) 9. The information processing device according to item 8, wherein the specific emotion includes at least one of the user's fear, sadness, anger, contempt, and disgust. (Item 10) 6. The information processing device according to item 5, wherein the prompt is generated by inserting the extracted keyword into a template according to the state of the conversation. (Item 11) 2. The information processing device according to item 1, wherein the acquisition means acquires emotion information of the user based on voice in a conversation between the user and another user different from the user. (Item 12) 2. The information processing device according to item 1, wherein the prompt generating means generates a prompt based on the content of the answer output by the output means and keywords extracted from the voice. (Item 13) 13. The information processing device according to any one of items 1 to 12, wherein the output means executes at least one of display control for displaying the answer on a display screen and voice control for outputting the answer by voice. (Item 14) A method executed by an information processing device, comprising: an acquisition step of acquiring emotion information of the user based on at least one of a captured image and a voice of the user; a prompt generating step of generating a prompt based on the emotion information and keywords extracted from the speech; and outputting an answer obtained using the prompt generated by the prompt generating step. (Item 15) A program for causing a computer to function as each of the means of the information processing device according to any one of items 1 to 13.

[0083] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0084] 101 Input section 102 Display section 103 Network I / F section 104 CPU 105 RAM 106 ROM 107 HDD 108 Data Bus

Claims

1. an acquisition means for acquiring emotion information of a user based on at least one of a captured image and a voice of the user; a prompt generating means for generating a prompt based on the emotion information and keywords extracted from the speech; and an output unit that outputs an answer obtained by using the prompt generated by the prompt generating unit.

2. The information processing apparatus according to claim 1 , wherein the prompt generating means generates a prompt based on one or more keywords identified based on the emotion information from a plurality of keywords extracted from the voice.

3. The information processing apparatus according to claim 2 , wherein the prompt generating means generates the prompt based on one or more keywords associated with a specific type of emotion from among a plurality of keywords extracted from the voice.

4. 4. The information processing device according to claim 3, wherein the prompt generation means generates a prompt based on one or more keywords among a plurality of keywords extracted from the voice, the one or more keywords being associated with the specific emotion type and having a likelihood of being the emotion equal to or greater than a threshold value.

5. A correction means for correcting a state of conversation based on keywords extracted from the voice is further provided, The information processing apparatus according to claim 1 , wherein the prompt generating means generates a prompt according to a state of a conversation with the user.

6. a correction means for correcting, when emotion information associated with a keyword at a first time and emotion information associated with a keyword that is the same as the keyword at a second time that is earlier than the first time, the emotion information of the keyword at the first time based on the emotion information of the keyword at the second time, The prompt generating means generates a prompt based on the one or more keywords identified based on the emotion information corrected by the correcting means.

5. The information processing device according to claim 2.

7. the correction means changes the state of the conversation in response to receiving a user operation to erase one or more keywords from a display screen among the keywords extracted from the voice. The information processing device according to claim 5 .

8. the correction means changes the conversation state to another conversation state different from the conversation state based on a ratio of the emotion information having a specific emotion. The information processing device according to claim 5 .

9. The particular emotion includes at least one of the user's fear, sadness, anger, contempt, and disgust; The information processing device according to claim 8.

10. The prompt is generated by inserting the extracted keywords into a template corresponding to the state of the conversation. The information processing device according to claim 5 .

11. The acquisition means acquires emotion information of the user based on a voice in a conversation between the user and another user different from the user. The information processing device according to claim 1 .

12. the prompt generating means generates a prompt based on the content of the answer output by the output means and keywords extracted from the voice. The information processing device according to claim 1 .

13. The information processing apparatus according to claim 1 , wherein the output means executes at least one of a display control for displaying the answer on a display screen and a voice control for outputting the answer by voice.

14. A method executed by an information processing device, comprising: an acquisition step of acquiring emotion information of the user based on at least one of a captured image and a voice of the user; a prompt generating step of generating a prompt based on the emotion information and keywords extracted from the speech; and outputting an answer obtained using the prompt generated by the prompt generating step.

15. On the computer, an acquisition step of acquiring emotion information of a user based on at least one of a captured image and a voice of the user; a prompt generation step of generating a prompt based on the emotion information and keywords extracted from the speech; an output step for outputting an answer obtained by using the prompt generated by the prompt generation step; and

Citation Information

Patent Citations

  • Prompt information generation method and voice robot thereof

    CN112185422A

  • System and method for inferring scenes based on visual context-free grammar model

    US20190251350A1

  • Chatbot device, chatbot system, and method of supporting selection by a plurality of users

    JP2021103411A