Information processing device, method, and program
The information processing apparatus addresses the challenge of improving generative AI response accuracy by using emotion and keyword analysis from user input to generate targeted prompts, resulting in more precise and relevant outputs.
Patent Information
- Application Number
- PCT/JP2024/037266
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-10-18
- Publication Date
- 2025-06-05
AI Technical Summary
Existing generative AI systems struggle with accurately processing user input due to the inclusion of unnecessary words and modifiers, particularly when voice input is used, making it difficult to distinguish important keywords from irrelevant ones.
An information processing apparatus that acquires user emotion information from captured images and voice, generates a prompt based on this information and extracted keywords from the voice input, and outputs answers generated using this prompt to improve the accuracy of responses.
The proposed solution enhances the accuracy of answers generated in response to user queries by effectively filtering out unnecessary information and focusing on key emotional and contextual cues.
Smart Images

Figure JP2024037266_05062025_PF_FP_ABST
Abstract
Description
Information processing device, method, and program
[0001] The present invention relates to an information processing device, method, and program.
[0002] Chatbots are being used to reduce labor costs and provide 24-hour support when responding to corporate inquiries and creating product plans. A chatbot is a program and / or device that automatically converses with a user via text and / or voice. Patent Document 1 discloses a method for conversing with a user based on the user's emotional information and browsing information estimated from keywords extracted from the user's voice information. Meanwhile, the spread of generative AI has made it possible to create more advanced suggestions than conventional chatbots. For example, if you input "Please tell me about famous tourist spots in Tokyo" into Chat-GPT, it will output information associating the names of several representative tourist spots with brief descriptions of the tourist spots. It is known that the quality of input information (also known as prompts) is important for generative AI.
[0003] Japanese Patent Application Laid-Open No. 2021-103411
[0004] M. Sarma, "Emotion identificationfrom raw speech signals using DNNs," Proc. Interspeech 2018S. Li, "Deep Facial Expression Recognition: A Survey", IEEE transactions on affective computing, 2020Sharma, P., & Li, Y. , "Self-Supervised Contextual Keyword and Keyphrase Retrieval with Self-Labelling", 2019
[0005] However, if the input information to the generative AI contains unnecessary words and / or modifiers, the expected output may not be obtained. In particular, when speech is used as input, it is difficult to distinguish between important and unnecessary keywords contained in the user's utterances, resulting in a decrease in the quality of the input information to the generative AI. Furthermore, Non-Patent Document 1 does not consider determining the importance of keywords contained in the user's utterances.
[0006] Therefore, an object of the present invention is to improve the accuracy of answers created in response to questions from users.
[0007] In order to achieve the object of the present invention, an information processing device according to one embodiment of the present invention comprises the following configuration: acquisition means for acquiring emotional information of a user based on at least one of a captured image and audio of the user, prompt generation means for generating a prompt based on the emotional information and keywords extracted from the audio, and output means for outputting an answer obtained using the prompt generated by the prompt generation means.
[0008] According to the present invention, it is possible to improve the accuracy of answers created in response to questions from users.
[0009] Other features and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings, in which the same or similar elements are designated by the same reference numerals.
[0010] The accompanying drawings are included in the specification, constitute a part thereof, illustrate embodiments of the present invention, and are used, together with the description thereof, to explain the principles of the present invention. A block diagram showing the hardware configuration of an information processing device according to a first embodiment. (Example 1) A block diagram showing the functional configuration of an information processing device according to a first embodiment. (Example 1) A diagram showing emotion information according to a first embodiment. (Example 1) A diagram showing keyword information according to a first embodiment. (Example 1) A diagram showing a prompt according to a first embodiment. (Example 1) A diagram showing an answer created by the generative AI according to the first embodiment. (Example 1) A flowchart explaining the process of creating an answer to a conversation according to the first embodiment. (Example 1) A diagram showing a UI displayed on a display unit according to the first embodiment. (Example 1) An enlarged view of a portion of the UI in FIG. 8A. (Example 1) A diagram explaining template information according to the first embodiment. (Example 1) A diagram showing a prompt according to a second embodiment. (Example 2) A diagram showing customer service support information created by the generative AI according to the second embodiment. (Example 2) A flowchart explaining the process of creating customer service support information according to the second embodiment. (Example 2)
[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0012] First Embodiment An information processing device 100 recognizes the emotion of a subject from a video in which the subject appears and / or the subject's voice. The information processing device 100 also creates information (also called a prompt) to be input to a generative AI based on the emotion and speech content of the subject. The information processing device 100 then inputs the prompt to the generative AI to create a response to a conversation with the subject (e.g., a customer visiting a store). In this embodiment, a system that provides customer service in a real store using a chatbot will be described as an example. Note that this embodiment can also be applied to situations where customers are served in a virtual store.
[0013] FIG. 1 is a block diagram showing the hardware configuration of an information processing apparatus according to the first embodiment.
[0014] The information processing device 100 includes an input unit 101, a display unit 102, a network I / F unit 103, a CPU 104, a RAM 105, a ROM 106, a HDD 107, and a data bus 108.
[0015] The input unit 101 includes, for example, at least one of a keyboard, a mouse, a touch panel, and a microphone, and receives input from a user.
[0016] The display unit 102 is, for example, a liquid crystal display, and displays information such as the results of various processes. The input unit 101 and the display unit 102 are connected to other functional units via a data bus 108 so as to be able to communicate with each other.
[0017] The network I / F unit 103 connects to an external device (not shown) via the Internet so that various types of information can be transmitted and received.
[0018] The CPU 104 reads out a control computer program stored in the ROM 106, loads it into the RAM 105, and executes various control processes. The CPU 104 executes an image processing program stored in the ROM 106 or the HDD 107, thereby realizing image processing of image data.
[0019] The RAM 105 is used as a temporary storage area for storing programs executed by the CPU 104, work memory, and the like.
[0020] The HDD 107 stores various types of information such as image data, setting parameters, various programs, etc. The HDD 107 can also receive data input from an external device (not shown) via the network I / F unit 103.
[0021] Image data and the like received from an external device (not shown) via the network I / F unit 103 is transmitted to and received from the CPU 104, RAM 105, and ROM 106 via the data bus 108.
[0022] FIG. 2 is a block diagram showing the functional configuration of the information processing apparatus according to the first embodiment.
[0023] The information processing device 100 includes an audio acquisition unit 201 , a video acquisition unit 202 , an emotion determination unit 203 , an extraction unit 204 , a generation unit 205 , and a correction unit 206 .
[0024] The voice acquisition unit 201 acquires the user's voice via the input unit 101 (for example, a microphone). The voice acquisition unit 201 may acquire the user's voice stored in the HDD 107, or may acquire the user's voice stored in an external device (for example, a server).
[0025] The video acquisition unit 202 acquires a video of the user captured by an external camera via the network I / F unit 103. The video acquisition unit 202 may also acquire a video of the user stored in the HDD 107, or may acquire a video of the user stored in an external device (for example, a server).
[0026] The emotion determination unit 203 generates emotion information of the person (i.e., the user) with whom the information processing device 100 is conversing, based on the audio acquired by the audio acquisition unit 201 and / or the video acquired by the video acquisition unit 202.
[0027] Here, FIG. 3 is a diagram showing emotion information according to the first embodiment.
[0028] The emotion information 300 includes emotion labels 310 that represent the user's emotions, such as "happiness," "surprise," "fear," "sadness," "anger," "contempt," "disgust," and "neutral," and emotion scores 320 for each emotion label 310. The emotion scores 320 are expressed as scalar values ranging from 0 to 1. The closer the black bar for an emotion label 310 is to 1, the higher the likelihood that the user has the emotion of the emotion label 310. Now, let's return to the description of FIG. 2.
[0029] The extraction unit 204 determines the importance of keywords included in the immediately preceding answer and / or the user's voice, based on the immediately preceding answer generated by the generation unit 205, the user's voice acquired from the voice acquisition unit 201, and emotion information 300 acquired from the emotion determination unit 203. The extraction unit 204 stores keyword information, in which keywords and emotion information 300 are linked for each time series of the conversation with the user, in a database server (not shown) and / or HDD 107.
[0030] Here, FIG. 4 is a diagram showing keyword information according to the first embodiment.
[0031] The keyword information 400 includes a date and time 410, a keyword 420, an emotion label 430, and an emotion score 440. The date and time 410 represents the time when the extraction unit 204 extracted the keyword 420 from the most recent answer and / or the user's voice. The keyword 420 is linked to the user's emotion label 430 and emotion score 440. For example, the keyword 420 "Sports day record" is linked to the emotion label 430 "Happiness" and the emotion score 440 "0.9." Note that a symbol representing the emotion is displayed to the right of the emotion in the emotion label 430, but this may not be displayed if necessary. Returning now to the description of FIG. 2 .
[0032] The extraction unit 204 creates a prompt based on the speech acquired from the speech acquisition unit 201, keyword information from the database server, conversation state information, and correction information (described below) acquired from the correction unit 206. Here, conversation state information refers to information that represents the purpose of the current conversation with the user. For example, assuming customer service in a store, the conversation state information includes information that represents a series of customer service flows, such as "initial state," "confirming the purpose of the visit," "proposing a product," and "explaining the details of the product." At this time, the extraction unit 204 internally stores information on product categories and / or product names according to the conversation state information. Product categories refer to the types of products sold in the store, such as cameras, vacuum cleaners, and refrigerators. Product names refer to specific product names, such as the model name of a camera.
[0033] Here, FIG. 5 is a diagram showing a prompt according to the first embodiment.
[0034] The prompt 500 is input information to the generation AI of the generation unit 205, which will be described later. The prompt 500 in FIG. 5 shows input information when "Explain the product in detail" is selected as the conversation state information. For example, the prompt 500 includes a message such as "Please keep the following points in mind and write an introduction to ABC in 200 characters or less" as instructions to the generation AI when creating an answer, as well as keywords 510. The keywords 510 are keywords used when creating an answer, and include, for example, "excellent autofocus, lightweight body, good image quality." Now, let us return to the explanation of FIG. 2.
[0035] The generation unit 205 generates a response by inputting the prompt 500 of FIG. 5 generated by the extraction unit 204 into a generative AI (for example, Chat-GPT).
[0036] FIG. 6 is a diagram showing an answer created by the generative AI according to the first embodiment.
[0037] Answer 600 is an answer created by the generative AI based on prompt 500 in Figure 5. Answer 600 includes a product description of ABC. In this case, answer 600 includes only a text description of the product, but may also include diagrams, tables, images, etc. depending on the instructions of prompt 500. Now, let us return to the explanation of Figure 2.
[0038] The display unit 102 presents the answer 600 (see FIG. 6) created by the generation unit 205 using the generative AI to the user.
[0039] The correction unit 206 corrects the emotion labels and / or emotion scores of the keyword information 400 and / or corrects the conversation state information based on the keyword information 400 in the database server (not shown) and / or the HDD 107 and the transition of the emotion information.
[0040] 7 is a flowchart illustrating the process of creating a response to a conversation according to the first embodiment. The process in FIG. 7 is started, for example, when the information processing device 100 receives an instruction to create a response from the user via the input unit 101.
[0041] In S701, the display unit 102 displays the answer (shown in FIG. 6 ) created by the generation unit 205 using generative AI. When the conversation state is in the initial state, the display unit 102 displays "What are you looking for?" as a general question to the user (e.g., a customer searching for a desired product). In this embodiment, the display of the answer portion of the display unit 102 is described as being updated based on the answer created in the processes of S702 to S709. At this time, the information processing device 100 may display, on the display unit 102, data obtained by converting the voice obtained in the conversation with the user into text and / or extracted keyword information 400.
[0042] Here, Fig. 8A is a diagram showing a UI displayed on the display unit according to the first embodiment, and Fig. 8B is a diagram showing an enlarged view of a part of the UI in Fig. 8A.
[0043] A region 810 on the left side of the screen 800 of the display unit 102 displays the contents of past conversations (i.e., conversation history) with a user (here, a customer searching for a desired camera). A region 820 in the upper right corner of the screen 800 displays the product category ("Mirrorless Camera") and the product name ("ABC"). A region 830 located below the region 820 displays keywords extracted from the conversation in the region 810. A position below the region 830 displays a store clerk avatar 840 serving the user. For example, the contents of the conversation in the region 810 corresponding to the keywords in the region 830 ("Good Image Quality," "Lightweight," "Autofocus") may be highlighted by bolding and / or coloring. Furthermore, a product description may be displayed on the product introduction page, with the hand 850 of the store clerk avatar 840 positioned in a position that characterizes the product (see FIG. 8B ). Returning now to the description of FIG. 7 .
[0044] In S702, the voice acquisition unit 201 acquires the voice of the user who is having a conversation, and the video acquisition unit 202 acquires a video showing the user who is having a conversation.
[0045] In S703, the emotion determination unit 203 analyzes the user's emotion during the conversation. The emotion determination unit 203 analyzes the user's emotion based on the user's voice and / or video of the user using a trained model that has undergone deep learning. Here, the emotion analysis process by the emotion determination unit 203 can be performed using known technology for analyzing the user's emotion from voice and / or video (images). The emotion determination unit 203 can acquire the user's emotion information from the voice by using the method described in Non-Patent Document 1. Furthermore, the emotion determination unit 203 can acquire the user's emotion information from video (specifically, images that constitute the video) by using the method described in Non-Patent Document 2.
[0046] In S704, the extraction unit 204 extracts keywords based on the most recent answer and the user's current voice. That is, the extraction unit 204 extracts one or more keywords from the content of answers previously presented to the user and the user's voice. These keywords are used to generate prompts. The extraction unit 204 extracts keywords from the user's voice using a known AI service. For example, keyword extraction may be performed based on a known technology for converting voice data to text and a known technology for automatically extracting characteristic phrases and / or words from the text. The extraction unit 204 converts the user's voice into text using an API for converting voice data to text, such as the Google Cloud Speech-to-Text API. Furthermore, the extraction unit 204 extracts keywords from the text using an API for automatically extracting characteristic phrases and / or words from the text, similar to the method described in Non-Patent Document 3.
[0047] In S705, the extraction unit 204 stores the keyword information 400, in which the extracted keywords are linked to the emotion information analyzed by the emotion determination unit 203, in a database server (not shown) and / or the HDD 107.
[0048] In S706, the extraction unit 204 creates a prompt by inserting the keywords from the keyword information 400 into a template (described below). The extraction unit 204 acquires template information for creating a prompt based on the state of the conversation and the correction information. Furthermore, the extraction unit 204 corrects the content of the prompt based on the keyword information 400 and the correction information.
[0049] The extraction unit 204 acquires "template information" that is stored in advance in a database server (not shown) or the like, based on the state of the conversation.
[0050] Here, FIG. 9 is a diagram illustrating template information according to the first embodiment.
[0051] The template information 900 is information that associates a conversation state 910 with a template 920 for creating a prompt. The conversation state 910 includes conversation states that represent a series of customer service flows, such as "initial state (state 1)," "confirming the purpose of the visit (state 2)," "proposing a product (state 3)," and "explaining the details of the product (state 4)." Note that the conversation state 910 may include fewer than four or five or more conversation states. The template 920 represents a standard prompt format for providing the optimal answer to the user depending on the conversation state 910. For example, if the conversation state 910 is the initial state (state 1), the template 920 corresponding to state 1 is "Are you looking for something?"
[0052] Next, the extraction unit 204 selects keywords to be applied to the template 920 based on the keyword information 400. The keyword selection may be performed based on the emotion information (emotion label 430 and emotion score 440) included in the keyword information 400 and predetermined keyword selection conditions. For example, the keyword selection conditions include "keywords for which the emotion is happiness and the emotion score is 0.8 or more" and "keywords for which the emotion is surprise and the emotion score is 0.8 or more."
[0053] The extraction unit 204 can select "record of athletic meet," "photo," "good image quality," "mirrorless camera," and "autofocus" from the keyword information 400 of FIG. 4 based on predetermined keyword selection conditions. The keyword selection conditions may also include "antonyms of keywords whose emotion is anger and whose emotion score is 0.9 or greater." For example, the extraction unit 204 can select "light," which is the antonym of "heavy," from the keyword information 400 of FIG. 4 based on "antonyms of keywords whose emotion is anger and whose emotion score is 0.9 or greater." That is, the extraction unit 204 can create a prompt based on one or more keywords associated with a specific emotion type from among multiple keywords extracted from speech. The extraction unit 204 can also generate a prompt based on one or more keywords associated with a specific emotion type from among the multiple keywords and whose likelihood of being the emotion is equal to or greater than a threshold.
[0054] Next, the extraction unit 204 creates a prompt by inserting the product category, product name, and selected keywords, which are managed together with the conversation state, into the corresponding locations in the template 920. Specifically, the extraction unit 204 creates a prompt by inserting each keyword selected (extracted) from the keyword information 400 into the "product category," "product name," and "keyword" in the template 920 in Fig. 9. Now, we return to the description of Fig. 7.
[0055] In S707, the generation unit 205 generates an answer to be presented to the user based on the prompt generated by the extraction unit 204. The answer generated in this manner is displayed on the display unit 102 in S701. While FIG. 2 shows an example in which the information processing device 100 includes the display unit 102, the information processing device 100 and the display unit 102 may be connected via a wired or wireless connection medium. In either case, the CPU 104 of the information processing device 100 executes display control for displaying the answer obtained based on the prompt on the display screen. Furthermore, the information processing device 100 of this embodiment can output the answer as audio instead of or in addition to displaying the answer.
[0056] In S708, the generation unit 205 determines whether the conversation with the user has ended. If the generation unit 205 determines that the conversation with the user has ended (Yes in S708), the process proceeds to S710. On the other hand, if the generation unit 205 determines that the conversation with the user has not ended (No in S708), the process proceeds to S709.
[0057] In S709, the correction unit 206 corrects the emotion information and conversation state based on the following conditions 1 to 4. Here, the corrected information refers to information obtained by correcting the emotion information (emotion label 430, emotion score 440) of the keyword information 400 or the conversation state 910 of the template information 900, which the correction unit 206 has corrected.
[0058] (Method of correcting emotional information based on condition 1) A method will be described with reference to FIG. 4 whereby correction section 206 modifies current emotional information based on past emotional information when past emotional information and current emotional information corresponding to the same keyword in keyword information 400 differ.
[0059] Based on the keyword 420 ("SLR camera") in the current (first time) keyword information 460, the correction unit 206 acquires past (second time) keyword information 450 containing "SLR camera" from the database server (not shown) and / or HDD 107. Here, keyword information 450 and keyword information 460 contain the keyword 420 ("SLR camera"). However, the emotion label 430 ("anger") in keyword information 460 differs from the emotion label 430 ("happiness") in keyword information 450. In this case, the correction unit 206 compares keyword information 450 and keyword information 460 and determines the emotion label 430 with the higher emotion score 440 from both pieces of keyword information. The emotion score 440 ("0.6") in keyword information 450 is higher than the emotion score 440 ("0.4") in keyword information 460. Therefore, correction unit 206 corrects "anger" in emotion label 430 of keyword information 460 to "happiness," and corrects emotion score 440 from "0.4" to "0.6." Note that in the correction method described here, the oldest emotion information in chronological order (emotion label 430 of keyword information 450) may be used as the current emotion information (emotion label 430 of keyword information 460, etc.), emotion information at the most recent time may be used, or emotion information based on the average value of all past emotion scores may be used.
[0060] (Method of correcting the conversation state based on condition 2) The following describes a method by which the correction unit 206 changes the conversation state to the next conversation state when it obtains information necessary for the conversation state transition (product category, product name, etc.). The correction unit 206 transitions the current conversation state to the next conversation state based on a predetermined condition for the conversation state transition.
[0061] For example, the conversation state transitions are defined in the following order: "initial state (hereinafter, "state 1")," "confirm purpose of visit (hereinafter, "state 2")," "suggest product (hereinafter, "state 3")," and "explain product details (hereinafter, "state 4"). Regarding states 1 to 4, the transition conditions from one conversation state to another are explained below. First, the transition condition from state 1 to state 2 includes the start of a conversation. The transition condition from state 2 to state 3 includes the acquisition of a product category. The transition condition from state 3 to state 4 includes the acquisition of a product name. If the current keyword information 400 includes a predetermined product category and / or product name, the correction unit 206 transitions the current conversation state (e.g., state 2) to the next conversation state (e.g., state 3).
[0062] (Method of Correcting the Conversation State Based on Condition 3) When a specific phrase is obtained during a conversation with a user, the correction unit 206 performs correction to return the current conversation state (e.g., State 2) to the previous conversation state (e.g., State 1). For example, phrases that return the conversation state to the previous state, such as "I want to see other products" or "I want to find products other than cameras," are set in advance. Then, when the correction unit 206 obtains the phrase during a conversation with a user, it returns the conversation state to the previous conversation state. Furthermore, when the user operates the product category and / or product name on the screen 800 of the display unit 102 to clear (delete) or change the product category and / or product name, the current conversation state may be returned to the previous conversation state. That is, the correction unit 206 can change the conversation state in response to receiving a user operation to delete one or more keywords extracted from the speech from the display screen.
[0063] (Method of Correcting the Conversation State Based on Condition 4) If the emotion information obtained for the most recent response indicates a negative emotion, the correction unit 206 corrects the conversation state to return it to the previous state. Here, negative emotions refer to "fear," "sadness," "anger," "contempt," and "disgust" in the emotion labels 310 of FIG. 3 . First, the correction unit 206 acquires keyword information from the keyword information 400 stored in the database server (not shown) and / or the HDD 107 for a predetermined period, for example, from one minute ago to the current time. If the number of negative emotions in the emotion information (emotion labels 430) included in the acquired keyword information exceeds a predetermined percentage, the correction unit 206 restores the current conversation state to the previous conversation state. For example, if the correction unit 206 determines to restore the current conversation state (State 2) to the previous conversation state (State 1), the correction unit 206 causes the generation unit 205 to create answers such as "Do you want to search for another product category?" and "Shall I suggest another product?" and displays one of the created answers on the display unit 102. Alternatively, the correction unit 206 may select an answer that is appropriate for the state of the conversation after correction from among the answers created by the generation unit 205 from the past to the present, and display the selected answer on the display unit 102.
[0064] As described above, according to the first embodiment, appropriate answers can be presented to the user in response to changes in the user's emotions, which can reduce stress felt by the answers presented by the information processing device 100 (specifically, the chatbot), thereby increasing the user's motivation to purchase.
[0065] Second Embodiment An information processing device 100 according to this embodiment recognizes the emotions of a subject from a video in which the subject appears and / or the subject's voice, and evaluates the importance of keywords in a conversation between the subject (customer) and a store clerk. Next, the information processing device 100 creates information (also called a prompt) to be input to a generative AI based on the content of the remarks and the importance of the keywords. The information processing device 100 then inputs the prompt to the generative AI to create customer service support information for the store clerk. This embodiment describes a system that supports store clerks who serve customers in a store. However, the system may also be applied to, for example, a system that supports doctors in questioning patients in a medical setting. The second embodiment will explain the differences from the first embodiment.
[0066] The extraction unit 204 determines the importance of keywords in the speech based on the speech of the conversation between the store clerk and the user (customer) acquired from the speech acquisition unit 201 and the emotion information acquired from the emotion determination unit 203. The extraction unit 204 also stores keyword information in which the keywords are linked to their importance in a database server (not shown) and / or the HDD 107. The extraction unit 204 then creates prompts to be input to the generative AI using the same procedure as in the first embodiment.
[0067] Here, FIG. 10 is a diagram showing a prompt according to the second embodiment.
[0068] Prompt 1000 is input information for the generative AI of generation unit 205. Prompt 1000 includes keywords 1010 as instructions to the generative AI when creating materials, such as "Please keep the following points in mind when creating ABC promotional materials." Keywords 1010 are keywords used when creating materials, and include, for example, "excellent autofocus, lightweight body, good image quality." Unlike prompt 500, prompt 1000 does not limit the number of characters in the output (deliverable) by the generative AI and allows expressions other than text responses (e.g., diagrams, tables, and images). Therefore, the deliverable (i.e., product promotional materials) generated by the generative AI based on prompt 1000 differs from the deliverable generated by the generative AI based on prompt 500.
[0069] Next, the generation unit 205 generates customer service support information by inputting the prompt 1000 of FIG. 10 generated by the extraction unit 204 into a generation AI (for example, Chat-GPT).
[0070] Here, FIG. 11 is a diagram showing customer service support information created by the generative AI according to the second embodiment.
[0071] The customer service support information 1100 includes areas 1110, 1120, and 1130. Area 1110 includes information for explaining the features of a product to a user in an easy-to-understand manner. Area 1110 includes, for example, a photo of the product and a catchphrase (e.g., "creative," "full size," "ABC"). Areas 1120 and 1130 include product descriptions corresponding to keyword 1010 in FIG. 10 . For example, area 1120 includes a product description corresponding to keyword 1010, "lightweight body, good image quality." Area 1130 includes a product description corresponding to keyword 1010, "excellent autofocus."
[0072] Fig. 12 is a flowchart illustrating the process of creating customer service assistance information according to the second embodiment. The process in Fig. 12 is started, for example, when the information processing device 100 receives an instruction from a store clerk via the input unit 101 to start creating materials.
[0073] Since steps S1201 to S1204 and S1206 are similar to steps S702 to S705 and S709 in the first embodiment, the description thereof will be omitted.
[0074] In S1205, the input unit 101 determines whether or not there is an instruction to create customer service assistance information. If the input unit 101 determines that there is an instruction to create customer service assistance information (Yes in S1205), the process proceeds to S1207. On the other hand, if the input unit 101 determines that there is no instruction to create customer service assistance information (No in S1205), the process proceeds to S1206.
[0075] In S1207, the extraction unit 204 creates the prompt 1000 by inserting the keywords in the keyword information 400 into the template 920. Specifically, the extraction unit 204 acquires a specific template 920 from the template information 900 for creating the prompt 1000, based on the state of the conversation and the correction information. Furthermore, the extraction unit 204 corrects the content of the prompt 1000 based on the keyword information 400 and the correction information.
[0076] The extraction unit 204 retrieves a specific template 920 from template information 900 stored in advance in a database server (not shown) or the like based on the state of the conversation. The template 920 includes, for example, "Please create catalog materials for 'product name' based on 'keywords'." Next, the extraction unit 204 creates a prompt 1000 by inserting the product category, product name, and selected keywords, which are managed together with the state of the conversation, into the template 920.
[0077] In S1208, the generation unit 205 generates customer service assistance information 1100 based on the prompt 1000 generated by the extraction unit 204. Then, the generation unit 205 displays the generated customer service assistance information 1100 on the display unit 102, and ends the process.
[0078] In this embodiment, the creation of a product catalog has been described as an example of customer service support information 1100, but other customer service support information may include the creation of product introduction videos and / or the creation of customer service messages.
[0079] As described above, according to the second embodiment, customer service support information (i.e., a product catalog to be presented to the user) that supports the sales clerk in serving customers can be efficiently created based on the conversation between the sales clerk and the user (customer). This customer service support information can reduce the sales clerk's burden of serving customers and further increase the user's desire to purchase.
[0080] <Other Embodiments> Although the exemplary embodiments have been described above in detail, the present invention can be embodied, for example, as a system, device, method, program, recording medium (storage medium), etc. Specifically, the present invention may be applied to a system made up of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus made up of a single device.
[0081] Needless to say, the object of the present invention can be achieved by the following: A recording medium (or storage medium) on which software program code (computer program) that realizes the functions of the above-described embodiments is recorded is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the recording medium on which the program code is recorded constitutes the present invention.
[0082] The present invention can also be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more of the functions.
[0083] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention.
[0084] This application claims priority based on Japanese Patent Application No. 2023-200138, filed November 27, 2023, the entire contents of which are incorporated herein by reference.
Claims
1. An information processing device comprising: an acquisition means for acquiring emotional information of a user based on at least one of a captured image and voice of the user; a prompt generation means for generating a prompt based on the emotional information and keywords extracted from the voice; and an output means for outputting an answer obtained using the prompt generated by the prompt generation means.
2. An information processing apparatus according to claim 1, wherein said prompt generating means generates a prompt based on one or more keywords identified based on said emotion information from a plurality of keywords extracted from said speech.
3. An information processing device according to claim 2, wherein said prompt generating means generates a prompt based on one or more keywords associated with a specific type of emotion from among a plurality of keywords extracted from said voice.
4. An information processing device as described in claim 3, wherein the prompt generation means generates a prompt based on one or more keywords out of a plurality of keywords extracted from the voice, which are associated with the specific type of emotion and have a likelihood of being that emotion equal to or greater than a threshold value.
5. An information processing device according to claim 1, further comprising a correction means for correcting a state of a conversation based on keywords extracted from the voice, wherein the prompt generation means generates a prompt according to the state of the conversation with the user.
6. An information processing device as described in any one of claims 2 to 4, further comprising a correction means for correcting the emotion information of the keyword at the first time based on the emotion information of the keyword at the second time when emotion information associated with a keyword at a first time differs from emotion information associated with the same keyword at a second time that is earlier than the first time, and the prompt generation means generates a prompt based on the one or more keywords identified based on the emotion information corrected by the correction means.
7. The information processing device according to claim 5, wherein the correction means changes the state of the conversation in response to receiving a user operation to erase one or more keywords from the keywords extracted from the voice from a display screen.
8. The information processing device according to claim 5, wherein said correction means changes said conversation state to another conversation state different from said conversation state based on a proportion of said emotion information having a specific emotion.
9. The information processing device according to claim 8, wherein the specific emotion includes at least one of the user's fear, sadness, anger, contempt, and disgust.
10. The information processing device according to claim 5, wherein the prompt is generated by inserting the extracted keyword into a template corresponding to the state of the conversation.
11. The information processing device according to claim 1, wherein the acquisition means acquires emotion information of the user based on voice in a conversation between the user and another user different from the user.
12. The information processing device according to claim 1, wherein said prompt generating means generates a prompt based on the content of the answer output by said output means and keywords extracted from the voice.
13. The information processing device according to claim 1, wherein said output means executes at least one of display control for displaying said answer on a display screen and voice control for outputting said answer as voice.
14. A method executed by an information processing device, comprising: an acquisition step of acquiring emotional information of a user based on at least one of a captured image and voice of the user; a prompt generation step of generating a prompt based on the emotional information and keywords extracted from the voice; and an output step of outputting an answer obtained using the prompt generated by the prompt generation step.
15. A program that causes a computer to execute an acquisition step of acquiring emotional information of a user based on at least one of a captured image and voice of the user; a prompt generation step of generating a prompt based on the emotional information and keywords extracted from the voice; and an output step of outputting an answer obtained using the prompt generated by the prompt generation step.
Citation Information
Patent Citations
Chatbot device, chatbot system, and method of supporting selection by a plurality of users
JP2021103411A
Information Processing Apparatus, Method, and Program
JP7706526B2
Prompt information generation method and voice robot thereof
CN112185422A