Information Processing Apparatus, Method, and Program
The information processing apparatus improves generative AI responses by using emotion and keyword analysis to filter and prioritize input, addressing the challenge of inaccurate outputs from unnecessary words in user input.
Patent Information
- Application Number
- JP2023200138
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-11-27
AI Technical Summary
Existing generative AI systems struggle with inaccurate outputs due to the inclusion of unnecessary words and modifiers in user input, particularly when distinguishing between important and unimportant keywords in voice-based interactions, which affects the quality of responses.
An information processing apparatus that acquires user emotion information from images and voice, determines the importance of keywords, and generates prompts for generative AI based on this information to improve response accuracy.
Enhances the accuracy of generative AI responses by filtering and prioritizing relevant keywords based on user emotion and conversation context, leading to improved interaction quality.
Smart Images

Figure 0007706526000001 
Figure 0007706526000002 
Figure 0007706526000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, method, and program.
Background Art
[0002] Inquiries to companies and creation of product plans, chatbots are utilized to reduce labor costs and provide 24-hour support. A chatbot is a program and / or device that automatically conducts conversations with users via text and / or voice. Patent Document 1 discloses a method of conducting conversations with users based on user emotion information estimated from keywords extracted from user voice information and browsing information. On the other hand, with the spread of generative AI, it has become possible to create more advanced proposals than conventional chatbots. For example, when "Please tell me famous tourist spots in Tokyo" is input to Chat-GPT, information associating the names of several representative tourist spots with a brief description of the tourist spots is output. Here, it is known that the quality of input information (also called a prompt) is important in generative AI.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
[0005] However, if the input information for generative AI contains unnecessary words and / or modifiers, the expected output may not be obtained. In particular, when inputting voice, it is difficult to distinguish between important keywords and unnecessary keywords included in the user's speech, and the quality of the input information for generative AI deteriorates. Also, Non-Patent Document 1 does not consider determining the importance of keywords included in the user's speech.
[0006] Therefore, an object of the present invention is to improve the accuracy of answers created in response to questions from users. [Means for Solving the Problems]
[0007] In order to achieve the object of the present invention, an information processing apparatus according to an embodiment of the present invention includes the following configuration. That is, acquisition means for acquiring the user's emotion information based on at least one of the user's photographed image and voice, and the emotion information Identified based on , keywords extracted from the voice, and , the state of the conversation with the user, and prompt generation means for generating a prompt based on, and output means for outputting an answer obtained using the prompt generated by the prompt generation means. [Effects of the Invention]
[0008] According to the present invention, the accuracy of answers created in response to questions from users can be improved.
Brief Description of Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Modes for Carrying Out the Invention
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are given the same reference numerals, and redundant descriptions are omitted.
[0011] <First Embodiment> The information processing apparatus 100 recognizes the emotion of a subject from an image in which the subject appears and / or the voice of the subject. Further, the information processing apparatus 100 creates information (also referred to as a prompt) to be input to the generative AI based on the emotion of the subject and the speech content. Then, the information processing apparatus 100 creates a response to a conversation with the subject (for example, a customer who has visited a store) by inputting the prompt to the generative AI. In this embodiment, a system for performing customer service in a physical store in the form of a chatbot will be described as an example. Note that this embodiment is also applicable to a scenario where customer service is provided to customers in a virtual store.
[0012] FIG. 1 is a block diagram showing the hardware configuration of the information processing apparatus according to the first embodiment.
[0013] The information processing apparatus 100 includes an input unit 101, a display unit 102, a network I / F unit 103, a CPU 104, a RAM 105, a ROM 106, an HDD 107, and a data bus 108.
[0014] The input unit 101 includes, for example, at least one of a keyboard, a mouse, a touch panel, and a microphone, and receives user input.
[0015] The display unit 102 is, for example, a liquid crystal display or the like, and displays information such as the results of various processes. The input unit 101 and the display unit 102 are communicably connected to other functional units via the data bus 108.
[0016] The network I / F unit 103 is connected to an external device (not shown) via the Internet so that various types of information can be transmitted and received.
[0017] The CPU 104 reads out the control computer program stored in the ROM 106, loads it into the RAM 105, and executes various control processes. By executing the image processing program stored in the ROM 106 or the HDD 107, image processing for the image data is realized.
[0018] The RAM 105 is used as a temporary storage area for storing programs or work memories that the CPU 104 executes.
[0019] The HDD 107 stores various types of information such as image data, setting parameters, or various programs. Also, the HDD 107 can receive data input from an external device (not shown) via the network I / F unit 103.
[0020] The image data etc. received from an external device (not shown) via the network I / F unit 103 are transmitted and received to and from the CPU 104, the RAM 105, and the ROM 106 via the data bus 108.
[0021] Figure 2 is a block diagram showing the functional configuration of the information processing apparatus according to the first embodiment.
[0022] The information processing apparatus 100 includes a voice acquisition unit 201, a video acquisition unit 202, an emotion determination unit 203, an extraction unit 204, a generation unit 205, and a correction unit 206.
[0023] The voice acquisition unit 201 acquires the user's voice via the input unit 101 (for example, a microphone). Also, the voice acquisition unit 201 may acquire the user's voice stored in the HDD 107, or may acquire the user's voice stored in an external device (for example, a server).
[0024] The video acquisition unit 202 acquires a video in which the user appears, which is taken by an external camera, via the network I / F unit 103. Further, the video acquisition unit 202 may acquire a video in which the user appears and is stored in the HDD 107, or may acquire a video in which the user appears and is stored in an external device (for example, a server).
[0025] Based on the voice acquired by the voice acquisition unit 201 and / or the video acquired by the video acquisition unit 202, the emotion determination unit 203 generates emotion information of the person (that is, the user) with whom the information processing apparatus 100 converses.
[0026] Here, FIG. 3 is a diagram showing the emotion information according to the first embodiment.
[0027] The emotion information 300 includes an emotion label 310 representing the user's emotions such as "happiness", "surprise", "fear", "sadness", "anger", "contempt", "disgust", "apathy", etc., and an emotion score 320 for each emotion label 310. The emotion score 320 is represented by a scalar value in the range of 0 to 1. The closer the black bar of the emotion label 310 is to 1, the higher the likelihood that the user has the emotion of the emotion label 310. Here, return to the description of FIG. 2.
[0028] Based on the immediately preceding answer generated by the generation unit 205, the user's voice acquired from the voice acquisition unit 201, and the emotion information 300 acquired from the emotion determination unit 203, the extraction unit 204 determines the importance of the keywords included in the immediately preceding answer and / or the user's voice. The extraction unit 204 stores keyword information in which keywords and the emotion information 300 are associated for each time series of conversations with the user in a database server (not shown) and / or the HDD 107.
[0029] Here, FIG. 4 is a diagram showing the keyword information according to the first embodiment.
[0030] The keyword information 400 includes a date and time 410, a keyword 420, an emotion label 430, and an emotion score 440. The date and time 410 represents the time when the extraction unit 204 extracted the keyword 420 from the most recent response and / or the user's voice. And the keyword 420 is associated with the user's emotion label 430 and emotion score 440. For example, the keyword "sports meet record" of the keyword 420 is associated with the emotion label "happiness" of the emotion label 430 and the emotion score "0.9" of the emotion score 440. Note that a symbol representing the emotion is displayed to the right of the emotion of the emotion label 430, but it may not be displayed if not necessary. Here, return to the description of FIG. 2.
[0031] The extraction unit 204 creates a prompt based on the voice obtained from the voice acquisition unit 201, the keyword information of the database server, the conversation state information, and the correction information (described later) obtained from the correction unit 206. Here, the conversation state information is information representing the purpose of the conversation with the current user. For example, assuming customer service in a store, the conversation state information includes information representing a series of processes of customer service such as "initial state", "confirming the purpose of coming to the store", "proposing products", and "explaining the details of the products". At this time, the extraction unit 204 internally holds information on product categories and / or product names according to the conversation state information. The product category means, for example, the types of products sold in the store such as cameras, vacuum cleaners, and refrigerators. The product name means, for example, a specific product name such as the model name of a camera.
[0032] Here, FIG. 5 is a diagram showing the prompt according to the first embodiment.
[0033] The prompt 500 is input information for the generative AI of the generation unit 205 described later. The prompt 500 in FIG. 5 shows the input information when "explain the details of the product" is selected as the conversation state information. For example, the prompt 500 includes a message such as "Please create an introduction to ABC within 200 characters while grasping the following points." as an instruction to the generative AI at the time of answer creation, and the keyword 510. The keyword 510 is a keyword at the time of answer creation and includes, for example, "excellent autofocus, light body, good image quality". Here, return to the description of FIG. 2.
[0034] The generation unit 205 creates an answer by inputting the prompt 500 in FIG. 5 created by the extraction unit 204 into a generative AI (for example, Chat-GPT).
[0035] FIG. 6 is a diagram showing an answer created by the generative AI according to the first embodiment.
[0036] The answer 600 is an answer created by the generative AI based on the prompt 500 in FIG. 5. The answer 600 includes a product description of ABC. At this time, the answer 600 only includes a literal description of the product, but may include figures, tables, images, etc. according to the instructions of the prompt 500. Here, return to the description of FIG. 2.
[0037] The display unit 102 presents the answer 600 (see FIG. 6) created by the generation unit 205 using the generative AI to the user.
[0038] The correction unit 206 corrects the sentiment label and / or sentiment score of the keyword information 400 and / or corrects the conversation state information based on the keyword information 400 of the database server (not shown) and / or the HDD 107 and the transition of the sentiment information.
[0039] FIG. 7 is a flowchart for explaining the answer creation process of the conversation according to the first embodiment. The process in FIG. 7 is started, for example, when the information processing device 100 receives a user's answer creation instruction via the input unit 101.
[0040] In S701, the display unit 102 displays the answer (illustrated in FIG. 6) created by the generation unit 205 using the generative AI. When the conversation state is the initial state, the display unit 102 displays "Are you looking for something?" as a general query to the user (for example, a customer looking for a desired product). In this embodiment, it will be described that the display of the answer part of the display unit 102 is updated based on the answers created in the processes of S702 to S709. At this time, the information processing apparatus 100 may display the data obtained by converting the voice obtained in the conversation with the user into text and / or the extracted keyword information 400 on the display unit 102.
[0041] Here, FIG. 8A is a diagram showing the UI displayed on the display unit according to the first embodiment. FIG. 8B is an enlarged view of a part of the UI of FIG. 8A.
[0042] In the left area 810 of the screen 800 of the display unit 102, the content of the past conversation with the user (here, a customer looking for a desired camera), that is, the conversation history, is displayed. In the upper right area 820 of the screen 800, the product category ("mirrorless camera") and the product name ("ABC") are displayed. In the area 830 located below the area 820, the keywords extracted from the conversation in the area 810 are displayed. Below the area 830, a clerk avatar 840 that conducts customer service for the user is displayed. For example, the content of the conversation in the area 810 corresponding to the keywords ("good image quality", "lightweight", "autofocus") in the area 830 may be highlighted by boldface and / or coloring. Also, a product description display may be performed by placing the hand 850 of the clerk avatar 840 at a position that is a feature of the product on the product introduction page (see FIG. 8B). Here, return to the description of FIG. 7.
[0043] In S702, the voice acquisition unit 201 acquires the voice of the user during the conversation. The video acquisition unit 202 acquires the video of the user during the conversation.
[0044] In S703, the emotion determination unit 203 analyzes the emotion of the user during the conversation. The emotion determination unit 203 performs user emotion analysis based on the user's voice and / or the video in which the user appears, using a trained model that has undergone deep learning. Here, the emotion analysis process by the emotion determination unit 203 can be executed by a known technique for analyzing the user's emotion from voice and / or video (image). The emotion determination unit 203 can obtain the user's emotion information from the voice by using the method of Non-Patent Document 1. Also, the emotion determination unit 203 can obtain the user's emotion information from the video (specifically, the images constituting the video) by using the method of Non-Patent Document 2.
[0045] In S704, the extraction unit 204 extracts keywords based on the previous answer and the current user's voice. That is, the extraction unit 204 extracts one or more keywords from the content of the answer presented to the user in the past and the user's voice. These keywords are used for the generation of prompts. The extraction unit 204 performs keyword extraction from the user's voice using a known AI service. For example, keyword extraction can be executed based on a known technique for converting voice data into text and a known technique for automatically extracting characteristic phrases and / or words from the text. The extraction unit 204 converts the user's voice into text by using an API such as the Google Cloud Speech-to-Text API for converting voice data into text. Also, the extraction unit 204 extracts keywords from the text by using an API for automatically extracting characteristic phrases and / or words from the text, similar to the method of Non-Patent Document 3.
[0046] In S705, the extraction unit 204 stores the keyword information 400 in which the extracted keywords are associated with the emotion information analyzed by the emotion determination unit 203 in a database server (not shown) and / or the HDD 107.
[0047] At S706, the extraction unit 204 creates a prompt by inserting the keywords of the keyword information 400 into the template described later. The extraction unit 204 acquires template information for creating a prompt based on the conversation state and correction information. Further, the extraction unit 204 corrects the content of the prompt based on the keyword information 400 and the correction information.
[0048] Based on the conversation state, the extraction unit 204 acquires "template information" that is pre-stored in a database server (not shown) or the like.
[0049] Here, FIG. 9 is a diagram for explaining the template information according to the first embodiment.
[0050] The template information 900 is information that associates the conversation state 910 with a template 920 for creating a prompt. The conversation state 910 includes conversation states representing a series of customer service processes such as "initial state (state 1)", "confirming the purpose of visiting the store (state 2)", "proposing a product (state 3)", and "explaining the details of the product (state 4)". Note that the conversation state 910 may include less than 4 or more than 5 conversation states. The template 920 represents a stereotype of a prompt for providing an optimal answer to the user according to the conversation state 910. For example, when the conversation state 910 is the initial state (state 1), the template 920 corresponding to state 1 is "Are you looking for something?"
[0051] Next, the extraction unit 204 selects keywords to be applied to the template 920 based on the keyword information 400. The selection of keywords can be executed based on the sentiment information (sentiment label 430 and sentiment score 440) included in the keyword information 400 and the predetermined keyword selection conditions. For example, the keyword selection conditions include "keywords with a sentiment of happiness and a sentiment score of 0.8 or higher", and "keywords with a sentiment of surprise and a sentiment score of 0.8 or higher".
[0052] The extraction unit 204 can select "sports meet records", "photos", "good image quality", "mirrorless camera", and "autofocus" from the keyword information 400 in FIG. 4 based on the established keyword selection conditions. Also, the keyword selection conditions may include "the antonyms of keywords with an anger emotion and an emotion score of 0.9 or higher". For example, the extraction unit 204 can select "light", which is the antonym of "heavy", from the keyword information 400 in FIG. 4 based on "the antonyms of keywords with an anger emotion and an emotion score of 0.9 or higher". That is, the extraction unit 204 can create a prompt based on one or more keywords associated with a specific type of emotion among the multiple keywords extracted from the voice. Also, the extraction unit 204 can generate a prompt based on one or more keywords that are associated with a specific type of emotion and have a likelihood of being that emotion equal to or higher than a threshold value among the multiple keywords.
[0053] Next, the extraction unit 204 creates a prompt by inserting the product category, product name, and selected keywords, which are managed together with the conversation state, into the corresponding locations in the template 920. Specifically, the extraction unit 204 inserts each keyword selected (extracted) from the keyword information 400 into the "product category", "product name", and "keyword" of the template 920 in FIG. 9 to create a prompt. Here, return to the description of FIG. 7.
[0054] In S707, the generation unit 205 creates a response to be presented to the user based on the prompt created by the extraction unit 204. The response generated in this way is displayed by the display unit 102 in S701. Although FIG. 2 shows an example where the information processing apparatus 100 has the display unit 102, the information processing apparatus 100 and the display unit 102 may be connected by a wired or wireless connection medium. In any case, the CPU 104 of the information processing apparatus 100 executes display control for displaying the response obtained based on the prompt on the display screen. Also, the information processing apparatus 100 of the present embodiment can output the response as voice instead of or in addition to the display of the response.
[0055] In S708, the generation unit 205 determines whether the conversation with the user has ended. If the generation unit 205 determines that the conversation with the user has ended (Yes in S708), the process proceeds to S710. On the other hand, if the generation unit 205 determines that the conversation with the user has not ended (No in S708), the process proceeds to S709.
[0056] In S709, the correction unit 206 corrects the emotion information and the conversation state based on the following conditions 1 to 4. Here, the correction information refers to the information obtained by the correction unit 206 correcting the emotion information (emotion label 430, emotion score 440) of the keyword information 400 or the conversation state 910 of the template information 900.
[0057] (Method of correcting emotion information based on condition 1) The method of correcting the current emotion information based on the past emotion information when the past emotion information corresponding to the same keyword in the keyword information 400 is different from the current emotion information by the correction unit 206 will be described with reference to FIG. 4.
[0058] Based on the keyword 420 (“single-lens reflex camera”) of the keyword information 460 at the current (first time), the correction unit 206 acquires the keyword information 450 of the past (second time) including “single-lens reflex camera” from a database server (not shown) and / or the HDD 107. Here, the keyword information 450 and the keyword information 460 include the keyword 420 (“single-lens reflex camera”). However, the emotion label 430 (“anger”) of the keyword information 460 is different from the emotion label 430 (“happiness”) of the keyword information 450. In this case, the correction unit 206 compares the keyword information 450 and the keyword information 460, and determines the emotion label 430 with the higher numerical value of the emotion score 440 from among the keyword information of both. The emotion score 440 (“0.6”) of the keyword information 450 is higher than the emotion score 440 (“0.4”) of the keyword information 460. Therefore, the correction unit 206 corrects the “anger” of the emotion label 430 of the keyword information 460 to “happiness” and corrects the “0.4” of the emotion score 440 to “0.6”. In the correction method described here, as the current emotion information (such as the emotion label 430 of the keyword information 460), the oldest emotion information (the emotion label 430 of the keyword information 450) in time series may be adopted, the emotion information at the latest time may be adopted, or the emotion information based on the average value of all past emotion scores may be adopted.
[0059] (Method for correcting the conversation state based on condition 2) When the correction unit 206 obtains information (product category, product name, etc.) necessary for the conversation state transition, a method for changing the conversation state to the next conversation state will be described. The correction unit 206 transitions the current conversation state to the next conversation state based on the conditions of the default conversation state transition.
[0060] For example, it is defined that the conversation state transitions in the order of "initial state (hereinafter referred to as state 1)", "confirm the purpose of visiting the store (hereinafter referred to as state 2)", "propose a product (hereinafter referred to as state 3)", and "explain the details of the product (hereinafter referred to as state 4)". Regarding the transition conditions from one conversation state to another in states 1 to 4, an explanation is provided. First, the transition condition from state 1 to state 2 includes that the conversation has started. The transition condition from state 2 to state 3 includes that a product category has been obtained. The transition condition from state 3 to state 4 includes that a product name has been obtained. And when a predetermined product category and / or product name is included in the current keyword information 400, the correction unit 206 transitions the current conversation state (for example, state 2) to the next conversation state (for example, state 3).
[0061] (Method for correcting the conversation state based on condition 3) When a specific phrase is obtained during the conversation with the user, the correction unit 206 performs a correction to return the current conversation state (for example, state 2) to the previous conversation state (for example, state 1). For example, phrases for returning the conversation state such as "want to see other products" and "want to search for products other than cameras" are preset. And when the correction unit 206 obtains the above phrases during the conversation with the user, it returns the conversation state to the previous conversation state. Also, when the user operates on the product category and / or product name on the screen 800 of the display unit 102 and clears (erases) or changes the product category and / or product name, the current conversation state may be returned to the previous conversation state. That is, the correction unit 206 can change the conversation state in response to receiving a user operation to delete from the display screen one or more keywords among the keywords extracted from the voice.
[0062] (Method for correcting the conversation state based on condition 4) If the sentiment information obtained for the most recent response is negative sentiment, the correction unit 206 performs a correction to return the conversation state to the previous state. Here, negative sentiment refers to "fear", "sadness", "anger", "contempt", and "disgust" in the sentiment label 310 of FIG. 3. First, the correction unit 206 acquires keyword information within a preset period, for example, from one minute before the current time to the current time, among the keyword information 400 in the database server (not shown) and / or the HDD 107. If the number of negative sentiments in the sentiment information (sentiment label 430) included in the acquired keyword information exceeds a preset ratio, the correction unit 206 returns the current conversation state to the previous conversation state. For example, when the correction unit 206 determines to return from the current conversation state (state 2) to the previous conversation state (state 1), it causes the generation unit 205 to create responses such as "Do you want to search for other product categories?" and "Shall I propose other products?", and displays any of the created responses on the display unit 102. Alternatively, the correction unit 206 may select a response suitable for the corrected conversation state from among the responses created by the generation unit 205 from the past to the present, and display the selected response on the display unit 102.
[0063] As described above, according to the first embodiment, an appropriate response can be presented to the user according to the change in the user's sentiment. Thereby, it is possible to prevent the user from feeling stress with respect to the response presented by the information processing apparatus 100 (specifically, the chatbot), and it is possible to enhance the user's purchasing desire.
[0064] <Second Embodiment> The information processing apparatus 100 according to this embodiment recognizes the emotion of a subject from the video in which the subject appears and / or the voice of the subject, and evaluates the importance of keywords in the conversation between the subject (customer) and the store clerk. Next, the information processing apparatus 100 creates information (also referred to as a prompt) to be input to the generative AI based on the utterance content and the importance of the keywords. Then, the information processing apparatus 100 creates customer service support information for the store clerk by inputting the prompt to the generative AI. In this embodiment, a system for supporting a store clerk who provides customer service in a store will be described. However, for example, it may be applied to a system for supporting a doctor's interview with a patient in a medical field. Note that in the second embodiment, the differences from the first embodiment will be described.
[0065] The extraction unit 204 determines the importance of keywords in the voice based on the voice of the conversation between the store clerk and the user (customer) acquired from the voice acquisition unit 201 and the emotion information acquired from the emotion determination unit 203. In addition, the extraction unit 204 stores the keyword information in which the keyword and the importance are associated in a database server (not shown) and / or the HDD 107. Then, the extraction unit 204 creates a prompt to be input to the generative AI in the same procedure as in the first embodiment.
[0066] Here, FIG. 10 is a diagram showing a prompt according to the second embodiment.
[0067] The prompt 1000 is input information for the generative AI of the generation unit 205. The prompt 1000 includes the keyword 1010 as an instruction to the generative AI when creating a document, such as "Please create an introduction document for ABC while emphasizing the following points." The keyword 1010 is a keyword when creating a document and includes, for example, "excellent autofocus, light body, good image quality". Here, the prompt 1000 does not limit the number of characters of the output (work product) by the generative AI compared to the prompt 500, and allows expressions other than text answers (for example, figures, tables, and images). Therefore, the work product (that is, the product introduction document) generated by the generative AI based on the prompt 1000 is different from the work product generated by the generative AI based on the prompt 500.
[0068] Next, the generation unit 205 creates customer service support information by inputting the prompt 1000 in FIG. 10 created by the extraction unit 204 into a generative AI (e.g., Chat-GPT).
[0069] Here, FIG. 11 is a diagram showing customer service support information created by the generative AI according to the second embodiment.
[0070] The customer service support information 1100 includes an area 1110, an area 1120, and an area 1130. The area 1110 includes information for clearly explaining the features of the product to the user. The area 1110 includes, for example, a product photo and catchphrases (e.g., "Creative", "Full size", "ABC"). The areas 1120 and 1130 include product descriptions corresponding to the keyword 1010 in FIG. 10. For example, the area 1120 includes a product description corresponding to "The body is light and the image quality is good" of the keyword 1010. The area 1130 includes a product description corresponding to "The autofocus is excellent" of the keyword 1010.
[0071] FIG. 12 is a flowchart for explaining the creation process of the customer service support information according to the second embodiment. The process in FIG. 12 is started, for example, when the information processing apparatus 100 receives an instruction to start creating materials for a store clerk via the input unit 101.
[0072] Since S1201 to S1204 and S1206 are the same as S702 to S705 and S709 of the first embodiment, the description thereof is omitted.
[0073] In S1205, the input unit 101 determines whether there is an instruction to create customer service support information. When the input unit 101 determines that there is an instruction to create customer service support information (Yes in S1205), the process proceeds to S1207. On the other hand, when the input unit 101 determines that there is no instruction to create customer service support information (No in S1205), the process proceeds to S1206.
[0074] In S1207, the extraction unit 204 creates the prompt 1000 by inserting the keywords of the keyword information 400 into the template 920. Specifically, the extraction unit 204 obtains a specific template 920 from among the template information 900 for creating the prompt 1000 based on the conversation state and the correction information. Further, the extraction unit 204 corrects the content of the prompt 1000 based on the keyword information 400 and the correction information.
[0075] Based on the conversation state, the extraction unit 204 obtains a specific template 920 from among the template information 900 stored in advance in a database server (not shown) or the like. The template 920 includes, for example, "Please create a catalog document of the 'product name' based on the 'keyword'." Next, the extraction unit 204 creates the prompt 1000 by inserting the product category, product name, and selected keywords managed together with the conversation state into the template 920.
[0076] In S1208, the generation unit 205 creates the customer service support information 1100 based on the prompt 1000 created by the extraction unit 204. Then, the generation unit 205 displays the created customer service support information 1100 on the display unit 102 and ends the process.
[0077] In this embodiment, product catalog creation has been described as an example of the customer service support information 1100, but creation of an introduction video of the product and / or creation of customer service phrases may be performed as other customer service support information.
[0078] As described above, according to the second embodiment, customer service support information (that is, a product catalog presented to the user) for supporting the customer service of the store clerk can be efficiently created based on the conversation between the store clerk and the user (customer). This customer service support information can reduce the customer service burden on the store clerk and further enhance the purchasing desire of the user.
[0079] <Other Embodiments> Although the above embodiments have been described in detail, the present invention can be implemented, for example, as an embodiment such as a system, apparatus, method, program, or recording medium (storage medium). Specifically, it may be applied to a system composed of a plurality of devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or it may also be applied to an apparatus composed of a single device.
[0080] Needless to say, the object of the present invention is achieved by the following means. That is, a recording medium (or storage medium) storing a program code (computer program) of software that realizes the functions of the above-described embodiments is supplied to a system or an apparatus. Needless to say, such a storage medium is a computer-readable storage medium. Then, a computer (or a CPU or MPU) of the system or apparatus reads and executes the program code stored in the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the recording medium storing the program code constitutes the present invention.
[0081] The present invention can also be realized by a process in which a program that realizes one or more functions of the above-described embodiments is supplied to a system or an apparatus via a network or a storage medium, and one or more processors in the computer of the system or apparatus read and execute the program. It can also be realized by a circuit (for example, an ASIC) that realizes one or more functions.
[0082] The disclosure of this specification includes the following information processing apparatus, method, and program. (Item 1) An acquisition means for acquiring the user's emotion information based on at least one of the user's captured image and voice; A prompt generation means for generating a prompt based on the emotion information and a keyword extracted from the voice; An information processing apparatus comprising an output means for outputting an answer obtained by using the prompt generated by the prompt generation means. (Configuration 2) The prompt generation means generates a prompt based on one or more keywords specified based on the emotion information from among a plurality of keywords extracted from the voice, according to the information processing apparatus of Item 1. (Item 3) The prompt generation means generates a prompt based on one or more keywords associated with a specific emotion type among the plurality of keywords extracted from the voice, according to the information processing apparatus of Item 1 or 2. (Item 4) The prompt generation means generates a prompt based on one or more keywords associated with the specific emotion type and having a likelihood of being that emotion equal to or greater than a threshold value among the plurality of keywords extracted from the voice, according to the information processing apparatus of any one of Items 1 to 3. (Item 5) The information processing apparatus further includes correction means for correcting the state of the conversation based on the keywords extracted from the voice. The prompt generation means generates a prompt according to the state of the conversation with the user, according to the information processing apparatus of any one of Items 1 to 4. (Item 6) The information processing apparatus further includes correction means for correcting the emotion information of the keyword at the first time based on the emotion information associated with the keyword at the second time, which is earlier than the first time and is the same as the keyword, when the emotion information associated with the keyword at the first time is different from the emotion information associated with the keyword at the second time. The prompt generation means generates a prompt based on the one or more keywords specified based on the emotion information corrected by the correction means. The information processing apparatus according to any one of Items 1 to 5. (Item 7) The correction means changes the state of the conversation in response to receiving a user operation to delete one or more keywords extracted from the voice from the display screen, according to the information processing apparatus of Item 5. (Item 8) The correction means changes the state of the conversation to a state of another conversation different from the state of the conversation based on the ratio of the emotional information having a specific emotion, according to the information processing apparatus of item 5. (Item 9) The specific emotion includes at least any one of fear, sadness, anger, contempt, and disgust of the user, according to the information processing apparatus of item 8. (Item 10) The prompt is generated by inserting the extracted keyword into a template according to the state of the conversation, according to the information processing apparatus of item 5. (Item 11) The acquisition means acquires the emotional information of the user based on the voice in the conversation between the user and another user different from the user, according to the information processing apparatus of item 1. (Item 12) The prompt generation means generates a prompt based on the content of the answer output by the output means and the keyword extracted from the voice, according to the information processing apparatus of item 1. (Item 13) The output means executes at least one of display control for displaying the answer on a display screen and voice control for voice-outputting the answer, according to the information processing apparatus of any one of items 1 to 12. (Item 14) A method executed by an information processing apparatus, comprising: an acquisition step of acquiring the emotional information of the user based on at least one of a captured image and voice of the user; a prompt generation step of generating a prompt based on the emotional information and a keyword extracted from the voice; an output step of outputting an answer obtained by using the prompt generated in the prompt generation step. (Item 15) A program for causing a computer to function as each means of the information processing apparatus according to any one of items 1 to 13.
[0083] The invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the claims are appended to disclose the scope of the invention.
Explanation of Signs
[0084] 101 Input section 102 Display section 103 Network I / F section 104 CPU 105 RAM 106 ROM 107 HDD 108 Data bus
Claims
1. An acquisition means for acquiring the user's emotional information based on at least one of the user's photographed image and voice; A prompt generation means for generating a prompt based on the keyword extracted from the voice and the conversation state with the user, which are specified based on the emotional information; An output means for outputting an answer obtained using the prompt generated by the prompt generation means, comprising: An information processing apparatus.
2. The prompt generation means generates a prompt based on one or more keywords specified based on the emotional information from among a plurality of keywords extracted from the voice. The information processing apparatus according to claim 1.
3. The prompt generation means generates a prompt based on one or more keywords associated with a specific type of emotion among a plurality of keywords extracted from the voice. The information processing apparatus according to claim 2.
4. The prompt generation means generates a prompt based on one or more keywords associated with the specific type of emotion and having a likelihood of being that emotion equal to or greater than a threshold value among a plurality of keywords extracted from the voice. The information processing apparatus according to claim 3.
5. Further comprising a correction means for correcting the conversation state based on the keyword extracted from the voice, The prompt generation means generates a prompt according to the conversation state. The information processing apparatus according to claim 1.
6. When the emotional information associated with the keyword at the first time is different from the emotional information associated with the same keyword as the keyword at the second time, which is earlier than the first time, further comprising a correction means for correcting the emotional information of the keyword at the first time based on the emotional information of the keyword at the second time, The prompt generation means generates a prompt based on one or more keywords specified based on the emotional information corrected by the correction means. The information processing apparatus according to any one of claims 2 to 4.
7. The correction means changes the conversation state in response to receiving a user operation to delete one or more keywords from among the keywords extracted from the voice from the display screen. The information processing apparatus according to claim 5.
8. The correction means changes the state of the conversation to a state of another conversation different from the state of the conversation based on the ratio of the emotion information having a specific emotion. The information processing apparatus according to claim 5.
9. The specific emotion includes at least any one of fear, sadness, anger, contempt, and disgust of the user. The information processing apparatus according to claim 8.
10. The prompt is generated by inserting the extracted keyword into a template according to the state of the conversation. The information processing apparatus according to claim 1.
11. The acquisition means acquires the emotion information of the user based on the voice in the conversation between the user and another user different from the user. The information processing apparatus according to claim 1.
12. The prompt generation means generates a prompt based on the content of the answer output by the output means and the keyword extracted from the voice. The information processing apparatus according to claim 1.
13. The output means executes at least one of display control for displaying the answer on a display screen and voice control for outputting the answer as voice. The information processing apparatus according to claim 1.
14. A method executed by an information processing apparatus, an acquisition step of acquiring the emotion information of the user based on at least any one of a photographed image and voice of the user; a prompt generation step of generating a prompt based on the keyword extracted from the voice and the state of the conversation between the user, which are specified based on the emotion information; an output step of outputting an answer obtained by using the prompt generated in the prompt generation step. Method.
15. Causing a computer to an acquisition procedure of acquiring the emotion information of the user based on at least any one of a photographed image and voice of the user; a prompt generation procedure of generating a prompt based on the keyword extracted from the voice and the state of the conversation between the user, which are specified based on the emotion information; a program for executing an output procedure of outputting an answer obtained by using the prompt generated in the prompt generation procedure.
Citation Information
Patent Citations
Prompt information generation method and voice robot thereof
CN112185422A
Chatbot device, chatbot system, and method of supporting selection by a plurality of users
JP2021103411A
System and method for inferring scenes based on visual context-free grammar model
US20190251350A1
Cited By
Information processing device, method, and program
WO2025115447A1