Information terminal and user assistance information presentation method
The information terminal with generative AI provides real-time user assistance, addressing gaps in decision-making and preventing fraud by generating tailored support in daily life scenarios.
Patent Information
- Application Number
- PCT/JP2024/022412
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-12-26
AI Technical Summary
Existing systems fail to provide comprehensive user assistance in various daily life situations beyond responding to fraudulent phone calls, leaving individuals vulnerable to fraud and poor decision-making.
An information terminal equipped with a processor, input and output devices, and a generative AI that generates user assistance information based on input data, providing real-time support through audio and visual outputs.
Enhances user awareness of information gaps and decision-making shortcomings, preventing fraud and improving decision-making in daily activities.
Smart Images

Figure JP2024022412_26122025_PF_FP_ABST
Abstract
Description
Information terminal and user support information presentation method
[0001] The present invention relates to an information terminal and a method for presenting user assistance information.
[0002] Patent Document 1 discloses that "a host system, a user terminal, and a telephone used by the user are freely connected to each other via a network, and a transfer fraud prevention system is provided which identifies calls to the telephone that are suspected to be transfer fraud. The host system comprises a means for converting the voice of the caller calling into text, a means for comparing keywords in the converted text with keywords in a keyword database provided in the host system, a means for comparing the telephone number of the caller to the telephone with a fraudulent phone number in a fraudulent phone number database provided in the host system, and a means for sending warning information to the telephone if it is determined to be fraudulent (summary excerpt)."
[0003] Japanese Patent Application Laid-Open No. 2007-323107
[0004] When a user performs daily life or work tasks, advice from others different from the user can help the user avoid becoming a victim of crime, receive shopping assistance, or smoothly perform high-value-added tasks. As such, situations in which advice from others different from the user is needed are diverse and not limited to phone calls. According to Patent Document 1, there is a problem in that it is not possible to provide support for the user's daily life or work in a wide variety of situations other than responding to fraudulent phone calls.
[0005] This technology was developed in light of this situation, and aims to provide a means of alerting users to information assistance information that complements information that the user is not good at or shortcomings in thinking and decision-making, for example, in everyday activities (shopping, consultations, phone calls, etc.), thereby preventing users from falling victim to malicious solicitations and fraud, as well as unnecessary shopping.
[0006] In order to achieve the above object, the present invention has the configurations described in the claims, for example, an information terminal including a processor, an input device, and an output device that performs at least one of audio and video display, wherein the processor generates question data for obtaining user assistance information to be presented to a user based on information input from the input device, obtains answer data to the question data generated by inputting the question data into a trained model, generates the user assistance information based on the answer data, and performs at least one of audio output and visual output of the user assistance information.
[0007] According to the present invention, it is possible to draw attention to information that a user is not good at or to shortcomings in thinking and decision-making by providing user support information that complements the information that the user is not good at, and to provide a means for preventing damage from malicious solicitations, fraud, etc., wasteful shopping, etc. Objects, configurations, and effects other than those described above will be made clear in the embodiments described below.
[0008] 1 is a schematic configuration diagram of a user assistance information presentation system according to a first embodiment; FIG. 2 is a hardware configuration diagram of an information terminal (smartphone); FIG. 3 is a hardware configuration diagram of an information terminal (smart glasses); FIG. 4 is a block diagram showing the functional configuration of the information terminal according to the first embodiment; FIG. 5 is a flowchart showing the processing flow of the user assistance information presentation system according to the first embodiment; FIG. 6 is an explanatory diagram of the operation of the information terminal according to the first embodiment; FIG. 7 is a diagram showing an example of a setting screen of a user assistance information provision app displayed on the information terminal; FIG. 8 is a diagram showing example training data of a trained model used in the user assistance information presentation system; FIG. 9 is a functional block diagram of an information terminal according to a second embodiment; FIG. 10 is a diagram showing a first example of a character setting screen; FIG. 11 is a diagram showing a second example of an assistant character setting screen; FIG. 12 is a main flowchart of a first example of an information terminal according to the second embodiment; FIG. 13 is a main flowchart explaining an example of the operation of providing user assistance information of the assistant function of the information terminal according to the second embodiment; FIG. 14 is a hardware configuration diagram of a head-mounted display information terminal (HMD) according to a third embodiment; FIG. 15 is a functional block diagram of an information terminal (HMD) according to the third embodiment; FIG. 16 is an external view of the information terminal (HMD) according to the third embodiment; FIG. 17 is a diagram showing an example display of the third embodiment; FIG. 18 is a flowchart showing the processing flow of an avatar setting unit in the information terminal according to the third embodiment; FIG. 1 is a second system overview diagram of an application example of the head-mounted information terminal of the third embodiment. FIG. 2 is a hardware configuration diagram of an information terminal of the fourth embodiment. FIG. 3 is a functional block diagram of the information terminal of the fourth embodiment. FIG. 4 is a diagram showing an example of video information visible from the information terminal of the fourth embodiment. FIG. 5 is a flowchart showing the flow of processing video information of the assistant function in augmented reality using the information terminal of the fourth embodiment. FIG. 6 is a flowchart showing the flow of processing audio information of the assistant function in augmented reality using the information terminal of the fourth embodiment. FIG. 7 is a functional block diagram of an information terminal of the fifth embodiment. FIG. 8 is a schematic diagram of operation using the information terminal of the fifth embodiment. FIG. 9 is a functional block diagram of an information terminal of the sixth embodiment. FIG. 10 is a flowchart showing the flow of information between the generation AI, the user, and the interlocutor, focusing on the information terminal of the sixth embodiment. FIG. 11 is a diagram showing an example of presenting user assistance information in the sixth embodiment.Fig. 13 is a flowchart showing a flow of processing by a supplemental information setting unit of an information terminal according to the sixth embodiment Fig. 14 is a flowchart showing another example (character determination of a user, etc.) of the information terminal according to the sixth embodiment.
[0009] The information terminal according to the present invention can provide user assistance information for supplementing information, thoughts, and decisions that the user is weak at for each scene of daily life activities in the virtual world or the real world. Therefore, the present invention can increase the commercial value of information processing devices and information processing systems to which the present invention is applied, and is therefore expected to contribute to Goal 8.2 of the Sustainable Development Goals (SDGs) advocated by the United Nations (increasing economic productivity through diversification, technological improvement, and innovation, particularly in industries that increase the value of goods and services and labor-intensive industries).
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The same components are designated by the same reference numerals throughout the drawings, and duplicated explanations will be omitted.
[0011] First Embodiment FIG. 1 is a schematic diagram of a user assistance information presentation system 100 according to a first embodiment.
[0012] The user assistance information presentation system 100 shown in FIG. 1 is configured by connecting an information terminal 10 to a server 1 including a generation AI 3 via a network 2. The information terminal 10 generates a question (corresponding to question data) for the generation AI 3 based on input information directly input by a user 5 or input information collected by the information terminal 10. A user assistance information presentation program 300 summarizes the answer (corresponding to answer data) obtained from the generation AI 3 to generate user assistance information for the user 5 and provides it to the user 5. The user assistance information presentation system 100 can provide user assistance information in real time in response to input information input in real time. While FIG. 1 illustrates one server 1 including the generation AI 3, a server group including multiple servers may also be used. In this case, each server may include a different generation AI 3, and the information terminal 10 may select a server to use depending on the application. The above phrase "each server includes a different generation AI 3" may include generation AIs that use the same algorithm but have trained models trained using different training data, and may include generation AIs that have trained models trained using different training data for different algorithms or the same training data.
[0013] Furthermore, wearable devices such as smart glasses 21 and a smart watch 30 may be communicatively connected to the information terminal 10, and user assistance information may be output to these wearable devices. The wearable device and the information terminal 10 may be connected wirelessly via a wireless router 4, or may be connected via short-range communication such as Bluetooth (registered trademark). While a smartphone is illustrated as an example of the information terminal 10 in FIG. 1, a personal computer or a tablet may also be used. The smart glasses 21 are an example of a head-mounted display type information terminal.
[0014] The generative AI 3 may include, for example, large language models (LLMs). Alternatively, the generative AI 3 may be a combination of natural language processing (NLP; Neuro Linguistic Programming) and machine learning. These NLP processes can process, understand, and generate responses in a natural way, with machine learning processes continuously improving the AI algorithms. The generative AI 3 uses natural language understanding (NLU) to decipher meaning and understand intent from input, either speech or text, and natural language generation (NLG) and NLP components to formulate responses. Finally, machine learning algorithms can be used to refine and improve the accuracy of responses over time.
[0015] 1 shows a configuration example in which the generation AI 3 is provided on a server 1 different from the information terminal 10, but as another example, the generation AI 3 may be installed on the information terminal 10. This makes it possible to provide user assistance information using the information terminal 10 even in an offline environment. This other example is common to all embodiments.
[0016] FIG. 2 is a diagram showing the hardware configuration of the information terminal 10 (smartphone).
[0017] The smartphone serving as the information terminal 10 includes an outer camera 111, an inner camera 112, a ranging sensor 113, a real-time clock (RTC) 114, an acceleration sensor 115, a gyro sensor 116, a geomagnetic sensor 117, a GPS receiver 118, a display 119, a communication interface (I / F) 120, a microphone 121, a speaker 122, a processor 125, a memory 128, and a telephone network communication device 131, all of which are connected to one another via a bus 140 that connects the various components. The communication interface (I / F) 120 is connected to an antenna 123 that transmits and receives network communication signals. The display 119 is an example of an output device capable of visually displaying and outputting user assistance information, and corresponds to a non-transparent display. As will be described later, the output device for user assistance information can also be configured as a transparent display 219 mounted on the smart glasses 21. Various sensors such as the microphone 121 , the outer camera 111 , the inner camera 112 , and the distance measurement sensor 113 are input devices that acquire external information about the information terminal 10 .
[0018] The communication interface (I / F) 120 is a communication interface that performs wireless communication between at least the information terminal 10 and the server 1 by short-range wireless communication, wireless LAN, or base station communication, and includes a communication processing circuit corresponding to various predetermined communication interfaces, and is connected to an antenna 123. The short-range wireless communication is performed using a wireless LAN such as Bluetooth (registered trademark), IrDA (Infrared Data Association, registered trademark), Zigbee (registered trademark), HomeRF (Home Radio Frequency, registered trademark), or Wi-Fi (registered trademark). Furthermore, the base station communication may be performed using long-range wireless communication such as 4G, 5G, or LTE.
[0019] A touch panel 130 is stacked on the display 119. The touch panel 130 functions as an input device that accepts operations from the user.
[0020] The processor 125 is configured by, for example, a CPU (Central Processing Unit). Alternatively, the processor 125 may be equipped with a GPU (Graphics Processing Unit) suitable for AI processing.
[0021] The memory 128 is configured by a flash memory and a nonvolatile memory, and stores programs 126 such as an OS (Operating System) and user support applications, and data 127 used by the processor 125.
[0022] The processor 125 loads a program 126 into the memory 128, executes the program, and reads data 127 as needed.
[0023] FIG. 3 is a hardware configuration diagram of the information terminal (smart glasses 21).
[0024] As an example of a head-mounted information terminal 20, the smart glasses 21 include an RTC 214, an acceleration sensor 215, a gyro sensor 216, a geomagnetic sensor 217, a GPS receiver 218, a transparent display 219, a communication interface (I / F) 220 and an antenna 223, a microphone 221, a speaker 222, a processor 225, a memory 228, and a camera 235, which are connected to each other via a bus 240.
[0025] FIG. 4 is a block diagram showing the functional configuration of the information terminal 10 according to the first embodiment.
[0026] The information terminal 10 includes, as components contributing to the user assistance information providing function, an input detection unit 301, a voice information processing unit 322, a question generation unit 302, a user assistance condition setting unit 303, a question transmission control unit 304, a summarization unit 305, a user assistance information presentation control unit 306, a voice utterance unit 307, and a text display unit 308. The information terminal 10 may further include a self-learning unit 350 that trains the LLMs of multiple generation AIs (referred to as "generation AI groups") used in the user assistance information presentation system 100. The self-learning unit 350 includes a machine learning data collection unit 351 and a machine learning execution unit 352. The functions of each unit will be described later with reference to flowcharts.
[0027] The RAW data storage unit 390 associates and stores as machine learning data voice data indicating the content of telephone conversations between customers and counselors, or voice-converted text data obtained by converting voice data into text, and customer satisfaction survey result data to which customers respond to the telephone conversation. The self-learning unit 350 stores the collected voice-converted text data and survey result data in the RAW data storage unit 390 as machine learning data, reads this data, and uses it for machine learning of the generated AI group. Details will be described later. The self-learning unit 350 is not essential to the information terminal 10; the functions of the self-learning unit 350 may be implemented in another information processing device, and the generated AI group may be accessed from that information processing device.
[0028] In the first embodiment, an example will be described in which a bank employee (counselor) uses the information terminal 10 for "telephone consultation on deposits, financial products, etc." as a use case of the information terminal 10.
[0029] Bank counselors who handle "telephone consultations on deposits, financial products, etc." from customers are required to have the ability to understand the content of the customer's consultation, the ability to infer the customer's financial literacy from the way the customer speaks, and the ability to explain financial products in accordance with the inferred customer's level of understanding. However, because junior counselors have less experience than veteran counselors, they may not be able to provide explanations and answers that satisfy the customer, which can lead to a decrease in customer satisfaction. Therefore, in the first embodiment, the know-how of veteran counselors is fed back to junior counselors in real time using an information terminal 10 while they are handling a consultation. Specifically, a summary of the veteran counselor's model answer is presented to the junior counselor in real time using audio, text, or both.
[0030] Fig. 5 is a flowchart showing the processing flow of the user assistance information presentation system of the first embodiment. Fig. 6 is a diagram illustrating the operation of the information terminal 10 of the first embodiment. Fig. 7 is a diagram showing an example of a setting screen for a user assistance information provision app displayed on the information terminal 10.
[0031] First, the user sets the user assistance information presentation conditions and the viewpoint for which assistance is required (hereinafter referred to as "advice viewpoint") on the setting screen 400 of the user assistance information providing application shown in FIG. 7 (FIG. 5: S101).
[0032] The user displays the setting screen 400 of the user assistance information providing app shown in Figure 7 and sets the user assistance information presentation method 410 and the advice perspective 420, which specifies the perspective from which advice (corresponding to user assistance information) the user wants to receive.
[0033] The setting screen 400 of the user assistance information providing app is configured to include a "screen" display ON / OFF switch 411 and an "audio (wireless earphone)" ON / OFF switch 412 as a user assistance information presentation method 410. A "product benefit explanation" selection button 421, a "risk explanation" selection button 422, a "wide range of options presented" selection button 423, and a "polite service" selection button 424 are displayed as advice viewpoints 420. The "product benefit explanation" selection button 421, the "risk explanation" selection button 422, the "wide range of options presented" selection button 423, and the "polite service" selection button 424 are configured as radio buttons. The items configured on the setting screen 400 of the user assistance information providing app become one of the presentation conditions for user assistance information.
[0034] The user assistance condition setting unit 303 sets server selection information 425 (details will be described later) and user assistance information presentation conditions based on the information entered on the setting screen 400. The counselor displays the setting screen 400 on the information terminal 10, performs initial settings, and waits for a telephone consultation.
[0035] The counselor waits for an incoming call from the customer (S102: No). When the telephone consultation starts (S102: Yes), the microphone 121 of the information terminal 10 picks up the voice of the telephone consultation conversation, or the telephone consultation is conducted using the information terminal 10 and the voice data of the conversation is recorded.
[0036] The input detection unit 301 detects conversational voice data picked up by the microphone 121 as external world information, and the voice information processing unit 322 performs voice recognition on the voice Au1 spoken by the customer and outputs voice-converted text data Tx1 (S103). In the example of Fig. 6, when the voice Au1 is converted into voice-converted text data Tx1, the text data generated is "Thanks to you, I was able to pay off my mortgage this time. However, I have a little bit of money left over that I saved for a house, and I wanted to hear more about the new investment trust product X that was mentioned the other day, so I called you." In addition to conversational voice, the external world information detected by the input detection unit 301 also includes video captured by the camera of the information terminal 10 and sensor information from various sensors.
[0037] The question generation unit 302 summarizes the speech-converted text data Tx1 and creates a question to be input to the generation AI 3 to obtain advice (equivalent to user support information) from the advice perspective set by the counselor (S104).
[0038] The question generation unit 302 may be configured by installing a local machine learning model (hereinafter referred to as "local AI") in the information terminal 10 and inputting the speech-converted text data Tx1 into the local AI to generate a question. For this purpose, an LLM may be used as the learning model for the local AI, and machine learning (training) may be performed by inputting a machine learning dataset into the LLM, in which a long sentence is input as input data and a question containing a summary of the long sentence is output data. The question generation unit 302 may also be configured using a text creation algorithm that does not use an AI engine, for example, a natural language processing program using a Markov chain.
[0039] In the example of FIG. 6 , the question generation unit 302 extracts "excess funds," "management," and "new investment trust product X" as keywords from the speech-converted text data Tx1, and generates question text data Tx2 as "Please explain investment trust X because it is a surplus fund management system." When local AI is used in the question generation unit 302, keyword extraction is not essential, and the speech-converted text data Tx1 may be input to the local AI to output question text data Tx2. When local AI is not used, for example, keywords may be registered in advance, extracted from the speech-converted text data Tx1, and then a natural language processing program using a Markov chain may be used to predict the next word based on the appearance pattern of the keywords and create a question.
[0040] The question transmission control unit 304 refers to the server selection information 425 and selects the advice perspective previously set by the junior counselor and a generation AI3 suitable for generating an answer in accordance with that advice perspective (S105).
[0041] As shown in FIG. 6 , the question transmission control unit 304 has a function of transmitting questions to each server A, B, C, and D included in the generation AI (LLM) group 1a on the network 2. Server A is a server including a generation AI 3 equipped with a trained model (e.g., LLM) suitable for inquiries about "product benefits." Server B is a server including a generation AI 3 equipped with a trained model (e.g., LLM) suitable for inquiries about "risk explanation." Server C is a server including a generation AI 3 equipped with a trained model (e.g., LLM) suitable for inquiries about "multiple choice." Server D is a server including a generation AI 3 equipped with a trained model (e.g., LLM) suitable for inquiries about "careful response." Examples of training data will be described later with reference to FIG. 7 . As described above, each of servers A, B, C, and D is equipped with one generation AI, and therefore, in this example, selecting a generation AI and selecting servers A, B, C, and D are described as synonymous.
[0042] In the server selection information 425 in Fig. 6, "merit of the product," "explanation of the risks," "wide range of options," "careful response," etc. are set as "advice viewpoints" 426 desired by the user, and "destination servers" 427 of the question text are associated with "(server) A," "(server) B," "(server) C," and "(server) D" corresponding to each advice viewpoint. When an advice viewpoint is set on the setting screen 400 of the user assistance information providing app in Fig. 7, a check box 428 in the server selection information 425 is checked (a flag is set) to indicate that the advice viewpoint has been selected.
[0043] Because the counselor has previously selected "risk explanation" as the advice viewpoint in the server selection information 425, the user assistance condition setting unit 303 refers to the server selection information 425, selects server B as the destination server for the question, and notifies the question transmission control unit 304. The advice viewpoint has already been set on the setting screen 400.
[0044] This time, the question transmission control unit 304 transmits the question text data Tx2 to Server B in accordance with the server selection information 425 (S106).
[0045] When the question text data Tx2 is input, the server B generates answer text data Tx3 in response to the question text data Tx2 and returns it to the information terminal 10, and the summarizing unit 305 acquires the answer text (S107).
[0046] An example of response text data Tx3 is output as follows: "Investment Trust X is a system that can handle all your asset management needs, and it will even research and trade the most suitable financial products on your behalf, so it is good for busy people or those with little knowledge. However, first of all, you should know that Investment Trust X is an investment product. Unlike savings and deposits, you must be aware that there is a risk of losing your principal due to fluctuations in stock prices and exchange rates at the time of purchase. On the other hand, X is a savings investment that can take advantage of a loss of principal. When purchasing the same amount each time, a lower unit price means you can purchase more shares, so a fall in the unit price can actually be a good thing. Conversely, if the unit price increases, the number of shares you can purchase for the same savings amount will decrease. This not only prevents you from "buying at too high a price," but is also expected to result in greater profits when the unit price rises in the future."
[0047] The response from Server B is sufficient as a risk explanation, but it is relatively long. Therefore, even if it is provided to a junior counselor as is, if the counselor tries to read the entire response from Server B while dealing with a customer, the customer consultation will temporarily stop while reading it, or the counselor will not be able to listen to what the customer is saying, so it is insufficient as real-time user support.
[0048] Therefore, the summarizing unit 305 receives notification from the user assistance condition setting unit 303 that "risk explanation" has been selected as the advice viewpoint, and extracts the portion relating to the risk explanation from the response text data Tx3, for example, "But first, we want you to know that investment trust X is a financial product with investment potential. Unlike savings and deposits, you must be aware that there is a risk of "loss of principal" due to fluctuations in stock prices and exchange rates at the time of purchase.", and generates summary text data Tx4 (corresponding to user assistance information) that matches the advice viewpoint desired by the counselor (S108).
[0049] For example, summary text data Tx4 could be, "X is a lump-sum asset management agency product, but please be aware that there is a risk of 'loss of principal'. Furthermore, the loss of principal will be used to prevent 'buying at an excessive price' in the long term."
[0050] If a local AI is used in the summarization unit 305, the answer sentence text data Tx3 may be input to the local AI (e.g., equipped with an LLM) to output summary text data Tx4. If a local AI is not used, a summary generation program using TF-IDF (Term Frequency-Inverse Document Frequency) may be used. TF-IDF calculates the importance of words in a document and summarizes the document based on that importance. TF (Term Frequency) represents the frequency of occurrence of each word in a document, so frequently occurring words in a document are considered to be highly important. On the other hand, IDF (Inverse Document Frequency) is an index indicating the rarity of a word, and therefore takes the inverse of the frequency of occurrence of a word in the entire document collection. Therefore, rare words are considered to be more important than frequently occurring words. By calculating the product of TF and IDF, the importance of each word is obtained, sentences containing highly important words are selected, and these are combined to create a summary.Instead of using TF-IDF to determine the importance of words, keywords that are frequently output in inquiries about financial products or that have high importance may be registered in advance, and a summary containing those words may be generated in the answer sentence text data Tx3.
[0051] The summarizing unit 305 outputs the summary text data Tx4 to the user assistance information presentation control unit 306. The user assistance information presentation control unit 306 refers to the user assistance information presentation conditions set in advance by the junior counselor, and presents the summary text data Tx4 in a manner that matches the presentation conditions (S109).
[0052] If text display is specified as the user assistance information presentation condition, the user assistance information presentation control unit 306 outputs the summary text data Tx4 to the text display unit 308, and the text display unit 308 displays the summary text data Tx4 in accordance with the presentation conditions on the display 119 of the information terminal 10, the transmissive display 219 of the smart glasses 21, the display of the smart watch 30, or the like. Alternatively, if voice notification is specified as the user assistance information presentation condition, the user assistance information presentation control unit 306 outputs the summary text data Tx4 to the voice utterance unit 307, and outputs voice data Au2 in which the voice utterance unit 307 reads out the summary text data Tx4 to the speaker 122 or the speaker 222 of the smart glasses 21.
[0053] 8 is a diagram showing an example of training data for a trained model used in the user assistance information presentation system. In FIG. 8, server B, which is suitable for "risk explanation," will be used as an example for explanation.
[0054] In a use case of the user assistance information presentation system of the first embodiment, advice is provided to a junior counselor as user assistance information from the perspective of what a veteran counselor would pay attention to when dealing with customers. Therefore, the generation AI 3 generates a machine learning dataset using the customer's utterances as input data and the client's utterances as output data, and uses this as training data. Recently, in order to improve the quality of telephone responses, the content of customer and client utterances during telephone responses is recorded, and this recorded voice data can be used to collect a machine learning dataset.
[0055] Therefore, the machine learning data collection unit 351 acquires recorded voice data of telephone conversations and classifies the recorded voice data from the start (incoming call) to the end of the conversation into customer speech and counselor speech. As an example of classification, the machine learning data collection unit 351 may perform a voice recognition process on the recorded voice data, recognize an utterance such as "This is XX from XX Bank" as the counselor's voice, and classify the recorded voice data into a voice that converses with the counselor as the customer. Furthermore, after customer interaction, a customer satisfaction survey may be conducted on the customer, and the survey results may be used as material for determining whether the recorded voice data can be used as machine learning data.
[0056] In Figure 8, a machine learning dataset D10 is created in which customer speech text data Tx11, which is obtained by recognizing and converting recorded voice data of a customer's speech into text, is used as input data, and counselor text data Tx12, which is obtained by recognizing and converting recorded voice data of a counselor into text, is used as output data.
[0057] An example of customer utterance text data Tx11 is, "Interest rates have been low recently and the yen has been weakening, so I'm thinking of investing in foreign currencies such as US dollars. Could you please tell me what methods are available?"; an example of counselor text data Tx12 is, "First of all, with foreign currency products, there is a risk that the principal will be lost due to the appreciation of the yen, in other words, if the exchange rate at the time of withdrawal is stronger than at the time of deposit, there is a risk that the amount will be less than the amount deposited."
[0058] Next, the machine learning data collection unit 351 refers to the results of a questionnaire survey of the client and classifies and determines which of the four advice perspectives the collected machine learning data set will be used as training data for the generation AI. As an example of a method for determining this, the unit may refer to the results of a questionnaire survey conducted by a counselor after a client has responded to the client to investigate the client's satisfaction with the telephone response. For example, in the questionnaire shown in FIG. 8 , the client is asked to rate each of the four advice perspectives—the appropriateness of the explanation of the "product benefits," the appropriateness of the "risk explanation," the appropriateness of the "multiple choice presentation" (corresponding to an evaluation of whether the advice perspective "wide range of choices" was provided), and the "polite response" (corresponding to an evaluation of whether the advice perspective "polite response" was provided)—on a three-point scale: "yes" (high rating), "average," and "no" (low rating). If any advice perspective is highly rated, the unit determines that the counselor's explanation was appropriate for that advice perspective, and selects a server equipped with a generation AI that generates an answer based on the highly rated advice perspective as the server for performing machine learning using the collected machine learning data.
[0059] The machine learning execution unit 352 inputs the machine learning dataset D10 collected by the machine learning data collection unit 351 into the generated AI 3 installed on the selected server (server B in the above example), and performs machine learning on the generated AI 3. As a result, the information terminal 10 collects machine learning data using recorded voice data from the customer's telephone consultation and the results of a questionnaire collected from the customer thereafter, and uses this to select a server to perform machine learning, allowing the machine learning on the generated AI 3 to proceed without human intervention.
[0060] According to this embodiment, the information terminal 10 extracts keywords from a real-time dialogue to generate a question, and then transmits the question to a generation AI that matches the advice viewpoint specified in advance by the user (counselor) to obtain an answer. Furthermore, the information terminal 10 summarizes the answer according to the advice viewpoint and then presents it to the user. Therefore, the user can ask a question to the generation AI during a real-time dialogue without having to create a question.
[0061] Furthermore, because the information terminal 10 presents user assistance information that summarizes the answer sentence, the user can quickly understand the user assistance information that matches the advice viewpoint that the user wants from the answer sentence. As a result, the user can obtain the user assistance information that the user wants during a real-time dialogue and use it in a way that suits the user's use case.
[0062] Second Embodiment The second embodiment is an embodiment in which user assistance information is provided using an assistant character. Fig. 9 is a functional block diagram of the information terminal 10 of the second embodiment. The difference from the first embodiment is that the user assistance condition setting unit 303 and the user assistance information presentation control unit 306 are configured as a character control unit 310 for use with an assistant character.
[0063] Character control unit 310 includes a character setting unit 311 that sets character attribute information including the character's personality, appearance, etc., and a character operation unit 312 that displays the character based on the set character attribute information and causes the character to present user assistance information through speech, etc. In another embodiment, when a conversation takes place between a character and another avatar in the virtual reality space, character operation unit 312 also controls the conversation with the other avatar via the character.
[0064] 10 is a diagram showing a first example of an assistant character setting screen.
[0065] The character setting screen 500a displayed on the display 119 of the information terminal 10 allows the user to select the assistant character's personality 501 (e.g., "shopping expert," "investment expert," "family member, etc."), assistant appearance 502, and voice quality 503 (e.g., "thick male voice," "thin female voice," "low male voice") for the assistant character. The selection operation may be performed using a touch panel 130 on the displayed display 119, or a non-contact input unit may be provided on the display 119 to accept contact or non-contact operations. While three options are shown here, more options may be displayed. Furthermore, by setting the personality 501 of a "family member, etc." as an example of an assistant, the user can, for example, hear user assistance information (comments) from an assistant with the "family member, etc." personality traits (which may include characteristics such as timing of greetings and speaking in a more intimate tone rather than a businesslike tone) even when shopping or investing alone.
[0066] The character setting screen 500a also displays an automatic button 504 and a setting button 505. Selecting the automatic button 504 with the finger 6 sets an automatic mode in which the situation in which the user is placed is determined and the most suitable character is automatically selected. Selecting the setting button 505 with the finger 6 sets a fixed mode in which the user is supported by a character using the personality 501, appearance 502, and voice quality 503 that have been set.
[0067] 11 is a diagram showing a second example of an assistant character setting screen 500b. In FIG. 11, the same processes or functions as those in FIG. 10 are denoted by the same reference numerals, and the description thereof will be omitted.
[0068] The personality of the assistant displayed on the display 119 of the information terminal 10 can be set to a preferred personality 512 (such as "safety," "economy," "universality," "necessity," or "convenience") according to each usage scene 511 (such as "shopping," "leisure," or "negotiation"). (Although not shown, appearance 502 and voice quality 503 in the first example may also be set.) These personality parameters may be set based on general personality traits such as extroversion or introversion, sensation or intuition, logic or emotion, and judgment or perception as classified by the MBTI (Myers-Briggs Type Indicator), to support the user's tendency to fall into shortcomings in the way they think about things and make decisions.
[0069] Fig. 12 is a main flowchart of a first example of the information terminal 10 according to the second embodiment. Fig. 13 is a main flowchart of a second example of providing user assistance information for the assistant function of the information terminal 10 according to the second embodiment.
[0070] When the information terminal 10 starts (initiates) the processing flow of FIG. 12, the character setting unit 311 selects an assistant character and determines whether the setting button 505 or the automatic button 504 has been pressed (S201).
[0071] If the character setting unit 311 determines that no character has been set (S201: No), the process ends.
[0072] On the other hand, if the character setting unit 311 determines that the character attribute information has already been set (S201: Yes), the character operation unit 312 sets an assistant character based on the character attribute information (S202), and the process ends. Note that the process flow in Fig. 12 is repeated in a cycle in which the approach of the user's finger 6 to the non-contact input unit configured on the surface of the display 119 is detected, or in a cycle in which the display 119 displays (for example, 60, 120, or 240 frames per second).
[0073] FIG. 13 is a main flowchart illustrating an example of an operation of the assistant function of the information terminal 10 according to the second embodiment to provide user assistance information.
[0074] 13, the input detection unit 301 detects whether the voice of the user and the person who is having a conversation with the user is input (S211). If the input detection unit 301 does not detect the voice of the user and the person who is having a conversation with the user (S211: No), the process ends.
[0075] On the other hand, if the input detection unit 301 detects a series of conversational voices between the user and the user's interlocutor (S211: Yes), the voice processing unit 183 analyzes the conversational information according to the voices detected by the input detection unit 301 and the character attribute information recognized by the character operation unit 312. Then, in order to obtain information on which the user is weak or information to provide support to complement thinking and decision-making, the question generation unit 302 determines the content of the question and generates a question. The question transmission control unit 304 selects a suitable generation AI from the group of generation AIs and transmits the question (S212).
[0076] The summarizing unit 305 receives the reply from the generation AI 3 and generates user assistance information based on this reply (S213).
[0077] The character operation unit 312 generates an assistant character based on the character attribute information and displays it on the display 119. Then, the assistant character is made to speak user assistance information (S214). Here, by changing the settings of the character setting unit 311, the information, thoughts, and decisions that the user is not good at are interpreted as different, and the content of the complementary user assistance information also changes.
[0078] In order to reduce the processing load on the information terminal 10, the character operation unit 312 does not display the assistant character, but outputs audio information from the speaker 122 in the voice of the personality 501 and voice quality 503 set in the character attribute information. If necessary, display output or audio output may be performed so that the information is conveyed not only to the user but also to the interlocutor. In this case, the messages for the user and the messages for the interlocutor may be different. Furthermore, the character operation unit 312 may present user assistance information on the display 119 in text format, without displaying the assistant character.
[0079] According to this embodiment, the user can set the character attributes, such as the personality and voice quality, of the character that presents the user assistance information. Therefore, the user can hear the user assistance information from a character that the user finds friendly, and therefore, even if comments that may seem unpleasant to the user at first glance, such as warnings or urging the user to reconsider a purchase, are presented, the user can feel less psychological resistance when listening to these comments.
[0080] Third Embodiment FIG. 14 is a hardware configuration diagram of a head-mounted display information terminal 60 (HMD) according to a third embodiment.
[0081] The third embodiment is an embodiment in which the display 119 of the first embodiment is configured as a head-mounted display and an assistant function in a virtual world is realized. The third embodiment differs from the information terminal 10 of the first embodiment in that it includes a virtual reality processing unit 320 that processes audio information and video information from a virtual reality providing server 61 to realize virtual reality, an avatar setting unit 331 that creates an avatar that is an alter ego of the user 5 in the virtual world, and BT (Bluetooth) controllers 694 a and 694 b that have a short-range wireless communication interface.
[0082] 14 includes an outer camera 611, an inner camera 612, a distance measurement sensor 613, a real-time clock (RTC) 614, an acceleration sensor 615, a gyro sensor 616, a geomagnetic sensor 617, a GPS receiver 618, a display 619, a communication interface (I / F) 620, a microphone 621, a speaker 622, a processor 625, a memory 628, and a battery 690, which are connected to one another via a bus 640 that connects the various components. The communication interface (I / F) 620 is connected to an antenna 623 that transmits and receives network communication signals.
[0083] Furthermore, the communication interface (I / F) 620 can be connected to the generation AI 3, other information terminals 65, and the virtual reality providing server 61 via the network 2.
[0084] FIG. 15 is a functional block diagram of an information terminal 60 (HMD) according to the third embodiment.
[0085] The functional blocks of the information terminal 10 of the second embodiment differ from those of the information terminal 10 of the second embodiment in that it includes a virtual reality processing unit 320, a video information processing unit 321, an audio information processing unit 322, a character control unit 310, and an avatar control unit 330. The avatar control unit 330 includes an avatar setting unit 331 that sets the appearance, gender, etc. of the user's alter ego (avatar) that the user places in the virtual reality space, and an avatar operation unit 332 that moves the avatar in accordance with the user's movements.
[0086] The character control unit 310 places an assistant character, which is separate from the avatar, in the virtual reality space.
[0087] The video information processing unit 321 and the audio information processing unit 322 handle as input information not only the video signal captured by the outer camera 611 and the audio information collected by the microphone 621, but also the video information of the virtual world obtained by the virtual reality processing unit 320 from the virtual reality providing server 61 and the audio information of the conversation between the avatar (and assistant) and the virtual store clerk in the virtual world.
[0088] FIG. 16 is an external view of an information terminal 60 (HMD) according to the third embodiment.
[0089] As shown in Figure 16, an information terminal 60 is worn on the head of a user 5, and content from a virtual reality providing server 61 is provided to the user's visual field, immersing the user 5 in a virtual world. When the information terminal 60 is worn on the user's head, a display 619 is positioned in the user's line of sight, and a left speaker 662L (right speaker not shown) is positioned at each ear. The user 5 also uses BT controllers 694a and 694b, which are held in each hand and have a short-range wireless communication interface that controls the motion tracking of the avatar's hands in the virtual reality space and the operation of various buttons and trackpads.
[0090] FIG. 17 is a diagram showing a display example of the third embodiment.
[0091] Figure 17 shows a virtual reality space of an apparel (clothing) store as an example of content of the virtual reality providing server 61. Figure 17 illustrates a state in which a virtual reality space 62 of an apparel (clothing) store is displayed on the display 619 as an example of a virtual store. In the virtual store, new products from various brands can be unveiled to people all over the world, and an avatar 63, which is the alter ego of the user 5 in the virtual reality space 62, and an assistant character (hereinafter abbreviated as "assistant") 64, which is set by the character control unit 310 to support the avatar 63, are displayed participating in the apparel (clothing) store.
[0092] Here, the assistant 64 can be set as an accessory to the avatar 63, and can move in accordance with the movements of the avatar 63. Thus, the user 5 can converse with a virtual salesperson (interlocutor) 65 at a virtual apparel (clothing) store and have the avatar 63 try on clothes of their choice. If they like the clothes, they can purchase them at a physical store or on an e-commerce site. Here, the assistant 64 performs an assistant function of providing information to support the conversation between the user 5 and the virtual salesperson (interlocutor) 65.
[0093] FIG. 18 is a flowchart showing the flow of processing by the avatar setting unit in the information terminal 60 according to the third embodiment.
[0094] The avatar setting unit 331 checks whether or not there is avatar attribute information (S301). If there is no avatar attribute information (S301: No), the process ends without setting an avatar.
[0095] On the other hand, if the avatar setting unit 331 can confirm the avatar attribute information (S301: Yes), it creates an avatar 63 according to the avatar attribute information (S302).
[0096] The avatar setting unit 331 places the avatar 63 in the virtual reality space 62, and the avatar operation unit 332 causes the avatar 63 to move in accordance with the instructions and movements of the user 5 (S303). The process returns to S301 and repeats until an instruction to end the placement of the avatar 63 in the virtual reality space is received (S304: No). If an instruction to end the placement of the avatar 63 in the virtual reality space is received (S304: Yes), the processes of the avatar setting unit 331 and the avatar operation unit 332 are terminated.
[0097] The avatar setting unit 331 displays a setting means similar to that of the assistant character described in the second embodiment on the display 619, and can set the avatar using the BT controllers 694a, 694b, etc. The processing flow in Fig. 19 is repeated at a cycle of detecting the selection of various buttons displayed on the display 619 or at a cycle of display on the display 619 (for example, 60, 120, 240 frames per second, etc.).
[0098] FIG. 19 is a flowchart showing the flow of processing of the assistant function in a virtual reality space using the information terminal 60 of the third embodiment.
[0099] It is detected whether the head-mounted information terminal 60 is connected to the virtual reality providing server 61 via the network 2 from the communication interface (I / F) 220 (S311). If the head-mounted information terminal 60 is not connected to the virtual reality providing server 61 (S311: No), the process ends.
[0100] On the other hand, if the head-mounted information terminal 60 is connected to the virtual reality providing server 61 (S311: Yes), it is detected whether or not the voice of the conversation between the avatar 63 of the user 5 and the virtual store clerk (interlocutor) 65 is input within the virtual reality providing server 61 (S312). If the voice information processing unit 322 does not detect the input of the voice of the conversation between the avatar 63 (and assistant 64) and the virtual store clerk 66 (S312: No), the processing ends.
[0101] On the other hand, if the voice information processing unit 322 detects input of conversational voice (S312: Yes), the question generation unit 302 analyzes the conversational information in the voice information processing unit 322 according to the character's personality based on the conversational voice and character attribute information, and generates a question to obtain information that the user 5 is not good at or information to provide support to complement their thinking and decision-making, and the question transmission control unit 304 sends the question to the generation AI 3 via the communication interface (I / F) 220 and network 2 (S313).
[0102] The summarizing unit 305 receives the response content from the generation AI 3 via the network 2 and the communication interface (I / F) 220, and generates user assistance information required for the user 5 (S314). Then, the assistant 64 in the virtual reality service presents the user assistance information by voice to the avatar 63 and the virtual store clerk (interlocutor) 65 (if necessary, the user assistance information may be displayed as text information on the display 619 or output as audio information from the speaker 622) (S315).
[0103] By changing the settings of the character setting unit 311, the information, thoughts, and decisions that the user 5 is weak at are interpreted as different, and the content of the complementary user support information also changes. The processing flow in Figure 19 is repeated in cycles for detecting the selection of various buttons displayed on the display 619 or in cycles for displaying the display 619 (for example, 60, 120, 240 frames per second, etc.).
[0104] FIG. 20 is a second system outline diagram of an application example of the head-mounted information terminal 60 of the third embodiment.
[0105] As a provision condition of the virtual reality providing server 61, there may be a case where two types of connection input, such as avatar 63 and assistant 64, from the same device (head-mounted information terminal 60) are rejected. To deal with this, a method will be described in which user 5 connects and inputs avatar 63 on the head-mounted information terminal 60 and connects and inputs assistant 64 on the smartphone-type information terminal 10 owned by user 5.
[0106] 20, as an initial setting, the avatar 63 is initially set using the avatar setting unit 331 of the information terminal 60, and the assistant 64 is initially set using the character setting unit 311 of the information terminal 10 (S331). For example, the setting is performed using the setting method shown in FIG. 10 or FIG. 11.
[0107] Thereafter, the avatar 63 and the assistant 64 are logged in using their respective IDs (identifications) (S332). Specifically, the avatar 63 of the user 5 is logged in to the virtual reality providing server 61 from the head-mounted information terminal 60 using an ID for identifying the avatar 63. Then, the assistant 64 is logged in to the virtual reality providing server 61 from the information terminal 10 using an ID for the assistant that is different from the ID for identifying the avatar 63.
[0108] The avatar operation unit 332 places the avatar 63 in the virtual reality space provided by the virtual reality providing server 61 and causes the avatar 63 to move in accordance with the movements of the user 5 (S333).
[0109] The character operation unit 312 places the assistant 64 in the virtual reality space provided by the virtual reality providing server 61, listens to the conversation between the avatar 63 and the virtual store clerk (interlocutor) 65 taking place in the virtual reality space, outputs the information to the audio information processing unit 322, and provides the avatar 63 with utterances based on user assistance information using a set character generated by the same processing as in the first example of the third embodiment (S334). For example, the assistant 64 listens to the conversation between the avatar 63, which is the alter ego of the user 5, and the virtual store clerk (interlocutor) 65 ((1) in the figure), asks a question to the generation AI 3 according to the attributes of the assistant 64 set by the character setting unit 311 ((2) in the figure), receives a response from the generation AI 3 ((3) in the figure), and utters user assistance information ((4) in the figure). In this way, the content of the assistant 64's utterances is linked to the generation AI 3.
[0110] In this embodiment, when the avatar 63 is moved using the head-mounted information terminal 60, the avatar operation unit 332 transmits the movement information (or the location information within the virtual reality providing server 61) to the information terminal 10 to share the information, thereby performing movement operation linked communication that synchronizes the movements of the avatar 63 and the assistant 64 within the virtual reality space.
[0111] In the third embodiment, in addition to the second embodiment, the user 5 can obtain user support information that complements information, thoughts, and decisions that the user 5 is not good at, through an avatar 63 that is an alter ego of the user 5, even within the content of the virtual world. Also, an assistant 64 can be presented to the virtual store clerk (interlocutor) 65, and by making the virtual store clerk (interlocutor) 65 aware that the avatar 63 has a formidable assistant 64, it is possible to prevent damage such as malicious misleading and fraud.
[0112] <Fourth embodiment> Fig. 21 is a hardware configuration diagram of an information terminal 20 according to a fourth embodiment. Fig. 22 is a functional block diagram of the information terminal 20 according to the fourth embodiment.
[0113] The fourth embodiment relates to a technology for applying an assistant function in augmented reality, in which the display 119 of the first embodiment is configured as a transmissive display 219. The head-mounted transmissive display 219 for realizing augmented reality has a function for displaying augmented information and simultaneously allowing the real world to be seen through. The transmissive display 219 displays augmented video information using a transmissive organic EL display, a transmissive inorganic EL display, a transmissive LCD display, or a projector type display.
[0114] The information terminal 20 is also connected to Bluetooth (Bluetooth) earphones 495a and 495b, which have a short-range wireless communication interface. User assistance information is output from a speaker 222 integrated with the information terminal 20 and the Bluetooth earphones 495a and 495b. The Bluetooth earphones 495a and 495b are an example of an output device for user assistance information; other examples of output devices include Bluetooth wireless headphones, Bluetooth bone conduction earphones, and Bluetooth bone conduction headphones. The communication connection with the information terminal 20 is also not limited to Bluetooth; it may be a Wi-Fi (registered trademark) communication connection or a wired connection.
[0115] Furthermore, as shown in FIG. 22, the information terminal 20 includes an augmented reality processing unit 340 that processes augmented information.
[0116] FIG. 23 is a diagram showing an example of video information that can be viewed from the information terminal 20 of the fourth embodiment.
[0117] Through the transmissive display 219, the user 5 sees the real world, including a physical apparel (clothing) store 71 existing in the real world and a physical store clerk (interlocutor) 74. User assistance information 75 from the virtual world is displayed as augmented information on the video information of the real world (an assistant 73 may also be displayed if necessary).
[0118] FIG. 24 is a flowchart showing the flow of processing video information of the assistant function in augmented reality using the information terminal 20 of the fourth embodiment.
[0119] The input detection unit 301 detects whether or not a conversational voice between the user 5 and the actual store clerk (interlocutor) 74 is input (S341). If the input detection unit 301 does not detect a conversational voice between the user 5 and the actual store clerk (interlocutor) 74 (S341: No), the process ends.
[0120] On the other hand, when the input detection unit 301 detects the conversational voice between the user 5 and the actual store clerk (interlocutor) 74 (S341: Yes), the voice information processing unit 322 analyzes the conversational voice detected by the input detection unit 301 and the character attribute information that defines the character's personality, etc., recognized by the character setting unit 311. Then, the question generation unit 302 generates a question to obtain information on which the user is weak or support information to complement thinking and decision-making, and the question transmission control unit 304 transmits the question to the generation AI 3 via the communication interface (I / F) 220 and the network 2 (S342).
[0121] The summarizing unit 305 receives the response content from the generation AI 3 via the network 2 and the communication interface (I / F) 220, and generates the necessary user assistance information (S343).
[0122] The character operation unit 312 displays the user assistance information 75 as text (character) information on the transmissive display 219 (S344). If the setting of the character setting unit 311 is changed, it is interpreted that the information, thoughts, and decisions that the user 5 is not good at are different, and the content of the complementary user assistance information 75 also changes. The processing flow of FIG. 24 is repeated in a cycle for detecting the setting of the character setting unit 311 or in a cycle for displaying the transmissive display 219 (for example, 60, 120, or 240 frames per second).
[0123] FIG. 25 is a flowchart showing the flow of processing voice information of the assistant function in augmented reality using the information terminal 20 of the fourth embodiment.
[0124] Steps S351 to S354 are the same as steps S341 to S344 in FIG. 24, and the user assistance information 75 is output from the head-mounted display information terminal 20.
[0125] In parallel with steps S351 to S354, it is detected whether or not the short-range wireless communication interface (IF) of the head-mounted display information terminal 20 and the short-range wireless communication of the BT earphones 495a and 495b are already connected (S355). If they are not connected (S355: No), the process proceeds to step S357, where the user assistance information 75 is output as audio from the speaker 222 that is integrally configured with the transmissive display 219 and the information terminal 20 (S357).
[0126] On the other hand, if connected (S355: Yes), the user assistance information 75 is output as audio from the BT earphones 495a and 495b separated from the information terminal 20, and is also output as video on the transmissive display 219 (S356).
[0127] By outputting the audio from the speaker 222 so that it can be heard by the actual store clerk (interlocutor) 74, it is possible to inform the actual store clerk (interlocutor) 74 that the user 5 has an excellent assistant 73. On the other hand, the audio output from the BT earphones 495a and 495b can be made inaudible to the actual store clerk (interlocutor) 74, so that the maliciousness or non-maliciousness of the information transmitted from the user assistance information 75 can be conveyed to the user 5 using more direct expressions. Here, by changing the setting of the character setting unit 311, the content of the user assistance information 75 that is interpreted as information that the user 5 is not good at, or thoughts and decisions that are different, also changes. Note that the processing flow of FIG. 25 is repeated at a cycle for detecting the setting of the character setting unit 311 or a cycle for displaying the transmissive display 219 (e.g., 60, 120, 240 frames per second, etc.).
[0128] In addition to the first embodiment, in the fourth embodiment, user support information 75 that complements information, thoughts, and decisions that the user 5 is not good at is provided as extended information in augmented reality, either only to the user 5, or deliberately so that the interlocutor 74 can also hear it, thereby attracting attention and preventing malicious solicitations, fraud, etc., and preventing unnecessary shopping, etc.
[0129] Fifth Embodiment FIG. 26 is a functional block diagram of a portable information terminal 10 according to a fifth embodiment.
[0130] The fifth embodiment differs from the information terminal 10 of the first embodiment in that it includes Bluetooth earphones 495 a and 495 b having a short-range wireless communication interface capable of communicating with the communication interface (I / F) 120 .
[0131] FIG. 27 is a schematic diagram of an operation using the information terminal 10 of the fifth embodiment.
[0132] The portable information terminal 10 listens to the voice of the conversation between the user 5, who is active in real space, and the interlocutor 80 using the input detection unit 301 ((1) in the figure), analyzes the content of the conversation using the voice information processing unit 322 based on the personality of the assistant 73, etc., which corresponds to the character attribute information set in the character setting unit 311, creates a question for the generated AI 3 using the question generation unit 302, and asks the question to the generated AI 3 via the communication interface (I / F) 120 and the network 2 ((2) in the figure).
[0133] Next, the answer from the generation AI 3 is received by the question generation unit 302 via the network 2 and the communication interface (I / F) 440 ((3) in the figure). Based on the received answer, the question generation unit 302 and the speech information processing unit 322 generate user assistance information in the summary unit 305, which is then spoken as speech information from the speaker 122 or the Bluetooth earphones 495a, 495b ((4) in the figure). The user assistance information 75 may also be displayed as text information on the display 119 (and the assistant 73 may be displayed if necessary).
[0134] Also, as shown in FIG. 27, the display 119 displays the user assistance information 75 generated by the summary unit 305 as text information together with the assistant 73 set in the character setting unit 311 (if necessary, the user 5 may also present the portable information terminal 10 to the interlocutor 80).
[0135] Furthermore, a touch panel 130 may be provided on the display 119 to display an automatic button 76 and an advice button 77 as an input interface (I / F). When the automatic button 76 is turned on, the information terminal 10 can be operated in an automatic mode in which the situation of the user 5 is automatically determined and text information of the user assistance information 75 (and the assistant 73, if necessary) is output as video to the display 119 only when necessary.
[0136] On the other hand, if the automatic button 76 is turned OFF, when the user 5 determines that user assistance information is necessary, the user can manually touch the advice button 77 to operate the information terminal 10 in a manual mode in which the text information of the user assistance information 75 and the assistant 73 are output as images on the display 119.
[0137] The user 5 can hold the information terminal 10 in his / her hand and present it to the interlocutor 80, so that the user assistance information 75 of the assistant 73 can also be presented to the interlocutor 80. Alternatively, the user assistance information 75 can be provided only to the user 5 from BT earphones 495a, 495b worn on the ears of the user 5, without being displayed on the display 119. Of course, the user assistance information 75 will not be output unless the automatic mode is turned off and the advice button 77 is touched, and the advice function can also be turned off.
[0138] Furthermore, the assistant function may be used for a conversation during a call on the portable information terminal 10, to provide user assistance information 75 for a call on a mobile phone. In this case, listening to the conversation ((1) in the figure) can be realized by inputting not only the microphone 121 but also the caller's voice from the communication interface (I / F) 120 to the voice information processing unit 322.
[0139] The processing flow for presenting user assistance information 75 for the conversation content between a user 5 and an interlocutor 80 using a portable information terminal 10 is such that the processing of the processor 225 of the smart glasses 21 in the fourth embodiment is executed by the processor 125 of the information terminal 10, and therefore a redundant explanation will be omitted.
[0140] In addition to the first embodiment, in the fifth embodiment, it is possible to provide user support information 75 to only user 5 in the real world, which complements information, thoughts, and decisions that user 5 is not good at, thereby drawing attention to the user, and it is also possible to warn the interlocutor 80 by presenting the portable information terminal 10, which has the effect of preventing fraud and other damage, wasteful shopping, etc.
[0141] <Sixth embodiment> The sixth embodiment is an embodiment in which an information terminal is equipped with a local AI 364, which cooperates with a generation AI 3 connected via a network 2 and supplements the answers of the generation AI 3. A first example of the sixth embodiment is an embodiment in which a supplemental information setting unit 363 is used to determine the situation in which a user 5 is placed and present more appropriate user assistance information 75 according to the scene. Figure 28 is a functional block diagram of an information terminal 10 of the sixth embodiment.
[0142] The database 362 is provided with a conversation keyword table (TBL) 3621. The conversation keyword table (TBL) 3621 is a table in which keywords obtained from the conversation are registered. The voice information processing unit 322 recognizes the conversation voice and converts it into text, and the text analysis unit 361 analyzes the text of the conversation voice by referring to the conversation keyword table (TBL) 3621.
[0143] The conversation keyword table (TBL) 3621 assumes, for example, names of clothing such as "coat," "jacket," or "suit," as well as words of a store clerk (interlocutor) such as "it suits you," "the last one," or "a popular item," and compares these with the database of the conversation keyword table (TBL) 3621 to determine the environment in which the user 5 is about to shop at a clothing store. If the character setting unit 311 is operating in the automatic mode described above, it automatically selects "shopping expert" as the character's personality. Then, to prevent the assistant 73 of the "shopping expert" character from being misled by the store clerk's (interlocutor) 74's (see FIG. 23 ) common phrases such as "it suits you," "the last one," or "a popular item," the text analysis unit 361 analyzes important missing information, such as the average usage period and average price of the product.
[0144] In addition, the supplemental information setting unit 363 analyzes the user 5's usual living environment (family composition, lifestyle, etc.) from the video information and audio information obtained from the video information processing unit 321 and audio information processing unit 322, and supplements the analysis of missing information to obtain user support information 75 that is more suitable for the user 5.
[0145] Furthermore, the current position may be obtained from the GPS receiver 118 and used to supplement the analysis of missing information in the user assistance information 75 .
[0146] The local AI (LLM) 364 answers any missing information from accumulated information such as past interactions, and queries the generation AI 3 for any information that is insufficient for the local AI (LLM) 364. If the local AI (LLM) 364 cannot connect to the network 2 due to a poor communication environment, for example, the user assistance information 75 is provided by the local AI 364 rather than the generation AI 3.
[0147] The generation AI 3 also converses with other information terminals 65 via the network 2, and by utilizing a large amount of information from similar cases, can provide more sophisticated responses than the local AI (LLM) 364. For example, the conversations of malicious and fraudulent interlocutors 80 who contacted the user 5 are often manualized, so it is highly likely that owners of other information terminals 15 have already experienced similar opportunities. If this information is stored in the generation AI 3, it can be immediately reflected in the responses to the information terminal 10 of the user 5, preventing damage to the user 5.
[0148] The answer obtained from the generated AI3 in this manner is summarized by the summarizing unit 305. The summary is then input to the local AI (LLM) 364, which reinforces the summary and outputs it from the user assistance information presentation control unit 306. The output may be made by the character operating unit 312 speaking through a character, by displaying text from the text display unit 308, or by outputting audio data from the audio speaking unit 307. In this case, the audio information output to the BT earphones 495a and 495b is not necessarily the same audio information, and the essential meaning of the statement or a warning such as "There is a possibility of fraud" may be conveyed only to the user 5.
[0149] Next, FIG. 29 is a flowchart showing the flow of information between the generation AI 3, the user 5, and the interlocutor 80, focusing on the information terminal 10 of the sixth embodiment.
[0150] When a conversation between the user 5 and the interlocutor 80 begins (S401, S402), the input detection unit 301 of the information terminal 10 detects the conversation voice (S403). The input detection unit 301 recognizes the conversation voice and converts it into text (characters) (S404). The text analysis unit 361 analyzes the converted text of the conversation using keywords and the like in the conversation keyword table (TBL) 3621 (S405). Set character attribute information may also be used as needed.
[0151] The question generation unit 302 generates a question including the keyword based on the analysis result and the character attribute information (S406).
[0152] The question transmission control unit 304 transmits the question to the generation AI 3 (S407).
[0153] The generation AI3 receives the question, analyzes it, and uses a large-scale language model (LLM) to decipher the meaning, understand the intention, etc. (S408). Then, it creates a response in natural language based on the analysis results and returns the answer (S409).
[0154] When the summarizing unit 305 receives the answer sentence from the generation AI 3 (S410), it inputs it to the local AI 364. The local AI 364 creates user assistance information 75 related to the keyword based on the answer sentence and character attribute information (S411).
[0155] The user assistance information presentation control unit 306 presents the user assistance information 75 by outputting it via characters, displaying it as text, outputting audio data, etc. (S412). The information is output as audio and video from the output unit 560 (or, if necessary, via the Bluetooth earphones 495a and 495b via the communication interface (I / F) 540). This allows the user 5 to converse with the interlocutor 80 while taking the user assistance information 75 into consideration (S413).
[0156] In addition to the first and fourth embodiments, in the sixth embodiment, even when the network 2 cannot be connected, the local AI (LLM) 364 installed in the information terminal 10 can provide user support information 75 that complements the information, thoughts, and decisions that the user 5 is not good at, thereby drawing attention to the user, thereby preventing fraud and other damage, wasteful shopping, etc.
[0157] As an example use case of the sixth embodiment, an example will be described in which the supplemental information setting unit 363 functions as a scene determination unit that determines the scene the user is in. The supplemental information setting unit 363 determines the situation of the user 5 from the video signal captured by the outer camera 111 of the information terminal 10 and the audio signal collected by the microphone 121, sets an assistant character that is optimal for the conversation scene in the character setting unit 311, and provides optimal user assistance information 75.
[0158] FIG. 30 is a diagram showing an example of user assistance information presentation in the sixth embodiment.
[0159] The display 119 displays the image captured by the outer camera 111, and the supplemental information setting unit 363 analyzes the situation around the user 5. For example, in FIG. 30 , the conversation situation at an apparel (clothing) store can be acquired from the video signal from the outer camera 111 and the audio signal from the microphone 121. The supplemental information setting unit 363 determines the situation in which the user 5 is shopping from the video signal and audio signal. Furthermore, if necessary, questions can be asked to the generation AI 3 from this information to determine a conversation scene with higher accuracy, and the character setting unit 311 can be changed to optimal conditions based on the results, thereby improving the performance of the user assistance information 75.
[0160] FIG. 31 is a flowchart showing the flow of processing by the supplementary information setting unit of the information terminal in the sixth embodiment.
[0161] The input detection unit 301 detects the voice of the conversation between the user 5 and the interlocutor 80 (S421). If the input detection unit 301 does not detect the voice of the conversation between the user 5 and the interlocutor 80 (S421: No), the process ends.
[0162] On the other hand, if the input detection unit 301 detects the voice of a conversation between the user 5 and the interlocutor 80 (S421: Yes), the text analysis unit 361 analyzes the detected voice of the conversation, and the supplemental information setting unit 363 selects one of a plurality of conversation scenes preset (S422). Also, the video captured by the outer camera 111 or the inner camera 112 is analyzed, and the supplemental information setting unit 363 selects one of a plurality of conversation scenes preset (S423). The order of steps S422 and S423 may be reversed. Alternatively, the supplemental information setting unit 363 may comprehensively determine the results of the two steps and select an optimal assistant from a plurality of conversation scenes preset by the character setting unit 311.
[0163] The character setting unit 311 sets the most suitable character for the assistant in accordance with the conversation scene selected by the supplemental information setting unit 363 (S424), and the process ends.
[0164] In the sixth embodiment, the supplemental information setting unit 363 determines the scene based on the audio signal and the video signal, but the input detection unit 301 may determine the current location based on input from various sensors (for example, the GPS receiver 118 shown in FIG. 1 ) and use that location as information for determining whether the location is an apparel store, a securities company, etc. The processing flow in FIG. 30 is repeated at a cycle of detecting the selection of various buttons configured on the surface of the display 119, or at a display cycle of the display 119 (for example, 60, 120, 240 frames per second, etc.).
[0165] In the sixth embodiment, the current conversation scene of the user 5 can be automatically determined, which has the effect of providing user support information 75 that optimally complements the information, thoughts, and decisions that the user 5 is not good at for each conversation scene, thereby attracting attention.
[0166] As a second example of the sixth embodiment, the supplemental information setting unit 363 may be used to determine the user's personality and present more suitable user assistance information 75. The supplemental information setting unit 363 determines the personality of the user, etc. from the video signal and audio signal from the input detection unit 301, and automatically sets in the character setting unit 311 the personality of an assistant that complements the information, thoughts, and decisions that the user 5 is not good at, based on the personality of the user 5, for example, to generate more suitable user assistance information.
[0167] FIG. 32 is a flowchart showing another example (character determination of a user, etc.) of the information terminal according to the sixth embodiment.
[0168] When this processing flow is activated (started), the input detection unit 301 detects whether or not the voice of the conversation between the user 5 and the interlocutor 80 is input (S431). If the input detection unit 301 does not detect the voice of the conversation between the user 5 and the interlocutor 80 (S431: No), the processing ends.
[0169] On the other hand, if the input detection unit 301 detects the conversational voice between the user 5 and the interlocutor 80 (S431: Yes), the voice information processing unit 322 analyzes the conversational voice detected by the input detection unit 301, and the supplementary information setting unit 363 determines the personalities of the user 5 and the interlocutor 80 (S432).
[0170] The character setting unit 311 automatically sets at least one assistant character based on the information about the personality of the user, etc., determined by the supplemental information setting unit 363 in step S432 (S433).
[0171] For example, in combination with the scene determination process executed by the supplemental information setting unit 363 in the first example of the sixth embodiment, the character setting unit 311 automatically sets the personality of the assistant that complements the information, thoughts, and decisions that the user 5 is not good at, based on the personality of the user 5 himself / herself for each usage scene 511 (shopping, leisure, negotiation, etc.) shown in Fig. 11, to generate more optimal user assistance information 75. In particular, a machine learning algorithm may be used based on information about the user 5's past unsuccessful experiences to prevent them from being repeated (this can be achieved by having the local AI 364 perform machine learning, for example), thereby refining the user assistance information response over time and increasing its accuracy.
[0172] Similarly, the personalities of family members and others who frequently converse with the user 5 can be analyzed and the personality of "family members and others" in personality 501 shown in Fig. 10 can be automatically set. This allows the user 5 to obtain user assistance information 75 from the assistant obtained by setting the "family members and others" even when shopping alone.
[0173] The processing flow of FIG. 32 is repeated in a cycle of detecting the selection of various buttons arranged on the surface of the display 119 or in a cycle of display on the display 119 (for example, 60, 120, 240 frames per second, etc.).
[0174] According to the sixth embodiment, the situation of the user 5 can be automatically determined from the conversational voice and the video of the surroundings, and the personality of the user 5 can be automatically determined from the user's everyday conversations, etc., and the determination results can be used to further modify the response of the generation AI 3 to suit the user 5, thereby providing the user assistance information 75. Furthermore, by installing the local AI 364 in the information terminal and having the local AI 364 perform machine learning of the user's conversational voice, etc. as needed, it is possible to automatically and repeatedly learn the user 5's personality weaknesses, etc., without having to set them in advance, and to provide the latest, improved, optimal user assistance information 75 to draw attention.
[0175] If the information terminals 10 and 20 equipped with the local AI 364 are equipped with microphones 121 and 221, and acquire the voice picked up by the microphones 121 and 221 as external world information, and input the speech-converted text data obtained by converting the voice into text into the local AI 364 to generate a question, a question that better fits the user's thinking tendencies can be generated. Furthermore, by inputting an answer from the generation AI 3 into the local AI 364, which has learned the user's thinking tendencies through machine learning using the user's conversational voice, and complementing the answer, user assistance information that is more suited to the user's thinking tendencies can be presented. Furthermore, if the information terminals 10 and 20 cannot connect to a server equipped with the generation AI 3 or if it is desired to present user assistance information in an on-premises environment, inputting a question into the local AI 364 to obtain an answer can also aim to improve response and confidentiality in poor communication environments.
[0176] The first to sixth embodiments can be interchanged and combined to the extent that no technical contradiction occurs. For example, the processing described using the portable information terminal 10 as an example may be executed by the head-mounted transparent display information terminal 20, and the processing described using the head-mounted transparent display information terminal 20 as an example may be executed by the portable information terminal 10. Furthermore, when using the head-mounted transparent display information terminal 20, user assistance information 75 may be obtained from a portable information terminal 10 that is also owned. Furthermore, when making a call using the portable information terminal 10, user assistance information 75 may be obtained from the head-mounted transparent display information terminal 20. Furthermore, the head-mounted transparent display information terminal 20 may be configured to realize the mobile phone functions of the portable information terminal 10 (for example, a SIM card may be loaded into the head-mounted transparent display information terminal 20, and the telephone functions may be implemented in the head-mounted transparent display information terminal 20).
[0177] The present invention is not limited to the above-described embodiments, and it is possible to replace part of the configuration of one embodiment with another embodiment. It is also possible to add the configuration of another embodiment to the configuration of one embodiment. These all fall within the scope of the present invention, and the numerical values, messages, etc. appearing in the text and figures are merely examples, and the use of different ones does not impair the effects of the present invention.
[0178] Furthermore, some or all of the functions of the invention may be implemented in hardware, for example, by designing an integrated circuit, a general-purpose processor, or an application-specific processor. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. Alternatively, the invention may be implemented in software by a microprocessor unit, processor, etc., interpreting and executing an operating program. Furthermore, the scope of software implementation is not limited, and hardware and software may be used together.
[0179] The above embodiments include the following inventions: (Supplementary Note 1) An information terminal comprising: a processor; an input device; and an output device for providing at least one of audio and video display, wherein the processor generates question data for obtaining user assistance information to be presented to a user based on information input from the input device, obtains answer data to the question data generated by inputting the question data into a trained model, generates the user assistance information based on the answer data, and performs at least one of audio output and visual output of the user assistance information. (Supplementary Note 2) A method for providing user assistance information using an information terminal, comprising: a processor provided in the information terminal receives information input via the input device, generates question data for obtaining user assistance information to be presented to a user based on the information input from the input device, obtains answer data to the question data generated by inputting the question data into a trained model, and summarizes the answer data to generate and output the user assistance information.
[0180] 1: Server, 2: Network, 3: Generating AI, 4: Wireless Router, 5: User, 6: Finger, 10: Information Terminal, 15: Other Information Terminal, 20: Information Terminal, 21: Smart Glasses, 30: Smart Watch, 60: Information Terminal, 61: Virtual Reality Providing Server, 62: Virtual Reality Space, 63: Avatar, 64: Assistant, 65: Information Terminal, 66: Virtual Clerk, 71: Physical Store, 73: Assistant, 74: Interlocutor, 75: User Assistance Information, 76: Automatic Button, 77: Advice Button, 80: Interlocutor, 100: User Assistance Information Presentation System, 111: Out-Camera, 112: In-Camera Camera, 113: Distance measurement sensor, 115: Acceleration sensor, 116: Gyro sensor, 117: Geomagnetic sensor, 118: GPS receiver, 119: Display, 121: Microphone, 122: Speaker, 123: Antenna, 125: Processor, 126: Program, 127: Data, 128: Memory, 130: Touch panel, 131: Telephone network communication device, 140: Bus, 183: Audio processing unit, 214: RTC, 215: Acceleration sensor, 216: Gyro sensor, 217: Geomagnetic sensor, 218: GPS receiver, 219: Transparent display, 221: Microphone, 22 2: Speaker, 223: Antenna, 225: Processor, 228: Memory, 235: Camera, 240: Bus, 300: User assistance information presentation program, 301: Input detection unit, 302: Question generation unit, 303: User assistance condition setting unit, 304: Question transmission control unit, 305: Summary unit, 306: User assistance information presentation control unit, 307: Voice utterance unit, 308: Text display unit, 310: Character control unit, 311: Character setting unit, 312: Character operation unit, 320: Virtual reality processing unit, 321: Video information processing unit, 322: Voice information processing unit, 330: Avatar control Control unit, 331: avatar setting unit, 332: avatar operation unit, 340: augmented reality processing unit, 350: self-learning unit, 351: machine learning data collection unit, 352: machine learning execution unit, 361: text analysis unit, 362: database, 3621: conversation keyword table (TBL), 363: supplementary information setting unit, 364: local AI, 390: RAW data storage unit, 400: setting screen, 410: user assistance information presentation method, 411: ON / OFF switching, 412: ON / OFF switching, 420: advice viewpoint, 421: selection button, 422: selection button, 423: selection button,424: Selection button, 425: Server selection information, 428: Check box, 495a: Bluetooth earphone, 495b: Bluetooth earphone, 500a: Character setting screen, 500b: Character setting screen, 501: Personality, 502: Appearance, 503: Voice quality, 504: Automatic button, 505: Setting button, 511: Usage scene, 512: Personality, 560: Output unit, 611: Out camera, 612: In camera camera, 613: distance measurement sensor, 615: acceleration sensor, 616: gyro sensor, 617: geomagnetic sensor, 618: GPS receiver, 619: display, 621: microphone, 622: speaker, 623: antenna, 625: processor, 628: memory, 640: bus, 662L: left speaker, 690: battery, 694a: Bluetooth controller, 694b: Bluetooth controller,
Claims
1. An information terminal comprising: a processor; an input device; and an output device that performs at least one of audio and video display, wherein the processor generates question data for obtaining user assistance information to be presented to a user based on information input from the input device; obtains answer data to the question data that is generated by inputting the question data into a trained model; generates the user assistance information based on the answer data; and performs at least one of audio output of the user assistance information and visual output of the user assistance information.
2. An information terminal according to claim 1, wherein the input device is at least one of a microphone, a camera, and a communication interface for receiving telephone calls.
3. An information terminal according to claim 1, wherein the output device is at least one of a speaker, a wireless earphone communicatively connected to the information terminal, a wireless headphone communicatively connected to the information terminal, a non-transparent display, and a transparent display.
4. An information terminal according to claim 1, wherein the processor collects conversations between people other than the user as machine learning data, and uses the machine learning data to machine-train the trained model.
5. An information terminal as described in claim 1, wherein the trained model includes a plurality of trained models trained using machine learning data in line with each of a plurality of advice perspectives in order to generate answer data for the question data from each of the plurality of advice perspectives, and the processor accepts a selection of any one of the plurality of advice perspectives, and inputs the question data to the trained model corresponding to the selected advice perspective.
6. An information terminal as described in claim 1, wherein the processor accepts settings of presentation conditions for the user assistance information, selects the output device corresponding to the presentation conditions, and performs at least one of audio output of the user assistance information and display output of the user assistance information.
7. An information terminal as described in claim 1, further comprising a display, wherein the processor accepts settings of character attribute information including at least one of the appearance, personality, and attributes of an assistant character that presents the user assistance information, sets an assistant character corresponding to the set character attribute information, displays the assistant character on the display, and causes the assistant character to speak the user assistance information.
8. An information terminal according to claim 7, further comprising a speaker, wherein the processor causes the speaker to speak the user assistance information.
9. An information terminal according to claim 8, further comprising wireless earphones or wireless headphones communicatively connected to the information terminal, wherein the processor outputs a voice message to the user based on the user assistance information from the wireless earphones or wireless headphones, and outputs a voice message to the interlocutor based on the user assistance information from the speaker.
10. An information terminal as described in claim 1, further comprising a display, wherein the processor: accepts settings of character attribute information including at least one of the appearance, personality, and attributes of an assistant character that presents the user assistance information; places an avatar representing the user and an assistant character corresponding to the character attribute information in the virtual reality space; displays on the display an image of the avatar and the assistant character placed in the virtual reality space; accepts input of conversational audio of the user's avatar conversing with another person's avatar in the virtual reality space; and causes the assistant character to speak the user assistance information generated based on the conversational audio.
11. An information terminal as described in claim 1, further comprising a sensor for detecting the current position of the information terminal, wherein the processor acquires the current position of the information terminal from the sensor, determines the situation in which the user is placed, and generates the question data for acquiring the user assistance information required in the situation in which the user is placed.
12. An information terminal as described in claim 1, further comprising a microphone and a local machine learning model, wherein the processor acquires speech picked up by the microphone as information input from the input device, converts the speech into text, obtains converted speech text data, and inputs the converted text data into the local machine learning model to generate the question data.
13. An information terminal as described in claim 1, further comprising a microphone and a local machine learning model, wherein the processor acquires speech picked up by the microphone as information input from the input device, converts the speech into text, obtains converted text data, and inputs the converted text data into the local machine learning model to learn the user's thinking tendencies, and inputs the answer data into the local machine learning model and presents the user assistance information interpolated from the answer data.
14. An information terminal as described in claim 1, further comprising a local machine learning model, wherein the trained model is installed on a server communicatively connected to the information terminal, and wherein when the processor is unable to communicate with the server, it inputs the question data into the local machine learning model to obtain the answer data, and generates the user assistance information based on the answer data.
15. A method for presenting user assistance information using an information terminal, comprising: a step in which a processor provided in the information terminal accepts input of information via an input device; a step in which, based on the information input from the input device, a processor generates question data for obtaining user assistance information to be presented to a user; a step in which answer data to the question data is generated by inputting the question data into a trained model; and a step in which the processor summarizes the answer data to generate and output the user assistance information.
Citation Information
Patent Citations
Dfstination setting device and agent device
JP2000266551A
Telephone set and system
JP2011130337A
Conversation facilitating apparatus, conversation facilitating method, and conversation facilitating program
JP2024017074A