system

The system addresses the issue of poor chatbot quality by allowing voice input, partner selection, and personalized responses, ensuring quick and accurate answers through a multi-unit approach, enhancing user trust and conversational effectiveness.

JP2026073049APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing chatbots lack the ability to provide high-quality information that meets user needs, making it difficult for users to obtain relevant and accurate responses.

Method used

A system that allows users to input questions by voice, select a conversation partner's voice and age group, and receive personalized, appropriate answers through a reception unit, generation unit, and advice unit, utilizing speech recognition, natural language processing, and generation AI to analyze and generate responses.

Benefits of technology

Enables users to receive quick, accurate, and personalized answers from a conversation partner of their choice, enhancing trust and persuasiveness by providing tailored advice in a familiar tone, thus improving the conversational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073049000001_ABST
    Figure 2026073049000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to allow a user to input a question by voice, select a conversation partner, and obtain an appropriate answer. [Solution] The system according to the embodiment comprises a reception unit, a generation unit, a selection unit, and an advice unit. The reception unit receives a question from the user via voice input. The generation unit analyzes the question received by the reception unit and provides an appropriate answer. The selection unit allows the user to choose the voice and age group of the person they are talking to. The advice unit provides the answer generated by the generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that the quality of chatbots is poor and it is impossible to reach the information required by users.

[0005] The system according to the embodiment aims to enable a user to input a question by voice, select an interlocutor, and obtain an appropriate answer.

Means for Solving the Problems

[0006] The system according to the embodiment includes a reception unit, a generation unit, a selection unit, and an advice unit. The reception unit receives a question input by a user by voice. The generation unit analyzes the question received by the reception unit and provides an appropriate answer. The selection unit allows the user to select the voice and age group of the interlocutor. The advice unit provides the answer generated by the generation unit. [Effects of the Invention]

[0007] The system according to this embodiment allows the user to input a question by voice, select a conversation partner, and obtain an appropriate answer. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The voice-interactive support tool according to an embodiment of the present invention is a system in which a user inputs a question by voice, and a generating AI analyzes the question and provides an appropriate answer. The voice-interactive support tool allows the user to choose a conversation partner, analyze their thinking, and provide advice in an appropriate tone. This is thought to create a sense of trust in the information obtained through the conversation, making it easier for the user to accept and leading to subsequent actions. For example, the user inputs a question by voice, and the generating AI analyzes the question and provides an appropriate answer. The user can choose the voice and age group of the conversation partner, and can receive advice in the voices of different characters, such as a doctor version, a young woman version, or a grandmother version. This allows the user to receive advice from a conversation partner that suits them, increasing familiarity and persuasiveness. In addition, the generating AI has learned from a vast amount of data and examples, and can provide accurate support. For example, if a user asks, "Should I replace the company's PCs with Windows or Mac?", the generating AI analyzes the question and provides appropriate advice. The doctoral version provides advice such as, "In Japan, the overwhelming majority of users prefer Windows. However, if your work is creative, a Mac can also be effective. It's best to choose an OS that supports the software you want to use, taking cost-effectiveness into consideration." In this way, the voice-interactive support tool of the present invention enhances user trust and provides familiar and persuasive support by having the user's chosen conversation partner provide accurate advice. As a result, the voice-interactive support tool can provide quick and accurate answers to user questions.

[0029] The voice-interactive support tool according to this embodiment comprises a reception unit, a generation unit, a selection unit, and an advice unit. The reception unit receives questions from the user via voice. For example, the user can input questions via voice using a microphone. The reception unit can also convert voice data into text data using speech recognition technology. For example, the reception unit uses speech recognition software to convert the user's voice into text in real time. Furthermore, the reception unit can save the user's voice data for later analysis. The generation unit uses a generation AI to analyze the questions received by the reception unit and provide appropriate answers. For example, the generation AI learns from a vast amount of data and examples to generate the optimal answer to the user's question. The generation unit can also use natural language processing technology to enable the generation AI to understand the intent of the user's question and provide relevant information. For example, the generation AI analyzes the user's question, retrieves information from a relevant database, and generates an answer. The selection unit allows the user to choose the voice and age range of the conversation partner. For example, the selection unit allows the user to choose the tone and accent of the conversation partner's voice. Furthermore, the selection unit allows the user to choose the age group of the conversation partner. For example, the selection unit allows the user to choose the voice of different characters, such as a doctor version, a young woman version, or a grandmother version. The advice unit provides the answer generated by the generation unit. The advice unit can, for example, provide the answer in the voice of the conversation partner selected by the user. The advice unit can also provide the answer in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advice unit can provide an answer in a professional tone for the doctor version and in a casual tone for the young woman version. As a result, the voice-interactive support tool according to the embodiment allows the user to input a question by voice and obtain an appropriate answer.

[0030] The reception desk accepts user input via voice. For example, users can input questions using a microphone. Specifically, when a user speaks a question into the microphone, the voice data is immediately transmitted to the system. The reception desk can also convert the voice data into text data using speech recognition technology. For example, the reception desk uses speech recognition software to convert the user's voice into text in real time. The speech recognition software uses an algorithm that analyzes the voice waveform, identifies phonemes and words, and converts them into text. Furthermore, the reception desk can save the user's voice data for later analysis. The saved voice data is used to analyze the user's speech patterns and question tendencies. This allows the system to learn the user's preferences and specific needs, enabling it to provide more personalized services in the future. For example, if a user frequently asks questions about a particular topic, the system can be adjusted to prioritize providing information related to that topic. Additionally, the saved voice data can be used to refer to past questions and answers. This allows users to easily review their past conversation history and efficiently obtain information without repeating the same questions.

[0031] The generation unit uses a generation AI to analyze questions received by the reception unit and provide appropriate answers. For example, the generation unit's generation AI learns from a vast amount of data and examples to generate the optimal answer to a user's question. The generation AI utilizes natural language processing technology to understand the intent of the user's question and provide relevant information. Specifically, the generation AI analyzes the context and keywords of the question and extracts information from relevant databases and knowledge bases. For example, if a user asks, "What's the weather like lately?", the generation AI accesses a weather forecast database, retrieves the latest weather information, and generates an answer. The generation AI can also consider past conversation history and user preferences when generating answers to user questions. This makes it possible to provide more personalized answers. Furthermore, the generation AI has an algorithm that generates multiple answer candidates and selects the most appropriate one. For example, the generation AI generates multiple answers to a user's question, evaluates the relevance and reliability of each answer, and selects the most appropriate one. This allows the generation unit to provide quick and accurate answers to user questions.

[0032] The selection feature allows users to choose the voice and age range of their conversation partner. For example, users can choose the tone and accent of their conversation partner's voice. Specifically, users can select the type of voice for their conversation partner in the system settings screen. For example, multiple options are provided, such as a calm voice, an energetic voice, or a voice with a specific regional accent. The selection feature also allows users to choose the age range of their conversation partner. For example, users can choose from different character voices, such as a doctor version, a young woman version, or a grandmother version. This allows users to choose the best conversation partner according to their preferences and situation. Furthermore, the selection feature can save the user's selection history and automatically apply the same settings in subsequent conversations. This saves users the trouble of changing settings each time. For example, if a user selects the doctor version once, the doctor version will be automatically applied in subsequent conversations. In addition, the selection feature can collect user feedback and continuously improve the voice and character options. This allows the selection feature to provide users with a more satisfying conversation experience.

[0033] The advisory unit provides the answers generated by the generation unit. For example, the advisory unit can provide answers in the voice of the conversation partner selected by the user. Specifically, the advisory unit converts the text data received from the generation unit into speech using a speech synthesis engine selected by the user. The speech synthesis engine uses the voice profile of the conversation partner selected by the user to generate natural-sounding speech. The advisory unit can also provide answers in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advisory unit might provide answers in a professional tone for the "doctor" version and in a casual tone for the "gal" version. This allows the user to receive appropriate answers according to the character of the conversation partner. Furthermore, the advisory unit can collect user feedback and continuously improve the quality of the answers and the naturalness of the voice. For example, the system can improve based on user evaluations of the content and quality of the answers. The advisory unit can also present multiple answer options, allowing the user to select the most appropriate answer. This ensures the user receives answers that best suit their needs. The advisory unit can provide users with quick and appropriate answers, improving the conversational experience.

[0034] The generation unit can learn from a vast amount of data and examples. For example, the generation unit's generation AI can learn from various types of data, such as text data, audio data, and image data. The generation unit can also have the generation AI learn from past examples and provide the best possible answers to user questions. For example, the generation unit can have the generation AI learn from past question and answer data and generate appropriate answers to similar questions. Furthermore, the generation unit can have the generation AI continuously learn from new data and provide answers based on the latest information. For example, the generation unit can have the generation AI learn from the latest news articles and academic papers on the internet and provide the latest information to user questions. In this way, the generation unit can provide appropriate answers by learning from a vast amount of data and examples.

[0035] The selection section allows the user to choose the voice and age range of their conversation partner. For example, the selection section allows the user to choose the tone and accent of the conversation partner's voice. The selection section also allows the user to choose the age range of the conversation partner. For example, the selection section allows the user to choose voices of different characters, such as a doctor version, a young woman version, or a grandmother version. This allows the user to receive advice from a conversation partner that suits them, increasing familiarity and persuasiveness. Some or all of the above processing in the selection section may be performed using AI, for example, or not using AI. For example, the selection section can input the user's selection history into the AI ​​and suggest the most suitable conversation partner.

[0036] The advisory unit can provide accurate advice from a conversation partner selected by the user. For example, the advisory unit can provide answers in the voice of the conversation partner selected by the user. The advisory unit can also provide answers in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advisory unit can provide answers in a professional tone in the doctoral version and in a casual tone in the young woman version. This increases user trust and provides support that is both approachable and persuasive by ensuring that the conversation partner selected by the user provides accurate advice. Some or all of the above processing in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's selection history into AI and provide optimal advice.

[0037] The reception desk can analyze the user's past question history and select the optimal reception method. For example, the reception desk may prioritize suggesting question formats that the user has frequently used in the past. It can also select a reception method suitable for a specific time of day based on the user's past question history. Furthermore, the reception desk can suggest the optimal reception method based on the user's preferred conversational format. This allows the reception desk to provide the most suitable reception method by analyzing the user's past question history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's past question history into AI to select the optimal reception method.

[0038] The reception desk can filter questions based on the user's current situation and areas of interest when receiving them. For example, the reception desk can accept only questions relevant to the user's current situation. It can also prioritize questions on specific topics based on the user's areas of interest. Furthermore, the reception desk can filter and accept appropriate questions according to the user's current activity. This allows the reception desk to receive highly relevant questions by filtering them based on the user's current situation and areas of interest. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input data on the user's current situation and areas of interest into an AI and use that to filter questions.

[0039] The reception desk can prioritize receiving questions that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the reception desk can prioritize questions related to that region. It can also prioritize questions related to the user's travel destination if the user is traveling. Furthermore, if the user is at home, the reception desk can prioritize questions related to their home. This allows for the prioritization of highly relevant questions by considering the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's geographical location into AI and filter for highly relevant questions.

[0040] The reception desk can analyze the user's social media activity when receiving a question and accept relevant questions. For example, the reception desk can prioritize questions related to topics the user is discussing on social media. It can also suggest questions that the user might be interested in based on their social media activity history. Furthermore, the reception desk can accept relevant questions based on the content of posts from accounts the user follows. In this way, by analyzing the user's social media activity, it is possible to prioritize the acceptance of relevant questions. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media activity data into AI and filter relevant questions.

[0041] The generation unit can adjust the level of detail in the answers based on the importance of the questions when generating responses. For example, the generation unit can generate detailed answers for high-importance questions. It can also generate concise answers for low-importance questions. Furthermore, the generation unit can generate answers with an appropriate level of detail depending on the importance of the questions. In this way, by adjusting the level of detail in the answers based on the importance of the questions, it is possible to provide answers with an appropriate level of detail. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question importance data into AI and adjust the level of detail in the answers.

[0042] The generation unit can apply different generation algorithms depending on the question category when generating answers. For example, the generation unit can apply a specialized generation algorithm to technical questions. It can also apply a general-purpose generation algorithm to general questions. Furthermore, the generation unit can select the optimal generation algorithm for each category and generate answers. This allows for the provision of appropriate answers by applying different generation algorithms depending on the question category. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question category data into AI and select the optimal generation algorithm.

[0043] The generation unit can determine the priority of answers based on when the questions were submitted when generating answers. For example, the generation unit can determine the priority of answers based on the time period in which the questions were submitted. The generation unit can also generate answers at an appropriate time depending on when the questions were submitted. Furthermore, the generation unit can determine the optimal order of answers, taking into account when the questions were submitted. This allows for timely provision of answers by determining the priority of answers based on when the questions were submitted. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question submission time data into AI to determine the priority of answers.

[0044] The generation unit can adjust the order of answers based on the relevance of the questions when generating answers. For example, the generation unit can prioritize generating the most relevant answers based on the relevance of the questions. The generation unit can also generate answers in an appropriate order, taking into account the relevance of the questions. Furthermore, the generation unit can adjust the order of answers according to the relevance of the questions. This allows for the priority provision of highly relevant answers by adjusting the order of answers based on the relevance of the questions. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question relevance data into AI and adjust the order of answers.

[0045] The selection unit can provide the optimal choice when selecting a conversation partner by referring to the user's past selection history. For example, the selection unit can suggest the optimal choice based on the conversation partners the user has previously selected. The selection unit can also suggest a conversation partner suitable for a specific time of day based on the user's past selection history. Furthermore, the selection unit can provide the optimal choice based on the conversation partners the user has previously preferred to select. In this way, the optimal conversation partner can be provided by referring to the user's past selection history. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's past selection history data into AI and suggest the optimal conversation partner.

[0046] The selection unit can customize the options based on the user's current situation when selecting a conversation partner. For example, the selection unit can suggest the most suitable conversation partner based on the user's current situation. The selection unit can also customize and suggest an appropriate conversation partner according to the user's current activity status. Furthermore, the selection unit can select the most suitable conversation partner considering the user's current situation. In this way, an appropriate conversation partner can be provided by customizing the conversation partner based on the user's current situation. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's current situation data into AI and customize the conversation partner options.

[0047] The selection unit can provide the optimal choice when selecting a conversation partner, taking into account the user's geographical location. For example, if the user is in a specific region, the selection unit can suggest a conversation partner related to that region. Furthermore, if the user is traveling, the selection unit can suggest a conversation partner related to their travel destination. Additionally, if the user is at home, the selection unit can suggest a conversation partner related to their home. This allows the selection unit to provide the optimal conversation partner by considering the user's geographical location. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's geographical location data into AI to suggest the optimal conversation partner.

[0048] The selection unit can analyze the user's social media activity and suggest options when selecting a conversation partner. For example, the selection unit can suggest conversation partners related to topics the user is discussing on social media. It can also suggest conversation partners that the user might be interested in based on their social media activity history. Furthermore, the selection unit can suggest relevant conversation partners based on the content of posts from accounts the user follows. In this way, relevant conversation partners can be provided by analyzing the user's social media activity. Some or all of the above processing in the selection unit may be performed using AI, for example, or not. For example, the selection unit can input the user's social media activity data into AI and suggest the most suitable conversation partner.

[0049] The advisory unit can provide optimal advice by referring to the user's past question history when providing advice. For example, the advisory unit can provide relevant advice based on the content of questions the user has asked in the past. The advisory unit can also provide advice on specific topics from the user's past question history. Furthermore, the advisory unit can analyze the user's past question history and provide the most appropriate advice. This allows the advisory unit to provide optimal advice by referring to the user's past question history. Some or all of the above processes in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's past question history data into AI and provide optimal advice.

[0050] The advisory unit can customize the content of the advice based on the user's current situation when providing advice. For example, the advisory unit can provide the best advice based on the user's current situation. The advisory unit can also customize and provide appropriate advice according to the user's current activity status. Furthermore, the advisory unit can provide the best advice considering the user's current situation. In this way, appropriate advice can be provided by customizing the content of the advice based on the user's current situation. Some or all of the above processing in the advisory unit may be performed using AI, for example, or without using AI. For example, the advisory unit can input the user's current situation data into AI and customize the content of the advice.

[0051] The advisory unit can provide optimal advice by considering the user's geographical location when providing advice. For example, if the user is in a specific region, the advisory unit can provide advice related to that region. Furthermore, if the user is traveling, the advisory unit can provide advice related to the travel destination. In addition, if the user is at home, the advisory unit can provide advice related to the home. This allows the advisory unit to provide optimal advice by considering the user's geographical location. Some or all of the above processing in the advisory unit may be performed using AI, for example, or without AI. For example, the advisory unit can input the user's geographical location data into AI to provide optimal advice.

[0052] The advisory unit can analyze the user's social media activity and propose advice when providing it. For example, the advisory unit can provide advice related to topics the user is discussing on social media. It can also suggest advice that the user might be interested in based on their social media activity history. Furthermore, the advisory unit can provide relevant advice based on the content of posts from accounts the user follows. In this way, relevant advice can be provided by analyzing the user's social media activity. Some or all of the above processing in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's social media activity data into AI and propose the most suitable advice.

[0053] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0054] The generation unit can learn the user's preferred answer style based on their past question history and provide appropriate answers. For example, if the user previously preferred detailed explanations, it will generate detailed answers. It can also generate concise answers if the user previously preferred brief answers. Furthermore, if the user has shown interest in a particular topic in the past, it can prioritize providing information related to that topic. This allows the system to leverage the user's past question history to provide the most suitable answers.

[0055] The selection unit can monitor the user's current activity level in real time and suggest the most suitable conversation partner. For example, if the user is exercising, it can suggest an energetic conversation partner. If the user is working, it can suggest a professional conversation partner. Furthermore, if the user is relaxed, it can suggest a friendly conversation partner. In this way, by suggesting a conversation partner based on the user's current activity level, it can provide the user with the most optimal support.

[0056] The reception desk can prioritize region-specific questions based on the user's geographical location. For example, if a user is in a specific city, it will prioritize questions related to that city. Similarly, if a user is traveling, it can prioritize questions related to their travel destination. Furthermore, if a user is at home, it can prioritize questions related to their home. This allows the system to prioritize highly relevant questions by considering the user's geographical location.

[0057] The reception desk can analyze users' social media activity and prioritize relevant questions. For example, it can prioritize questions related to topics users are discussing on social media. It can also suggest questions that users might be interested in based on their social media activity history. Furthermore, it can accept relevant questions based on the content of posts from accounts users follow. In this way, by analyzing users' social media activity, it is possible to prioritize the acceptance of relevant questions.

[0058] The selection function can suggest the most suitable conversation partner based on the user's past selection history. For example, it can suggest the best options based on conversation partners the user has previously selected. It can also suggest conversation partners suitable for a specific time of day based on the user's past selection history. Furthermore, it can provide the best options based on conversation partners the user has previously preferred. In this way, the system can provide the most suitable conversation partner by referring to the user's past selection history.

[0059] The following briefly describes the processing flow for example form 1.

[0060] Step 1: The reception desk receives the user's question via voice input. For example, the user can use a microphone to input their question by voice, and speech recognition technology can be used to convert the voice data into text data. Furthermore, the reception desk can save the user's voice data and analyze it later. Step 2: The generation unit analyzes the questions received by the reception unit and provides appropriate answers. The generation unit learns from a vast amount of data and examples using generation AI, understands the intent of the user's questions using natural language processing technology, and provides relevant information. Step 3: The selection section allows the user to choose the voice and age range of the person they are talking to. For example, the user can choose the tone and accent of the person they are talking to, as well as their age range (e.g., doctor version, young woman version, grandmother version). Step 4: The advice unit provides the answer generated by the generation unit. For example, it provides the answer in the voice of the conversation partner selected by the user, and in a tone appropriate to the selected age group.

[0061] (Example of form 2) The voice-interactive support tool according to an embodiment of the present invention is a system in which a user inputs a question by voice, and a generating AI analyzes the question and provides an appropriate answer. The voice-interactive support tool allows the user to choose a conversation partner, analyze their thinking, and provide advice in an appropriate tone. This is thought to create a sense of trust in the information obtained through the conversation, making it easier for the user to accept and leading to subsequent actions. For example, the user inputs a question by voice, and the generating AI analyzes the question and provides an appropriate answer. The user can choose the voice and age group of the conversation partner, and can receive advice in the voices of different characters, such as a doctor version, a young woman version, or a grandmother version. This allows the user to receive advice from a conversation partner that suits them, increasing familiarity and persuasiveness. In addition, the generating AI has learned from a vast amount of data and examples, and can provide accurate support. For example, if a user asks, "Should I replace the company's PCs with Windows or Mac?", the generating AI analyzes the question and provides appropriate advice. The doctoral version provides advice such as, "In Japan, the overwhelming majority of users prefer Windows. However, if your work is creative, a Mac can also be effective. It's best to choose an OS that supports the software you want to use, taking cost-effectiveness into consideration." In this way, the voice-interactive support tool of the present invention enhances user trust and provides familiar and persuasive support by having the user's chosen conversation partner provide accurate advice. As a result, the voice-interactive support tool can provide quick and accurate answers to user questions.

[0062] The voice-interactive support tool according to this embodiment comprises a reception unit, a generation unit, a selection unit, and an advice unit. The reception unit receives questions from the user via voice. For example, the user can input questions via voice using a microphone. The reception unit can also convert voice data into text data using speech recognition technology. For example, the reception unit uses speech recognition software to convert the user's voice into text in real time. Furthermore, the reception unit can save the user's voice data for later analysis. The generation unit uses a generation AI to analyze the questions received by the reception unit and provide appropriate answers. For example, the generation AI learns from a vast amount of data and examples to generate the optimal answer to the user's question. The generation unit can also use natural language processing technology to enable the generation AI to understand the intent of the user's question and provide relevant information. For example, the generation AI analyzes the user's question, retrieves information from a relevant database, and generates an answer. The selection unit allows the user to choose the voice and age range of the conversation partner. For example, the selection unit allows the user to choose the tone and accent of the conversation partner's voice. Furthermore, the selection unit allows the user to choose the age group of the conversation partner. For example, the selection unit allows the user to choose the voice of different characters, such as a doctor version, a young woman version, or a grandmother version. The advice unit provides the answer generated by the generation unit. The advice unit can, for example, provide the answer in the voice of the conversation partner selected by the user. The advice unit can also provide the answer in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advice unit can provide an answer in a professional tone for the doctor version and in a casual tone for the young woman version. As a result, the voice-interactive support tool according to the embodiment allows the user to input a question by voice and obtain an appropriate answer.

[0063] The reception desk accepts user input via voice. For example, users can input questions using a microphone. Specifically, when a user speaks a question into the microphone, the voice data is immediately transmitted to the system. The reception desk can also convert the voice data into text data using speech recognition technology. For example, the reception desk uses speech recognition software to convert the user's voice into text in real time. The speech recognition software uses an algorithm that analyzes the voice waveform, identifies phonemes and words, and converts them into text. Furthermore, the reception desk can save the user's voice data for later analysis. The saved voice data is used to analyze the user's speech patterns and question tendencies. This allows the system to learn the user's preferences and specific needs, enabling it to provide more personalized services in the future. For example, if a user frequently asks questions about a particular topic, the system can be adjusted to prioritize providing information related to that topic. Additionally, the saved voice data can be used to refer to past questions and answers. This allows users to easily review their past conversation history and efficiently obtain information without repeating the same questions.

[0064] The generation unit uses a generation AI to analyze questions received by the reception unit and provide appropriate answers. For example, the generation unit's generation AI learns from a vast amount of data and examples to generate the optimal answer to a user's question. The generation AI utilizes natural language processing technology to understand the intent of the user's question and provide relevant information. Specifically, the generation AI analyzes the context and keywords of the question and extracts information from relevant databases and knowledge bases. For example, if a user asks, "What's the weather like lately?", the generation AI accesses a weather forecast database, retrieves the latest weather information, and generates an answer. The generation AI can also consider past conversation history and user preferences when generating answers to user questions. This makes it possible to provide more personalized answers. Furthermore, the generation AI has an algorithm that generates multiple answer candidates and selects the most appropriate one. For example, the generation AI generates multiple answers to a user's question, evaluates the relevance and reliability of each answer, and selects the most appropriate one. This allows the generation unit to provide quick and accurate answers to user questions.

[0065] The selection feature allows users to choose the voice and age range of their conversation partner. For example, users can choose the tone and accent of their conversation partner's voice. Specifically, users can select the type of voice for their conversation partner in the system settings screen. For example, multiple options are provided, such as a calm voice, an energetic voice, or a voice with a specific regional accent. The selection feature also allows users to choose the age range of their conversation partner. For example, users can choose from different character voices, such as a doctor version, a young woman version, or a grandmother version. This allows users to choose the best conversation partner according to their preferences and situation. Furthermore, the selection feature can save the user's selection history and automatically apply the same settings in subsequent conversations. This saves users the trouble of changing settings each time. For example, if a user selects the doctor version once, the doctor version will be automatically applied in subsequent conversations. In addition, the selection feature can collect user feedback and continuously improve the voice and character options. This allows the selection feature to provide users with a more satisfying conversation experience.

[0066] The advisory unit provides the answers generated by the generation unit. For example, the advisory unit can provide answers in the voice of the conversation partner selected by the user. Specifically, the advisory unit converts the text data received from the generation unit into speech using a speech synthesis engine selected by the user. The speech synthesis engine uses the voice profile of the conversation partner selected by the user to generate natural-sounding speech. The advisory unit can also provide answers in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advisory unit might provide answers in a professional tone for the "doctor" version and in a casual tone for the "gal" version. This allows the user to receive appropriate answers according to the character of the conversation partner. Furthermore, the advisory unit can collect user feedback and continuously improve the quality of the answers and the naturalness of the voice. For example, the system can improve based on user evaluations of the content and quality of the answers. The advisory unit can also present multiple answer options, allowing the user to select the most appropriate answer. This ensures the user receives answers that best suit their needs. The advisory unit can provide users with quick and appropriate answers, improving the conversational experience.

[0067] The generation unit can learn from a vast amount of data and examples. For example, the generation unit's generation AI can learn from various types of data, such as text data, audio data, and image data. The generation unit can also have the generation AI learn from past examples and provide the best possible answers to user questions. For example, the generation unit can have the generation AI learn from past question and answer data and generate appropriate answers to similar questions. Furthermore, the generation unit can have the generation AI continuously learn from new data and provide answers based on the latest information. For example, the generation unit can have the generation AI learn from the latest news articles and academic papers on the internet and provide the latest information to user questions. In this way, the generation unit can provide appropriate answers by learning from a vast amount of data and examples.

[0068] The selection section allows the user to choose the voice and age range of their conversation partner. For example, the selection section allows the user to choose the tone and accent of the conversation partner's voice. The selection section also allows the user to choose the age range of the conversation partner. For example, the selection section allows the user to choose voices of different characters, such as a doctor version, a young woman version, or a grandmother version. This allows the user to receive advice from a conversation partner that suits them, increasing familiarity and persuasiveness. Some or all of the above processing in the selection section may be performed using AI, for example, or not using AI. For example, the selection section can input the user's selection history into the AI ​​and suggest the most suitable conversation partner.

[0069] The advisory unit can provide accurate advice from a conversation partner selected by the user. For example, the advisory unit can provide answers in the voice of the conversation partner selected by the user. The advisory unit can also provide answers in a tone appropriate to the age group of the conversation partner selected by the user. For example, the advisory unit can provide answers in a professional tone in the doctoral version and in a casual tone in the young woman version. This increases user trust and provides support that is both approachable and persuasive by ensuring that the conversation partner selected by the user provides accurate advice. Some or all of the above processing in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's selection history into AI and provide optimal advice.

[0070] The reception desk can estimate the user's emotions and adjust the timing of question acceptance based on the estimated emotions. For example, if the user is stressed, the reception desk can temporarily delay question acceptance to give the user time to relax. Conversely, if the user is agitated, the reception desk can accept questions immediately and respond quickly. Furthermore, if the user is tired, the reception desk can simplify the question acceptance process and adjust it to be completed in a short time. In this way, the user's stress is reduced by adjusting the timing of question acceptance based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI, for example, or not using AI. For example, the reception desk can input user emotion data into AI and adjust the timing of question acceptance.

[0071] The reception desk can analyze the user's past question history and select the optimal reception method. For example, the reception desk may prioritize suggesting question formats that the user has frequently used in the past. It can also select a reception method suitable for a specific time of day based on the user's past question history. Furthermore, the reception desk can suggest the optimal reception method based on the user's preferred conversational format. This allows the reception desk to provide the most suitable reception method by analyzing the user's past question history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's past question history into AI to select the optimal reception method.

[0072] The reception desk can filter questions based on the user's current situation and areas of interest when receiving them. For example, the reception desk can accept only questions relevant to the user's current situation. It can also prioritize questions on specific topics based on the user's areas of interest. Furthermore, the reception desk can filter and accept appropriate questions according to the user's current activity. This allows the reception desk to receive highly relevant questions by filtering them based on the user's current situation and areas of interest. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input data on the user's current situation and areas of interest into an AI and use that to filter questions.

[0073] The reception desk can estimate the user's emotions and determine the priority of questions to accept based on the estimated emotions. For example, if the user is feeling anxious, the reception desk will prioritize accepting urgent questions. If the user is relaxed, the reception desk may also prioritize accepting normal questions. Furthermore, if the user is excited, the reception desk may prioritize accepting interesting questions. This allows for prioritizing urgent questions based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input user emotion data into an AI to determine the priority of questions.

[0074] The reception desk can prioritize receiving questions that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific region, the reception desk can prioritize questions related to that region. It can also prioritize questions related to the user's travel destination if the user is traveling. Furthermore, if the user is at home, the reception desk can prioritize questions related to their home. This allows for the prioritization of highly relevant questions by considering the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can input the user's geographical location into AI and filter for highly relevant questions.

[0075] The reception desk can analyze the user's social media activity when receiving a question and accept relevant questions. For example, the reception desk can prioritize questions related to topics the user is discussing on social media. It can also suggest questions that the user might be interested in based on their social media activity history. Furthermore, the reception desk can accept relevant questions based on the content of posts from accounts the user follows. In this way, by analyzing the user's social media activity, it is possible to prioritize the acceptance of relevant questions. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media activity data into AI and filter relevant questions.

[0076] The generation unit can estimate the user's emotions and adjust the way the response is expressed based on the estimated emotions. For example, if the user is feeling anxious, the generation unit can generate a response using reassuring language. It can also generate a response using friendly language if the user is relaxed. Furthermore, if the user is excited, the generation unit can generate a response using energetic language. This allows for the provision of appropriate responses to the user by adjusting the way the response is expressed based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, or not. For example, the generation unit can input user emotion data into an AI and adjust the way the response is expressed.

[0077] The generation unit can adjust the level of detail in the answers based on the importance of the questions when generating responses. For example, the generation unit can generate detailed answers for high-importance questions. It can also generate concise answers for low-importance questions. Furthermore, the generation unit can generate answers with an appropriate level of detail depending on the importance of the questions. In this way, by adjusting the level of detail in the answers based on the importance of the questions, it is possible to provide answers with an appropriate level of detail. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question importance data into AI and adjust the level of detail in the answers.

[0078] The generation unit can apply different generation algorithms depending on the question category when generating answers. For example, the generation unit can apply a specialized generation algorithm to technical questions. It can also apply a general-purpose generation algorithm to general questions. Furthermore, the generation unit can select the optimal generation algorithm for each category and generate answers. This allows for the provision of appropriate answers by applying different generation algorithms depending on the question category. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question category data into AI and select the optimal generation algorithm.

[0079] The generation unit can estimate the user's emotions and adjust the length of the response based on the estimated emotions. For example, if the user is in a hurry, the generation unit can generate a short, concise response. If the user is relaxed, the generation unit can also generate a longer response with detailed explanations. Furthermore, if the user is excited, the generation unit can generate a response with visually stimulating effects. This allows the system to provide responses of an appropriate length for the user by adjusting the response length based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI or not. For example, the generation unit can input user emotion data into the AI ​​and adjust the response length.

[0080] The generation unit can determine the priority of answers based on when the questions were submitted when generating answers. For example, the generation unit can determine the priority of answers based on the time period in which the questions were submitted. The generation unit can also generate answers at an appropriate time depending on when the questions were submitted. Furthermore, the generation unit can determine the optimal order of answers, taking into account when the questions were submitted. This allows for timely provision of answers by determining the priority of answers based on when the questions were submitted. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question submission time data into AI to determine the priority of answers.

[0081] The generation unit can adjust the order of answers based on the relevance of the questions when generating answers. For example, the generation unit can prioritize generating the most relevant answers based on the relevance of the questions. The generation unit can also generate answers in an appropriate order, taking into account the relevance of the questions. Furthermore, the generation unit can adjust the order of answers according to the relevance of the questions. This allows for the priority provision of highly relevant answers by adjusting the order of answers based on the relevance of the questions. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input question relevance data into AI and adjust the order of answers.

[0082] The selection unit can estimate the user's emotions and adjust the method of selecting a conversation partner based on the estimated emotions. For example, if the user is feeling anxious, the selection unit will prioritize suggesting a conversation partner that provides a sense of security. It can also suggest a friendly conversation partner if the user is relaxed. Furthermore, if the user is excited, it can suggest an energetic conversation partner. This allows the system to provide the user with an appropriate conversation partner by adjusting the selection method based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the selection unit may be performed using AI, or not. For example, the selection unit can input user emotion data into an AI to adjust the method of selecting a conversation partner.

[0083] The selection unit can provide the optimal choice when selecting a conversation partner by referring to the user's past selection history. For example, the selection unit can suggest the optimal choice based on the conversation partners the user has previously selected. The selection unit can also suggest a conversation partner suitable for a specific time of day based on the user's past selection history. Furthermore, the selection unit can provide the optimal choice based on the conversation partners the user has previously preferred to select. In this way, the optimal conversation partner can be provided by referring to the user's past selection history. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's past selection history data into AI and suggest the optimal conversation partner.

[0084] The selection unit can customize the options based on the user's current situation when selecting a conversation partner. For example, the selection unit can suggest the most suitable conversation partner based on the user's current situation. The selection unit can also customize and suggest an appropriate conversation partner according to the user's current activity status. Furthermore, the selection unit can select the most suitable conversation partner considering the user's current situation. In this way, an appropriate conversation partner can be provided by customizing the conversation partner based on the user's current situation. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's current situation data into AI and customize the conversation partner options.

[0085] The selection unit can estimate the user's emotions and determine the priority of conversation partners based on the estimated emotions. For example, if the user is feeling anxious, the selection unit will prioritize selecting a conversation partner who provides a sense of security. Similarly, if the user is relaxed, the selection unit may prioritize selecting a friendly conversation partner. Furthermore, if the user is excited, the selection unit may prioritize selecting an energetic conversation partner. This allows the system to provide appropriate conversation partners by prioritizing them based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the selection unit may be performed using AI, or not. For example, the selection unit can input user emotion data into an AI to determine the priority of conversation partners.

[0086] The selection unit can provide the optimal choice when selecting a conversation partner, taking into account the user's geographical location. For example, if the user is in a specific region, the selection unit can suggest a conversation partner related to that region. Furthermore, if the user is traveling, the selection unit can suggest a conversation partner related to their travel destination. Additionally, if the user is at home, the selection unit can suggest a conversation partner related to their home. This allows the selection unit to provide the optimal conversation partner by considering the user's geographical location. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's geographical location data into AI to suggest the optimal conversation partner.

[0087] The selection unit can analyze the user's social media activity and suggest options when selecting a conversation partner. For example, the selection unit can suggest conversation partners related to topics the user is discussing on social media. It can also suggest conversation partners that the user might be interested in based on their social media activity history. Furthermore, the selection unit can suggest relevant conversation partners based on the content of posts from accounts the user follows. In this way, relevant conversation partners can be provided by analyzing the user's social media activity. Some or all of the above processing in the selection unit may be performed using AI, for example, or not. For example, the selection unit can input the user's social media activity data into AI and suggest the most suitable conversation partner.

[0088] The advice unit can estimate the user's emotions and adjust the way it expresses advice based on those emotions. For example, if the user is feeling anxious, the advice unit will provide advice in a reassuring way. It can also provide advice in a friendly way if the user is relaxed. Furthermore, if the user is excited, it can provide advice in an energetic way. This allows the advice unit to provide appropriate advice to the user by adjusting the way it expresses advice based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the advice unit may be performed using AI, or not. For example, the advice unit can input user emotion data into an AI and adjust the way it expresses advice.

[0089] The advisory unit can provide optimal advice by referring to the user's past question history when providing advice. For example, the advisory unit can provide relevant advice based on the content of questions the user has asked in the past. The advisory unit can also provide advice on specific topics from the user's past question history. Furthermore, the advisory unit can analyze the user's past question history and provide the most appropriate advice. This allows the advisory unit to provide optimal advice by referring to the user's past question history. Some or all of the above processes in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's past question history data into AI and provide optimal advice.

[0090] The advisory unit can customize the content of the advice based on the user's current situation when providing advice. For example, the advisory unit can provide the best advice based on the user's current situation. The advisory unit can also customize and provide appropriate advice according to the user's current activity status. Furthermore, the advisory unit can provide the best advice considering the user's current situation. In this way, appropriate advice can be provided by customizing the content of the advice based on the user's current situation. Some or all of the above processing in the advisory unit may be performed using AI, for example, or without using AI. For example, the advisory unit can input the user's current situation data into AI and customize the content of the advice.

[0091] The advisory unit can estimate the user's emotions and determine the priority of advice based on the estimated emotions. For example, if the user is feeling anxious, the advisory unit will prioritize providing urgent advice. It can also prioritize providing normal advice if the user is relaxed. Furthermore, if the user is excited, it can prioritize providing interesting advice. This allows for the prioritization of urgent advice based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the advisory unit may be performed using AI, or not. For example, the advisory unit can input user emotion data into an AI to determine the priority of advice.

[0092] The advisory unit can provide optimal advice by considering the user's geographical location when providing advice. For example, if the user is in a specific region, the advisory unit can provide advice related to that region. Furthermore, if the user is traveling, the advisory unit can provide advice related to the travel destination. In addition, if the user is at home, the advisory unit can provide advice related to the home. This allows the advisory unit to provide optimal advice by considering the user's geographical location. Some or all of the above processing in the advisory unit may be performed using AI, for example, or without AI. For example, the advisory unit can input the user's geographical location data into AI to provide optimal advice.

[0093] The advisory unit can analyze the user's social media activity and propose advice when providing it. For example, the advisory unit can provide advice related to topics the user is discussing on social media. It can also suggest advice that the user might be interested in based on their social media activity history. Furthermore, the advisory unit can provide relevant advice based on the content of posts from accounts the user follows. In this way, relevant advice can be provided by analyzing the user's social media activity. Some or all of the above processing in the advisory unit may be performed using AI, for example, or not using AI. For example, the advisory unit can input the user's social media activity data into AI and propose the most suitable advice.

[0094] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0095] The reception desk can analyze the user's tone and speed of voice to estimate their emotional state. For example, if a user speaks quickly, it can estimate they are excited and immediately accept a question. If a user speaks slowly, it can estimate they are relaxed and accept a more detailed question. Furthermore, if a user speaks in a low tone, it can estimate they are tired and accept a simplified question. By adjusting the questioning method based on the user's tone and speed of voice, the system can provide optimal support to the user.

[0096] The generation unit can learn the user's preferred answer style based on their past question history and provide appropriate answers. For example, if the user previously preferred detailed explanations, it will generate detailed answers. It can also generate concise answers if the user previously preferred brief answers. Furthermore, if the user has shown interest in a particular topic in the past, it can prioritize providing information related to that topic. This allows the system to leverage the user's past question history to provide the most suitable answers.

[0097] The selection unit can monitor the user's current activity level in real time and suggest the most suitable conversation partner. For example, if the user is exercising, it can suggest an energetic conversation partner. If the user is working, it can suggest a professional conversation partner. Furthermore, if the user is relaxed, it can suggest a friendly conversation partner. In this way, by suggesting a conversation partner based on the user's current activity level, it can provide the user with the most optimal support.

[0098] The advice function can estimate the user's emotions and adjust the advice based on those emotions. For example, if the user is feeling anxious, it can provide reassuring advice. If the user is relaxed, it can provide detailed advice. Furthermore, if the user is excited, it can provide energetic advice. In this way, by adjusting the advice based on the user's emotions, it can provide advice that is appropriate for the user.

[0099] The reception desk can prioritize region-specific questions based on the user's geographical location. For example, if a user is in a specific city, it will prioritize questions related to that city. Similarly, if a user is traveling, it can prioritize questions related to their travel destination. Furthermore, if a user is at home, it can prioritize questions related to their home. This allows the system to prioritize highly relevant questions by considering the user's geographical location.

[0100] The generation unit can estimate the user's emotions and adjust the tone of the response based on those emotions. For example, if the user is feeling anxious, it can generate a response in a reassuring tone. If the user is relaxed, it can generate a response in a friendly tone. Furthermore, if the user is excited, it can generate a response in an energetic tone. By adjusting the tone of the response based on the user's emotions, it can provide the user with an appropriate response.

[0101] The reception desk can analyze users' social media activity and prioritize relevant questions. For example, it can prioritize questions related to topics users are discussing on social media. It can also suggest questions that users might be interested in based on their social media activity history. Furthermore, it can accept relevant questions based on the content of posts from accounts users follow. In this way, by analyzing users' social media activity, it is possible to prioritize the acceptance of relevant questions.

[0102] The generation unit can estimate the user's emotions and adjust the level of detail in the response based on that estimation. For example, if the user is in a hurry, it can generate a concise response. If the user is relaxed, it can generate a detailed response. Furthermore, if the user is excited, it can generate a response with visually stimulating effects. By adjusting the level of detail in the response based on the user's emotions, it can provide the user with an appropriate response.

[0103] The selection function can suggest the most suitable conversation partner based on the user's past selection history. For example, it can suggest the best options based on conversation partners the user has previously selected. It can also suggest conversation partners suitable for a specific time of day based on the user's past selection history. Furthermore, it can provide the best options based on conversation partners the user has previously preferred. In this way, the system can provide the most suitable conversation partner by referring to the user's past selection history.

[0104] The advice unit can estimate the user's emotions and prioritize advice based on those emotions. For example, if the user is feeling anxious, it can prioritize urgent advice. If the user is relaxed, it can prioritize normal advice. Furthermore, if the user is excited, it can prioritize interesting advice. By prioritizing advice based on the user's emotions, it can prioritize urgent advice.

[0105] The following briefly describes the processing flow for example form 2.

[0106] Step 1: The reception desk receives the user's question via voice input. For example, the user can use a microphone to input their question by voice, and speech recognition technology can be used to convert the voice data into text data. Furthermore, the reception desk can save the user's voice data and analyze it later. Step 2: The generation unit analyzes the questions received by the reception unit and provides appropriate answers. The generation unit learns from a vast amount of data and examples using generation AI, understands the intent of the user's questions using natural language processing technology, and provides relevant information. Step 3: The selection section allows the user to choose the voice and age range of the person they are talking to. For example, the user can choose the tone and accent of the person they are talking to, as well as their age range (e.g., doctor version, young woman version, grandmother version). Step 4: The advice unit provides the answer generated by the generation unit. For example, it provides the answer in the voice of the conversation partner selected by the user, and in a tone appropriate to the selected age group.

[0107] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0108] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0109] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0110] Each of the multiple elements described above, including the reception unit, generation unit, selection unit, and advice unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives the user's voice using the microphone 38B of the smart device 14 and converts the voice data into text data using the control unit 46A. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and analyzes the user's question using generation AI and generates an appropriate answer. The selection unit is implemented in the control unit 46A of the smart device 14, for example, and allows the user to select the voice and age group of the person they are talking to. The advice unit provides the user with an answer using the speaker 40B of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0111] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0112] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0113] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0114] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0115] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0116] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0117] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0118] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0119] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0120] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0121] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0122] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0123] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0124] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0125] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0126] Each of the multiple elements described above, including the reception unit, generation unit, selection unit, and advice unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives the user's voice using the microphone 238 of the smart glasses 214 and converts the voice data into text data using the control unit 46A. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the user's question using generation AI and generates an appropriate answer. The selection unit is implemented, for example, by the control unit 46A of the smart glasses 214, which allows the user to select the voice and age group of the person they are talking to. The advice unit provides the user with an answer using, for example, the speaker 240 of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0127] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0128] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0129] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0130] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0131] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0132] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0133] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0134] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0135] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0136] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0137] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0138] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0139] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0140] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0141] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0142] Each of the multiple elements described above, including the reception unit, generation unit, selection unit, and advice unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives the user's voice using the microphone 238 of the headset terminal 314 and converts the voice data into text data using the control unit 46A. The generation unit is implemented in the identification processing unit 290 of the data processing unit 12, for example, and analyzes the user's question using a generation AI and generates an appropriate answer. The selection unit is implemented in the control unit 46A of the headset terminal 314, for example, and allows the user to select the voice and age group of the person they are talking to. The advice unit provides answers to the user using the speaker 240 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0143] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0144] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0145] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0146] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0147] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0149] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0150] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0151] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0152] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0153] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0154] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0155] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0156] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0157] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0158] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0159] Each of the multiple elements described above, including the reception unit, generation unit, selection unit, and advice unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives the user's voice using the microphone 238 of the robot 414 and converts the voice data into text data using the control unit 46A. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the user's question using a generation AI and generates an appropriate answer. The selection unit is implemented, for example, by the control unit 46A of the robot 414, allowing the user to select the voice and age group of the person they are talking to. The advice unit provides answers to the user using, for example, the speaker 240 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0160] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0161] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0162] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0163] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0164] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0165] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0166] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0167] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0168] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0169] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0170] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0171] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0172] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0173] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0174] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0175] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0176] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0177] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0178] (Note 1) A reception desk where users input questions by voice, A generation unit analyzes the questions received by the reception unit and provides appropriate answers, A selection section where the user can choose the voice and age range of the person they are talking to, The system includes an advisory unit that provides the answer generated by the generation unit. A system characterized by the following features. (Note 2) The generating unit is Learn from a vast amount of data and case studies. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned selection unit is Users can choose the voice and age range of their conversation partner. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned advisory unit, The conversation partner chosen by the user provides accurate advice. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of question submissions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is Analyze the user's past question history and select the most suitable method of handling inquiries. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is When receiving a question, filtering is performed based on the user's current situation and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is It estimates the user's emotions and determines the priority of questions to ask based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving questions, the system prioritizes accepting questions that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is When receiving a question, the system analyzes the user's social media activity and accepts relevant questions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The generating unit is It estimates the user's emotions and adjusts the way responses are expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The generating unit is When generating answers, adjust the level of detail in the answers based on the importance of the question. The system described in Appendix 1, characterized by the features described herein. (Note 13) The generating unit is When generating answers, different generation algorithms are applied depending on the question category. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is It estimates the user's emotions and adjusts the length of the response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is When generating answers, the system prioritizes answers based on when the questions were submitted. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is When generating answers, the order of answers is adjusted based on the relevance of the questions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned selection unit is It estimates the user's emotions and adjusts how the conversation partner is selected based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned selection unit is When selecting a conversation partner, the system provides the most suitable option by referring to the user's past selection history. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned selection unit is When selecting a conversation partner, customize the options based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned selection unit is It estimates the user's emotions and determines the priority of the conversation partner based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned selection unit is When selecting a conversation partner, the system takes the user's geographical location into consideration to provide the most suitable option. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned selection unit is When selecting a conversation partner, the system analyzes the user's social media activity to suggest options. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned advisory unit, It estimates the user's emotions and adjusts the way advice is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned advisory unit, When providing advice, we refer to the user's past question history to provide the most appropriate advice. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned advisory unit, When providing advice, customize the content of the advice based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned advisory unit, It estimates the user's emotions and determines the priority of advice based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned advisory unit, When providing advice, we take the user's geographical location into consideration to provide the most appropriate advice. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned advisory unit, When providing advice, we analyze the user's social media activity and propose content for the advice. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0179] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception desk where users input questions by voice, A generation unit analyzes the questions received by the reception unit and provides appropriate answers, A selection section where the user can choose the voice and age range of the person they are talking to, The system includes an advisory unit that provides the answer generated by the generation unit. A system characterized by the following features.

2. The generating unit is Learn from a vast amount of data and case studies. The system according to feature 1.

3. The aforementioned selection unit is Users can choose the voice and age range of their conversation partner. The system according to feature 1.

4. The aforementioned advisory unit, The conversation partner chosen by the user provides accurate advice. The system according to feature 1.

5. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of question submissions based on those estimated emotions. The system according to feature 1.

6. The aforementioned reception unit is Analyze the user's past question history and select the most suitable method of handling inquiries. The system according to feature 1.

7. The aforementioned reception unit is When receiving a question, filtering is performed based on the user's current situation and areas of interest. The system according to feature 1.

8. The aforementioned reception unit is The system estimates the user's emotions and prioritizes the questions to be asked based on those estimated emotions. The system according to feature 1.

9. The aforementioned reception unit is When receiving questions, the system prioritizes accepting questions that are highly relevant, taking into account the user's geographical location. The system according to feature 1.

10. The aforementioned reception unit is When receiving a question, the system analyzes the user's social media activity and accepts relevant questions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A