System
The integration of speech recognition and generative AI in robot systems addresses the limitations of conventional robots by enabling diverse and personalized conversations through advanced language recognition and response generation.
Patent Information
- Application Number
- JP2024132363
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Conventional robot systems are limited in their conversational capabilities, primarily relying on programmed content, which restricts the range and depth of interactions.
A system combining speech recognition technology and generative AI to recognize user utterances and generate appropriate responses, including multilingual support, emotion analysis, and personalized interactions.
Enables a wide variety of conversations by recognizing multiple languages and dialects, analyzing user emotions and urgency, and generating personalized responses, thereby enhancing interaction flexibility and effectiveness.
Smart Images

Figure 2026029514000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With conventional technology, the robot's responses were limited to programmed content, making it difficult to have a wide range of conversations.
[0005] The system according to the embodiment aims to realize a variety of conversations by combining speech recognition and generation AI. [Means for solving the problem]
[0006] The system according to the embodiment includes a speech recognition unit, a generation unit, and a response unit. The speech recognition unit recognizes a user's speech using speech recognition technology. The generation unit generates an appropriate response based on the speech content recognized by the speech recognition unit. The response unit provides the user with the response generated by the generation unit. [Effects of the Invention]
[0007] The system according to the embodiment can realize a variety of conversations by combining voice recognition and generation AI. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION
[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0010] First, the terms used in the following description will be explained.
[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).
[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.
[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example 1) A robot system according to an embodiment of the present invention is a system that recognizes user utterances and engages in a variety of conversations by combining speech recognition technology and generative AI. This allows the robot system to recognize user utterances and generate and provide appropriate responses.
[0029] A robot system according to an embodiment includes a speech recognition unit, a generation unit, and a response unit. The speech recognition unit recognizes a user's speech. For example, the speech recognition unit converts the user's speech into text data using deep learning-based speech recognition technology. The speech recognition unit can also analyze the speech content using a hidden Markov model (HMM). The speech recognition unit can also perform phonemic analysis to analyze the speech content in detail. The generation unit generates an appropriate response based on the speech content recognized by the speech recognition unit. For example, the generation unit can generate a natural response based on the user's speech content using a text generation AI (e.g., GPT-3). The generation unit can also generate a response based on a predefined template. The generation unit can also generate a response based on the user's intention. The response unit provides the response generated by the generation unit to the user. For example, the response unit can provide a voice response using a speaker. The response unit can also display a text response using a display. The response unit can also send a notification to the user's device. This allows the robot system according to an embodiment to recognize a user's speech and generate and provide an appropriate response.
[0030] The speech recognition unit can simultaneously recognize multiple languages and dialects and generate responses according to the user's language selection. For example, the speech recognition unit enhances speech recognition technology to simultaneously recognize multiple languages and dialects. For example, if a user speaks a mixture of English and Spanish, the robot recognizes both languages and generates a response in the appropriate language. The speech recognition unit also generates a response according to the user's language selection. For example, if a user says, "Bonjour, comment ca va?" in French, the robot responds in French with, "Bonjour, je vais bien, merci. Et vous?" The speech recognition unit also simultaneously recognizes multiple languages and dialects and generates a response according to the user's language selection. For example, if a user says, "How is it today?" in Japanese, the robot responds in Japanese with, "The weather is very nice today." This enables multilingual support by recognizing multiple languages and dialects and generating a response according to the user's language selection.
[0031] The speech recognition unit can analyze the user's speech rate and tone, determine the user's level of urgency and importance, and adjust the response accordingly. The speech recognition unit, for example, uses speech recognition technology to analyze the user's speech rate and tone. For example, if the user speaks in a hurry, the robot recognizes the level of urgency and generates a quick response. The speech recognition unit also analyzes the user's speech rate and tone to determine the level of urgency and importance. For example, if the user speaks in a calm tone, the robot determines the level of importance as low and generates a normal response. The speech recognition unit also analyzes the speech rate and tone to determine the user's level of urgency and importance and adjust the response accordingly. For example, if the user speaks in a high tone, the robot determines the level of importance as high and generates a quick and detailed response. This makes it possible to respond appropriately by analyzing the user's speech rate and tone and generating a response according to the level of urgency and importance.
[0032] The speech recognition unit can translate what a user says in real time and support communication between users who speak different languages. The speech recognition unit translates what a user says in real time, for example, using speech recognition technology. For example, if a user speaks in English, the robot translates that content into Japanese and responds in Japanese. The speech recognition unit also supports communication between users who speak different languages. For example, when an English-speaking user and a Japanese-speaking user are conversing, the robot stands between them and translates in real time. The speech recognition unit also translates what a user says in real time and supports communication between users who speak different languages. For example, when a French-speaking user and a Spanish-speaking user are conversing, the robot stands between them and translates in real time. In this way, the speech recognition unit can translate what a user says in real time and support communication between users who speak different languages, thereby enabling smooth multilingual communication.
[0033] The speech recognition unit can save the user's speech content as text data and build a database that can be searched and analyzed later. The speech recognition unit, for example, uses speech recognition technology to save the user's speech content as text data. For example, if a user says, "What's the weather like today?", the content is saved as text data. The speech recognition unit also registers the saved text data in a database that can be searched and analyzed later. For example, it searches past conversations and extracts conversations related to specific keywords. The speech recognition unit also saves the user's speech content as text data and builds a database that can be searched and analyzed later. For example, it saves the content of customer support conversations and analyzes them later to solve problems. In this way, the content of user's speech is saved as text data and a database that can be searched and analyzed later is built, making it easier to manage and utilize information.
[0034] The generation unit can generate more personalized responses by referring to the user's past conversation history. The generation unit, for example, strengthens the generation AI and refers to the user's past conversation history. For example, if the user previously said, "I have a dog," the generation unit generates a response such as, "How is your dog doing?" in the next conversation. The generation unit also generates personalized responses based on the user's past conversation history. For example, if the user said, "I like traveling," the generation unit generates a response such as, "Have you traveled anywhere recently?" The generation unit also refers to the past conversation history and generates responses based on the user's preferences and interests. For example, if the user said, "I like movies," the generation unit generates a response such as, "What movie have you seen recently?" This makes it possible to generate more personalized responses by referring to the user's past conversation history.
[0035] The generation unit can automatically search for related information based on the user's utterance and include it in the response. The generation unit, for example, uses a generation AI to automatically search for related information based on the user's utterance. For example, if a user asks, "What's the weather like today?", the generation unit searches for weather information and generates a response such as, "It's sunny today. The temperature is about 25 degrees." The generation unit also searches for related information based on the user's utterance and includes it in the response. For example, if a user asks, "What's the latest news?", the generation unit searches for the latest news and generates a response such as, "The latest news is ____." The generation unit also automatically searches for related information based on the user's utterance and includes it in the response. For example, if a user says, "Tell me about a restaurant nearby," the generation unit searches for nearby restaurant information and generates a response such as, "There's a restaurant called ____ nearby." This automatically searches for related information based on the user's utterance and includes it in the response, enabling more informative conversations.
[0036] The generation unit generates a story based on the user's speech, which can be used for entertainment or educational purposes. The generation unit generates a story based on the user's speech, for example, using a generation AI. For example, if the user starts by saying, "Once upon a time," the robot generates the rest of the story and creates a story. The generation unit also generates stories for entertainment purposes based on the user's speech. For example, if the user says, "Tell me an adventure story," the robot generates an adventure story and begins to tell it. The generation unit also generates stories for educational purposes based on the user's speech. For example, if the user says, "Tell me a history story," the robot generates a story based on historical events and begins to tell it. In this way, generating stories based on the user's speech can be used for entertainment or educational purposes.
[0037] The generation unit can automatically generate emails and messages based on the content of the user's utterances to support communication. The generation unit, for example, uses a generation AI to automatically generate emails based on the content of the user's utterances. For example, if the user says, "Let me know about my meeting schedule," the robot generates and sends an email based on that content. The generation unit also automatically generates messages based on the content of the user's utterances. For example, if the user says, "Tell my friend thank you," the robot generates and sends a message based on that content. The generation unit also automatically generates emails and messages to support communication based on the content of the utterances. For example, if the user says, "Send the report to my boss," the robot generates and sends an email based on that content. In this way, communication can be supported by automatically generating emails and messages based on the content of the user's utterances.
[0038] A robot can upgrade its hardware and equip itself with more advanced sensors and cameras to analyze the user's movements and facial expressions. For example, a robot can upgrade its hardware and equip itself with more advanced sensors. For example, a motion sensor can be added to analyze the user's movements. The robot can also be equipped with an advanced camera to analyze the user's facial expressions. For example, if the user speaks to the robot with a smile, the robot can recognize that expression and generate a positive response. The robot can also use sensors and cameras to analyze the user's movements and facial expressions in real time. For example, if the user waves their hand, the robot can recognize that movement and generate a response such as "Hello!" This enables more advanced interactions by analyzing the user's movements and facial expressions.
[0039] A robot can update its software and combine multiple generation AIs to realize a wider variety of conversation patterns. For example, a robot can update its software and combine multiple generation AIs. For example, it can combine AIs with different conversation styles to generate a wider variety of responses. A robot can also use multiple generation AIs to generate the optimal response based on what the user says. For example, if a user asks, "What's the weather like today?", it can combine an AI that provides weather information with an AI that continues the conversation. A robot can also update its software to optimize the combination of generation AIs. For example, it can select the optimal combination of AIs based on the user's past conversation history and generate a response. In this way, by combining multiple generation AIs, a wider variety of conversation patterns can be realized.
[0040] Robots can be repurposed for different industries and applications, and can be used to support communication in medical settings and nursing homes. For example, robots can be repurposed in medical settings to support communication with patients. For example, they can talk to patients, asking, "How are you feeling today?" and listen to their symptoms. Robots can also be used to support communication in nursing homes. For example, they can talk to elderly people, asking, "What did you do today?" and enjoy everyday conversation. Robots can also be repurposed in different industries to provide communication support. For example, in education, they can support students by asking, "What subject will you study today?" This means that robots can be repurposed for different industries and applications, and can be used to support communication in medical settings and nursing homes.
[0041] Multiple robots can be linked together to build a system in which they work together as a team to accomplish tasks. For example, multiple robots can work together to accomplish tasks as a team. For example, multiple robots can work together to assist customers in customer service. Robots can also develop systems in which robots work together to efficiently share tasks. For example, multiple robots can work together to support classes in educational settings. Robots can also build robot collaboration systems to accomplish tasks as a team. For example, multiple robots can work together to care for patients in medical settings. In this way, by linking multiple robots together, it is possible to build a system in which they work together as a team to accomplish tasks.
[0042] The robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes learning support functions for educational settings. Furthermore, the robot can perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes customer support functions for customer service work. Furthermore, the robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes communication support functions for use in the home. In this way, by performing customized demonstrations tailored to its target customer base, it is possible to increase the marketing effect.
[0043] Robots can approach more customers by making their pricing flexible and introducing subscription or leasing models. For example, Robots can make their pricing flexible and introduce a subscription model. For example, they can offer a plan that allows customers to use the robot for a monthly fee. Robots can also introduce a leasing model that allows customers to use the robot for a certain period of time. For example, they can offer a one-year lease contract with an option to purchase thereafter. Robots can also make their pricing flexible and introduce subscription or leasing models. For example, they can offer a monthly usage plan for customers who want to use the robot for a short period of time. By making their pricing flexible and introducing subscription or leasing models, they can approach more customers.
[0044] Robots can be deployed in different markets to acquire new customers in the education and entertainment markets. For example, robots can be deployed in the education market to acquire new customers. For example, their learning support functions can be enhanced to promote their use in schools and cram schools. Robots can also be deployed in the entertainment market to acquire new customers. For example, their performance functions at events and shows can be enhanced. Robots can also be deployed in different markets to acquire new customers. For example, their patient support functions in the medical market can be enhanced to promote their use in hospitals and clinics. In this way, new customers can be acquired by deploying robots in different markets.
[0045] A robot can be developed in collaboration with a partner company and conduct a joint marketing campaign to acquire new customers. For example, a robot may develop a robot in collaboration with a partner company and conduct a joint marketing campaign. For example, a robot may partner with an educational institution to jointly develop a learning assistance robot. A robot may also conduct a joint marketing campaign to acquire new customers. For example, a robot may partner with an entertainment company to jointly promote a robot performance at an event. A robot may also develop a robot in collaboration with a partner company and conduct a marketing campaign. For example, a robot may partner with a medical institution to jointly develop a patient assistance robot and promote its use in hospitals. In this way, a robot can acquire new customers by developing a robot in collaboration with a partner company and conducting a joint marketing campaign.
[0046] A robot can analyze a student's learning progress in an educational setting and provide individually customized learning support. For example, in an educational setting, a robot can analyze a student's learning progress and provide individually customized learning support. For example, if a student says, "I'm not good at math," the robot analyzes that progress and suggests an appropriate learning plan. The robot can also analyze a student's learning progress in real time and provide individually customized learning support. For example, if a student says, "English grammar is difficult," the robot analyzes that progress and provides grammar practice questions. The robot can also analyze a student's learning progress and provide individually customized learning support. For example, if a student says, "My history test is coming up," the robot analyzes that progress and suggests a study plan for test preparation. In this way, by analyzing a student's learning progress in an educational setting and providing individually customized learning support, it is possible to improve learning effectiveness.
[0047] A robot can manage family schedules within a home and provide reminders at appropriate times. For example, a robot can manage family schedules within a home and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "You have a doctor's appointment tomorrow." A robot can also manage family schedules in real time and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "We have plans for a family trip this weekend." A robot can also manage schedules and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "Your child has a school event today." In this way, by managing family schedules within a home and providing reminders at appropriate times, communication within the home can be facilitated.
[0048] A robot can analyze a customer's purchase history during customer service and make personalized product suggestions. For example, a robot can analyze a customer's purchase history during customer service and make personalized product suggestions. For example, a robot can suggest, "I also recommend this product," based on products the customer has purchased in the past. A robot can also analyze a customer's purchase history in real time and make personalized product suggestions. For example, if a customer says, "The shampoo I bought last time was great," the robot can suggest, "I also recommend conditioner from the same brand." A robot can also analyze a customer's purchase history and make personalized product suggestions. For example, if a customer says, "The wine I recently purchased was delicious," the robot can suggest, "I also recommend other wines from the same winery." In this way, customer satisfaction can be improved by analyzing a customer's purchase history during customer service and making personalized product suggestions.
[0049] Robots can converse with multiple students simultaneously in educational settings and promote group discussions. For example, a robot can speak to multiple students at the same time, asking, "What do you all think about this topic?". A robot can also converse with multiple students at the same time and promote group discussions. For example, a robot can ask, "What do you think about this problem?", encouraging students to exchange opinions among themselves. A robot can also converse with multiple students to promote group discussions. For example, a robot can say, "Let's discuss this topic," encouraging students to exchange opinions among themselves. In this way, a robot can converse with multiple students simultaneously in educational settings and promote group discussions, thereby stimulating the exchange of opinions among students.
[0050] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0051] The robot system may further include a health management unit that monitors the user's health condition. For example, the health management unit may measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. The health management unit may also record the user's amount of exercise and generate a message encouraging exercise if it detects a lack of exercise. Furthermore, the health management unit may record the user's diet and provide advice on nutritional balance. This allows the robot system to support the user's health management and contribute to maintaining good health.
[0052] The robot system can further include an information providing unit that provides customized information based on the user's hobbies and interests. For example, if the user says, "I like traveling," the latest travel information can be provided. If the user says, "I like cooking," new recipes can be suggested. If the user says, "I like sports," the latest sports news can be provided. This allows the robot system to provide more personalized services by providing information based on the user's hobbies and interests.
[0053] The robot system can further include a learning support unit that analyzes the user's learning progress and provides an individually customized learning plan. For example, if the user says, "I'm not good at math," the robot system can analyze the user's learning progress and suggest an appropriate learning plan. If the user says, "English grammar is difficult," the robot system can provide grammar practice questions. If the user says, "My history test is coming up," the robot system can suggest a study plan for test preparation. In this way, the robot system can analyze the user's learning progress and provide individually customized learning support, thereby improving learning effectiveness.
[0054] The robot system can further include a purchasing support unit that analyzes the user's purchasing history and makes personalized product suggestions. For example, the robot system can suggest "I also recommend this product" based on products the user has purchased in the past. If the user says, "The shampoo I bought last time was good," the robot system can suggest, "I also recommend the conditioner from the same brand." If the user says, "The wine I recently purchased was delicious," the robot system can suggest, "I also recommend other wines from the same winery." In this way, the robot system can improve customer satisfaction by analyzing the user's purchasing history and making personalized product suggestions.
[0055] The robot system can further include a schedule management unit that manages the user's schedule and provides reminders at appropriate times. For example, it can provide a reminder such as "You have a doctor's appointment tomorrow." It can also provide a reminder such as "You have a family trip planned this weekend." It can also provide a reminder such as "Your child has a school event today." In this way, the robot system can support the user's daily life by managing the user's schedule and providing reminders at appropriate times.
[0056] The processing flow of the first embodiment will be briefly explained below.
[0057] Step 1: The speech recognition unit recognizes the user's speech. For example, the speech recognition unit converts the user's speech into text data using deep learning-based speech recognition technology. The speech recognition unit can also analyze the content of the speech using an HMM (hidden Markov model). Furthermore, the speech recognition unit can perform phonemic analysis to analyze the content of the speech in detail. Step 2: The generator generates an appropriate response based on the utterance recognized by the speech recognizer. For example, the generator uses a text generation AI (e.g., GPT-3) to generate a natural response based on the user's utterance. The generator can also generate a response based on a predefined template. Furthermore, the generator can also generate a response based on the user's intention. Step 3: The response unit provides the response generated by the generation unit to the user. For example, the response unit may provide the response audibly using a speaker. The response unit may also display the response in text using a display. Furthermore, the response unit may also send a notification to the user's device.
[0058] (Example 2) A robot system according to an embodiment of the present invention is a system that recognizes user utterances and engages in a variety of conversations by combining speech recognition technology and generative AI. This allows the robot system to recognize user utterances and generate and provide appropriate responses.
[0059] A robot system according to an embodiment includes a speech recognition unit, a generation unit, and a response unit. The speech recognition unit recognizes a user's speech. For example, the speech recognition unit converts the user's speech into text data using deep learning-based speech recognition technology. The speech recognition unit can also analyze the speech content using a hidden Markov model (HMM). The speech recognition unit can also perform phonemic analysis to analyze the speech content in detail. The generation unit generates an appropriate response based on the speech content recognized by the speech recognition unit. For example, the generation unit can generate a natural response based on the user's speech content using a text generation AI (e.g., GPT-3). The generation unit can also generate a response based on a predefined template. The generation unit can also generate a response based on the user's intention. The response unit provides the response generated by the generation unit to the user. For example, the response unit can provide a voice response using a speaker. The response unit can also display a text response using a display. The response unit can also send a notification to the user's device. This allows the robot system according to an embodiment to recognize a user's speech and generate and provide an appropriate response.
[0060] The speech recognition unit can analyze a user's emotions in real time and generate a response that corresponds to the emotion. For example, the speech recognition unit incorporates an emotion estimation function into speech recognition technology to analyze not only the content of a user's speech but also their emotions. For example, if a user says, "I'm tired today," the robot recognizes that emotion and generates a response such as, "Thank you for your hard work. Shall I teach you how to relax?" The speech recognition unit also uses the emotion estimation function to generate a response that corresponds to the user's emotions. For example, if a user says, "I have some good news," the robot generates a response such as, "That's great! What's the news?" The speech recognition unit also analyzes a user's emotions in real time and generates a response that corresponds to the emotion. For example, if a user says, "I'm sad today," the robot generates a response such as, "What happened? Want to talk?" This allows for a more natural and friendly conversation by generating a response that corresponds to the user's emotions.
[0061] The speech recognition unit can simultaneously recognize multiple languages and dialects and generate responses according to the user's language selection. For example, the speech recognition unit enhances speech recognition technology to simultaneously recognize multiple languages and dialects. For example, if a user speaks a mixture of English and Spanish, the robot recognizes both languages and generates a response in the appropriate language. The speech recognition unit also generates a response according to the user's language selection. For example, if a user says, "Bonjour, comment ca va?" in French, the robot responds in French with, "Bonjour, je vais bien, merci. Et vous?" The speech recognition unit also simultaneously recognizes multiple languages and dialects and generates a response according to the user's language selection. For example, if a user says, "How is it today?" in Japanese, the robot responds in Japanese with, "The weather is very nice today." This enables multilingual support by recognizing multiple languages and dialects and generating a response according to the user's language selection.
[0062] The speech recognition unit can analyze the user's speech rate and tone, determine the user's level of urgency and importance, and adjust the response accordingly. The speech recognition unit, for example, uses speech recognition technology to analyze the user's speech rate and tone. For example, if the user speaks in a hurry, the robot recognizes the level of urgency and generates a quick response. The speech recognition unit also analyzes the user's speech rate and tone to determine the level of urgency and importance. For example, if the user speaks in a calm tone, the robot determines the level of importance as low and generates a normal response. The speech recognition unit also analyzes the speech rate and tone to determine the user's level of urgency and importance and adjust the response accordingly. For example, if the user speaks in a high tone, the robot determines the level of importance as high and generates a quick and detailed response. This makes it possible to respond appropriately by analyzing the user's speech rate and tone and generating a response according to the level of urgency and importance.
[0063] The speech recognition unit can translate what a user says in real time and support communication between users who speak different languages. The speech recognition unit translates what a user says in real time, for example, using speech recognition technology. For example, if a user speaks in English, the robot translates that content into Japanese and responds in Japanese. The speech recognition unit also supports communication between users who speak different languages. For example, when an English-speaking user and a Japanese-speaking user are conversing, the robot stands between them and translates in real time. The speech recognition unit also translates what a user says in real time and supports communication between users who speak different languages. For example, when a French-speaking user and a Spanish-speaking user are conversing, the robot stands between them and translates in real time. In this way, the speech recognition unit can translate what a user says in real time and support communication between users who speak different languages, thereby enabling smooth multilingual communication.
[0064] The speech recognition unit can save the user's speech content as text data and build a database that can be searched and analyzed later. The speech recognition unit, for example, uses speech recognition technology to save the user's speech content as text data. For example, if a user says, "What's the weather like today?", the content is saved as text data. The speech recognition unit also registers the saved text data in a database that can be searched and analyzed later. For example, it searches past conversations and extracts conversations related to specific keywords. The speech recognition unit also saves the user's speech content as text data and builds a database that can be searched and analyzed later. For example, it saves the content of customer support conversations and analyzes them later to solve problems. In this way, the content of user's speech is saved as text data and a database that can be searched and analyzed later is built, making it easier to manage and utilize information.
[0065] The speech recognition unit can analyze a user's emotions and generate a customized response based on the emotions. For example, the speech recognition unit adds an emotion estimation function to speech recognition technology to analyze the user's emotions. For example, if a user says, "I'm tired today," the speech recognition unit recognizes the emotion and generates a response such as, "Thank you for your hard work. Shall I teach you how to relax?" The speech recognition unit also uses the emotion estimation function to generate a customized response based on the user's emotions. For example, if a user says, "I have some good news," the speech recognition unit generates a response such as, "That's great! What's the news?" The speech recognition unit also analyzes a user's emotions in real time and generates a customized response based on the emotions. For example, if a user says, "I'm sad today," the speech recognition unit generates a response such as, "What happened? Want to talk?" This enables more personalized responses by generating customized responses based on the user's emotions.
[0066] The generation unit generates responses based on the user's emotions, enabling more natural conversations. For example, the generation unit incorporates an emotion estimation function into the generation AI to analyze the user's emotions. For example, if a user says, "I'm tired today," the generation unit recognizes that emotion and generates a response such as, "Thank you for your hard work. Shall I teach you how to relax?" The generation unit also uses the emotion estimation function to generate responses based on the user's emotions. For example, if a user says, "I have some good news," the generation unit generates a response such as, "That's great! What's the news?" The generation unit also analyzes the user's emotions in real time and generates a response based on the emotions. For example, if a user says, "I'm sad today," the generation unit generates a response such as, "What happened? Want to talk?" This allows for more natural and friendly conversations by generating responses based on the user's emotions.
[0067] The generation unit can generate more personalized responses by referring to the user's past conversation history. The generation unit, for example, strengthens the generation AI and refers to the user's past conversation history. For example, if the user previously said, "I have a dog," the generation unit generates a response such as, "How is your dog doing?" in the next conversation. The generation unit also generates personalized responses based on the user's past conversation history. For example, if the user said, "I like traveling," the generation unit generates a response such as, "Have you traveled anywhere recently?" The generation unit also refers to the past conversation history and generates responses based on the user's preferences and interests. For example, if the user said, "I like movies," the generation unit generates a response such as, "What movie have you seen recently?" This makes it possible to generate more personalized responses by referring to the user's past conversation history.
[0068] The generation unit can automatically search for related information based on the user's utterance and include it in the response. The generation unit, for example, uses a generation AI to automatically search for related information based on the user's utterance. For example, if a user asks, "What's the weather like today?", the generation unit searches for weather information and generates a response such as, "It's sunny today. The temperature is about 25 degrees." The generation unit also searches for related information based on the user's utterance and includes it in the response. For example, if a user asks, "What's the latest news?", the generation unit searches for the latest news and generates a response such as, "The latest news is ____." The generation unit also automatically searches for related information based on the user's utterance and includes it in the response. For example, if a user says, "Tell me about a restaurant nearby," the generation unit searches for nearby restaurant information and generates a response such as, "There's a restaurant called ____ nearby." This automatically searches for related information based on the user's utterance and includes it in the response, enabling more informative conversations.
[0069] The generation unit generates a story based on the user's speech, which can be used for entertainment or educational purposes. The generation unit generates a story based on the user's speech, for example, using a generation AI. For example, if the user starts by saying, "Once upon a time," the robot generates the rest of the story and creates a story. The generation unit also generates stories for entertainment purposes based on the user's speech. For example, if the user says, "Tell me an adventure story," the robot generates an adventure story and begins to tell it. The generation unit also generates stories for educational purposes based on the user's speech. For example, if the user says, "Tell me a history story," the robot generates a story based on historical events and begins to tell it. In this way, generating stories based on the user's speech can be used for entertainment or educational purposes.
[0070] The generation unit can automatically generate emails and messages based on the content of the user's utterances to support communication. The generation unit, for example, uses a generation AI to automatically generate emails based on the content of the user's utterances. For example, if the user says, "Let me know about my meeting schedule," the robot generates and sends an email based on that content. The generation unit also automatically generates messages based on the content of the user's utterances. For example, if the user says, "Tell my friend thank you," the robot generates and sends a message based on that content. The generation unit also automatically generates emails and messages to support communication based on the content of the utterances. For example, if the user says, "Send the report to my boss," the robot generates and sends an email based on that content. In this way, communication can be supported by automatically generating emails and messages based on the content of the user's utterances.
[0071] The generation unit generates a response based on the user's emotions and can provide emotional support. For example, the generation unit adds an emotion estimation function to the generation AI and analyzes the user's emotions. For example, if a user says, "I'm sad today," the generation unit recognizes that emotion and generates a response such as, "What happened? Wanna talk?" The generation unit also uses the emotion estimation function to generate a response based on the user's emotions. For example, if a user says, "I have some good news," the generation unit generates a response such as, "That's great! What's the news?" The generation unit also analyzes the user's emotions in real time and generates a response based on the emotion. For example, if a user says, "I'm tired today," the generation unit generates a response such as, "Thank you for your hard work. Shall I teach you how to relax?" This makes it possible to provide emotional support by generating a response based on the user's emotions.
[0072] A robot can incorporate an emotion estimation function and generate responses based on the user's emotions, making it more approachable. For example, a robot can incorporate an emotion estimation function and analyze the user's emotions. For example, if the user says, "I'm sad today," the robot can recognize that emotion and generate a response such as, "What happened? Wanna tell me?" The robot can also use the emotion estimation function to generate a response based on the user's emotions. For example, if the user says, "I have some good news," the robot can generate a response such as, "That's great! What's the news?" The robot can also analyze the user's emotions in real time and generate a response based on the emotions. For example, if the user says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Shall I teach you how to relax?" This allows the robot to generate responses based on the user's emotions, making it more approachable.
[0073] A robot can upgrade its hardware and equip itself with more advanced sensors and cameras to analyze the user's movements and facial expressions. For example, a robot can upgrade its hardware and equip itself with more advanced sensors. For example, a motion sensor can be added to analyze the user's movements. The robot can also be equipped with an advanced camera to analyze the user's facial expressions. For example, if the user speaks to the robot with a smile, the robot can recognize that expression and generate a positive response. The robot can also use sensors and cameras to analyze the user's movements and facial expressions in real time. For example, if the user waves their hand, the robot can recognize that movement and generate a response such as "Hello!" This enables more advanced interactions by analyzing the user's movements and facial expressions.
[0074] A robot can update its software and combine multiple generation AIs to realize a wider variety of conversation patterns. For example, a robot can update its software and combine multiple generation AIs. For example, it can combine AIs with different conversation styles to generate a wider variety of responses. A robot can also use multiple generation AIs to generate the optimal response based on what the user says. For example, if a user asks, "What's the weather like today?", it can combine an AI that provides weather information with an AI that continues the conversation. A robot can also update its software to optimize the combination of generation AIs. For example, it can select the optimal combination of AIs based on the user's past conversation history and generate a response. In this way, by combining multiple generation AIs, a wider variety of conversation patterns can be realized.
[0075] Robots can be repurposed for different industries and applications, and can be used to support communication in medical settings and nursing homes. For example, robots can be repurposed in medical settings to support communication with patients. For example, they can talk to patients, asking, "How are you feeling today?" and listen to their symptoms. Robots can also be used to support communication in nursing homes. For example, they can talk to elderly people, asking, "What did you do today?" and enjoy everyday conversation. Robots can also be repurposed in different industries to provide communication support. For example, in education, they can support students by asking, "What subject will you study today?" This means that robots can be repurposed for different industries and applications, and can be used to support communication in medical settings and nursing homes.
[0076] Multiple robots can be linked together to build a system in which they work together as a team to accomplish tasks. For example, multiple robots can work together to accomplish tasks as a team. For example, multiple robots can work together to assist customers in customer service. Robots can also develop systems in which robots work together to efficiently share tasks. For example, multiple robots can work together to support classes in educational settings. Robots can also build robot collaboration systems to accomplish tasks as a team. For example, multiple robots can work together to care for patients in medical settings. In this way, by linking multiple robots together, it is possible to build a system in which they work together as a team to accomplish tasks.
[0077] A robot can provide emotional support by adding an emotion estimation function and generating a response based on the user's emotions. For example, a robot can add an emotion estimation function and analyze the user's emotions. For example, if a user says, "I'm sad today," the robot recognizes the emotion and generates a response such as, "What happened? Wanna talk?" The robot can also use the emotion estimation function to generate a response based on the user's emotions. For example, if a user says, "I have some good news," the robot can generate a response such as, "That's great! What's the news?" The robot can also analyze the user's emotions in real time and generate a response based on the emotion. For example, if a user says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Shall I teach you how to relax?" This allows the robot to provide emotional support by generating a response based on the user's emotions.
[0078] A robot can incorporate an emotion estimation function to generate a response according to a user's emotions, thereby improving customer satisfaction. For example, a robot can incorporate an emotion estimation function and analyze a user's emotions. For example, if a user says, "I'm sad today," the robot recognizes the emotion and generates a response such as, "What happened? Wanna tell me?" The robot can also use the emotion estimation function to generate a response according to the user's emotions. For example, if a user says, "I have some good news," the robot can generate a response such as, "That's great! What's the news?" The robot can also analyze a user's emotions in real time and generate a response according to the emotion. For example, if a user says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Shall I teach you how to relax?" This allows the robot to generate a response according to the user's emotions, improving customer satisfaction.
[0079] The robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes learning support functions for educational settings. Furthermore, the robot can perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes customer support functions for customer service work. Furthermore, the robot can enhance its marketing strategy and perform customized demonstrations tailored to its target customer base. For example, the robot can perform a demonstration that emphasizes communication support functions for use in the home. In this way, by performing customized demonstrations tailored to its target customer base, it is possible to increase the marketing effect.
[0080] Robots can approach more customers by making their pricing flexible and introducing subscription or leasing models. For example, Robots can make their pricing flexible and introduce a subscription model. For example, they can offer a plan that allows customers to use the robot for a monthly fee. Robots can also introduce a leasing model that allows customers to use the robot for a certain period of time. For example, they can offer a one-year lease contract with an option to purchase thereafter. Robots can also make their pricing flexible and introduce subscription or leasing models. For example, they can offer a monthly usage plan for customers who want to use the robot for a short period of time. By making their pricing flexible and introducing subscription or leasing models, they can approach more customers.
[0081] Robots can be deployed in different markets to acquire new customers in the education and entertainment markets. For example, robots can be deployed in the education market to acquire new customers. For example, their learning support functions can be enhanced to promote their use in schools and cram schools. Robots can also be deployed in the entertainment market to acquire new customers. For example, their performance functions at events and shows can be enhanced. Robots can also be deployed in different markets to acquire new customers. For example, their patient support functions in the medical market can be enhanced to promote their use in hospitals and clinics. In this way, new customers can be acquired by deploying robots in different markets.
[0082] A robot can be developed in collaboration with a partner company and conduct a joint marketing campaign to acquire new customers. For example, a robot may develop a robot in collaboration with a partner company and conduct a joint marketing campaign. For example, a robot may partner with an educational institution to jointly develop a learning assistance robot. A robot may also conduct a joint marketing campaign to acquire new customers. For example, a robot may partner with an entertainment company to jointly promote a robot performance at an event. A robot may also develop a robot in collaboration with a partner company and conduct a marketing campaign. For example, a robot may partner with a medical institution to jointly develop a patient assistance robot and promote its use in hospitals. In this way, a robot can acquire new customers by developing a robot in collaboration with a partner company and conducting a joint marketing campaign.
[0083] A robot can provide emotional support by adding an emotion estimation function and generating a response based on the user's emotions. For example, a robot can add an emotion estimation function and analyze the user's emotions. For example, if a user says, "I'm sad today," the robot recognizes the emotion and generates a response such as, "What happened? Wanna talk?" The robot can also use the emotion estimation function to generate a response based on the user's emotions. For example, if a user says, "I have some good news," the robot can generate a response such as, "That's great! What's the news?" The robot can also analyze the user's emotions in real time and generate a response based on the emotion. For example, if a user says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Shall I teach you how to relax?" This allows the robot to provide emotional support by generating a response based on the user's emotions.
[0084] A robot can estimate a customer's emotions during customer service and generate a response based on the emotion, thereby improving customer satisfaction. For example, when a customer says, "I'm tired today," a robot can estimate a customer's emotions and generate a response based on the emotion. For example, if a customer says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Are you looking for a product that will help you relax?" The robot can also analyze a customer's emotions in real time and generate a response based on the emotion. For example, if a customer says, "I have some great news," the robot can generate a response such as, "That's great! What's the news?" The robot can also use its emotion estimation function to generate a response based on the customer's emotions, improving customer satisfaction. For example, if a customer says, "I'm sad today," the robot can generate a response such as, "What happened? Wanna talk?" This allows the robot to estimate a customer's emotions during customer service and generate a response based on the emotion, thereby improving customer satisfaction.
[0085] A robot can analyze a student's learning progress in an educational setting and provide individually customized learning support. For example, in an educational setting, a robot can analyze a student's learning progress and provide individually customized learning support. For example, if a student says, "I'm not good at math," the robot analyzes that progress and suggests an appropriate learning plan. The robot can also analyze a student's learning progress in real time and provide individually customized learning support. For example, if a student says, "English grammar is difficult," the robot analyzes that progress and provides grammar practice questions. The robot can also analyze a student's learning progress and provide individually customized learning support. For example, if a student says, "My history test is coming up," the robot analyzes that progress and suggests a study plan for test preparation. In this way, by analyzing a student's learning progress in an educational setting and providing individually customized learning support, it is possible to improve learning effectiveness.
[0086] A robot can manage family schedules within a home and provide reminders at appropriate times. For example, a robot can manage family schedules within a home and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "You have a doctor's appointment tomorrow." A robot can also manage family schedules in real time and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "We have plans for a family trip this weekend." A robot can also manage schedules and provide reminders at appropriate times. For example, a robot can provide a reminder such as, "Your child has a school event today." In this way, by managing family schedules within a home and providing reminders at appropriate times, communication within the home can be facilitated.
[0087] A robot can analyze a customer's purchase history during customer service and make personalized product suggestions. For example, a robot can analyze a customer's purchase history during customer service and make personalized product suggestions. For example, a robot can suggest, "I also recommend this product," based on products the customer has purchased in the past. A robot can also analyze a customer's purchase history in real time and make personalized product suggestions. For example, if a customer says, "The shampoo I bought last time was great," the robot can suggest, "I also recommend conditioner from the same brand." A robot can also analyze a customer's purchase history and make personalized product suggestions. For example, if a customer says, "The wine I recently purchased was delicious," the robot can suggest, "I also recommend other wines from the same winery." In this way, customer satisfaction can be improved by analyzing a customer's purchase history during customer service and making personalized product suggestions.
[0088] Robots can converse with multiple students simultaneously in educational settings and promote group discussions. For example, a robot can speak to multiple students at the same time, asking, "What do you all think about this topic?". A robot can also converse with multiple students at the same time and promote group discussions. For example, a robot can ask, "What do you think about this problem?", encouraging students to exchange opinions among themselves. A robot can also converse with multiple students to promote group discussions. For example, a robot can say, "Let's discuss this topic," encouraging students to exchange opinions among themselves. In this way, a robot can converse with multiple students simultaneously in educational settings and promote group discussions, thereby stimulating the exchange of opinions among students.
[0089] A robot can estimate the emotions of family members at home and generate responses based on those emotions, thereby facilitating communication within the home. For example, a robot can estimate the emotions of family members at home and generate responses based on those emotions. For example, if a family member says, "I'm tired today," the robot can generate a response such as, "Thank you for your hard work. Shall I teach you how to relax?" The robot can also analyze the emotions of family members in real time and generate responses based on those emotions. For example, if a family member says, "I have some good news," the robot can generate a response such as, "That's great! What's the news?" The robot can also use its emotion estimation function to generate responses based on the emotions of family members, thereby facilitating communication within the home. For example, if a family member says, "I'm sad today," the robot can generate a response such as, "What happened? Want to talk about it?" This allows the robot to estimate the emotions of family members at home and generate responses based on those emotions, thereby facilitating communication within the home.
[0090] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0091] The robot system may further include a health management unit that monitors the user's health condition. For example, the health management unit may measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. The health management unit may also record the user's amount of exercise and generate a message encouraging exercise if it detects a lack of exercise. Furthermore, the health management unit may record the user's diet and provide advice on nutritional balance. This allows the robot system to support the user's health management and contribute to maintaining good health.
[0092] The robot system may further include a music providing unit that estimates the user's emotions and selects music based on the estimated emotions. For example, if the user says, "I'm tired today," relaxing music may be provided. If the user says, "I have some good news," upbeat music may be provided. Furthermore, if the user says, "I'm sad today," music that soothes the mood may be provided. In this way, the robot system can provide emotional support by providing music that corresponds to the user's emotions.
[0093] The robot system can further include an information providing unit that provides customized information based on the user's hobbies and interests. For example, if the user says, "I like traveling," the latest travel information can be provided. If the user says, "I like cooking," new recipes can be suggested. If the user says, "I like sports," the latest sports news can be provided. This allows the robot system to provide more personalized services by providing information based on the user's hobbies and interests.
[0094] The robot system can further include a relaxation suggestion unit that estimates the user's emotions and suggests relaxation methods based on the estimated emotions. For example, if the user says, "I feel tired today," the robot system can suggest deep breathing or stretching methods. If the user says, "I feel stressed today," the robot system can suggest meditation or yoga methods. If the user says, "I want to relax today," the robot system can suggest aromatherapy or music therapy. In this way, the robot system can contribute to stress reduction by suggesting relaxation methods according to the user's emotions.
[0095] The robot system can further include a learning support unit that analyzes the user's learning progress and provides an individually customized learning plan. For example, if the user says, "I'm not good at math," the robot system can analyze the user's learning progress and suggest an appropriate learning plan. If the user says, "English grammar is difficult," the robot system can provide grammar practice questions. If the user says, "My history test is coming up," the robot system can suggest a study plan for test preparation. In this way, the robot system can analyze the user's learning progress and provide individually customized learning support, thereby improving learning effectiveness.
[0096] The robot system may further include an entertainment provider that estimates the user's emotions and provides entertainment based on the estimated emotions. For example, if the user says, "I'm tired today," the robot system may suggest relaxing movies or TV dramas. If the user says, "I'm in a good mood today," the robot system may suggest comedy movies or variety shows. If the user says, "I'm bored today," the robot system may suggest interesting games or activities. In this way, the robot system can improve the user's mood by providing entertainment that matches the user's emotions.
[0097] The robot system can further include a purchasing support unit that analyzes the user's purchasing history and makes personalized product suggestions. For example, the robot system can suggest "I also recommend this product" based on products the user has purchased in the past. If the user says, "The shampoo I bought last time was good," the robot system can suggest, "I also recommend the conditioner from the same brand." If the user says, "The wine I recently purchased was delicious," the robot system can suggest, "I also recommend other wines from the same winery." In this way, the robot system can improve customer satisfaction by analyzing the user's purchasing history and making personalized product suggestions.
[0098] The robot system may further include a message generation unit that estimates the user's emotions and generates a customized message based on the estimated emotions. For example, if the user says, "I'm sad today," the robot system may generate a message such as, "What happened? Wanna talk?". If the user says, "I have some good news," the robot system may generate a message such as, "That's great! What's the news?". If the user says, "I'm tired today," the robot system may generate a message such as, "Thank you for your hard work. Shall I teach you how to relax?". This allows the robot system to provide emotional support by generating customized messages based on the user's emotions.
[0099] The robot system can further include a schedule management unit that manages the user's schedule and provides reminders at appropriate times. For example, it can provide a reminder such as "You have a doctor's appointment tomorrow." It can also provide a reminder such as "You have a family trip planned this weekend." It can also provide a reminder such as "Your child has a school event today." In this way, the robot system can support the user's daily life by managing the user's schedule and providing reminders at appropriate times.
[0100] The robot system can further include an exercise support unit that estimates the user's emotions and provides a customized exercise plan based on the estimated emotions. For example, if the user says, "I feel tired today," the robot system can suggest light stretching or relaxation exercises. If the user says, "I feel energetic today," the robot system can suggest more strenuous exercises. Furthermore, if the user says, "I feel stressed today," the robot system can suggest exercises to relieve stress. In this way, the robot system can contribute to maintaining the user's health by providing an exercise plan that matches the user's emotions.
[0101] The processing flow of the second embodiment will be briefly explained below.
[0102] Step 1: The speech recognition unit recognizes the user's speech. For example, the speech recognition unit converts the user's speech into text data using deep learning-based speech recognition technology. The speech recognition unit can also analyze the content of the speech using an HMM (hidden Markov model). Furthermore, the speech recognition unit can perform phonemic analysis to analyze the content of the speech in detail. Step 2: The generator generates an appropriate response based on the utterance recognized by the speech recognizer. For example, the generator uses a text generation AI (e.g., GPT-3) to generate a natural response based on the user's utterance. The generator can also generate a response based on a predefined template. Furthermore, the generator can also generate a response based on the user's intention. Step 3: The response unit provides the response generated by the generation unit to the user. For example, the response unit may provide the response audibly using a speaker. The response unit may also display the response in text using a display. Furthermore, the response unit may also send a notification to the user's device.
[0103] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0104] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0105] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0106] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0107] 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0108] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0109] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0110] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0111] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0112] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0113] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0114] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0115] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0116] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0117] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0118] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0119] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0120] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0121] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0122] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0123] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0124] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0125] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0126] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0127] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0128] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0129] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0130] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0131] In the headset type terminal 314, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0132] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0133] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0134] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0135] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0136] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0137] 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0138] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0139] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0140] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0141] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0142] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0143] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0144] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0145] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0146] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0147] In the robot 414, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The robot 414 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0148] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0149] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0150] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0151] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0152] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0153] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[0154] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[0155] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[0156] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.
[0157] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[0158] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[0159] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.
[0160] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[0161] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0162] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[0163] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[0164] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[0165] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[0166] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[0167] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.
[0168] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0169] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]
[0170] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot
Claims
1. a speech recognition unit that recognizes a user's speech using speech recognition technology; a generation unit that generates an appropriate response based on the utterance content recognized by the speech recognition unit; a response unit that provides the response generated by the generation unit to the user. A system characterized by:
2. The voice recognition unit Analyzes user emotions in real time and generates responses according to those emotions 2. The system of claim 1.
3. The voice recognition unit Recognize multiple languages and dialects simultaneously and generate responses based on the user's language selection 2. The system of claim 1.
4. The voice recognition unit Analyzes the user's speech rate and tone, determines the urgency and importance of the user, and adjusts the response accordingly 2. The system of claim 1.
5. The voice recognition unit Translates user speech in real time to support communication between users who speak different languages 2. The system of claim 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A