system
The system effectively processes diverse media data to provide personalized and multilingual answers, enhancing user interaction by integrating natural language processing and emotional recognition for continuous improvement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
Smart Images

Figure 2026068424000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, a vast amount of knowledge is generated daily using various information media, but a system that efficiently organizes and utilizes this knowledge and accurately answers specific questions of users has not fully developed. Also, with the complexity of information increasing, the need for multilingual support and personalized information provision is growing, and a technical solution to address this is required.
Means for Solving the Problems
[0005] This invention provides a system that stores multiple types of media data, such as text, images, and audio, input via an information processing device into a database, and generates knowledge data corresponding to human questions by processing that data. It also provides a system that generates appropriate answers to questions using the generated knowledge data. Furthermore, it can acquire user feedback and use it to improve the accuracy of the knowledge data. In addition, by incorporating natural language processing technology and multilingual support capabilities, it can respond to a wide range of user questions and provide accurate and timely information.
[0006] An "information processing device" is a general term for electronic devices that have the functions of inputting, storing, processing, and outputting data.
[0007] "Media data" refers to information expressed in different formats, such as text, images, and audio.
[0008] A "database" is an information system that stores data in a structured format, allowing for efficient management and retrieval.
[0009] "Knowledge data" refers to a collection of information generated for a specific purpose that can be used to solve problems or answer questions.
[0010] "Natural language processing technology" is a collection of computational methods that enable computers to understand and process human language.
[0011] "Multilingual support" refers to the ability to handle input and output in multiple different languages, enabling appropriate dialogue.
[0012] "Feedback" refers to opinions and evaluations provided by users, which are used to improve and adjust the system. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] The system of the present invention effectively processes diverse forms of information input by a user and generates highly accurate answers to the user's questions. This system is implemented with a configuration including an information processing device, a database, an AI model, and a user interface.
[0035] First, the user provides information through their device. This information is acquired as text input, images, or audio data. This acquired information is formatted appropriately by the device and sent to the server. The server stores the received information in a database and makes it available for subsequent processing.
[0036] Next, the server analyzes the data and executes a process to generate knowledge data using an AI model. Natural language processing techniques are applied here to understand the meaning from the data and extract the information necessary to answer the user's questions. This AI model improves the accuracy of its knowledge through regular learning. Furthermore, the model has multilingual capabilities, ensuring smooth operation in different language environments.
[0037] When a user asks a question from their device, the device sends it to the server. The server interprets the question, retrieves relevant knowledge data from its database, and generates an appropriate answer. This answer is provided in a personalized format according to the user's needs.
[0038] The generated responses are sent to the device and displayed through the user interface. Furthermore, it is possible to have the responses read aloud using the voice output function, providing not only visual but also auditory information.
[0039] Finally, users can provide feedback on the information provided. The device sends this feedback to the server, which analyzes it and uses it to further improve the accuracy of the AI model. User feedback allows the system to continuously improve, enabling more natural and human-like conversations.
[0040] Thus, the system of the present invention provides quick and accurate answers to user questions by going through a variety of processes starting with information provided by the user. All elements supporting this process are closely coordinated to realize advanced information processing and dialogue functions.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The user inputs information through their device. This information is captured as text, audio, or images. For example, the user might input text about their travels or upload images of tourist destinations.
[0044] Step 2:
[0045] The device converts the acquired information into an appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from images using image analysis technology.
[0046] Step 3:
[0047] The terminal sends formatted data to the server. This data is transmitted according to the communication protocol because it will be used in subsequent processing.
[0048] Step 4:
[0049] The server stores the received data in a database. The data is structured so that it can be efficiently searched and used later.
[0050] Step 5:
[0051] The server analyzes the data and generates knowledge data using an AI model. Natural language processing is applied here to analyze the meaning of the data.
[0052] Step 6:
[0053] The user enters a question via their device. The question is specific and constitutes a request for information from the system.
[0054] Step 7:
[0055] The terminal sends the question received from the user to the server. The question is formatted and prepared for processing on the server side.
[0056] Step 8:
[0057] The server interprets the question and retrieves relevant knowledge data from the database. This forms the basis for the answer to the question.
[0058] Step 9:
[0059] The server uses an AI model to generate appropriate answers. The generated answers are presented in a format that is easy for the user to understand.
[0060] Step 10:
[0061] The server generates the response and sends it to the terminal. The response data is delivered to the terminal via communication.
[0062] Step 11:
[0063] The device displays the answer through the user interface. The answer can be presented in both text and audio formats.
[0064] Step 12:
[0065] Users provide feedback on the answers. This feedback is valuable information for improving the system.
[0066] Step 13:
[0067] The device sends user feedback to the server. This feedback data is managed as material for analysis.
[0068] Step 14:
[0069] The server analyzes the feedback and incorporates it into the AI model to improve accuracy. Based on the feedback, the system is improved.
[0070] (Example 1)
[0071] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0072] There was a need for an information system that could accept data in various formats from users, process it effectively, and provide precise answers. Furthermore, challenges existed in handling questions in different languages and improving system accuracy through user feedback. This project aims to solve these problems.
[0073] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0074] In this invention, the server includes means for converting data in multiple formats input via information processing means into a unified format and storing it in a database; means for analyzing the data and generating knowledge to respond to human questions using an artificial intelligence model; and means for deriving the optimal answer to the received question based on the knowledge generated by the artificial intelligence model. This enables multilingual support and highly accurate information provision that reflects feedback.
[0075] "Information processing means" refers to a device or process that has the function of receiving input data, converting it into a standardized format, and storing it in a database.
[0076] A "database" is a collection of information for efficient storage, management, and retrieval, and a means of quickly obtaining necessary data.
[0077] An "artificial intelligence model" is an algorithm or system used to analyze data and generate knowledge in response to human questions based on the results.
[0078] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language, and is used for questioning and data analysis.
[0079] "Feedback" refers to evaluations and opinions provided by users regarding the system's response, and is information used to improve and adjust the system.
[0080] "Multilingual support" refers to a system's ability to understand multiple different languages and provide appropriate responses in each language.
[0081] The system according to the present invention is an integrated information processing system for improving flexibility and accuracy in information handling. This system efficiently executes a series of processes, from information input to response generation and improvement using feedback.
[0082] First, the user inputs information through the terminal. The input information can take various forms, including text, images, and audio data. For example, audio data is converted into text data using speech recognition technology. Standard speech recognition software is used for this process. The formatted data is then sent to the server via a secure communication protocol.
[0083] The server stores the received data in a database. This database is designed to efficiently manage large amounts of information and retrieve it quickly when needed. Specifically, a database management system is commonly used.
[0084] Next, the server analyzes the data using an artificial intelligence model. This analysis process utilizes natural language processing techniques to understand the meaning of the information and generate knowledge in response to the user's questions. For example, open-source natural language processing frameworks and cloud-based AI services are commonly used as generative AI models.
[0085] When a user asks a question, the device sends the question to the server, which interprets the content of the question. Based on the generated knowledge, the server derives the best answer. This answer is personalized based on the user's past interactions and profile information. For example, in response to a prompt such as "I'm thinking of traveling, can you recommend some tourist spots?", the system will suggest tourist spots considering the user's hobbies and past visits.
[0086] Finally, users can provide feedback on their answers. This feedback is sent from the device to the server and used to train the AI model. This allows the system to continuously improve its accuracy and enhance the user experience.
[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0088] Step 1:
[0089] Users input information via a terminal. This information is collected in various forms, including text, images, and audio data. The terminal receives the input data and converts it into the appropriate format. Specifically, if audio data is input, the terminal uses speech recognition technology to convert it into text data. This conversion generates a unified data format that can be transmitted to the server.
[0090] Step 2:
[0091] The terminal sends formatted data to the server. The server stores the received data in a database. The database plays a role in enabling centralized data management and efficient searching. Once the input data is stored, it becomes available for use in subsequent analysis processes.
[0092] Step 3:
[0093] The server retrieves data from the database and performs analysis using a generative AI model. Here, natural language processing is used to interpret the meaning of the data and generate relevant knowledge. This process inputs data as prompts into the AI model, and the analysis results are obtained as output from the model. These results can be used to answer user questions.
[0094] Step 4:
[0095] When a user enters a question, the device sends this question to the server. The server analyzes the question and extracts relevant knowledge from its database. Using a generative AI model, it generates the most appropriate answer to the question and provides it to the user in a personalized manner. The output is a prompt-based answer.
[0096] Step 5:
[0097] The generated responses are sent to the device and displayed through the user interface. In addition, for users who have difficulty visualizing, the device provides the responses audibly using its voice output function. In this way, users can receive information through both visual and auditory means.
[0098] Step 6:
[0099] Users can input feedback on the answers they receive. The device sends the feedback to the server. The server analyzes the feedback and incorporates it into the AI model's learning process to improve the system's accuracy and responsiveness. This process continuously improves the user experience.
[0100] (Application Example 1)
[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0102] In today's information-saturated society, users demand quick and accurate information suggestions, but existing systems have limitations in the accuracy of personalized content recommendations. Therefore, there is a need for technology that effectively delivers appropriate content tailored to users' interests and preferences.
[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0104] In this invention, the server includes means for storing media data input via an information processing device in a data storage means, means for using technology to process the media data and generate knowledge data, and means for suggesting personalized content based on the user's interests through a user interface. This enables the suggestion of personalized content in response to the user's questions and interests.
[0105] An "information processing device" is an electronic device used to collect and analyze data and generate output according to a specific purpose.
[0106] "Media data" refers to information provided by users in various formats, such as text, audio, and images.
[0107] A "data storage means" is a mechanism for organizing and storing collected data.
[0108] "Knowledge data" refers to useful information in response to user questions, obtained by analyzing and processing collected data.
[0109] A "user interface" is a means of interaction that allows a user to exchange information with a system in a two-way manner.
[0110] "Personalization" is the process of adapting information and content to each user's interests and preferences.
[0111] A "proposal" is an action that presents users with options or information.
[0112] The system that realizes this application consists of an information processing device, a database, an AI model, and a user interface.
[0113] The server first collects media data such as text, images, and audio input from users via smartphones or smart glasses, formats it appropriately, and stores it in a database. This data is efficiently managed by the data storage system.
[0114] Next, the server uses Python to perform natural language processing and extract relevant information from the media data. This process utilizes AI models such as the BERT model from the Transformers library to achieve advanced data analysis. After the analysis, knowledge data is generated, which forms the basis for deriving information based on the user's questions and interests.
[0115] When a user asks a specific question through the user interface, the server quickly processes the question, retrieves relevant knowledge data from the database, and provides personalized content. This structure ensures that content most closely matches the user's interests is suggested.
[0116] Furthermore, the server receives user feedback and improves the performance of the AI model. This feedback loop is a crucial element in ensuring continuous improvement of the system.
[0117] As a concrete example, if a user asks a question about movie recommendations, the system will suggest the most suitable movie based on their past viewing history. An example of a prompt to the generating AI model might be, "What movies do you recommend? I used to like action movies." Based on this prompt, the AI model will extract and suggest the movie that best suits the user's preferences.
[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0119] Step 1:
[0120] Users input questions or data in text, image, or audio format through their device. This input data is formatted on the device and sent to the server over the network. The input data arrives at the server in its original format.
[0121] Step 2:
[0122] The server parses the received data and first stores it in a database. The data is efficiently stored using a database system (e.g., MySQL®). The stored data is structured with the assumption that it will be used for subsequent natural language processing.
[0123] Step 3:
[0124] The server executes Python scripts and analyzes the stored data using natural language processing. It utilizes tools such as the BERT model from the Transformers library to extract meaning and relationships from the data. Specifically, it converts the input data into feature vectors to understand context and related information. As a result, knowledge data is generated, which forms the basis for answering user questions.
[0125] Step 4:
[0126] When a user enters a specific question from their device, it is sent to the server. The server interprets this question and extracts the appropriate answer by referring to knowledge data generated from a database. By using prompts, the AI model quickly finds relevant information and suggests content based on the user's interests.
[0127] Step 5:
[0128] The device receives responses from the server and displays them to the user through the user interface. This display includes details of the suggested content and information. The device also provides the ability to read the responses aloud using the Google® Cloud Text-to-Speech API.
[0129] Step 6:
[0130] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and used to further train the AI model. The server analyzes this feedback to improve the AI model and increase the accuracy of its answers to future questions.
[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0132] The system of the present invention not only processes diverse forms of information input from users and generates highly accurate answers to questions, but also has the function of recognizing the user's emotional state and optimizing the response. This system is implemented with a configuration including an information processing device, a database, an AI model, an emotion engine, and a user interface.
[0133] First, the user inputs information via the device. This information can be obtained as text, images, or audio. For example, if the user inputs a voice message, the device uses speech recognition technology to convert it into text.
[0134] Next, the device sends the converted data to the server. The server stores this data in a database and uses an AI model to generate knowledge data. This process employs natural language processing techniques to analyze the meaning of the information and prepare to generate answers to questions.
[0135] When a user enters a question, the terminal transmits it to the server, which then compares it with knowledge data to interpret the question. Based on the user's question and related information, the server generates an appropriate answer.
[0136] The key here is the role of the emotion engine. The emotion engine analyzes the user's emotional state from their input. For example, it can analyze the content of the text and the tone of their voice to determine whether the user is angry or sad. Based on this emotional information, the server generates personalized responses that are appropriate to the user's state. For example, if the user is feeling stressed, the response will include content that promotes relaxation.
[0137] The generated response is sent to the device and displayed through the user interface. Furthermore, the device can use emotion recognition to provide text-to-speech that matches the tone and nuances the user is likely to prefer.
[0138] Finally, users provide feedback on the information provided. This feedback is crucial for the continuous improvement of the system, and the server uses it to adjust the AI model and emotion engine.
[0139] Thus, this system, which incorporates an emotion engine, provides not only the functionality of a conventional question-answering system but also emotion recognition and response optimization to improve the user experience. This enables more natural and adaptable dialogue.
[0140] The following describes the processing flow.
[0141] Step 1:
[0142] The user inputs information through their device. This information is obtained in text, audio, or image format. For example, the user might input a thank-you message via voice.
[0143] Step 2:
[0144] The device converts the information into the appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from image data as needed.
[0145] Step 3:
[0146] The terminal sends formatted data to the server. This data is sent according to a predetermined communication protocol because it will be used for subsequent processing.
[0147] Step 4:
[0148] The server stores the received data in a database. The data is structured for easy searching and use and stored in the database.
[0149] Step 5:
[0150] The server analyzes the data and generates knowledge data using an AI model. Natural language processing techniques are applied here to analyze the meaning and context of the data.
[0151] Step 6:
[0152] The server uses an emotion engine to analyze the user's emotional state. For example, it analyzes the tone of the input text or voice to identify the user's emotions.
[0153] Step 7:
[0154] The user enters a question through their device. The question seeks specific information or an answer.
[0155] Step 8:
[0156] The terminal sends the question received from the user to the server. The question data is sent in a predetermined format, preparing it for processing on the server side.
[0157] Step 9:
[0158] The server interprets the question and retrieves relevant knowledge data from the database. The server analyzes the context of the question and extracts relevant information.
[0159] Step 10:
[0160] The server uses an AI model to generate appropriate responses. Here, it considers the analysis results of the emotion engine to create personalized responses tailored to the user's emotional state.
[0161] Step 11:
[0162] The server generates the response and sends it to the device. The response data is sent in a format that the device can use.
[0163] Step 12:
[0164] The device displays the answer through its user interface. The answer is displayed on the screen as text and may also be played back using a text-to-speech function.
[0165] Step 13:
[0166] Users provide feedback on the information they receive. Users rate whether the answers were satisfactory.
[0167] Step 14:
[0168] The device sends user feedback to the server. This feedback information is managed for future system improvements.
[0169] Step 15:
[0170] The server analyzes the feedback and incorporates it into the AI model and emotion engine to improve accuracy. This allows the system to continuously improve the user experience.
[0171] (Example 2)
[0172] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0173] Traditional information processing systems focus on providing accurate answers to user questions, but they fail to adequately optimize responses based on the user's emotional state. As a result, the user experience is uniform, making it difficult to provide individualized interactions. Furthermore, there are challenges in handling different languages, highlighting the need for high-quality question-and-answer systems in multiple languages.
[0174] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0175] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a data storage area, means for generating knowledge data to respond to human questions using natural language processing technology, and means for analyzing the user's emotional state using emotion recognition technology and optimizing responses according to that state. This enables personalized dialogue that takes the user's emotions into consideration, and further enables the provision of high-quality responses that support multiple languages.
[0176] An "information processing device" is a general term for hardware and software that receives data in various formats from users and performs the necessary processing.
[0177] "Media data" refers to various information formats that users can input, such as text, images, and audio.
[0178] "Data storage area" refers to a storage area for saving received data, and includes databases and other storage methods.
[0179] "Natural language processing technology" is a technology that enables computers to understand, interpret, and respond to human language, and is used to analyze user questions.
[0180] "Knowledge data" refers to a dataset that systematically organizes the information necessary to generate answers to user questions.
[0181] "Emotion recognition technology" is a technology that analyzes the emotional state from the user's input and identifies a specific emotion.
[0182] "Multilingual support" refers to a system's ability to understand inquiries in different languages and generate responses in the appropriate language accordingly.
[0183] The system of the present invention includes an information processing device, a database, a generative AI model, an emotion recognition engine, and a user interface. These components process diverse forms of information input from the user, generate highly accurate answers to questions, and recognize the user's emotional state to optimize responses.
[0184] The user provides input data via the device in the form of text, images, or audio. In the case of audio input, the device uses speech recognition software to convert the audio data into text data. A common cloud-based speech recognition service is used for this conversion. The device then sends the converted data to the server.
[0185] The server stores the received data in data storage and generates knowledge data from the information using a generative AI model. This process employs natural language processing techniques to analyze the user's questions and prepare the optimal answers. An advanced AI framework is used to generate the knowledge data.
[0186] Furthermore, the server uses an emotion recognition engine to analyze the user's emotional state from their input. This emotion analysis is based on the user's text content and tone of voice. For example, if a user inputs "I've been feeling tired lately," the emotion recognition engine detects fatigue and generates a personalized response accordingly.
[0187] The generated responses are sent to the device and displayed or read aloud through the user interface. In the case of read-aloud responses, the tone is tailored to the user's emotional state.
[0188] For example, if a user inputs "I've been feeling stressed lately," the emotion recognition engine detects the stress, the server generates a response such as "Why don't you try listening to some relaxing music?", and the device displays or announces this message.
[0189] An example of a prompt in this invention is, "If the user is feeling stressed, how do you generate a response that promotes relaxation?"
[0190] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0191] Step 1:
[0192] The user inputs information using a device. The input format can be text, images, or audio. For example, by opening a smartphone application and speaking "What's the weather like today?", audio data is acquired. This input data is then converted into text data using speech recognition technology. Specifically, speech recognition software analyzes the audio data and outputs the result as the text "What's the weather like today?".
[0193] Step 2:
[0194] The terminal sends the converted text data to the server. This transmission uses an HTTP request with an internet connection. The server checks the received text data and saves it to its data storage area. Specifically, the server accesses the database system and saves the text "What's the weather like today?".
[0195] Step 3:
[0196] The server inputs the stored text data into a generative AI model, which uses natural language processing techniques to analyze the meaning of the question. In this process, for example, the generative AI model analyzes the text, identifies that the user is asking a question about "weather," and organizes the relevant information as knowledge data. The output at this stage is the analyzed intent of the question and related information.
[0197] Step 4:
[0198] The server uses an emotion recognition engine to analyze the user's emotional state from their input. It identifies the emotional state based on the user's text content and tone. For example, if the input includes the word "tired," it detects stress and fatigue. This emotional information is then incorporated into subsequent response generation.
[0199] Step 5:
[0200] The server integrates knowledge data and sentiment information using a generative AI model to generate the optimal response. Specifically, based on information such as "the user is asking about the weather and is feeling stressed," it constructs a response like "It's sunny today. How about going for a walk?" This response is then sent to the device.
[0201] Step 6:
[0202] The terminal displays the response received from the server in its user interface. It also uses speech synthesis technology to read the response aloud as needed. Taking emotional elements into consideration, it provides guidance in a gentle tone, such as, "It's sunny today. Why not go for a walk?"
[0203] Step 7:
[0204] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and stored to improve the system's accuracy. This feedback is used to fine-tune the AI model and emotion recognition engine, helping to improve the system.
[0205] (Application Example 2)
[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0207] When users inquire about information, they are required to receive appropriate and prompt answers to their questions, and furthermore, to have those answers optimized to suit their emotional state. In particular, in the context of electronic payments, users often experience anxiety and stress, and it is necessary for the system to respond appropriately to these feelings.
[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0209] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a database, means for processing the media data and generating knowledge data to respond to human questions, and means for analyzing the user's emotional state and optimizing the response. This enables the provision of a fast, appropriate, and emotionally-driven response to the user.
[0210] An "information processing device" is a device that receives data input from a user and processes it in a predetermined format.
[0211] "Media data" refers to information expressed in multiple formats, such as text, audio, and images.
[0212] A "database" is a digital storage system used to store and manage information efficiently and systematically.
[0213] "Knowledge data" refers to information generated by processing raw data in order to produce appropriate answers to user input questions.
[0214] "Means for analyzing emotional states" refers to functions that recognize and evaluate emotions based on user input information and adjust responses accordingly.
[0215] "Means of optimizing responses" refer to methods for personalizing the content of responses according to the user's emotional state and needs, and providing them in the most optimal way.
[0216] "Feedback" refers to the collection of evaluations and opinions from users that are used to improve future processes.
[0217] The system implementing this invention comprises an information processing device, a database, multiple AI models, and a user interface. The server processes data input from the user in multiple formats, such as voice, text, and images. This includes transcribing voice data into text using a speech recognition library (e.g., speech recognition technology), analyzing text data using natural language processing technology, and storing it in the database.
[0218] Based on the information stored in the database, the server uses an AI model (e.g., a generative AI model) to generate answers to questions. During this process, an emotion analysis engine (e.g., emotion recognition technology) is used to analyze the user's emotional state and adjust the content and tone of the answers to match the user's emotions.
[0219] On the device, the generated response is displayed through the user interface and can also be read aloud. This allows users to obtain information through both sight and sound. User feedback is collected and recorded in a database to help improve the system. Based on this feedback, the accuracy of the AI model and sentiment analysis engine is improved.
[0220] As a concrete example, in an electronic payment scenario, if a user asks, "Is this payment secure?", the system would return the message, "You can use it with confidence. This transaction is highly secure." An example of a prompt used here would be, "Generate an appropriate response to the user's question and optimize the tone based on the emotion detected by emotion recognition technology."
[0221] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0222] Step 1:
[0223] The terminal receives the voice data entered by the user. This voice data is then converted into text data using speech recognition technology. This converted text data becomes the input for the next step.
[0224] Step 2:
[0225] The server receives text data sent from the terminal. Natural language processing techniques are used to analyze the meaning of the question. Based on the text data, relevant information is retrieved from the database to generate knowledge data. This knowledge data serves as input for the next step.
[0226] Step 3:
[0227] The server uses sentiment analysis technology to analyze the user's emotional state. This analysis includes calculations to identify emotions from the text content. Based on the analysis results, it generates personalized responses tailored to the user's emotional state. This generated response data becomes the input for the next step.
[0228] Step 4:
[0229] The server uses a generative AI model to concretize appropriate responses for the user based on knowledge data and sentiment analysis results. Here, prompts are used to instruct the AI model. Based on these instructions, the final response is output.
[0230] Step 5:
[0231] The device receives personalized responses sent from the server. These responses are displayed through the user interface and read aloud using speech synthesis technology. This process allows the user to obtain information both visually and aurally.
[0232] Step 6:
[0233] Users provide feedback on the answers provided by the system. The device collects this feedback and sends it to the server. This feedback data is used for the continuous improvement of the system and contributes to improving the accuracy of the generative AI model and sentiment analysis engine.
[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0237] [Second Embodiment]
[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0250] The system of the present invention effectively processes diverse forms of information input by a user and generates highly accurate answers to the user's questions. This system is implemented with a configuration including an information processing device, a database, an AI model, and a user interface.
[0251] First, the user provides information through their device. This information is acquired as text input, images, or audio data. This acquired information is formatted appropriately by the device and sent to the server. The server stores the received information in a database and makes it available for subsequent processing.
[0252] Next, the server analyzes the data and executes a process to generate knowledge data using an AI model. Natural language processing techniques are applied here to understand the meaning from the data and extract the information necessary to answer the user's questions. This AI model improves the accuracy of its knowledge through regular learning. Furthermore, the model has multilingual capabilities, ensuring smooth operation in different language environments.
[0253] When a user asks a question from their device, the device sends it to the server. The server interprets the question, retrieves relevant knowledge data from its database, and generates an appropriate answer. This answer is provided in a personalized format according to the user's needs.
[0254] The generated responses are sent to the device and displayed through the user interface. Furthermore, it is possible to have the responses read aloud using the voice output function, providing not only visual but also auditory information.
[0255] Finally, users can provide feedback on the information provided. The device sends this feedback to the server, which analyzes it and uses it to further improve the accuracy of the AI model. User feedback allows the system to continuously improve, enabling more natural and human-like conversations.
[0256] Thus, the system of the present invention provides quick and accurate answers to user questions by going through a variety of processes starting with information provided by the user. All elements supporting this process are closely coordinated to realize advanced information processing and dialogue functions.
[0257] The following describes the processing flow.
[0258] Step 1:
[0259] The user inputs information through their device. This information is captured as text, audio, or images. For example, the user might input text about their travels or upload images of tourist destinations.
[0260] Step 2:
[0261] The device converts the acquired information into an appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from images using image analysis technology.
[0262] Step 3:
[0263] The terminal sends formatted data to the server. This data is transmitted according to the communication protocol because it will be used in subsequent processing.
[0264] Step 4:
[0265] The server stores the received data in a database. The data is structured so that it can be efficiently searched and used later.
[0266] Step 5:
[0267] The server analyzes the data and generates knowledge data using an AI model. Natural language processing is applied here to analyze the meaning of the data.
[0268] Step 6:
[0269] The user enters a question via their device. The question is specific and constitutes a request for information from the system.
[0270] Step 7:
[0271] The terminal sends the question received from the user to the server. The question is formatted and prepared for processing on the server side.
[0272] Step 8:
[0273] The server interprets the content of the question and extracts relevant knowledge data from the database. This forms the basis for answering the question.
[0274] Step 9:
[0275] The server uses an AI model to generate an appropriate answer. The generated answer is created in a format that is easy for the user to understand.
[0276] Step 10:
[0277] The server sends the generated answer to the terminal. The answer data is delivered to the terminal side through communication.
[0278] Step 11:
[0279] The terminal displays the answer through the user interface. The answer can be presented both in text and voice.
[0280] Step 12:
[0281] The user inputs feedback on the answer. This feedback is valuable information for system improvement.
[0282] Step 13:
[0283] The terminal sends the user's feedback to the server. The feedback data is managed as material for analysis.
[0284] Step 14:
[0285] The server analyzes the feedback and reflects it in the AI model to improve accuracy. Based on the feedback, the system is improved.
[0286] (Example 1)
[0287] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0288] There was a need for an information system that could accept data in various formats from users, process it effectively, and provide precise answers. Furthermore, challenges existed in handling questions in different languages and improving system accuracy through user feedback. This project aims to solve these problems.
[0289] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0290] In this invention, the server includes means for converting data in multiple formats input via information processing means into a unified format and storing it in a database; means for analyzing the data and generating knowledge to respond to human questions using an artificial intelligence model; and means for deriving the optimal answer to the received question based on the knowledge generated by the artificial intelligence model. This enables multilingual support and highly accurate information provision that reflects feedback.
[0291] "Information processing means" refers to a device or process that has the function of receiving input data, converting it into a standardized format, and storing it in a database.
[0292] A "database" is a collection of information for efficient storage, management, and retrieval, and a means of quickly obtaining necessary data.
[0293] An "artificial intelligence model" is an algorithm or system used to analyze data and generate knowledge in response to human questions based on the results.
[0294] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language, and is used for questioning and data analysis.
[0295] "Feedback" refers to evaluations and opinions provided by users regarding the system's response, and is information used to improve and adjust the system.
[0296] "Multilingual support" refers to a system's ability to understand multiple different languages and provide appropriate responses in each language.
[0297] The system according to the present invention is an integrated information processing system for improving flexibility and accuracy in information handling. This system efficiently executes a series of processes, from information input to response generation and improvement using feedback.
[0298] First, the user inputs information through the terminal. The input information can take various forms, including text, images, and audio data. For example, audio data is converted into text data using speech recognition technology. Standard speech recognition software is used for this process. The formatted data is then sent to the server via a secure communication protocol.
[0299] The server stores the received data in a database. This database is designed to efficiently manage large amounts of information and retrieve it quickly when needed. Specifically, a database management system is commonly used.
[0300] Next, the server analyzes the data using an artificial intelligence model. This analysis process utilizes natural language processing techniques to understand the meaning of the information and generate knowledge in response to the user's questions. For example, open-source natural language processing frameworks and cloud-based AI services are commonly used as generative AI models.
[0301] When a user asks a question, the terminal sends the question to the server, and the server interprets the content of the question. Based on the generated knowledge, the server derives an optimal answer. This answer is personalized based on the user's past interactions and profile information. For example, for a prompt sentence like "I'm thinking about traveling, so please tell me some recommended tourist destinations", the system makes a recommendation for tourist destinations considering the user's hobbies and past visited places.
[0302] Finally, the user can provide feedback on the answer. This feedback is sent from the terminal to the server and used for the learning of the AI model. As a result, it becomes possible for the system to continuously improve the accuracy and enhance the user experience.
[0303] The flow of the specific process in Example 1 will be described using FIG. 11.
[0304] Step 1:
[0305] The user inputs information via the terminal. This information is collected in forms such as text, images, audio data, etc. The terminal receives the input data and converts it into an appropriate format. Specifically, when audio data is input, the terminal uses speech recognition technology to convert it into text data. This conversion generates a unified data format that can be transmitted to the server.
[0306] Step 2:
[0307] The terminal sends the formatted data to the server. The server stores the received data in the database. The database plays a role in enabling unified management and efficient search of data. By storing the input data, it becomes available for subsequent analysis processes.
[0308] Step 3:
[0309] The server retrieves data from the database and performs analysis using a generative AI model. Here, natural language processing is used to interpret the meaning of the data and generate relevant knowledge. This process inputs data as prompts into the AI model, and the analysis results are obtained as output from the model. These results can be used to answer user questions.
[0310] Step 4:
[0311] When a user enters a question, the device sends this question to the server. The server analyzes the question and extracts relevant knowledge from its database. Using a generative AI model, it generates the most appropriate answer to the question and provides it to the user in a personalized manner. The output is a prompt-based answer.
[0312] Step 5:
[0313] The generated responses are sent to the device and displayed through the user interface. In addition, for users who have difficulty visualizing, the device provides the responses audibly using its voice output function. In this way, users can receive information through both visual and auditory means.
[0314] Step 6:
[0315] Users can input feedback on the answers they receive. The device sends the feedback to the server. The server analyzes the feedback and incorporates it into the AI model's learning process to improve the system's accuracy and responsiveness. This process continuously improves the user experience.
[0316] (Application Example 1)
[0317] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0318] In today's information-saturated society, users demand quick and accurate information suggestions, but existing systems have limitations in the accuracy of personalized content recommendations. Therefore, there is a need for technology that effectively delivers appropriate content tailored to users' interests and preferences.
[0319] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0320] In this invention, the server includes means for storing media data input via an information processing device in a data storage means, means for using technology to process the media data and generate knowledge data, and means for suggesting personalized content based on the user's interests through a user interface. This enables the suggestion of personalized content in response to the user's questions and interests.
[0321] An "information processing device" is an electronic device used to collect and analyze data and generate output according to a specific purpose.
[0322] "Media data" refers to information provided by users in various formats, such as text, audio, and images.
[0323] A "data storage means" is a mechanism for organizing and storing collected data.
[0324] "Knowledge data" refers to useful information in response to user questions, obtained by analyzing and processing collected data.
[0325] A "user interface" is a means of interaction that allows a user to exchange information with a system in a two-way manner.
[0326] "Personalization" is the process of adapting information and content to each user's interests and preferences.
[0327] A "proposal" is an action that presents users with options or information.
[0328] The system that realizes this application consists of an information processing device, a database, an AI model, and a user interface.
[0329] The server first collects media data such as text, images, and audio input from users via smartphones or smart glasses, formats it appropriately, and stores it in a database. This data is efficiently managed by the data storage system.
[0330] Next, the server uses Python to perform natural language processing and extract relevant information from the media data. This process utilizes AI models such as the BERT model from the Transformers library to achieve advanced data analysis. After the analysis, knowledge data is generated, which forms the basis for deriving information based on the user's questions and interests.
[0331] When a user asks a specific question through the user interface, the server quickly processes the question, retrieves relevant knowledge data from the database, and provides personalized content. This structure ensures that content most closely matches the user's interests is suggested.
[0332] Furthermore, the server receives user feedback and improves the performance of the AI model. This feedback loop is a crucial element in ensuring continuous improvement of the system.
[0333] As a concrete example, if a user asks a question about movie recommendations, the system will suggest the most suitable movie based on their past viewing history. An example of a prompt to the generating AI model might be, "What movies do you recommend? I used to like action movies." Based on this prompt, the AI model will extract and suggest the movie that best suits the user's preferences.
[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0335] Step 1:
[0336] Users input questions or data in text, image, or audio format through their device. This input data is formatted on the device and sent to the server over the network. The input data arrives at the server in its original format.
[0337] Step 2:
[0338] The server parses the received data and first stores it in a database. The data is efficiently stored using a database system (e.g., MySQL). The stored data is structured with the assumption that it will be used for subsequent natural language processing.
[0339] Step 3:
[0340] The server executes Python scripts and analyzes the stored data using natural language processing. It utilizes tools such as the BERT model from the Transformers library to extract meaning and relationships from the data. Specifically, it converts the input data into feature vectors to understand context and related information. As a result, knowledge data is generated, which forms the basis for answering user questions.
[0341] Step 4:
[0342] When a user enters a specific question from their device, it is sent to the server. The server interprets this question and extracts the appropriate answer by referring to knowledge data generated from a database. By using prompts, the AI model quickly finds relevant information and suggests content based on the user's interests.
[0343] Step 5:
[0344] The device receives responses from the server and displays them to the user through a user interface. This display includes details of the suggested content and information. The device also provides the ability to read the responses aloud using the Google Cloud Text-to-Speech API.
[0345] Step 6:
[0346] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and used to further train the AI model. The server analyzes this feedback to improve the AI model and increase the accuracy of its answers to future questions.
[0347] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0348] The system of the present invention not only processes diverse forms of information input from users and generates highly accurate answers to questions, but also has the function of recognizing the user's emotional state and optimizing the response. This system is implemented with a configuration including an information processing device, a database, an AI model, an emotion engine, and a user interface.
[0349] First, the user inputs information via the device. This information can be obtained as text, images, or audio. For example, if the user inputs a voice message, the device uses speech recognition technology to convert it into text.
[0350] Next, the device sends the converted data to the server. The server stores this data in a database and uses an AI model to generate knowledge data. This process employs natural language processing techniques to analyze the meaning of the information and prepare to generate answers to questions.
[0351] When a user enters a question, the terminal transmits it to the server, which then compares it with knowledge data to interpret the question. Based on the user's question and related information, the server generates an appropriate answer.
[0352] The key here is the role of the emotion engine. The emotion engine analyzes the user's emotional state from their input. For example, it can analyze the content of the text and the tone of their voice to determine whether the user is angry or sad. Based on this emotional information, the server generates personalized responses that are appropriate to the user's state. For example, if the user is feeling stressed, the response will include content that promotes relaxation.
[0353] The generated response is sent to the device and displayed through the user interface. Furthermore, the device can use emotion recognition to provide text-to-speech that matches the tone and nuances the user is likely to prefer.
[0354] Finally, users provide feedback on the information provided. This feedback is crucial for the continuous improvement of the system, and the server uses it to adjust the AI model and emotion engine.
[0355] Thus, this system, which incorporates an emotion engine, provides not only the functionality of a conventional question-answering system but also emotion recognition and response optimization to improve the user experience. This enables more natural and adaptable dialogue.
[0356] The following describes the processing flow.
[0357] Step 1:
[0358] The user inputs information through their device. This information is obtained in text, audio, or image format. For example, the user might input a thank-you message via voice.
[0359] Step 2:
[0360] The device converts the information into the appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from image data as needed.
[0361] Step 3:
[0362] The terminal sends formatted data to the server. This data is sent according to a predetermined communication protocol because it will be used for subsequent processing.
[0363] Step 4:
[0364] The server stores the received data in a database. The data is structured for easy searching and use and stored in the database.
[0365] Step 5:
[0366] The server analyzes the data and generates knowledge data using an AI model. Natural language processing techniques are applied here to analyze the meaning and context of the data.
[0367] Step 6:
[0368] The server uses an emotion engine to analyze the user's emotional state. For example, it analyzes the tone of the input text or voice to identify the user's emotions.
[0369] Step 7:
[0370] The user enters a question through their device. The question seeks specific information or an answer.
[0371] Step 8:
[0372] The terminal sends the question received from the user to the server. The question data is sent in a predetermined format, preparing it for processing on the server side.
[0373] Step 9:
[0374] The server interprets the question and retrieves relevant knowledge data from the database. The server analyzes the context of the question and extracts relevant information.
[0375] Step 10:
[0376] The server uses an AI model to generate appropriate responses. Here, it considers the analysis results of the emotion engine to create personalized responses tailored to the user's emotional state.
[0377] Step 11:
[0378] The server generates the response and sends it to the device. The response data is sent in a format that the device can use.
[0379] Step 12:
[0380] The device displays the answer through its user interface. The answer is displayed on the screen as text and may also be played back using a text-to-speech function.
[0381] Step 13:
[0382] Users provide feedback on the information they receive. Users rate whether the answers were satisfactory.
[0383] Step 14:
[0384] The device sends user feedback to the server. This feedback information is managed for future system improvements.
[0385] Step 15:
[0386] The server analyzes the feedback and incorporates it into the AI model and emotion engine to improve accuracy. This allows the system to continuously improve the user experience.
[0387] (Example 2)
[0388] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0389] Traditional information processing systems focus on providing accurate answers to user questions, but they fail to adequately optimize responses based on the user's emotional state. As a result, the user experience is uniform, making it difficult to provide individualized interactions. Furthermore, there are challenges in handling different languages, highlighting the need for high-quality question-and-answer systems in multiple languages.
[0390] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0391] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a data storage area, means for generating knowledge data to respond to human questions using natural language processing technology, and means for analyzing the user's emotional state using emotion recognition technology and optimizing responses according to that state. This enables personalized dialogue that takes the user's emotions into consideration, and further enables the provision of high-quality responses that support multiple languages.
[0392] An "information processing device" is a general term for hardware and software that receives data in various formats from users and performs the necessary processing.
[0393] "Media data" refers to various information formats that users can input, such as text, images, and audio.
[0394] "Data storage area" refers to a storage area for saving received data, and includes databases and other storage methods.
[0395] "Natural language processing technology" is a technology that enables computers to understand, interpret, and respond to human language, and is used to analyze user questions.
[0396] "Knowledge data" refers to a dataset that systematically organizes the information necessary to generate answers to user questions.
[0397] "Emotion recognition technology" is a technology that analyzes the emotional state from the user's input and identifies a specific emotion.
[0398] "Multilingual support" refers to a system's ability to understand inquiries in different languages and generate responses in the appropriate language accordingly.
[0399] The system of the present invention includes an information processing device, a database, a generative AI model, an emotion recognition engine, and a user interface. These components process diverse forms of information input from the user, generate highly accurate answers to questions, and recognize the user's emotional state to optimize responses.
[0400] The user provides input data via the device in the form of text, images, or audio. In the case of audio input, the device uses speech recognition software to convert the audio data into text data. A common cloud-based speech recognition service is used for this conversion. The device then sends the converted data to the server.
[0401] The server stores the received data in data storage and generates knowledge data from the information using a generative AI model. This process employs natural language processing techniques to analyze the user's questions and prepare the optimal answers. An advanced AI framework is used to generate the knowledge data.
[0402] Furthermore, the server uses an emotion recognition engine to analyze the user's emotional state from their input. This emotion analysis is based on the user's text content and tone of voice. For example, if a user inputs "I've been feeling tired lately," the emotion recognition engine detects fatigue and generates a personalized response accordingly.
[0403] The generated responses are sent to the device and displayed or read aloud through the user interface. In the case of read-aloud responses, the tone is tailored to the user's emotional state.
[0404] For example, if a user inputs "I've been feeling stressed lately," the emotion recognition engine detects the stress, the server generates a response such as "Why don't you try listening to some relaxing music?", and the device displays or announces this message.
[0405] An example of a prompt in this invention is, "If the user is feeling stressed, how do you generate a response that promotes relaxation?"
[0406] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0407] Step 1:
[0408] The user inputs information using a device. The input format can be text, images, or audio. For example, by opening a smartphone application and speaking "What's the weather like today?", audio data is acquired. This input data is then converted into text data using speech recognition technology. Specifically, speech recognition software analyzes the audio data and outputs the result as the text "What's the weather like today?".
[0409] Step 2:
[0410] The terminal sends the converted text data to the server. This transmission uses an HTTP request with an internet connection. The server checks the received text data and saves it to its data storage area. Specifically, the server accesses the database system and saves the text "What's the weather like today?".
[0411] Step 3:
[0412] The server inputs the stored text data into a generative AI model, which uses natural language processing techniques to analyze the meaning of the question. In this process, for example, the generative AI model analyzes the text, identifies that the user is asking a question about "weather," and organizes the relevant information as knowledge data. The output at this stage is the analyzed intent of the question and related information.
[0413] Step 4:
[0414] The server uses an emotion recognition engine to analyze the user's emotional state from their input. It identifies the emotional state based on the user's text content and tone. For example, if the input includes the word "tired," it detects stress and fatigue. This emotional information is then incorporated into subsequent response generation.
[0415] Step 5:
[0416] The server integrates knowledge data and sentiment information using a generative AI model to generate the optimal response. Specifically, based on information such as "the user is asking about the weather and is feeling stressed," it constructs a response like "It's sunny today. How about going for a walk?" This response is then sent to the device.
[0417] Step 6:
[0418] The terminal displays the response received from the server in its user interface. It also uses speech synthesis technology to read the response aloud as needed. Taking emotional elements into consideration, it provides guidance in a gentle tone, such as, "It's sunny today. Why not go for a walk?"
[0419] Step 7:
[0420] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and stored to improve the system's accuracy. This feedback is used to fine-tune the AI model and emotion recognition engine, helping to improve the system.
[0421] (Application Example 2)
[0422] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0423] When users inquire about information, they are required to receive appropriate and prompt answers to their questions, and furthermore, to have those answers optimized to suit their emotional state. In particular, in the context of electronic payments, users often experience anxiety and stress, and it is necessary for the system to respond appropriately to these feelings.
[0424] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0425] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a database, means for processing the media data and generating knowledge data to respond to human questions, and means for analyzing the user's emotional state and optimizing the response. This enables the provision of a fast, appropriate, and emotionally-driven response to the user.
[0426] An "information processing device" is a device that receives data input from a user and processes it in a predetermined format.
[0427] "Media data" refers to information expressed in multiple formats, such as text, audio, and images.
[0428] A "database" is a digital storage system used to store and manage information efficiently and systematically.
[0429] "Knowledge data" refers to information generated by processing raw data in order to produce appropriate answers to user input questions.
[0430] "Means for analyzing emotional states" refers to functions that recognize and evaluate emotions based on user input information and adjust responses accordingly.
[0431] "Means of optimizing responses" refer to methods for personalizing the content of responses according to the user's emotional state and needs, and providing them in the most optimal way.
[0432] "Feedback" refers to the collection of evaluations and opinions from users that are used to improve future processes.
[0433] The system implementing this invention comprises an information processing device, a database, multiple AI models, and a user interface. The server processes data input from the user in multiple formats, such as voice, text, and images. This includes transcribing voice data into text using a speech recognition library (e.g., speech recognition technology), analyzing text data using natural language processing technology, and storing it in the database.
[0434] Based on the information stored in the database, the server uses an AI model (e.g., a generative AI model) to generate answers to questions. During this process, an emotion analysis engine (e.g., emotion recognition technology) is used to analyze the user's emotional state and adjust the content and tone of the answers to match the user's emotions.
[0435] On the device, the generated response is displayed through the user interface and can also be read aloud. This allows users to obtain information through both sight and sound. User feedback is collected and recorded in a database to help improve the system. Based on this feedback, the accuracy of the AI model and sentiment analysis engine is improved.
[0436] As a concrete example, in an electronic payment scenario, if a user asks, "Is this payment secure?", the system would return the message, "You can use it with confidence. This transaction is highly secure." An example of a prompt used here would be, "Generate an appropriate response to the user's question and optimize the tone based on the emotion detected by emotion recognition technology."
[0437] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0438] Step 1:
[0439] The terminal receives the voice data entered by the user. This voice data is then converted into text data using speech recognition technology. This converted text data becomes the input for the next step.
[0440] Step 2:
[0441] The server receives text data sent from the terminal. Natural language processing techniques are used to analyze the meaning of the question. Based on the text data, relevant information is retrieved from the database to generate knowledge data. This knowledge data serves as input for the next step.
[0442] Step 3:
[0443] The server uses sentiment analysis technology to analyze the user's emotional state. This analysis includes calculations to identify emotions from the text content. Based on the analysis results, it generates personalized responses tailored to the user's emotional state. This generated response data becomes the input for the next step.
[0444] Step 4:
[0445] The server uses a generative AI model to concretize appropriate responses for the user based on knowledge data and sentiment analysis results. Here, prompts are used to instruct the AI model. Based on these instructions, the final response is output.
[0446] Step 5:
[0447] The device receives personalized responses sent from the server. These responses are displayed through the user interface and read aloud using speech synthesis technology. This process allows the user to obtain information both visually and aurally.
[0448] Step 6:
[0449] Users provide feedback on the answers provided by the system. The device collects this feedback and sends it to the server. This feedback data is used for the continuous improvement of the system and contributes to improving the accuracy of the generative AI model and sentiment analysis engine.
[0450] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0451] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0452] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0453] [Third Embodiment]
[0454] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0455] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0456] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0457] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0458] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0460] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0461] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0462] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0463] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0464] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0465] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0466] The system of the present invention effectively processes diverse forms of information input by a user and generates highly accurate answers to the user's questions. This system is implemented with a configuration including an information processing device, a database, an AI model, and a user interface.
[0467] First, the user provides information through their device. This information is acquired as text input, images, or audio data. This acquired information is formatted appropriately by the device and sent to the server. The server stores the received information in a database and makes it available for subsequent processing.
[0468] Next, the server analyzes the data and executes a process to generate knowledge data using an AI model. Natural language processing techniques are applied here to understand the meaning from the data and extract the information necessary to answer the user's questions. This AI model improves the accuracy of its knowledge through regular learning. Furthermore, the model has multilingual capabilities, ensuring smooth operation in different language environments.
[0469] When a user asks a question from their device, the device sends it to the server. The server interprets the question, retrieves relevant knowledge data from its database, and generates an appropriate answer. This answer is provided in a personalized format according to the user's needs.
[0470] The generated responses are sent to the device and displayed through the user interface. Furthermore, it is possible to have the responses read aloud using the voice output function, providing not only visual but also auditory information.
[0471] Finally, users can provide feedback on the information provided. The device sends this feedback to the server, which analyzes it and uses it to further improve the accuracy of the AI model. User feedback allows the system to continuously improve, enabling more natural and human-like conversations.
[0472] Thus, the system of the present invention provides quick and accurate answers to user questions by going through a variety of processes starting with information provided by the user. All elements supporting this process are closely coordinated to realize advanced information processing and dialogue functions.
[0473] The following describes the processing flow.
[0474] Step 1:
[0475] The user inputs information through their device. This information is captured as text, audio, or images. For example, the user might input text about their travels or upload images of tourist destinations.
[0476] Step 2:
[0477] The device converts the acquired information into an appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from images using image analysis technology.
[0478] Step 3:
[0479] The terminal sends formatted data to the server. This data is transmitted according to the communication protocol because it will be used in subsequent processing.
[0480] Step 4:
[0481] The server stores the received data in a database. The data is structured so that it can be efficiently searched and used later.
[0482] Step 5:
[0483] The server analyzes the data and generates knowledge data using an AI model. Natural language processing is applied here to analyze the meaning of the data.
[0484] Step 6:
[0485] The user enters a question via their device. The question is specific and constitutes a request for information from the system.
[0486] Step 7:
[0487] The terminal sends the question received from the user to the server. The question is formatted and prepared for processing on the server side.
[0488] Step 8:
[0489] The server interprets the question and retrieves relevant knowledge data from the database. This forms the basis for the answer to the question.
[0490] Step 9:
[0491] The server uses an AI model to generate appropriate answers. The generated answers are presented in a format that is easy for the user to understand.
[0492] Step 10:
[0493] The server generates the response and sends it to the terminal. The response data is delivered to the terminal via communication.
[0494] Step 11:
[0495] The device displays the answer through the user interface. The answer can be presented in both text and audio formats.
[0496] Step 12:
[0497] Users provide feedback on the answers. This feedback is valuable information for improving the system.
[0498] Step 13:
[0499] The device sends user feedback to the server. This feedback data is managed as material for analysis.
[0500] Step 14:
[0501] The server analyzes the feedback and incorporates it into the AI model to improve accuracy. Based on the feedback, the system is improved.
[0502] (Example 1)
[0503] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0504] There was a need for an information system that could accept data in various formats from users, process it effectively, and provide precise answers. Furthermore, challenges existed in handling questions in different languages and improving system accuracy through user feedback. This project aims to solve these problems.
[0505] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0506] In this invention, the server includes means for converting data in multiple formats input via information processing means into a unified format and storing it in a database; means for analyzing the data and generating knowledge to respond to human questions using an artificial intelligence model; and means for deriving the optimal answer to the received question based on the knowledge generated by the artificial intelligence model. This enables multilingual support and highly accurate information provision that reflects feedback.
[0507] "Information processing means" refers to a device or process that has the function of receiving input data, converting it into a standardized format, and storing it in a database.
[0508] A "database" is a collection of information for efficient storage, management, and retrieval, and a means of quickly obtaining necessary data.
[0509] An "artificial intelligence model" is an algorithm or system used to analyze data and generate knowledge in response to human questions based on the results.
[0510] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language, and is used for questioning and data analysis.
[0511] "Feedback" refers to evaluations and opinions provided by users regarding the system's response, and is information used to improve and adjust the system.
[0512] "Multilingual support" refers to a system's ability to understand multiple different languages and provide appropriate responses in each language.
[0513] The system according to the present invention is an integrated information processing system for improving flexibility and accuracy in information handling. This system efficiently executes a series of processes, from information input to response generation and improvement using feedback.
[0514] First, the user inputs information through the terminal. The input information can take various forms, including text, images, and audio data. For example, audio data is converted into text data using speech recognition technology. Standard speech recognition software is used for this process. The formatted data is then sent to the server via a secure communication protocol.
[0515] The server stores the received data in a database. This database is designed to efficiently manage large amounts of information and retrieve it quickly when needed. Specifically, a database management system is commonly used.
[0516] Next, the server analyzes the data using an artificial intelligence model. This analysis process utilizes natural language processing techniques to understand the meaning of the information and generate knowledge in response to the user's questions. For example, open-source natural language processing frameworks and cloud-based AI services are commonly used as generative AI models.
[0517] When a user asks a question, the device sends the question to the server, which interprets the content of the question. Based on the generated knowledge, the server derives the best answer. This answer is personalized based on the user's past interactions and profile information. For example, in response to a prompt such as "I'm thinking of traveling, can you recommend some tourist spots?", the system will suggest tourist spots considering the user's hobbies and past visits.
[0518] Finally, users can provide feedback on their answers. This feedback is sent from the device to the server and used to train the AI model. This allows the system to continuously improve its accuracy and enhance the user experience.
[0519] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0520] Step 1:
[0521] Users input information via a terminal. This information is collected in various forms, including text, images, and audio data. The terminal receives the input data and converts it into the appropriate format. Specifically, if audio data is input, the terminal uses speech recognition technology to convert it into text data. This conversion generates a unified data format that can be transmitted to the server.
[0522] Step 2:
[0523] The terminal sends formatted data to the server. The server stores the received data in a database. The database plays a role in enabling centralized data management and efficient searching. Once the input data is stored, it becomes available for use in subsequent analysis processes.
[0524] Step 3:
[0525] The server retrieves data from the database and performs analysis using a generative AI model. Here, natural language processing is used to interpret the meaning of the data and generate relevant knowledge. This process inputs data as prompts into the AI model, and the analysis results are obtained as output from the model. These results can be used to answer user questions.
[0526] Step 4:
[0527] When a user enters a question, the device sends this question to the server. The server analyzes the question and extracts relevant knowledge from its database. Using a generative AI model, it generates the most appropriate answer to the question and provides it to the user in a personalized manner. The output is a prompt-based answer.
[0528] Step 5:
[0529] The generated responses are sent to the device and displayed through the user interface. In addition, for users who have difficulty visualizing, the device provides the responses audibly using its voice output function. In this way, users can receive information through both visual and auditory means.
[0530] Step 6:
[0531] Users can input feedback on the answers they receive. The device sends the feedback to the server. The server analyzes the feedback and incorporates it into the AI model's learning process to improve the system's accuracy and responsiveness. This process continuously improves the user experience.
[0532] (Application Example 1)
[0533] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0534] In today's information-saturated society, users demand quick and accurate information suggestions, but existing systems have limitations in the accuracy of personalized content recommendations. Therefore, there is a need for technology that effectively delivers appropriate content tailored to users' interests and preferences.
[0535] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0536] In this invention, the server includes means for storing media data input via an information processing device in a data storage means, means for using technology to process the media data and generate knowledge data, and means for suggesting personalized content based on the user's interests through a user interface. This enables the suggestion of personalized content in response to the user's questions and interests.
[0537] An "information processing device" is an electronic device used to collect and analyze data and generate output according to a specific purpose.
[0538] "Media data" refers to information provided by users in various formats, such as text, audio, and images.
[0539] A "data storage means" is a mechanism for organizing and storing collected data.
[0540] "Knowledge data" refers to useful information in response to user questions, obtained by analyzing and processing collected data.
[0541] A "user interface" is a means of interaction that allows a user to exchange information with a system in a two-way manner.
[0542] "Personalization" is the process of adapting information and content to each user's interests and preferences.
[0543] A "proposal" is an action that presents users with options or information.
[0544] The system that realizes this application consists of an information processing device, a database, an AI model, and a user interface.
[0545] The server first collects media data such as text, images, and audio input from users via smartphones or smart glasses, formats it appropriately, and stores it in a database. This data is efficiently managed by the data storage system.
[0546] Next, the server uses Python to perform natural language processing and extract relevant information from the media data. This process utilizes AI models such as the BERT model from the Transformers library to achieve advanced data analysis. After the analysis, knowledge data is generated, which forms the basis for deriving information based on the user's questions and interests.
[0547] When a user asks a specific question through the user interface, the server quickly processes the question, retrieves relevant knowledge data from the database, and provides personalized content. This structure ensures that content most closely matches the user's interests is suggested.
[0548] Furthermore, the server receives user feedback and improves the performance of the AI model. This feedback loop is a crucial element in ensuring continuous improvement of the system.
[0549] As a concrete example, if a user asks a question about movie recommendations, the system will suggest the most suitable movie based on their past viewing history. An example of a prompt to the generating AI model might be, "What movies do you recommend? I used to like action movies." Based on this prompt, the AI model will extract and suggest the movie that best suits the user's preferences.
[0550] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0551] Step 1:
[0552] Users input questions or data in text, image, or audio format through their device. This input data is formatted on the device and sent to the server over the network. The input data arrives at the server in its original format.
[0553] Step 2:
[0554] The server parses the received data and first stores it in a database. The data is efficiently stored using a database system (e.g., MySQL). The stored data is structured with the assumption that it will be used for subsequent natural language processing.
[0555] Step 3:
[0556] The server executes Python scripts and analyzes the stored data using natural language processing. It utilizes tools such as the BERT model from the Transformers library to extract meaning and relationships from the data. Specifically, it converts the input data into feature vectors to understand context and related information. As a result, knowledge data is generated, which forms the basis for answering user questions.
[0557] Step 4:
[0558] When a user enters a specific question from their device, it is sent to the server. The server interprets this question and extracts the appropriate answer by referring to knowledge data generated from a database. By using prompts, the AI model quickly finds relevant information and suggests content based on the user's interests.
[0559] Step 5:
[0560] The device receives responses from the server and displays them to the user through a user interface. This display includes details of the suggested content and information. The device also provides the ability to read the responses aloud using the Google Cloud Text-to-Speech API.
[0561] Step 6:
[0562] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and used to further train the AI model. The server analyzes this feedback to improve the AI model and increase the accuracy of its answers to future questions.
[0563] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0564] The system of the present invention not only processes diverse forms of information input from users and generates highly accurate answers to questions, but also has the function of recognizing the user's emotional state and optimizing the response. This system is implemented with a configuration including an information processing device, a database, an AI model, an emotion engine, and a user interface.
[0565] First, the user inputs information via the device. This information can be obtained as text, images, or audio. For example, if the user inputs a voice message, the device uses speech recognition technology to convert it into text.
[0566] Next, the device sends the converted data to the server. The server stores this data in a database and uses an AI model to generate knowledge data. This process employs natural language processing techniques to analyze the meaning of the information and prepare to generate answers to questions.
[0567] When a user enters a question, the terminal transmits it to the server, which then compares it with knowledge data to interpret the question. Based on the user's question and related information, the server generates an appropriate answer.
[0568] The key here is the role of the emotion engine. The emotion engine analyzes the user's emotional state from their input. For example, it can analyze the content of the text and the tone of their voice to determine whether the user is angry or sad. Based on this emotional information, the server generates personalized responses that are appropriate to the user's state. For example, if the user is feeling stressed, the response will include content that promotes relaxation.
[0569] The generated response is sent to the device and displayed through the user interface. Furthermore, the device can use emotion recognition to provide text-to-speech that matches the tone and nuances the user is likely to prefer.
[0570] Finally, users provide feedback on the information provided. This feedback is crucial for the continuous improvement of the system, and the server uses it to adjust the AI model and emotion engine.
[0571] Thus, this system, which incorporates an emotion engine, provides not only the functionality of a conventional question-answering system but also emotion recognition and response optimization to improve the user experience. This enables more natural and adaptable dialogue.
[0572] The following describes the processing flow.
[0573] Step 1:
[0574] The user inputs information through their device. This information is obtained in text, audio, or image format. For example, the user might input a thank-you message via voice.
[0575] Step 2:
[0576] The device converts the information into the appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from image data as needed.
[0577] Step 3:
[0578] The terminal sends formatted data to the server. This data is sent according to a predetermined communication protocol because it will be used for subsequent processing.
[0579] Step 4:
[0580] The server stores the received data in a database. The data is structured for easy searching and use and stored in the database.
[0581] Step 5:
[0582] The server analyzes the data and generates knowledge data using an AI model. Natural language processing techniques are applied here to analyze the meaning and context of the data.
[0583] Step 6:
[0584] The server uses an emotion engine to analyze the user's emotional state. For example, it analyzes the tone of the input text or voice to identify the user's emotions.
[0585] Step 7:
[0586] The user enters a question through their device. The question seeks specific information or an answer.
[0587] Step 8:
[0588] The terminal sends the question received from the user to the server. The question data is sent in a predetermined format, preparing it for processing on the server side.
[0589] Step 9:
[0590] The server interprets the question and retrieves relevant knowledge data from the database. The server analyzes the context of the question and extracts relevant information.
[0591] Step 10:
[0592] The server uses an AI model to generate appropriate responses. Here, it considers the analysis results of the emotion engine to create personalized responses tailored to the user's emotional state.
[0593] Step 11:
[0594] The server generates the response and sends it to the device. The response data is sent in a format that the device can use.
[0595] Step 12:
[0596] The device displays the answer through its user interface. The answer is displayed on the screen as text and may also be played back using a text-to-speech function.
[0597] Step 13:
[0598] Users provide feedback on the information they receive. Users rate whether the answers were satisfactory.
[0599] Step 14:
[0600] The device sends user feedback to the server. This feedback information is managed for future system improvements.
[0601] Step 15:
[0602] The server analyzes the feedback and incorporates it into the AI model and emotion engine to improve accuracy. This allows the system to continuously improve the user experience.
[0603] (Example 2)
[0604] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0605] Traditional information processing systems focus on providing accurate answers to user questions, but they fail to adequately optimize responses based on the user's emotional state. As a result, the user experience is uniform, making it difficult to provide individualized interactions. Furthermore, there are challenges in handling different languages, highlighting the need for high-quality question-and-answer systems in multiple languages.
[0606] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0607] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a data storage area, means for generating knowledge data to respond to human questions using natural language processing technology, and means for analyzing the user's emotional state using emotion recognition technology and optimizing responses according to that state. This enables personalized dialogue that takes the user's emotions into consideration, and further enables the provision of high-quality responses that support multiple languages.
[0608] An "information processing device" is a general term for hardware and software that receives data in various formats from users and performs the necessary processing.
[0609] "Media data" refers to various information formats that users can input, such as text, images, and audio.
[0610] "Data storage area" refers to a storage area for saving received data, and includes databases and other storage methods.
[0611] "Natural language processing technology" is a technology that enables computers to understand, interpret, and respond to human language, and is used to analyze user questions.
[0612] "Knowledge data" refers to a dataset that systematically organizes the information necessary to generate answers to user questions.
[0613] "Emotion recognition technology" is a technology that analyzes the emotional state from the user's input and identifies a specific emotion.
[0614] "Multilingual support" refers to a system's ability to understand inquiries in different languages and generate responses in the appropriate language accordingly.
[0615] The system of the present invention includes an information processing device, a database, a generative AI model, an emotion recognition engine, and a user interface. These components process diverse forms of information input from the user, generate highly accurate answers to questions, and recognize the user's emotional state to optimize responses.
[0616] The user provides input data via the device in the form of text, images, or audio. In the case of audio input, the device uses speech recognition software to convert the audio data into text data. A common cloud-based speech recognition service is used for this conversion. The device then sends the converted data to the server.
[0617] The server stores the received data in data storage and generates knowledge data from the information using a generative AI model. This process employs natural language processing techniques to analyze the user's questions and prepare the optimal answers. An advanced AI framework is used to generate the knowledge data.
[0618] Furthermore, the server uses an emotion recognition engine to analyze the user's emotional state from their input. This emotion analysis is based on the user's text content and tone of voice. For example, if a user inputs "I've been feeling tired lately," the emotion recognition engine detects fatigue and generates a personalized response accordingly.
[0619] The generated responses are sent to the device and displayed or read aloud through the user interface. In the case of read-aloud responses, the tone is tailored to the user's emotional state.
[0620] For example, if a user inputs "I've been feeling stressed lately," the emotion recognition engine detects the stress, the server generates a response such as "Why don't you try listening to some relaxing music?", and the device displays or announces this message.
[0621] An example of a prompt in this invention is, "If the user is feeling stressed, how do you generate a response that promotes relaxation?"
[0622] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0623] Step 1:
[0624] The user inputs information using a device. The input format can be text, images, or audio. For example, by opening a smartphone application and speaking "What's the weather like today?", audio data is acquired. This input data is then converted into text data using speech recognition technology. Specifically, speech recognition software analyzes the audio data and outputs the result as the text "What's the weather like today?".
[0625] Step 2:
[0626] The terminal sends the converted text data to the server. This transmission uses an HTTP request with an internet connection. The server checks the received text data and saves it to its data storage area. Specifically, the server accesses the database system and saves the text "What's the weather like today?".
[0627] Step 3:
[0628] The server inputs the stored text data into a generative AI model, which uses natural language processing techniques to analyze the meaning of the question. In this process, for example, the generative AI model analyzes the text, identifies that the user is asking a question about "weather," and organizes the relevant information as knowledge data. The output at this stage is the analyzed intent of the question and related information.
[0629] Step 4:
[0630] The server uses an emotion recognition engine to analyze the user's emotional state from their input. It identifies the emotional state based on the user's text content and tone. For example, if the input includes the word "tired," it detects stress and fatigue. This emotional information is then incorporated into subsequent response generation.
[0631] Step 5:
[0632] The server integrates knowledge data and sentiment information using a generative AI model to generate the optimal response. Specifically, based on information such as "the user is asking about the weather and is feeling stressed," it constructs a response like "It's sunny today. How about going for a walk?" This response is then sent to the device.
[0633] Step 6:
[0634] The terminal displays the response received from the server in its user interface. It also uses speech synthesis technology to read the response aloud as needed. Taking emotional elements into consideration, it provides guidance in a gentle tone, such as, "It's sunny today. Why not go for a walk?"
[0635] Step 7:
[0636] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and stored to improve the system's accuracy. This feedback is used to fine-tune the AI model and emotion recognition engine, helping to improve the system.
[0637] (Application Example 2)
[0638] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0639] When users inquire about information, they are required to receive appropriate and prompt answers to their questions, and furthermore, to have those answers optimized to suit their emotional state. In particular, in the context of electronic payments, users often experience anxiety and stress, and it is necessary for the system to respond appropriately to these feelings.
[0640] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0641] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a database, means for processing the media data and generating knowledge data to respond to human questions, and means for analyzing the user's emotional state and optimizing the response. This enables the provision of a fast, appropriate, and emotionally-driven response to the user.
[0642] An "information processing device" is a device that receives data input from a user and processes it in a predetermined format.
[0643] "Media data" refers to information expressed in multiple formats, such as text, audio, and images.
[0644] A "database" is a digital storage system used to store and manage information efficiently and systematically.
[0645] "Knowledge data" refers to information generated by processing raw data in order to produce appropriate answers to user input questions.
[0646] "Means for analyzing emotional states" refers to functions that recognize and evaluate emotions based on user input information and adjust responses accordingly.
[0647] "Means of optimizing responses" refer to methods for personalizing the content of responses according to the user's emotional state and needs, and providing them in the most optimal way.
[0648] "Feedback" refers to the collection of evaluations and opinions from users that are used to improve future processes.
[0649] The system implementing this invention comprises an information processing device, a database, multiple AI models, and a user interface. The server processes data input from the user in multiple formats, such as voice, text, and images. This includes transcribing voice data into text using a speech recognition library (e.g., speech recognition technology), analyzing text data using natural language processing technology, and storing it in the database.
[0650] Based on the information stored in the database, the server uses an AI model (e.g., a generative AI model) to generate answers to questions. During this process, an emotion analysis engine (e.g., emotion recognition technology) is used to analyze the user's emotional state and adjust the content and tone of the answers to match the user's emotions.
[0651] On the device, the generated response is displayed through the user interface and can also be read aloud. This allows users to obtain information through both sight and sound. User feedback is collected and recorded in a database to help improve the system. Based on this feedback, the accuracy of the AI model and sentiment analysis engine is improved.
[0652] As a concrete example, in an electronic payment scenario, if a user asks, "Is this payment secure?", the system would return the message, "You can use it with confidence. This transaction is highly secure." An example of a prompt used here would be, "Generate an appropriate response to the user's question and optimize the tone based on the emotion detected by emotion recognition technology."
[0653] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0654] Step 1:
[0655] The terminal receives the voice data entered by the user. This voice data is then converted into text data using speech recognition technology. This converted text data becomes the input for the next step.
[0656] Step 2:
[0657] The server receives text data sent from the terminal. Natural language processing techniques are used to analyze the meaning of the question. Based on the text data, relevant information is retrieved from the database to generate knowledge data. This knowledge data serves as input for the next step.
[0658] Step 3:
[0659] The server uses sentiment analysis technology to analyze the user's emotional state. This analysis includes calculations to identify emotions from the text content. Based on the analysis results, it generates personalized responses tailored to the user's emotional state. This generated response data becomes the input for the next step.
[0660] Step 4:
[0661] The server uses a generative AI model to concretize appropriate responses for the user based on knowledge data and sentiment analysis results. Here, prompts are used to instruct the AI model. Based on these instructions, the final response is output.
[0662] Step 5:
[0663] The device receives personalized responses sent from the server. These responses are displayed through the user interface and read aloud using speech synthesis technology. This process allows the user to obtain information both visually and aurally.
[0664] Step 6:
[0665] Users provide feedback on the answers provided by the system. The device collects this feedback and sends it to the server. This feedback data is used for the continuous improvement of the system and contributes to improving the accuracy of the generative AI model and sentiment analysis engine.
[0666] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0667] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0669] [Fourth Embodiment]
[0670] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0671] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0673] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0677] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0678] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0679] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0680] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0681] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0682] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] The system of the present invention effectively processes diverse forms of information input by a user and generates highly accurate answers to the user's questions. This system is implemented with a configuration including an information processing device, a database, an AI model, and a user interface.
[0684] First, the user provides information through their device. This information is acquired as text input, images, or audio data. This acquired information is formatted appropriately by the device and sent to the server. The server stores the received information in a database and makes it available for subsequent processing.
[0685] Next, the server analyzes the data and executes a process to generate knowledge data using an AI model. Natural language processing techniques are applied here to understand the meaning from the data and extract the information necessary to answer the user's questions. This AI model improves the accuracy of its knowledge through regular learning. Furthermore, the model has multilingual capabilities, ensuring smooth operation in different language environments.
[0686] When a user asks a question from their device, the device sends it to the server. The server interprets the question, retrieves relevant knowledge data from its database, and generates an appropriate answer. This answer is provided in a personalized format according to the user's needs.
[0687] The generated responses are sent to the device and displayed through the user interface. Furthermore, it is possible to have the responses read aloud using the voice output function, providing not only visual but also auditory information.
[0688] Finally, users can provide feedback on the information provided. The device sends this feedback to the server, which analyzes it and uses it to further improve the accuracy of the AI model. User feedback allows the system to continuously improve, enabling more natural and human-like conversations.
[0689] Thus, the system of the present invention provides quick and accurate answers to user questions by going through a variety of processes starting with information provided by the user. All elements supporting this process are closely coordinated to realize advanced information processing and dialogue functions.
[0690] The following describes the processing flow.
[0691] Step 1:
[0692] The user inputs information through their device. This information is captured as text, audio, or images. For example, the user might input text about their travels or upload images of tourist destinations.
[0693] Step 2:
[0694] The device converts the acquired information into an appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from images using image analysis technology.
[0695] Step 3:
[0696] The terminal sends formatted data to the server. This data is transmitted according to the communication protocol because it will be used in subsequent processing.
[0697] Step 4:
[0698] The server stores the received data in a database. The data is structured so that it can be efficiently searched and used later.
[0699] Step 5:
[0700] The server analyzes the data and generates knowledge data using an AI model. Natural language processing is applied here to analyze the meaning of the data.
[0701] Step 6:
[0702] The user enters a question via their device. The question is specific and constitutes a request for information from the system.
[0703] Step 7:
[0704] The terminal sends the question received from the user to the server. The question is formatted and prepared for processing on the server side.
[0705] Step 8:
[0706] The server interprets the question and retrieves relevant knowledge data from the database. This forms the basis for the answer to the question.
[0707] Step 9:
[0708] The server uses an AI model to generate appropriate answers. The generated answers are presented in a format that is easy for the user to understand.
[0709] Step 10:
[0710] The server generates the response and sends it to the terminal. The response data is delivered to the terminal via communication.
[0711] Step 11:
[0712] The device displays the answer through the user interface. The answer can be presented in both text and audio formats.
[0713] Step 12:
[0714] Users provide feedback on the answers. This feedback is valuable information for improving the system.
[0715] Step 13:
[0716] The device sends user feedback to the server. This feedback data is managed as material for analysis.
[0717] Step 14:
[0718] The server analyzes the feedback and incorporates it into the AI model to improve accuracy. Based on the feedback, the system is improved.
[0719] (Example 1)
[0720] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0721] There was a need for an information system that could accept data in various formats from users, process it effectively, and provide precise answers. Furthermore, challenges existed in handling questions in different languages and improving system accuracy through user feedback. This project aims to solve these problems.
[0722] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0723] In this invention, the server includes means for converting data in multiple formats input via information processing means into a unified format and storing it in a database; means for analyzing the data and generating knowledge to respond to human questions using an artificial intelligence model; and means for deriving the optimal answer to the received question based on the knowledge generated by the artificial intelligence model. This enables multilingual support and highly accurate information provision that reflects feedback.
[0724] "Information processing means" refers to a device or process that has the function of receiving input data, converting it into a standardized format, and storing it in a database.
[0725] A "database" is a collection of information for efficient storage, management, and retrieval, and a means of quickly obtaining necessary data.
[0726] An "artificial intelligence model" is an algorithm or system used to analyze data and generate knowledge in response to human questions based on the results.
[0727] "Natural language processing technology" is a technology that enables computers to understand, interpret, and generate human language, and is used for questioning and data analysis.
[0728] "Feedback" refers to evaluations and opinions provided by users regarding the system's response, and is information used to improve and adjust the system.
[0729] "Multilingual support" refers to a system's ability to understand multiple different languages and provide appropriate responses in each language.
[0730] The system according to the present invention is an integrated information processing system for improving flexibility and accuracy in information handling. This system efficiently executes a series of processes, from information input to response generation and improvement using feedback.
[0731] First, the user inputs information through the terminal. The input information can take various forms, including text, images, and audio data. For example, audio data is converted into text data using speech recognition technology. Standard speech recognition software is used for this process. The formatted data is then sent to the server via a secure communication protocol.
[0732] The server stores the received data in a database. This database is designed to efficiently manage large amounts of information and retrieve it quickly when needed. Specifically, a database management system is commonly used.
[0733] Next, the server analyzes the data using an artificial intelligence model. This analysis process utilizes natural language processing techniques to understand the meaning of the information and generate knowledge in response to the user's questions. For example, open-source natural language processing frameworks and cloud-based AI services are commonly used as generative AI models.
[0734] When a user asks a question, the device sends the question to the server, which interprets the content of the question. Based on the generated knowledge, the server derives the best answer. This answer is personalized based on the user's past interactions and profile information. For example, in response to a prompt such as "I'm thinking of traveling, can you recommend some tourist spots?", the system will suggest tourist spots considering the user's hobbies and past visits.
[0735] Finally, users can provide feedback on their answers. This feedback is sent from the device to the server and used to train the AI model. This allows the system to continuously improve its accuracy and enhance the user experience.
[0736] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0737] Step 1:
[0738] Users input information via a terminal. This information is collected in various forms, including text, images, and audio data. The terminal receives the input data and converts it into the appropriate format. Specifically, if audio data is input, the terminal uses speech recognition technology to convert it into text data. This conversion generates a unified data format that can be transmitted to the server.
[0739] Step 2:
[0740] The terminal sends formatted data to the server. The server stores the received data in a database. The database plays a role in enabling centralized data management and efficient searching. Once the input data is stored, it becomes available for use in subsequent analysis processes.
[0741] Step 3:
[0742] The server retrieves data from the database and performs analysis using a generative AI model. Here, natural language processing is used to interpret the meaning of the data and generate relevant knowledge. This process inputs data as prompts into the AI model, and the analysis results are obtained as output from the model. These results can be used to answer user questions.
[0743] Step 4:
[0744] When a user enters a question, the device sends this question to the server. The server analyzes the question and extracts relevant knowledge from its database. Using a generative AI model, it generates the most appropriate answer to the question and provides it to the user in a personalized manner. The output is a prompt-based answer.
[0745] Step 5:
[0746] The generated responses are sent to the device and displayed through the user interface. In addition, for users who have difficulty visualizing, the device provides the responses audibly using its voice output function. In this way, users can receive information through both visual and auditory means.
[0747] Step 6:
[0748] Users can input feedback on the answers they receive. The device sends the feedback to the server. The server analyzes the feedback and incorporates it into the AI model's learning process to improve the system's accuracy and responsiveness. This process continuously improves the user experience.
[0749] (Application Example 1)
[0750] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0751] In today's information-saturated society, users demand quick and accurate information suggestions, but existing systems have limitations in the accuracy of personalized content recommendations. Therefore, there is a need for technology that effectively delivers appropriate content tailored to users' interests and preferences.
[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0753] In this invention, the server includes means for storing media data input via an information processing device in a data storage means, means for using technology to process the media data and generate knowledge data, and means for suggesting personalized content based on the user's interests through a user interface. This enables the suggestion of personalized content in response to the user's questions and interests.
[0754] An "information processing device" is an electronic device used to collect and analyze data and generate output according to a specific purpose.
[0755] "Media data" refers to information provided by users in various formats, such as text, audio, and images.
[0756] A "data storage means" is a mechanism for organizing and storing collected data.
[0757] "Knowledge data" refers to useful information in response to user questions, obtained by analyzing and processing collected data.
[0758] A "user interface" is a means of interaction that allows a user to exchange information with a system in a two-way manner.
[0759] "Personalization" is the process of adapting information and content to each user's interests and preferences.
[0760] A "proposal" is an action that presents users with options or information.
[0761] The system that realizes this application consists of an information processing device, a database, an AI model, and a user interface.
[0762] The server first collects media data such as text, images, and audio input from users via smartphones or smart glasses, formats it appropriately, and stores it in a database. This data is efficiently managed by the data storage system.
[0763] Next, the server uses Python to perform natural language processing and extract relevant information from the media data. This process utilizes AI models such as the BERT model from the Transformers library to achieve advanced data analysis. After the analysis, knowledge data is generated, which forms the basis for deriving information based on the user's questions and interests.
[0764] When a user asks a specific question through the user interface, the server quickly processes the question, retrieves relevant knowledge data from the database, and provides personalized content. This structure ensures that content most closely matches the user's interests is suggested.
[0765] Furthermore, the server receives user feedback and improves the performance of the AI model. This feedback loop is a crucial element in ensuring continuous improvement of the system.
[0766] As a concrete example, if a user asks a question about movie recommendations, the system will suggest the most suitable movie based on their past viewing history. An example of a prompt to the generating AI model might be, "What movies do you recommend? I used to like action movies." Based on this prompt, the AI model will extract and suggest the movie that best suits the user's preferences.
[0767] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0768] Step 1:
[0769] Users input questions or data in text, image, or audio format through their device. This input data is formatted on the device and sent to the server over the network. The input data arrives at the server in its original format.
[0770] Step 2:
[0771] The server parses the received data and first stores it in a database. The data is efficiently stored using a database system (e.g., MySQL). The stored data is structured with the assumption that it will be used for subsequent natural language processing.
[0772] Step 3:
[0773] The server executes Python scripts and analyzes the stored data using natural language processing. It utilizes tools such as the BERT model from the Transformers library to extract meaning and relationships from the data. Specifically, it converts the input data into feature vectors to understand context and related information. As a result, knowledge data is generated, which forms the basis for answering user questions.
[0774] Step 4:
[0775] When a user enters a specific question from their device, it is sent to the server. The server interprets this question and extracts the appropriate answer by referring to knowledge data generated from a database. By using prompts, the AI model quickly finds relevant information and suggests content based on the user's interests.
[0776] Step 5:
[0777] The device receives responses from the server and displays them to the user through a user interface. This display includes details of the suggested content and information. The device also provides the ability to read the responses aloud using the Google Cloud Text-to-Speech API.
[0778] Step 6:
[0779] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and used to further train the AI model. The server analyzes this feedback to improve the AI model and increase the accuracy of its answers to future questions.
[0780] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0781] The system of the present invention not only processes diverse forms of information input from users and generates highly accurate answers to questions, but also has the function of recognizing the user's emotional state and optimizing the response. This system is implemented with a configuration including an information processing device, a database, an AI model, an emotion engine, and a user interface.
[0782] First, the user inputs information via the device. This information can be obtained as text, images, or audio. For example, if the user inputs a voice message, the device uses speech recognition technology to convert it into text.
[0783] Next, the device sends the converted data to the server. The server stores this data in a database and uses an AI model to generate knowledge data. This process employs natural language processing techniques to analyze the meaning of the information and prepare to generate answers to questions.
[0784] When a user enters a question, the terminal transmits it to the server, which then compares it with knowledge data to interpret the question. Based on the user's question and related information, the server generates an appropriate answer.
[0785] The key here is the role of the emotion engine. The emotion engine analyzes the user's emotional state from their input. For example, it can analyze the content of the text and the tone of their voice to determine whether the user is angry or sad. Based on this emotional information, the server generates personalized responses that are appropriate to the user's state. For example, if the user is feeling stressed, the response will include content that promotes relaxation.
[0786] The generated response is sent to the device and displayed through the user interface. Furthermore, the device can use emotion recognition to provide text-to-speech that matches the tone and nuances the user is likely to prefer.
[0787] Finally, users provide feedback on the information provided. This feedback is crucial for the continuous improvement of the system, and the server uses it to adjust the AI model and emotion engine.
[0788] Thus, this system, which incorporates an emotion engine, provides not only the functionality of a conventional question-answering system but also emotion recognition and response optimization to improve the user experience. This enables more natural and adaptable dialogue.
[0789] The following describes the processing flow.
[0790] Step 1:
[0791] The user inputs information through their device. This information is obtained in text, audio, or image format. For example, the user might input a thank-you message via voice.
[0792] Step 2:
[0793] The device converts the information into the appropriate format. Audio data is converted to text using speech recognition technology, and text information is extracted from image data as needed.
[0794] Step 3:
[0795] The terminal sends formatted data to the server. This data is sent according to a predetermined communication protocol because it will be used for subsequent processing.
[0796] Step 4:
[0797] The server stores the received data in a database. The data is structured for easy searching and use and stored in the database.
[0798] Step 5:
[0799] The server analyzes the data and generates knowledge data using an AI model. Natural language processing techniques are applied here to analyze the meaning and context of the data.
[0800] Step 6:
[0801] The server uses an emotion engine to analyze the user's emotional state. For example, it analyzes the tone of the input text or voice to identify the user's emotions.
[0802] Step 7:
[0803] The user enters a question through their device. The question seeks specific information or an answer.
[0804] Step 8:
[0805] The terminal sends the question received from the user to the server. The question data is sent in a predetermined format, preparing it for processing on the server side.
[0806] Step 9:
[0807] The server interprets the question and retrieves relevant knowledge data from the database. The server analyzes the context of the question and extracts relevant information.
[0808] Step 10:
[0809] The server uses an AI model to generate appropriate responses. Here, it considers the analysis results of the emotion engine to create personalized responses tailored to the user's emotional state.
[0810] Step 11:
[0811] The server generates the response and sends it to the device. The response data is sent in a format that the device can use.
[0812] Step 12:
[0813] The device displays the answer through its user interface. The answer is displayed on the screen as text and may also be played back using a text-to-speech function.
[0814] Step 13:
[0815] Users provide feedback on the information they receive. Users rate whether the answers were satisfactory.
[0816] Step 14:
[0817] The device sends user feedback to the server. This feedback information is managed for future system improvements.
[0818] Step 15:
[0819] The server analyzes the feedback and incorporates it into the AI model and emotion engine to improve accuracy. This allows the system to continuously improve the user experience.
[0820] (Example 2)
[0821] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0822] Traditional information processing systems focus on providing accurate answers to user questions, but they fail to adequately optimize responses based on the user's emotional state. As a result, the user experience is uniform, making it difficult to provide individualized interactions. Furthermore, there are challenges in handling different languages, highlighting the need for high-quality question-and-answer systems in multiple languages.
[0823] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0824] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a data storage area, means for generating knowledge data to respond to human questions using natural language processing technology, and means for analyzing the user's emotional state using emotion recognition technology and optimizing responses according to that state. This enables personalized dialogue that takes the user's emotions into consideration, and further enables the provision of high-quality responses that support multiple languages.
[0825] An "information processing device" is a general term for hardware and software that receives data in various formats from users and performs the necessary processing.
[0826] "Media data" refers to various information formats that users can input, such as text, images, and audio.
[0827] "Data storage area" refers to a storage area for saving received data, and includes databases and other storage methods.
[0828] "Natural language processing technology" is a technology that enables computers to understand, interpret, and respond to human language, and is used to analyze user questions.
[0829] "Knowledge data" refers to a dataset that systematically organizes the information necessary to generate answers to user questions.
[0830] "Emotion recognition technology" is a technology that analyzes the emotional state from the user's input and identifies a specific emotion.
[0831] "Multilingual support" refers to a system's ability to understand inquiries in different languages and generate responses in the appropriate language accordingly.
[0832] The system of the present invention includes an information processing device, a database, a generative AI model, an emotion recognition engine, and a user interface. These components process diverse forms of information input from the user, generate highly accurate answers to questions, and recognize the user's emotional state to optimize responses.
[0833] The user provides input data via the device in the form of text, images, or audio. In the case of audio input, the device uses speech recognition software to convert the audio data into text data. A common cloud-based speech recognition service is used for this conversion. The device then sends the converted data to the server.
[0834] The server stores the received data in data storage and generates knowledge data from the information using a generative AI model. This process employs natural language processing techniques to analyze the user's questions and prepare the optimal answers. An advanced AI framework is used to generate the knowledge data.
[0835] Furthermore, the server uses an emotion recognition engine to analyze the user's emotional state from their input. This emotion analysis is based on the user's text content and tone of voice. For example, if a user inputs "I've been feeling tired lately," the emotion recognition engine detects fatigue and generates a personalized response accordingly.
[0836] The generated responses are sent to the device and displayed or read aloud through the user interface. In the case of read-aloud responses, the tone is tailored to the user's emotional state.
[0837] For example, if a user inputs "I've been feeling stressed lately," the emotion recognition engine detects the stress, the server generates a response such as "Why don't you try listening to some relaxing music?", and the device displays or announces this message.
[0838] An example of a prompt in this invention is, "If the user is feeling stressed, how do you generate a response that promotes relaxation?"
[0839] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0840] Step 1:
[0841] The user inputs information using a device. The input format can be text, images, or audio. For example, by opening a smartphone application and speaking "What's the weather like today?", audio data is acquired. This input data is then converted into text data using speech recognition technology. Specifically, speech recognition software analyzes the audio data and outputs the result as the text "What's the weather like today?".
[0842] Step 2:
[0843] The terminal sends the converted text data to the server. This transmission uses an HTTP request with an internet connection. The server checks the received text data and saves it to its data storage area. Specifically, the server accesses the database system and saves the text "What's the weather like today?".
[0844] Step 3:
[0845] The server inputs the stored text data into a generative AI model, which uses natural language processing techniques to analyze the meaning of the question. In this process, for example, the generative AI model analyzes the text, identifies that the user is asking a question about "weather," and organizes the relevant information as knowledge data. The output at this stage is the analyzed intent of the question and related information.
[0846] Step 4:
[0847] The server uses an emotion recognition engine to analyze the user's emotional state from their input. It identifies the emotional state based on the user's text content and tone. For example, if the input includes the word "tired," it detects stress and fatigue. This emotional information is then incorporated into subsequent response generation.
[0848] Step 5:
[0849] The server integrates knowledge data and sentiment information using a generative AI model to generate the optimal response. Specifically, based on information such as "the user is asking about the weather and is feeling stressed," it constructs a response like "It's sunny today. How about going for a walk?" This response is then sent to the device.
[0850] Step 6:
[0851] The terminal displays the response received from the server in its user interface. It also uses speech synthesis technology to read the response aloud as needed. Taking emotional elements into consideration, it provides guidance in a gentle tone, such as, "It's sunny today. Why not go for a walk?"
[0852] Step 7:
[0853] Users provide feedback on the answers they receive. This feedback is sent from the device to the server and stored to improve the system's accuracy. This feedback is used to fine-tune the AI model and emotion recognition engine, helping to improve the system.
[0854] (Application Example 2)
[0855] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0856] When users inquire about information, they are required to receive appropriate and prompt answers to their questions, and furthermore, to have those answers optimized to suit their emotional state. In particular, in the context of electronic payments, users often experience anxiety and stress, and it is necessary for the system to respond appropriately to these feelings.
[0857] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0858] In this invention, the server includes means for storing multiple types of media data input via an information processing device in a database, means for processing the media data and generating knowledge data to respond to human questions, and means for analyzing the user's emotional state and optimizing the response. This enables the provision of a fast, appropriate, and emotionally-driven response to the user.
[0859] An "information processing device" is a device that receives data input from a user and processes it in a predetermined format.
[0860] "Media data" refers to information expressed in multiple formats, such as text, audio, and images.
[0861] A "database" is a digital storage system used to store and manage information efficiently and systematically.
[0862] "Knowledge data" refers to information generated by processing raw data in order to produce appropriate answers to user input questions.
[0863] "Means for analyzing emotional states" refers to functions that recognize and evaluate emotions based on user input information and adjust responses accordingly.
[0864] "Means of optimizing responses" refer to methods for personalizing the content of responses according to the user's emotional state and needs, and providing them in the most optimal way.
[0865] "Feedback" refers to the collection of evaluations and opinions from users that are used to improve future processes.
[0866] The system implementing this invention comprises an information processing device, a database, multiple AI models, and a user interface. The server processes data input from the user in multiple formats, such as voice, text, and images. This includes transcribing voice data into text using a speech recognition library (e.g., speech recognition technology), analyzing text data using natural language processing technology, and storing it in the database.
[0867] Based on the information stored in the database, the server uses an AI model (e.g., a generative AI model) to generate answers to questions. During this process, an emotion analysis engine (e.g., emotion recognition technology) is used to analyze the user's emotional state and adjust the content and tone of the answers to match the user's emotions.
[0868] On the device, the generated response is displayed through the user interface and can also be read aloud. This allows users to obtain information through both sight and sound. User feedback is collected and recorded in a database to help improve the system. Based on this feedback, the accuracy of the AI model and sentiment analysis engine is improved.
[0869] As a concrete example, in an electronic payment scenario, if a user asks, "Is this payment secure?", the system would return the message, "You can use it with confidence. This transaction is highly secure." An example of a prompt used here would be, "Generate an appropriate response to the user's question and optimize the tone based on the emotion detected by emotion recognition technology."
[0870] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0871] Step 1:
[0872] The terminal receives the voice data entered by the user. This voice data is then converted into text data using speech recognition technology. This converted text data becomes the input for the next step.
[0873] Step 2:
[0874] The server receives text data sent from the terminal. Natural language processing techniques are used to analyze the meaning of the question. Based on the text data, relevant information is retrieved from the database to generate knowledge data. This knowledge data serves as input for the next step.
[0875] Step 3:
[0876] The server uses sentiment analysis technology to analyze the user's emotional state. This analysis includes calculations to identify emotions from the text content. Based on the analysis results, it generates personalized responses tailored to the user's emotional state. This generated response data becomes the input for the next step.
[0877] Step 4:
[0878] The server uses a generative AI model to concretize appropriate responses for the user based on knowledge data and sentiment analysis results. Here, prompts are used to instruct the AI model. Based on these instructions, the final response is output.
[0879] Step 5:
[0880] The device receives personalized responses sent from the server. These responses are displayed through the user interface and read aloud using speech synthesis technology. This process allows the user to obtain information both visually and aurally.
[0881] Step 6:
[0882] Users provide feedback on the answers provided by the system. The device collects this feedback and sends it to the server. This feedback data is used for the continuous improvement of the system and contributes to improving the accuracy of the generative AI model and sentiment analysis engine.
[0883] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0884] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0885] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0886] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0887] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0888] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0889] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0890] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0891] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0892] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0893] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0894] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0895] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0896] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0897] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0898] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0899] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0900] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0901] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0902] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0903] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0904] The following is further disclosed regarding the embodiments described above.
[0905] (Claim 1)
[0906] A means for storing multiple types of media data input via an information processing device into a database,
[0907] A means for processing the aforementioned media data and generating knowledge data to respond to human questions,
[0908] A means for generating an appropriate answer to a question received based on the generated knowledge data,
[0909] A means for obtaining user feedback and using it to improve the accuracy of the aforementioned knowledge data,
[0910] A system that includes this.
[0911] (Claim 2)
[0912] The system according to claim 1, comprising means for interpreting the question using natural language processing technology and retrieving relevant information from the media data.
[0913] (Claim 3)
[0914] The system according to claim 1, which has multilingual support capabilities and means for understanding questions in different languages and providing answers in the relevant languages.
[0915] "Example 1"
[0916] (Claim 1)
[0917] A means for converting data in multiple formats input via information processing means into a unified format and storing it in a database,
[0918] A means for analyzing the aforementioned data and generating knowledge to respond to human questions using an artificial intelligence model,
[0919] A means for deriving the optimal answer to a received question based on the knowledge generated by the aforementioned artificial intelligence model,
[0920] A means for obtaining feedback from users and sequentially improving the accuracy of the artificial intelligence model based on that feedback,
[0921] An information processing system that includes this.
[0922] (Claim 2)
[0923] The information processing system according to claim 1, comprising means for analyzing the question using natural language processing technology and extracting relevant knowledge from the database.
[0924] (Claim 3)
[0925] The information processing system according to claim 1, comprising means for supporting multiple languages, understanding questions in different languages, and providing answers in the corresponding languages.
[0926] "Application Example 1"
[0927] (Claim 1)
[0928] A means for storing multiple types of media data input via an information processing device in a data storage means,
[0929] Means for using techniques to process the aforementioned media data and generate knowledge data,
[0930] A means for generating an answer to a question received based on the generated knowledge data,
[0931] A means for obtaining user evaluations and using them to improve the accuracy of the aforementioned knowledge data,
[0932] A means of personalizedly suggesting content based on user interests through a user interface,
[0933] A system that includes this.
[0934] (Claim 2)
[0935] The system according to claim 1, comprising means for interpreting the question using natural language processing technology, searching for relevant information from the media data, and generating personalized suggestions.
[0936] (Claim 3)
[0937] The system according to claim 1, which has multilingual support capabilities and means for understanding questions in different languages and providing suggestions in the relevant languages.
[0938] "Example 2 of combining an emotion engine"
[0939] (Claim 1)
[0940] A means for storing multiple types of media data input via an information processing device in a data storage area,
[0941] A means for processing the aforementioned media data and generating knowledge data for responding to human questions using natural language processing technology,
[0942] A means for generating an appropriate answer to a question received based on the generated knowledge data,
[0943] A means for analyzing a user's emotional state using emotion recognition technology and optimizing responses according to that state,
[0944] A means for obtaining user feedback and using it to improve the accuracy of the aforementioned knowledge data and sentiment analysis,
[0945] A system that includes this.
[0946] (Claim 2)
[0947] The system according to claim 1, comprising means for analyzing user input using the aforementioned emotion recognition technology and personalizing the tone and content of the response to suit the user.
[0948] (Claim 3)
[0949] The system according to claim 1, which has multilingual support capabilities and means for understanding questions in different languages and providing answers in the relevant languages.
[0950] "Application example 2 when combining with an emotional engine"
[0951] (Claim 1)
[0952] A means for storing multiple types of media data input via an information processing device into a database,
[0953] A means for processing the aforementioned media data and generating knowledge data to respond to human questions,
[0954] A means for generating an appropriate answer to a question received based on the generated knowledge data,
[0955] A means to analyze the user's emotional state and optimize the response,
[0956] A means for obtaining user feedback and using it to improve the accuracy of the aforementioned knowledge data,
[0957] A system that includes this.
[0958] (Claim 2)
[0959] The system according to claim 1, comprising means for interpreting the question using natural language processing technology, retrieving relevant information from the media data, and further comprising means for personalizing the answer based on the user's emotions.
[0960] (Claim 3)
[0961] The system according to claim 1, which has multilingual support capabilities, means for understanding questions in different languages and providing answers in the relevant languages, and means for performing sentiment analysis on inputs in different languages. [Explanation of Symbols]
[0962] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for storing multiple types of media data input via an information processing device into a database, A means for processing the aforementioned media data and generating knowledge data to respond to human questions, A means for generating an appropriate answer to a question received based on the generated knowledge data, A means for obtaining user feedback and using it to improve the accuracy of the aforementioned knowledge data, A system that includes this.
2. The system according to claim 1, comprising means for interpreting the question using natural language processing technology and retrieving relevant information from the media data.
3. The system according to claim 1, which has multilingual support capabilities and means for understanding questions in different languages and providing answers in the relevant languages.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A