system

The system addresses the challenge of serving multilingual and technologically unfamiliar users by using natural language processing to generate personalized responses and collect feedback, enhancing user experience and business efficiency.

JP2026071725APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional systems struggle to provide personalized services to users who speak different languages or are unfamiliar with technology, leading to a decline in customer experience and business efficiency.

Method used

A system utilizing natural language processing to convert user input into speech or text data, support multiple languages, generate tailored responses, and collect feedback for continuous improvement.

Benefits of technology

Enables effective communication with diverse users, providing personalized and efficient services by understanding user intent, translating languages, and improving service quality through feedback analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071725000001_ABST
    Figure 2026071725000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for communicating using natural language processing, Means for converting user input into voice or text data, Means for supporting multiple languages, Means for retrieving information from a database, Means for generating a response based on the retrieved information, Means for presenting the generated response to the user, Means for collecting and recording user feedback, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, providing personalized services for diverse users is an important issue. In particular, the responses to users who speak different languages or elderly people who are unfamiliar with technology may not be adequately handled by conventional systems. This has led to a decline in the customer experience and a problem of limited business efficiency for service providers.

Means for Solving the Problems

[0005] This invention provides a system that uses natural language processing to communicate naturally with users and converts user input into speech or text data. This enables support for multiple languages ​​and allows the system to extract information from a database to generate responses tailored to user needs. Furthermore, by presenting the generated responses to the user in an intuitive manner and collecting and recording feedback, the system enables continuous improvement of service quality.

[0006] "Natural language processing" is a technology that enables computers to understand, generate, and respond to human language.

[0007] "User input" refers to data that indicates the information or instructions a user provides to the system.

[0008] "Audio or text data" refers to data obtained by converting an audio signal into a set of characters, or a data format composed of characters themselves.

[0009] "Supporting multiple languages" refers to having the ability to translate and understand different languages.

[0010] A "database" is a system that organizes and manages digital information, making it searchable and modifiable as needed.

[0011] "Searching for information" means extracting necessary data from a database according to specific criteria.

[0012] "Generating a response" refers to the process of creating an appropriate reply to input information.

[0013] "To present" means to provide information to the user visually or aurally.

[0014] "Collecting and recording feedback" refers to gathering user reactions and opinions and saving them for future evaluation and improvement.

Brief Description of the Drawings

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Modes for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a user-terminal-server type interactive system that utilizes natural language processing. This system is designed to provide comprehensive and personalized services to users who speak different languages ​​or who are unfamiliar with digital technologies.

[0037] The user inputs information into the device via voice or text. The device uses speech recognition technology to convert the input into text data and sends it to the server. The server analyzes this text data through a natural language processing engine to accurately understand the user's intent and requests. It also uses translation functions as needed to facilitate communication between different languages.

[0038] Based on the analysis results, the server searches the database for relevant information. For example, if a user is looking for nearby restaurants, the server collects local restaurant information based on their location. Next, it generates a response based on the searched information and sends it to the terminal.

[0039] The terminal receives this response and communicates it to the user through visual and auditory means. This allows the user to obtain the necessary information in real time. The user can also input feedback about the information provided. The terminal records this feedback and sends it to the server. The server records this as collected feedback data in a database and uses it to improve the system's services.

[0040] In this way, the present invention is a system that responds to the diverse needs of users and supports service providers in performing their duties more efficiently. For example, the present invention can be implemented in a wide range of applications, such as setting reminders to help elderly people remember their medication schedules, or providing interpretation services to travelers from different countries.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user enters information into the device via voice or text. In the case of voice input, the device uses speech recognition to convert the voice into text data.

[0044] Step 2:

[0045] The terminal sends user input data to the server. This data includes information in text format.

[0046] Step 3:

[0047] The server analyzes the received text data using a natural language processing engine to understand the user's intent and question. The data is translated as needed.

[0048] Step 4:

[0049] The server searches the database for relevant information based on the analysis results. For example, it collects information about the vicinity of the proposed location.

[0050] Step 5:

[0051] The server generates a response using the retrieved information and sends it to the terminal in text format.

[0052] Step 6:

[0053] The terminal presents the response received from the server to the user through visual or audio feedback.

[0054] Step 7:

[0055] The user generates feedback about the information provided and enters it into the device.

[0056] Step 8:

[0057] The device collects user feedback and sends it to the server.

[0058] Step 9:

[0059] The server records feedback in a database to help improve the service.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] The challenge lies in improving the efficiency and accuracy of information provision to multilingual users and users unfamiliar with digital technology, as well as effectively collecting and utilizing user feedback to improve system services.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for communicating using natural language processing, means for converting speech to text using speech recognition technology, and means for generating prompt sentences from the analysis results using a generative AI model. This enables efficient information exchange between diverse languages, accurate understanding of user intent, and effective collection of feedback for system improvement.

[0065] "Natural language processing" is a technology that uses computers to understand, analyze, and generate language that humans use on a daily basis.

[0066] "Means of communication" refers to a function that enables the sending and receiving of data and the exchange of information.

[0067] "Speech recognition technology" is a technology that converts voice input into text data.

[0068] "Means of converting to text data" refers to functions that convert audio or other data formats into textual information.

[0069] "Means of supporting multiple languages" refers to a function that has the ability to understand different languages ​​and translate / convert between them.

[0070] A "data set" is a general term for databases and records that collect and manage diverse information.

[0071] A "means of searching for information" is a function that finds necessary information from a data set based on specific conditions.

[0072] "Means for generating responses" refers to functions that create appropriate answers or information based on input data and search results.

[0073] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate and analyze text and data.

[0074] A "prompt" refers to an instruction or question given to a generative AI model.

[0075] "Means of ensuring security" refer to functions that maintain data safety and prevent unauthorized access and data leaks.

[0076] "Means for collecting and recording feedback" refers to a function that collects and saves user reactions and opinions to be used for future improvements.

[0077] This invention is an interactive information provision system that handles multiple languages ​​and is based on a three-tier structure consisting of a user, terminal, and server. Specific embodiments of this system are described below.

[0078] Terminal roles and functions

[0079] The user first inputs information into the terminal via voice or text. If voice input is used, the terminal receives the information through the microphone and utilizes speech recognition technology. This speech recognition includes the ability to convert speech into text data using a common API (e.g., a speech recognition API). If text input is provided, it is sent directly to the server. The terminal also displays the response from the server and provides the information to the user as voice through speech synthesis technology.

[0080] Server roles and functions

[0081] The server receives text data sent from the terminal and parses it via a natural language processing engine. During this process, a generative AI model is used to generate prompt sentences. For example, in response to a user requesting "I'm looking for a nearby Italian restaurant," the server might generate a prompt such as "Provide information on nearby Italian restaurants."

[0082] Based on the analysis, the server retrieves relevant information from a data set. This data set is managed in a common database system (e.g., a relational database), enabling efficient searching. Furthermore, the server supports multiple languages, utilizing translation functions as needed to tailor responses to each user's language. To ensure secure communication, security protocols (e.g., SSL / TLS) are used for data transmission during this process.

[0083] Gathering feedback and improving the system

[0084] Users can provide feedback on the information provided. The device receives the user's feedback and sends it to the server. The server collects this feedback data and records it in a data set. This recorded feedback is regularly analyzed to update the generative AI model and improve the quality of responses, leading to continuous improvement of the system.

[0085] This system will enable smoother information exchange and meet diverse user needs. For example, when tourists are looking for a restaurant at their travel destination, they can instantly obtain the most suitable information, overcoming language barriers.

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user inputs information into the device via voice or text. If the input is voice, they might speak into the microphone, for example, "Find a nearby Italian restaurant." The input data is either voice or text data. This data forms the basis for subsequent processing.

[0089] Step 2:

[0090] The terminal uses speech recognition technology to convert voice input into text data. Specifically, it uses a speech recognition API to analyze the audio waveform and generate the corresponding text. This output text is then sent to the server.

[0091] Step 3:

[0092] The terminal sends the converted text data to the server. The data is encrypted using a security protocol and transferred securely. This ensures that the user's request reaches the server accurately and securely.

[0093] Step 4:

[0094] The server inputs text data received from the terminal into a natural language processing engine. Using a generative AI model, it analyzes the user's intent from the text. This process generates prompt sentences, such as "Provide information on nearby Italian restaurants." Based on this analysis, a specific data search is performed.

[0095] Step 5:

[0096] The server searches the data set for relevant information based on the prompt message and the analysis results. It executes a database query to retrieve, for example, a list of Italian restaurants closest to the current location. This search result forms the core of the information to be provided to the user.

[0097] Step 6:

[0098] The server generates a response based on the retrieved information. Using a multilingual module, it formats the text to match the user's language and prepares it for visual and auditory presentation. This generated response is then sent to the terminal.

[0099] Step 7:

[0100] The terminal presents the received response to the user. It displays the information on the screen and plays the content aloud using speech synthesis technology. This allows the user to easily confirm the information.

[0101] Step 8:

[0102] Users provide feedback on the information provided. They input comments via text through their device, such as "This information was helpful" or "I'd like to know about other restaurants." This feedback becomes valuable data for improving the system.

[0103] Step 9:

[0104] The device sends user feedback to the server. The transmitted feedback data is recorded and analyzed on the server side for future improvements. Based on this feedback, the server takes measures to improve the quality of the service.

[0105] (Application Example 1)

[0106] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0107] In physical stores, customers often face challenges in quickly and appropriately obtaining information due to language barriers and unfamiliarity with technology. In particular, there is a lack of readily available means to guide foreign-speaking tourists and the elderly with product information and store services. Therefore, there is a need for ways to improve the customer experience in stores.

[0108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0109] In this invention, the server includes means for communicating using natural language processing, means for converting user input into voice or text data, and means for acquiring and providing product information in response to user questions using speech recognition in order to provide guidance within a physical store. This enables users to receive the information they need in real time, regardless of language, thereby improving the customer experience in physical stores.

[0110] "Natural language processing" is a technology that enables computers to understand, interpret, and generate human language.

[0111] "Speech recognition" is a technology that converts speech into digital data and interprets the content of that speech as textual information.

[0112] "Multilingual support" refers to the ability to process and provide information in multiple different languages.

[0113] A "database" is a structured digital collection of information designed to efficiently store, retrieve, and manage information.

[0114] "Feedback" refers to responses and evaluations from system users, and is data used to improve the service.

[0115] A "physical store" is a facility that exists in a physical location and is accessible to customers for sales or service provision.

[0116] "Product information" refers to data containing detailed descriptions and specifications about a product or service.

[0117] "Foreign language speakers" refer to people who speak a language other than their native language and receive information within that linguistic environment.

[0118] The system implementing this invention links a terminal installed in a physical store environment with the user's smart device and uses natural language processing technology to smoothly provide product information in response to the user's questions.

[0119] Applications on terminals or smart devices capture user voice input and convert it into text data using open-source or commercially available speech recognition software. Google® Cloud Speech-to-Text technology is preferably used for this purpose. The converted text data is transferred to a server in real time.

[0120] The server uses IBM Watson® Natural Language Understanding to analyze the content of the received text data and identify the user's intent. Based on the analysis results, it searches a database (e.g., Firebase Realtime Database) to retrieve relevant product information. The retrieved information is translated as needed via Microsoft® Translator Text API and adapted to the user's language.

[0121] Users can visually or audibly verify the information provided by their device. Multilingual support is available, allowing foreign language speakers to receive guidance in their own language. For example, if a tourist asks a question like, "I'm looking for traditional Japanese souvenirs," they will be provided with real-time information on store locations and products that match their intent.

[0122] Examples of prompt statements are as follows:

[0123] "The user has asked you for information about Japanese souvenirs in the store in English. Search the database for relevant information and answer in both English and Japanese."

[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0125] Step 1:

[0126] The user inputs a question by voice into the device's microphone. This voice input is converted into text data using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data. The audio waveform data is sampled, features are extracted, and the audio is documented based on the results.

[0127] Step 2:

[0128] The terminal sends the obtained text data to the server. The server receives this text data and processes it using a natural language processing engine (e.g., IBM Watson Natural Language Understanding) for analysis. The input is text data, and the output is structured data that indicates the user's intent. The server analyzes the document's structure, keywords, and context to interpret the user's intent.

[0129] Step 3:

[0130] The server searches the database based on the analysis results. The database contains product and store information, and extracts information that matches the user's intent. The input is structured intent data, and the output is related information data. The server generates queries and retrieves the necessary entries from the database.

[0131] Step 4:

[0132] If necessary, the server translates the extracted information into the user's native language. A translation API (e.g., Microsoft Translator Text API) is used for translation. The input is informational data, and the output is translated informational data. The server sends the informational data to the translation engine to generate text in the new language.

[0133] Step 5:

[0134] The server sends the translated information back to the terminal. After receiving this information, the terminal provides it to the user visually or audibly. The input is the translated information data, and the output is a display or audio output to the user. The terminal displays the information on the screen or plays synthesized speech through its speaker.

[0135] Step 6:

[0136] Users provide feedback on the information presented. This feedback is sent from the terminal to the server, which records its contents in a database. The input is the feedback data, and the output is the recorded feedback data. The server receives the feedback, stores it as evaluation information, and uses it to improve future processes.

[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0138] This invention is a user-terminal-server type interactive system that combines natural language processing and an emotion engine. This system is designed to understand not only the words but also the emotions behind them in response to user input, and to generate more appropriate and personalized responses.

[0139] Users input information into the device via voice or text. The device then uses speech recognition technology to convert the voice into text data, which is then sent to an emotion engine for analysis. The emotion engine extracts emotional information from the user's voice tone and text wording.

[0140] The terminal sends the converted and analyzed data to the server. The server analyzes the text data using a natural language processing engine to understand the user's intent. Based on the results, it searches the database and retrieves relevant information.

[0141] The server generates an appropriate response based on the acquired information and emotional information, and sends it to the terminal. For example, if the user is dissatisfied, it will generate a response that provides reassurance.

[0142] The terminal presents the response received from the server to the user using visual or audio feedback. In doing so, it adjusts the tone and expression of the feedback based on emotional information.

[0143] Furthermore, when users input feedback into their devices, data is collected for the purpose of improving the service. The server records this feedback in a database and uses it to continuously improve the system, including sentiment analysis.

[0144] In this way, the present invention is a system that achieves more human-like interaction by taking user emotions into consideration. For example, the present invention can be implemented in customer support, such as providing a friendly response that quickly resolves problems when it is determined that a customer is feeling stressed.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The user enters information into the device via voice or text. This input includes user questions and requests.

[0148] Step 2:

[0149] When the device receives voice input, it uses its speech recognition function to convert the speech into text data. At the same time, it also records characteristics such as the tone and speed of the speech.

[0150] Step 3:

[0151] The device passes the converted text data to an emotion engine, which then recognizes emotions from the input. For example, it can determine whether the user is angry or happy based on the wording and voice features of the input.

[0152] Step 4:

[0153] The terminal sends the sentiment analysis results along with text data to the server. The server receives this data and prepares it for analysis.

[0154] Step 5:

[0155] The server analyzes the received data using a natural language processing engine to understand the user's intent. For example, it identifies what kind of information is needed.

[0156] Step 6:

[0157] Based on the analysis results, the server searches the database for relevant information and extracts the necessary data.

[0158] Step 7:

[0159] The server uses the retrieved information and sentiment analysis results to generate a response for the user. For example, if it determines that the user is dissatisfied, it will include reassuring language.

[0160] Step 8:

[0161] The server sends the generated response to the terminal. The terminal receives it and prepares to present it to the user.

[0162] Step 9:

[0163] The device presents responses to the user visually or audibly. During presentation, it adjusts the tone and expression based on the sentiment analysis results.

[0164] Step 10:

[0165] Users can provide feedback on the information presented. The device receives this feedback and sends it to the server.

[0166] Step 11:

[0167] The server records feedback data in a database and uses it for analysis to improve the service. This includes sentiment data.

[0168] (Example 2)

[0169] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0170] Conventional technologies only provide linguistic responses to user input, making it difficult to deeply understand the user's emotional state and intentions. Furthermore, they lacked the means to generate responses tailored to the user's emotions, resulting in insufficient means of achieving personalized interaction. Additionally, the effective collection and utilization of user feedback was challenging, posing a challenge to continuous system improvement.

[0171] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0172] In this invention, the server includes means for analyzing information using natural language processing functions, means for extracting data from information resources, and a module for generating individual responses based on the extracted information. This makes it possible to generate more human-like and personalized responses that take into account the user's emotions and intentions, thereby achieving effective interaction with the user.

[0173] "Natural language processing" refers to the technology that enables computers to understand human language and analyze its meaning.

[0174] A "processing unit for converting to audio or text data" refers to hardware and software components for converting data into different formats, such as converting user voice input to text.

[0175] An "interface for identifying multiple languages" is a function that accurately identifies different language inputs and performs corresponding processing.

[0176] "Means of extracting data from information resources" refers to a function or process for obtaining necessary information from existing databases, etc.

[0177] A "module that generates individual responses" is a component that creates responses tailored to the user's specific conditions and needs based on the information it has acquired.

[0178] A "display device" is a device that presents a generated response to the user visually or audibly.

[0179] "Means for analyzing emotional information and reflecting it in response generation" refers to the processes and tools for identifying a user's emotions and reflecting the results in the response.

[0180] A "feedback collection and storage device" is a device designed to efficiently gather user reactions and opinions so that they can be analyzed later.

[0181] "Methods for analyzing saved feedback and using it to improve the system" refers to the process of analyzing recorded feedback data to improve the functionality and usability of the system.

[0182] The present invention is a system for interacting with a user through voice or text input, understanding the user's emotions, and generating responses. Its embodiments are described in detail below.

[0183] The process begins with the user inputting information into the device via voice or text. In the case of voice input, the device's built-in microphone is used to capture the voice. Next, the device applies speech recognition software to convert the voice data into text data. This process can utilize, for example, a common speech recognition API or open-source or commercial speech recognition technology.

[0184] The input text data is sent to an engine equipped with sentiment analysis capabilities. This engine uses natural language processing techniques to extract emotional information from the user's input. In this process, the terminal, for example, utilizes publicly available natural language processing libraries to analyze the user's intentions and emotions.

[0185] Sentimental information and text data are transferred from the terminal to the server, which then generates a response based on this data. The data is analyzed by a natural language processing model to identify the information and actions the user is seeking. In this process, for example, generative AI models are used to perform advanced language analysis.

[0186] The server then extracts relevant information from the database and generates a response based on the sentiment analysis results. The generated response is sent to the terminal and presented to the user as audio or text. The terminal adjusts the tone and expression presented to match the user's emotions, enabling more natural communication.

[0187] Such systems have diverse applications, including customer support. For example, they can provide appropriate responses to customers who are experiencing stress.

[0188] A concrete example of a prompt message is that it can be input into a generative AI model in the form of, "Please tell me how to generate an appropriate response when the user is excited."

[0189] Thus, the present invention is an efficient system that achieves human-like interaction by analyzing the user's emotions and reflecting them in the response.

[0190] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0191] Step 1:

[0192] The user inputs information into the device. Input methods include voice or text. In the case of voice input, the device captures the voice using a microphone. Speech recognition is performed based on the input voice data, and this is converted into text data. The output of speech recognition is text data.

[0193] Step 2:

[0194] The terminal sends the converted text data to an emotion analysis unit. This unit uses natural language processing techniques to extract emotions from the text data. The input data is text data, and the output is data representing the user's emotional state. This analysis is achieved by utilizing publicly available emotion analysis libraries.

[0195] Step 3:

[0196] The device sends the analyzed sentiment data and the original text data to the server. The server analyzes the received data using a natural language processing engine to recognize the user's intent. This analysis process identifies the information necessary to provide the desired response from the input text and sentiment information.

[0197] Step 4:

[0198] The server searches the database based on the recognized intent and retrieves relevant information. In this step, the input is data about the user's intent, and the output is relevant database information.

[0199] Step 5:

[0200] The server uses acquired information and sentiment information to generate an appropriate response. It performs advanced language generation using a generative AI model. The input data consists of relevant and sentiment information, and the output is the generated response.

[0201] Step 6:

[0202] The device presents the generated response to the user. Visual output is provided through a screen display, and auditory output through speech synthesis. The tone and expression of the response are adjusted based on the user's emotional information. The user then decides on their next action.

[0203] Step 7:

[0204] The device records the feedback provided by the user and sends it to the server. This feedback is stored as data used to improve the system in the future. The server analyzes this information and uses it to improve the system's functionality.

[0205] (Application Example 2)

[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0207] In current interactive systems, generating appropriate responses that take user emotions into account is difficult. Furthermore, there is a lack of adequate responses in emergencies, and the means to provide users with a sense of security are limited. As a result, users may experience inadequate communication and support. Therefore, there is a need to develop methods for generating responses based on the user's emotional state and providing rapid responses in emergencies.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0209] In this invention, the server includes means for communicating using natural language processing, means for analyzing the user's emotional state, and means for adjusting the tone of the response based on the user's emotional state. This makes it possible to provide the user with emotionally sensitive, personalized responses, and furthermore, to provide rapid support in emergencies.

[0210] "Natural language processing" is the technology that enables computers to understand and process human language.

[0211] "Means of communication" refers to technologies for sending and receiving information, such as the internet and wireless communication.

[0212] "Means of converting user input into speech or text data" refers to technology that converts information provided by a user into a format that can be processed by a machine.

[0213] "Means of supporting multiple languages" refers to technologies that enable information processing in different languages.

[0214] "Methods for searching information from a database" refer to technologies for finding relevant information from stored data.

[0215] "Means for generating responses based on retrieved information" refers to technologies that create responses for users based on acquired data.

[0216] "Means of presenting the generated response to the user" refers to technologies that display or communicate the answer generated by the system to the user.

[0217] "Methods for analyzing a user's emotional state" refer to technologies that identify emotions from a user's tone of voice and word choice.

[0218] "Means of adjusting the tone of response" refers to techniques that change the way a response is expressed to match the user's emotions.

[0219] "Means for collecting and recording user feedback" refers to technologies that collect user opinions and reactions and store them in a format that can be used later.

[0220] "Means of providing rapid support in emergencies" refers to technologies that enable a rapid response to emergency situations and provide necessary assistance.

[0221] This invention is an interactive system that takes user emotions into account and is comprised of a combination of an emotion analysis engine and a natural language processing engine. The system processes user voice and text input and provides appropriate responses.

[0222] First, the user inputs information into a device such as a smartphone via voice or text. The device has built-in speech recognition software, which uses libraries such as a "speech recognition API" to convert the voice into text data. This converted text data is then sent to a server located in the cloud.

[0223] The server uses an "emotion analysis engine" to analyze the emotional state from the wording and nuances of the voice in the received text data. For example, an emotion analysis tool like "Affectiva" is used in this process. Once the analysis results are obtained, the data is then sent to a "natural language processing engine" where the user's intentions are analyzed in detail.

[0224] Based on this analysis and emotional information, the server searches the database and retrieves relevant information. Based on the retrieved information, it generates a response tailored to the user's emotional state and sends it to the terminal. In this process, the "Natural Language Processing API" is used to generate the text.

[0225] For example, if a user is feeling anxious during an emergency, the server can quickly generate the most appropriate support information nearby and provide reassurance to the user.

[0226] A concrete example to understand how this system works is when a user mutters "I'm scared" while walking alone at night and enters it into the application. In response to this input, the system will present the user with information about safe places to alleviate their anxiety.

[0227] Examples of prompt statements include the following:

[0228] "Users feel uneasy walking alone at night. Please consider ways to reassure users who are expressing fear."

[0229] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0230] Step 1:

[0231] The device receives voice or text input from the user. Once this input is received, if it's voice, it's converted into text data using a "speech recognition API." This transforms the raw voice input into concrete text output.

[0232] Step 2:

[0233] The terminal sends the converted text data to the server. The server prepares to process the received data and passes it to the sentiment analysis engine for the next analysis step. It receives text data from the terminal as input and formats it into a parseable data format as output.

[0234] Step 3:

[0235] The server uses an emotion analysis engine to analyze the emotional state of text data. For example, it might use "Affectiva" to extract emotional information based on the nuances and tone of the input text. The input is text data, and the output is metadata indicating the user's emotional state.

[0236] Step 4:

[0237] The server uses a natural language processing engine to analyze text data containing emotional information in detail. Here, natural language processing APIs are used to process the data and extract the user's intent. Based on the text data and emotional metadata obtained as input, the server outputs the user's specific intent.

[0238] Step 5:

[0239] The server searches the database based on the analysis results and collects information related to the user's emotions and intentions. The input is the analysis results, and the output is the relevant information to be presented to the user. A rapid search algorithm works in conjunction with this process.

[0240] Step 6:

[0241] The server generates the optimal response based on the collected information and emotional state. By using a generative AI model, flexible responses are possible. The input consists of relevant information and emotional state, while the output is a response text that matches the emotional state.

[0242] Step 7:

[0243] The terminal presents the user with responses obtained from the server. It provides visual or auditory feedback and, as a termination process, offers information in a tone appropriate to the user's emotions. The input is response data from the server, and the output is the display or audio as feedback to the user.

[0244] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0245] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0246] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0247] [Second Embodiment]

[0248] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0249] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0250] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0251] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0252] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0253] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0254] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0255] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0256] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0257] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0258] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0259] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0260] This invention is a user-terminal-server type interactive system that utilizes natural language processing. This system is designed to provide comprehensive and personalized services to users who speak different languages ​​or who are unfamiliar with digital technologies.

[0261] The user inputs information into the device via voice or text. The device uses speech recognition technology to convert the input into text data and sends it to the server. The server analyzes this text data through a natural language processing engine to accurately understand the user's intent and requests. It also uses translation functions as needed to facilitate communication between different languages.

[0262] Based on the analysis results, the server searches the database for relevant information. For example, if a user is looking for nearby restaurants, the server collects local restaurant information based on their location. Next, it generates a response based on the searched information and sends it to the terminal.

[0263] The terminal receives this response and communicates it to the user through visual and auditory means. This allows the user to obtain the necessary information in real time. The user can also input feedback about the information provided. The terminal records this feedback and sends it to the server. The server records this as collected feedback data in a database and uses it to improve the system's services.

[0264] In this way, the present invention is a system that responds to the diverse needs of users and supports service providers in performing their duties more efficiently. For example, the present invention can be implemented in a wide range of applications, such as setting reminders to help elderly people remember their medication schedules, or providing interpretation services to travelers from different countries.

[0265] The following describes the processing flow.

[0266] Step 1:

[0267] The user enters information into the device via voice or text. In the case of voice input, the device uses speech recognition to convert the voice into text data.

[0268] Step 2:

[0269] The terminal sends the input data from the user to the server. This data includes information in text format.

[0270] Step 3:

[0271] The server analyzes the received text data with a natural language processing engine to understand the user's intention and the content of the question. If necessary, the data is translated.

[0272] Step 4:

[0273] The server searches the database for relevant information based on the analysis results. For example, it collects information about the vicinity of the proposed location.

[0274] Step 5:

[0275] The server uses the retrieved information to generate a response and sends it to the terminal in text format.

[0276] Step 6:

[0277] The terminal presents the response received from the server to the user through visual or audio feedback.

[0278] Step 7:

[0279] The user generates feedback about the provided information and inputs it into the terminal.

[0280] Step 8:

[0281] The terminal collects the feedback from the user and sends it to the server.

[0282] Step 9:

[0283] The server records the feedback in the database and uses it to improve the service.

[0284] (Example 1)

[0285] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0286] It is an issue to improve the efficiency and accuracy of information provision for users who speak multiple languages or are unfamiliar with digital technology, and to effectively collect and utilize user feedback to improve the system's services.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0288] In this invention, the server includes means for communicating using natural language processing, means for converting speech to text using speech recognition technology, and means for generating a prompt sentence from an analysis result using a generative AI model. Thereby, it becomes possible to improve the efficiency of information exchange between various languages, accurately understand user intentions, effectively collect feedback, and improve the system.

[0289] "Natural language processing" is a technology for a computer to understand, analyze, and generate the language that humans use in daily life.

[0290] "Means for communicating" is a function for transmitting and receiving data and enabling information exchange.

[0291] "Speech recognition technology" is a technology for converting speech input into text data.

[0292] "Means for converting to text data" is a function for converting speech and other forms of data into character information.

[0293] "Means for supporting multiple languages" is a function having the ability to understand different languages and perform translation and conversion.

[0294] "Data set" is a general term for databases and records that collect and manage various information.

[0295] A "means of searching for information" is a function that finds necessary information from a data set based on specific conditions.

[0296] "Means for generating responses" refers to functions that create appropriate answers or information based on input data and search results.

[0297] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate and analyze text and data.

[0298] A "prompt" refers to an instruction or question given to a generative AI model.

[0299] "Means of ensuring security" refer to functions that maintain data safety and prevent unauthorized access and data leaks.

[0300] "Means for collecting and recording feedback" refers to a function that collects and saves user reactions and opinions to be used for future improvements.

[0301] This invention is an interactive information provision system that handles multiple languages ​​and is based on a three-tier structure consisting of a user, terminal, and server. Specific embodiments of this system are described below.

[0302] Terminal roles and functions

[0303] The user first inputs information into the terminal via voice or text. If voice input is used, the terminal receives the information through the microphone and utilizes speech recognition technology. This speech recognition includes the ability to convert speech into text data using a common API (e.g., a speech recognition API). If text input is provided, it is sent directly to the server. The terminal also displays the response from the server and provides the information to the user as voice through speech synthesis technology.

[0304] Server roles and functions

[0305] The server receives the text data sent from the terminal and analyzes it via a natural language processing engine. At this time, a prompt sentence is generated using a generative AI model. For example, for a user request such as "looking for a nearby Italian restaurant", a prompt such as "Provide information on nearby Italian restaurants" is generated.

[0306] Based on the analysis, the server searches for relevant information from the data set. This data set is managed by a general database system (e.g., a relational database), enabling efficient searches. Also, for multi-language support, the server utilizes a translation function as needed to tailor the response to each user's language. In this process, a security protocol (e.g., SSL / TLS) is used for data transmission to ensure secure communication.

[0307] Collection of Feedback and Improvement of the System

[0308] Users can provide feedback on the information provided. The terminal receives the user's feedback and sends it to the server. The server collects this feedback data and records it in the data set. This recorded feedback is regularly analyzed to update the generative AI model and improve the quality of responses, leading to continuous improvement of the system.

[0309] This system can accommodate various user needs and enable smoother information exchange. For example, when a tourist is looking for a restaurant at their travel destination, they can instantly obtain optimal information across language barriers.

[0310] The flow of specific processing in Example 1 will be described using FIG. 11.

[0311] Step 1:

[0312] The user inputs information into the device via voice or text. If the input is voice, they might speak into the microphone, for example, "Find a nearby Italian restaurant." The input data is either voice or text data. This data forms the basis for subsequent processing.

[0313] Step 2:

[0314] The terminal uses speech recognition technology to convert voice input into text data. Specifically, it uses a speech recognition API to analyze the audio waveform and generate the corresponding text. This output text is then sent to the server.

[0315] Step 3:

[0316] The terminal sends the converted text data to the server. The data is encrypted using a security protocol and transferred securely. This ensures that the user's request reaches the server accurately and securely.

[0317] Step 4:

[0318] The server inputs text data received from the terminal into a natural language processing engine. Using a generative AI model, it analyzes the user's intent from the text. This process generates prompt sentences, such as "Provide information on nearby Italian restaurants." Based on this analysis, a specific data search is performed.

[0319] Step 5:

[0320] The server searches the data set for relevant information based on the prompt message and the analysis results. It executes a database query to retrieve, for example, a list of Italian restaurants closest to the current location. This search result forms the core of the information to be provided to the user.

[0321] Step 6:

[0322] The server generates a response based on the retrieved information. Using a multilingual module, it formats the text to match the user's language and prepares it for visual and auditory presentation. This generated response is then sent to the terminal.

[0323] Step 7:

[0324] The terminal presents the received response to the user. It displays the information on the screen and plays the content aloud using speech synthesis technology. This allows the user to easily confirm the information.

[0325] Step 8:

[0326] Users provide feedback on the information provided. They input comments via text through their device, such as "This information was helpful" or "I'd like to know about other restaurants." This feedback becomes valuable data for improving the system.

[0327] Step 9:

[0328] The device sends user feedback to the server. The transmitted feedback data is recorded and analyzed on the server side for future improvements. Based on this feedback, the server takes measures to improve the quality of the service.

[0329] (Application Example 1)

[0330] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0331] In physical stores, customers often face challenges in quickly and appropriately obtaining information due to language barriers and unfamiliarity with technology. In particular, there is a lack of readily available means to guide foreign-speaking tourists and the elderly with product information and store services. Therefore, there is a need for ways to improve the customer experience in stores.

[0332] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0333] In this invention, the server includes means for communicating using natural language processing, means for converting user input into voice or text data, and means for acquiring and providing product information in response to user questions using speech recognition in order to provide guidance within a physical store. This enables users to receive the information they need in real time, regardless of language, thereby improving the customer experience in physical stores.

[0334] "Natural language processing" is a technology that enables computers to understand, interpret, and generate human language.

[0335] "Speech recognition" is a technology that converts speech into digital data and interprets the content of that speech as textual information.

[0336] "Multilingual support" refers to the ability to process and provide information in multiple different languages.

[0337] A "database" is a structured digital collection of information designed to efficiently store, retrieve, and manage information.

[0338] "Feedback" refers to responses and evaluations from system users, and is data used to improve the service.

[0339] A "physical store" is a facility that exists in a physical location and is accessible to customers for sales or service provision.

[0340] "Product information" refers to data containing detailed descriptions and specifications about a product or service.

[0341] "Foreign language speakers" refer to people who speak a language other than their native language and receive information within that linguistic environment.

[0342] The system implementing this invention links a terminal installed in a physical store environment with the user's smart device and uses natural language processing technology to smoothly provide product information in response to the user's questions.

[0343] Applications on terminals or smart devices capture user voice input and convert it into text data using open-source or commercially available speech recognition software. Google Cloud Speech-to-Text technology is preferably used for this purpose. The converted text data is transferred to a server in real time.

[0344] The server uses IBM Watson Natural Language Understanding to analyze the content of the received text data and identify the user's intent. Based on the analysis results, it searches a database (e.g., Firebase Realtime Database) to retrieve relevant product information. The retrieved information is translated as needed via the Microsoft Translator Text API and adapted to the user's language.

[0345] Users can visually or audibly verify the information provided by their device. Multilingual support is available, allowing foreign language speakers to receive guidance in their own language. For example, if a tourist asks a question like, "I'm looking for traditional Japanese souvenirs," they will be provided with real-time information on store locations and products that match their intent.

[0346] Examples of prompt statements are as follows:

[0347] "The user has asked you for information about Japanese souvenirs in the store in English. Search the database for relevant information and answer in both English and Japanese."

[0348] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0349] Step 1:

[0350] The user inputs a question by voice into the device's microphone. This voice input is converted into text data using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data. The audio waveform data is sampled, features are extracted, and the audio is documented based on the results.

[0351] Step 2:

[0352] The terminal sends the obtained text data to the server. The server receives this text data and processes it using a natural language processing engine (e.g., IBM Watson Natural Language Understanding) for analysis. The input is text data, and the output is structured data that indicates the user's intent. The server analyzes the document's structure, keywords, and context to interpret the user's intent.

[0353] Step 3:

[0354] The server searches the database based on the analysis results. The database contains product and store information, and extracts information that matches the user's intent. The input is structured intent data, and the output is related information data. The server generates queries and retrieves the necessary entries from the database.

[0355] Step 4:

[0356] If necessary, the server translates the extracted information into the user's native language. A translation API (e.g., Microsoft Translator Text API) is used for translation. The input is informational data, and the output is translated informational data. The server sends the informational data to the translation engine to generate text in the new language.

[0357] Step 5:

[0358] The server sends the translated information back to the terminal. After receiving this information, the terminal provides it to the user visually or audibly. The input is the translated information data, and the output is a display or audio output to the user. The terminal displays the information on the screen or plays synthesized speech through its speaker.

[0359] Step 6:

[0360] Users provide feedback on the information presented. This feedback is sent from the terminal to the server, which records its contents in a database. The input is the feedback data, and the output is the recorded feedback data. The server receives the feedback, stores it as evaluation information, and uses it to improve future processes.

[0361] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0362] This invention is a user-terminal-server type interactive system that combines natural language processing and an emotion engine. This system is designed to understand not only the words but also the emotions behind them in response to user input, and to generate more appropriate and personalized responses.

[0363] Users input information into the device via voice or text. The device then uses speech recognition technology to convert the voice into text data, which is then sent to an emotion engine for analysis. The emotion engine extracts emotional information from the user's voice tone and text wording.

[0364] The terminal sends the converted and analyzed data to the server. The server analyzes the text data using a natural language processing engine to understand the user's intent. Based on the results, it searches the database and retrieves relevant information.

[0365] The server generates an appropriate response based on the acquired information and emotional information, and sends it to the terminal. For example, if the user is dissatisfied, it will generate a response that provides reassurance.

[0366] The terminal presents the response received from the server to the user using visual or audio feedback. In doing so, it adjusts the tone and expression of the feedback based on emotional information.

[0367] Furthermore, when users input feedback into their devices, data is collected for the purpose of improving the service. The server records this feedback in a database and uses it to continuously improve the system, including sentiment analysis.

[0368] In this way, the present invention is a system that achieves more human-like interaction by taking user emotions into consideration. For example, the present invention can be implemented in customer support, such as providing a friendly response that quickly resolves problems when it is determined that a customer is feeling stressed.

[0369] The following describes the processing flow.

[0370] Step 1:

[0371] The user enters information into the device via voice or text. This input includes user questions and requests.

[0372] Step 2:

[0373] When the device receives voice input, it uses its speech recognition function to convert the speech into text data. At the same time, it also records characteristics such as the tone and speed of the speech.

[0374] Step 3:

[0375] The device passes the converted text data to an emotion engine, which then recognizes emotions from the input. For example, it can determine whether the user is angry or happy based on the wording and voice features of the input.

[0376] Step 4:

[0377] The terminal sends the sentiment analysis results along with text data to the server. The server receives this data and prepares it for analysis.

[0378] Step 5:

[0379] The server analyzes the received data using a natural language processing engine to understand the user's intent. For example, it identifies what kind of information is needed.

[0380] Step 6:

[0381] Based on the analysis results, the server searches the database for relevant information and extracts the necessary data.

[0382] Step 7:

[0383] The server uses the retrieved information and sentiment analysis results to generate a response for the user. For example, if it determines that the user is dissatisfied, it will include reassuring language.

[0384] Step 8:

[0385] The server sends the generated response to the terminal. The terminal receives it and prepares to present it to the user.

[0386] Step 9:

[0387] The device presents responses to the user visually or audibly. During presentation, it adjusts the tone and expression based on the sentiment analysis results.

[0388] Step 10:

[0389] Users can provide feedback on the information presented. The device receives this feedback and sends it to the server.

[0390] Step 11:

[0391] The server records feedback data in a database and uses it for analysis to improve the service. This includes sentiment data.

[0392] (Example 2)

[0393] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0394] Conventional technologies only provide linguistic responses to user input, making it difficult to deeply understand the user's emotional state and intentions. Furthermore, they lacked the means to generate responses tailored to the user's emotions, resulting in insufficient means of achieving personalized interaction. Additionally, the effective collection and utilization of user feedback was challenging, posing a challenge to continuous system improvement.

[0395] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0396] In this invention, the server includes means for analyzing information using natural language processing functions, means for extracting data from information resources, and a module for generating individual responses based on the extracted information. This makes it possible to generate more human-like and personalized responses that take into account the user's emotions and intentions, thereby achieving effective interaction with the user.

[0397] "Natural language processing" refers to the technology that enables computers to understand human language and analyze its meaning.

[0398] A "processing unit for converting to audio or text data" refers to hardware and software components for converting data into different formats, such as converting user voice input to text.

[0399] An "interface for identifying multiple languages" is a function that accurately identifies different language inputs and performs corresponding processing.

[0400] "Means of extracting data from information resources" refers to a function or process for obtaining necessary information from existing databases, etc.

[0401] A "module that generates individual responses" is a component that creates responses tailored to the user's specific conditions and needs based on the information it has acquired.

[0402] A "display device" is a device that presents a generated response to the user visually or audibly.

[0403] "Means for analyzing emotional information and reflecting it in response generation" refers to the processes and tools for identifying a user's emotions and reflecting the results in the response.

[0404] A "feedback collection and storage device" is a device designed to efficiently gather user reactions and opinions so that they can be analyzed later.

[0405] "Methods for analyzing saved feedback and using it to improve the system" refers to the process of analyzing recorded feedback data to improve the functionality and usability of the system.

[0406] The present invention is a system for interacting with a user through voice or text input, understanding the user's emotions, and generating responses. Its embodiments are described in detail below.

[0407] The process begins with the user inputting information into the device via voice or text. In the case of voice input, the device's built-in microphone is used to capture the voice. Next, the device applies speech recognition software to convert the voice data into text data. This process can utilize, for example, a common speech recognition API or open-source or commercial speech recognition technology.

[0408] The input text data is sent to an engine equipped with sentiment analysis capabilities. This engine uses natural language processing techniques to extract emotional information from the user's input. In this process, the terminal, for example, utilizes publicly available natural language processing libraries to analyze the user's intentions and emotions.

[0409] Sentimental information and text data are transferred from the terminal to the server, which then generates a response based on this data. The data is analyzed by a natural language processing model to identify the information and actions the user is seeking. In this process, for example, generative AI models are used to perform advanced language analysis.

[0410] The server then extracts relevant information from the database and generates a response based on the sentiment analysis results. The generated response is sent to the terminal and presented to the user as audio or text. The terminal adjusts the tone and expression presented to match the user's emotions, enabling more natural communication.

[0411] Such systems have diverse applications, including customer support. For example, they can provide appropriate responses to customers who are experiencing stress.

[0412] A concrete example of a prompt message is that it can be input into a generative AI model in the form of, "Please tell me how to generate an appropriate response when the user is excited."

[0413] Thus, the present invention is an efficient system that achieves human-like interaction by analyzing the user's emotions and reflecting them in the response.

[0414] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0415] Step 1:

[0416] The user inputs information into the device. Input methods include voice or text. In the case of voice input, the device captures the voice using a microphone. Speech recognition is performed based on the input voice data, and this is converted into text data. The output of speech recognition is text data.

[0417] Step 2:

[0418] The terminal sends the converted text data to an emotion analysis unit. This unit uses natural language processing techniques to extract emotions from the text data. The input data is text data, and the output is data representing the user's emotional state. This analysis is achieved by utilizing publicly available emotion analysis libraries.

[0419] Step 3:

[0420] The device sends the analyzed sentiment data and the original text data to the server. The server analyzes the received data using a natural language processing engine to recognize the user's intent. This analysis process identifies the information necessary to provide the desired response from the input text and sentiment information.

[0421] Step 4:

[0422] The server searches the database based on the recognized intent and retrieves relevant information. In this step, the input is data about the user's intent, and the output is relevant database information.

[0423] Step 5:

[0424] The server uses acquired information and sentiment information to generate an appropriate response. It performs advanced language generation using a generative AI model. The input data consists of relevant and sentiment information, and the output is the generated response.

[0425] Step 6:

[0426] The device presents the generated response to the user. Visual output is provided through a screen display, and auditory output through speech synthesis. The tone and expression of the response are adjusted based on the user's emotional information. The user then decides on their next action.

[0427] Step 7:

[0428] The device records the feedback provided by the user and sends it to the server. This feedback is stored as data used to improve the system in the future. The server analyzes this information and uses it to improve the system's functionality.

[0429] (Application Example 2)

[0430] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0431] In current interactive systems, generating appropriate responses that take user emotions into account is difficult. Furthermore, there is a lack of adequate responses in emergencies, and the means to provide users with a sense of security are limited. As a result, users may experience inadequate communication and support. Therefore, there is a need to develop methods for generating responses based on the user's emotional state and providing rapid responses in emergencies.

[0432] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0433] In this invention, the server includes means for communicating using natural language processing, means for analyzing the user's emotional state, and means for adjusting the tone of the response based on the user's emotional state. This makes it possible to provide the user with emotionally sensitive, personalized responses, and furthermore, to provide rapid support in emergencies.

[0434] "Natural language processing" is the technology that enables computers to understand and process human language.

[0435] "Means of communication" refers to technologies for sending and receiving information, such as the internet and wireless communication.

[0436] "Means of converting user input into speech or text data" refers to technology that converts information provided by a user into a format that can be processed by a machine.

[0437] "Means of supporting multiple languages" refers to technologies that enable information processing in different languages.

[0438] "Methods for searching information from a database" refer to technologies for finding relevant information from stored data.

[0439] "Means for generating responses based on retrieved information" refers to technologies that create responses for users based on acquired data.

[0440] "Means of presenting the generated response to the user" refers to technologies that display or communicate the answer generated by the system to the user.

[0441] "Methods for analyzing a user's emotional state" refer to technologies that identify emotions from a user's tone of voice and word choice.

[0442] "Means of adjusting the tone of response" refers to techniques that change the way a response is expressed to match the user's emotions.

[0443] "Means for collecting and recording user feedback" refers to technologies that collect user opinions and reactions and store them in a format that can be used later.

[0444] "Means of providing rapid support in emergencies" refers to technologies that enable a rapid response to emergency situations and provide necessary assistance.

[0445] This invention is an interactive system that takes user emotions into account and is comprised of a combination of an emotion analysis engine and a natural language processing engine. The system processes user voice and text input and provides appropriate responses.

[0446] First, the user inputs information into a device such as a smartphone via voice or text. The device has built-in speech recognition software, which uses libraries such as a "speech recognition API" to convert the voice into text data. This converted text data is then sent to a server located in the cloud.

[0447] The server uses an "emotion analysis engine" to analyze the emotional state from the wording and nuances of the voice in the received text data. For example, an emotion analysis tool like "Affectiva" is used in this process. Once the analysis results are obtained, the data is then sent to a "natural language processing engine" where the user's intentions are analyzed in detail.

[0448] Based on this analysis and emotional information, the server searches the database and retrieves relevant information. Based on the retrieved information, it generates a response tailored to the user's emotional state and sends it to the terminal. In this process, the "Natural Language Processing API" is used to generate the text.

[0449] For example, if a user is feeling anxious during an emergency, the server can quickly generate the most appropriate support information nearby and provide reassurance to the user.

[0450] A concrete example to understand how this system works is when a user mutters "I'm scared" while walking alone at night and enters it into the application. In response to this input, the system will present the user with information about safe places to alleviate their anxiety.

[0451] Examples of prompt statements include the following:

[0452] "Users feel uneasy walking alone at night. Please consider ways to reassure users who are expressing fear."

[0453] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0454] Step 1:

[0455] The device receives voice or text input from the user. Once this input is received, if it's voice, it's converted into text data using a "speech recognition API." This transforms the raw voice input into concrete text output.

[0456] Step 2:

[0457] The terminal sends the converted text data to the server. The server prepares to process the received data and passes it to the sentiment analysis engine for the next analysis step. It receives text data from the terminal as input and formats it into a parseable data format as output.

[0458] Step 3:

[0459] The server uses an emotion analysis engine to analyze the emotional state of text data. For example, it might use "Affectiva" to extract emotional information based on the nuances and tone of the input text. The input is text data, and the output is metadata indicating the user's emotional state.

[0460] Step 4:

[0461] The server uses a natural language processing engine to analyze text data containing emotional information in detail. Here, natural language processing APIs are used to process the data and extract the user's intent. Based on the text data and emotional metadata obtained as input, the server outputs the user's specific intent.

[0462] Step 5:

[0463] The server searches the database based on the analysis results and collects information related to the user's emotions and intentions. The input is the analysis results, and the output is the relevant information to be presented to the user. A rapid search algorithm works in conjunction with this process.

[0464] Step 6:

[0465] The server generates the optimal response based on the collected information and emotional state. By using a generative AI model, flexible responses are possible. The input consists of relevant information and emotional state, while the output is a response text that matches the emotional state.

[0466] Step 7:

[0467] The terminal presents the user with responses obtained from the server. It provides visual or auditory feedback and, as a termination process, offers information in a tone appropriate to the user's emotions. The input is response data from the server, and the output is the display or audio as feedback to the user.

[0468] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0469] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0470] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0471] [Third Embodiment]

[0472] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0473] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0474] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0475] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0476] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0477] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0478] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0479] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0480] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0481] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0482] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0483] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0484] This invention is a user-terminal-server type interactive system that utilizes natural language processing. This system is designed to provide comprehensive and personalized services to users who speak different languages ​​or who are unfamiliar with digital technologies.

[0485] The user inputs information into the device via voice or text. The device uses speech recognition technology to convert the input into text data and sends it to the server. The server analyzes this text data through a natural language processing engine to accurately understand the user's intent and requests. It also uses translation functions as needed to facilitate communication between different languages.

[0486] Based on the analysis results, the server searches the database for relevant information. For example, if a user is looking for nearby restaurants, the server collects local restaurant information based on their location. Next, it generates a response based on the searched information and sends it to the terminal.

[0487] The terminal receives this response and communicates it to the user through visual and auditory means. This allows the user to obtain the necessary information in real time. The user can also input feedback about the information provided. The terminal records this feedback and sends it to the server. The server records this as collected feedback data in a database and uses it to improve the system's services.

[0488] In this way, the present invention is a system that responds to the diverse needs of users and supports service providers in performing their duties more efficiently. For example, the present invention can be implemented in a wide range of applications, such as setting reminders to help elderly people remember their medication schedules, or providing interpretation services to travelers from different countries.

[0489] The following describes the processing flow.

[0490] Step 1:

[0491] The user enters information into the device via voice or text. In the case of voice input, the device uses speech recognition to convert the voice into text data.

[0492] Step 2:

[0493] The terminal sends user input data to the server. This data includes information in text format.

[0494] Step 3:

[0495] The server analyzes the received text data using a natural language processing engine to understand the user's intent and question. The data is translated as needed.

[0496] Step 4:

[0497] The server searches the database for relevant information based on the analysis results. For example, it collects information about the vicinity of the proposed location.

[0498] Step 5:

[0499] The server generates a response using the retrieved information and sends it to the terminal in text format.

[0500] Step 6:

[0501] The terminal presents the response received from the server to the user through visual or audio feedback.

[0502] Step 7:

[0503] The user generates feedback about the information provided and enters it into the device.

[0504] Step 8:

[0505] The device collects user feedback and sends it to the server.

[0506] Step 9:

[0507] The server records feedback in a database to help improve the service.

[0508] (Example 1)

[0509] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] The challenge lies in improving the efficiency and accuracy of information provision to multilingual users and users unfamiliar with digital technology, as well as effectively collecting and utilizing user feedback to improve system services.

[0511] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0512] In this invention, the server includes means for communicating using natural language processing, means for converting speech to text using speech recognition technology, and means for generating prompt sentences from the analysis results using a generative AI model. This enables efficient information exchange between diverse languages, accurate understanding of user intent, and effective collection of feedback for system improvement.

[0513] "Natural language processing" is a technology that uses computers to understand, analyze, and generate language that humans use on a daily basis.

[0514] "Means of communication" refers to a function that enables the sending and receiving of data and the exchange of information.

[0515] "Speech recognition technology" is a technology that converts voice input into text data.

[0516] "Means of converting to text data" refers to functions that convert audio or other data formats into textual information.

[0517] "Means of supporting multiple languages" refers to a function that has the ability to understand different languages ​​and translate / convert between them.

[0518] A "data set" is a general term for databases and records that collect and manage diverse information.

[0519] A "means of searching for information" is a function that finds necessary information from a data set based on specific conditions.

[0520] "Means for generating responses" refers to functions that create appropriate answers or information based on input data and search results.

[0521] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate and analyze text and data.

[0522] A "prompt" refers to an instruction or question given to a generative AI model.

[0523] "Means of ensuring security" refer to functions that maintain data safety and prevent unauthorized access and data leaks.

[0524] "Means for collecting and recording feedback" refers to a function that collects and saves user reactions and opinions to be used for future improvements.

[0525] This invention is an interactive information provision system that handles multiple languages ​​and is based on a three-tier structure consisting of a user, terminal, and server. Specific embodiments of this system are described below.

[0526] Terminal roles and functions

[0527] The user first inputs information into the terminal via voice or text. If voice input is used, the terminal receives the information through the microphone and utilizes speech recognition technology. This speech recognition includes the ability to convert speech into text data using a common API (e.g., a speech recognition API). If text input is provided, it is sent directly to the server. The terminal also displays the response from the server and provides the information to the user as voice through speech synthesis technology.

[0528] Server roles and functions

[0529] The server receives text data sent from the terminal and parses it via a natural language processing engine. During this process, a generative AI model is used to generate prompt sentences. For example, in response to a user requesting "I'm looking for a nearby Italian restaurant," the server might generate a prompt such as "Provide information on nearby Italian restaurants."

[0530] Based on the analysis, the server retrieves relevant information from a data set. This data set is managed in a common database system (e.g., a relational database), enabling efficient searching. Furthermore, the server supports multiple languages, utilizing translation functions as needed to tailor responses to each user's language. To ensure secure communication, security protocols (e.g., SSL / TLS) are used for data transmission during this process.

[0531] Gathering feedback and improving the system

[0532] Users can provide feedback on the information provided. The device receives the user's feedback and sends it to the server. The server collects this feedback data and records it in a data set. This recorded feedback is regularly analyzed to update the generative AI model and improve the quality of responses, leading to continuous improvement of the system.

[0533] This system will enable smoother information exchange and meet diverse user needs. For example, when tourists are looking for a restaurant at their travel destination, they can instantly obtain the most suitable information, overcoming language barriers.

[0534] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0535] Step 1:

[0536] The user inputs information into the device via voice or text. If the input is voice, they might speak into the microphone, for example, "Find a nearby Italian restaurant." The input data is either voice or text data. This data forms the basis for subsequent processing.

[0537] Step 2:

[0538] The terminal uses speech recognition technology to convert voice input into text data. Specifically, it uses a speech recognition API to analyze the audio waveform and generate the corresponding text. This output text is then sent to the server.

[0539] Step 3:

[0540] The terminal sends the converted text data to the server. The data is encrypted using a security protocol and transferred securely. This ensures that the user's request reaches the server accurately and securely.

[0541] Step 4:

[0542] The server inputs text data received from the terminal into a natural language processing engine. Using a generative AI model, it analyzes the user's intent from the text. This process generates prompt sentences, such as "Provide information on nearby Italian restaurants." Based on this analysis, a specific data search is performed.

[0543] Step 5:

[0544] The server searches the data set for relevant information based on the prompt message and the analysis results. It executes a database query to retrieve, for example, a list of Italian restaurants closest to the current location. This search result forms the core of the information to be provided to the user.

[0545] Step 6:

[0546] The server generates a response based on the retrieved information. Using a multilingual module, it formats the text to match the user's language and prepares it for visual and auditory presentation. This generated response is then sent to the terminal.

[0547] Step 7:

[0548] The terminal presents the received response to the user. It displays the information on the screen and plays the content aloud using speech synthesis technology. This allows the user to easily confirm the information.

[0549] Step 8:

[0550] Users provide feedback on the information provided. They input comments via text through their device, such as "This information was helpful" or "I'd like to know about other restaurants." This feedback becomes valuable data for improving the system.

[0551] Step 9:

[0552] The device sends user feedback to the server. The transmitted feedback data is recorded and analyzed on the server side for future improvements. Based on this feedback, the server takes measures to improve the quality of the service.

[0553] (Application Example 1)

[0554] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] In physical stores, customers often face challenges in quickly and appropriately obtaining information due to language barriers and unfamiliarity with technology. In particular, there is a lack of readily available means to guide foreign-speaking tourists and the elderly with product information and store services. Therefore, there is a need for ways to improve the customer experience in stores.

[0556] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0557] In this invention, the server includes means for communicating using natural language processing, means for converting user input into voice or text data, and means for acquiring and providing product information in response to user questions using speech recognition in order to provide guidance within a physical store. This enables users to receive the information they need in real time, regardless of language, thereby improving the customer experience in physical stores.

[0558] "Natural language processing" is a technology that enables computers to understand, interpret, and generate human language.

[0559] "Speech recognition" is a technology that converts speech into digital data and interprets the content of that speech as textual information.

[0560] "Multilingual support" refers to the ability to process and provide information in multiple different languages.

[0561] A "database" is a structured digital collection of information designed to efficiently store, retrieve, and manage information.

[0562] "Feedback" refers to responses and evaluations from system users, and is data used to improve the service.

[0563] A "physical store" is a facility that exists in a physical location and is accessible to customers for sales or service provision.

[0564] "Product information" refers to data containing detailed descriptions and specifications about a product or service.

[0565] "Foreign language speakers" refer to people who speak a language other than their native language and receive information within that linguistic environment.

[0566] The system implementing this invention links a terminal installed in a physical store environment with the user's smart device and uses natural language processing technology to smoothly provide product information in response to the user's questions.

[0567] Applications on terminals or smart devices capture user voice input and convert it into text data using open-source or commercially available speech recognition software. Google Cloud Speech-to-Text technology is preferably used for this purpose. The converted text data is transferred to a server in real time.

[0568] The server uses IBM Watson Natural Language Understanding to analyze the content of the received text data and identify the user's intent. Based on the analysis results, it searches a database (e.g., Firebase Realtime Database) to retrieve relevant product information. The retrieved information is translated as needed via the Microsoft Translator Text API and adapted to the user's language.

[0569] Users can visually or audibly verify the information provided by their device. Multilingual support is available, allowing foreign language speakers to receive guidance in their own language. For example, if a tourist asks a question like, "I'm looking for traditional Japanese souvenirs," they will be provided with real-time information on store locations and products that match their intent.

[0570] Examples of prompt statements are as follows:

[0571] "The user has asked you for information about Japanese souvenirs in the store in English. Search the database for relevant information and answer in both English and Japanese."

[0572] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0573] Step 1:

[0574] The user inputs a question by voice into the device's microphone. This voice input is converted into text data using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data. The audio waveform data is sampled, features are extracted, and the audio is documented based on the results.

[0575] Step 2:

[0576] The terminal sends the obtained text data to the server. The server receives this text data and processes it using a natural language processing engine (e.g., IBM Watson Natural Language Understanding) for analysis. The input is text data, and the output is structured data that indicates the user's intent. The server analyzes the document's structure, keywords, and context to interpret the user's intent.

[0577] Step 3:

[0578] The server searches the database based on the analysis results. The database contains product and store information, and extracts information that matches the user's intent. The input is structured intent data, and the output is related information data. The server generates queries and retrieves the necessary entries from the database.

[0579] Step 4:

[0580] If necessary, the server translates the extracted information into the user's native language. A translation API (e.g., Microsoft Translator Text API) is used for translation. The input is informational data, and the output is translated informational data. The server sends the informational data to the translation engine to generate text in the new language.

[0581] Step 5:

[0582] The server sends the translated information back to the terminal. After receiving this information, the terminal provides it to the user visually or audibly. The input is the translated information data, and the output is a display or audio output to the user. The terminal displays the information on the screen or plays synthesized speech through its speaker.

[0583] Step 6:

[0584] Users provide feedback on the information presented. This feedback is sent from the terminal to the server, which records its contents in a database. The input is the feedback data, and the output is the recorded feedback data. The server receives the feedback, stores it as evaluation information, and uses it to improve future processes.

[0585] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0586] This invention is a user-terminal-server type interactive system that combines natural language processing and an emotion engine. This system is designed to understand not only the words but also the emotions behind them in response to user input, and to generate more appropriate and personalized responses.

[0587] Users input information into the device via voice or text. The device then uses speech recognition technology to convert the voice into text data, which is then sent to an emotion engine for analysis. The emotion engine extracts emotional information from the user's voice tone and text wording.

[0588] The terminal sends the converted and analyzed data to the server. The server analyzes the text data using a natural language processing engine to understand the user's intent. Based on the results, it searches the database and retrieves relevant information.

[0589] The server generates an appropriate response based on the acquired information and emotional information, and sends it to the terminal. For example, if the user is dissatisfied, it will generate a response that provides reassurance.

[0590] The terminal presents the response received from the server to the user using visual or audio feedback. In doing so, it adjusts the tone and expression of the feedback based on emotional information.

[0591] Furthermore, when users input feedback into their devices, data is collected for the purpose of improving the service. The server records this feedback in a database and uses it to continuously improve the system, including sentiment analysis.

[0592] In this way, the present invention is a system that achieves more human-like interaction by taking user emotions into consideration. For example, the present invention can be implemented in customer support, such as providing a friendly response that quickly resolves problems when it is determined that a customer is feeling stressed.

[0593] The following describes the processing flow.

[0594] Step 1:

[0595] The user enters information into the device via voice or text. This input includes user questions and requests.

[0596] Step 2:

[0597] When the device receives voice input, it uses its speech recognition function to convert the speech into text data. At the same time, it also records characteristics such as the tone and speed of the speech.

[0598] Step 3:

[0599] The device passes the converted text data to an emotion engine, which then recognizes emotions from the input. For example, it can determine whether the user is angry or happy based on the wording and voice features of the input.

[0600] Step 4:

[0601] The terminal sends the sentiment analysis results along with text data to the server. The server receives this data and prepares it for analysis.

[0602] Step 5:

[0603] The server analyzes the received data using a natural language processing engine to understand the user's intent. For example, it identifies what kind of information is needed.

[0604] Step 6:

[0605] Based on the analysis results, the server searches the database for relevant information and extracts the necessary data.

[0606] Step 7:

[0607] The server uses the retrieved information and sentiment analysis results to generate a response for the user. For example, if it determines that the user is dissatisfied, it will include reassuring language.

[0608] Step 8:

[0609] The server sends the generated response to the terminal. The terminal receives it and prepares to present it to the user.

[0610] Step 9:

[0611] The device presents responses to the user visually or audibly. During presentation, it adjusts the tone and expression based on the sentiment analysis results.

[0612] Step 10:

[0613] Users can provide feedback on the information presented. The device receives this feedback and sends it to the server.

[0614] Step 11:

[0615] The server records feedback data in a database and uses it for analysis to improve the service. This includes sentiment data.

[0616] (Example 2)

[0617] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0618] Conventional technologies only provide linguistic responses to user input, making it difficult to deeply understand the user's emotional state and intentions. Furthermore, they lacked the means to generate responses tailored to the user's emotions, resulting in insufficient means of achieving personalized interaction. Additionally, the effective collection and utilization of user feedback was challenging, posing a challenge to continuous system improvement.

[0619] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0620] In this invention, the server includes means for analyzing information using natural language processing functions, means for extracting data from information resources, and a module for generating individual responses based on the extracted information. This makes it possible to generate more human-like and personalized responses that take into account the user's emotions and intentions, thereby achieving effective interaction with the user.

[0621] "Natural language processing" refers to the technology that enables computers to understand human language and analyze its meaning.

[0622] A "processing unit for converting to audio or text data" refers to hardware and software components for converting data into different formats, such as converting user voice input to text.

[0623] An "interface for identifying multiple languages" is a function that accurately identifies different language inputs and performs corresponding processing.

[0624] "Means of extracting data from information resources" refers to a function or process for obtaining necessary information from existing databases, etc.

[0625] A "module that generates individual responses" is a component that creates responses tailored to the user's specific conditions and needs based on the information it has acquired.

[0626] A "display device" is a device that presents a generated response to the user visually or audibly.

[0627] "Means for analyzing emotional information and reflecting it in response generation" refers to the processes and tools for identifying a user's emotions and reflecting the results in the response.

[0628] A "feedback collection and storage device" is a device designed to efficiently gather user reactions and opinions so that they can be analyzed later.

[0629] "Methods for analyzing saved feedback and using it to improve the system" refers to the process of analyzing recorded feedback data to improve the functionality and usability of the system.

[0630] The present invention is a system for interacting with a user through voice or text input, understanding the user's emotions, and generating responses. Its embodiments are described in detail below.

[0631] The process begins with the user inputting information into the device via voice or text. In the case of voice input, the device's built-in microphone is used to capture the voice. Next, the device applies speech recognition software to convert the voice data into text data. This process can utilize, for example, a common speech recognition API or open-source or commercial speech recognition technology.

[0632] The input text data is sent to an engine equipped with sentiment analysis capabilities. This engine uses natural language processing techniques to extract emotional information from the user's input. In this process, the terminal, for example, utilizes publicly available natural language processing libraries to analyze the user's intentions and emotions.

[0633] Sentimental information and text data are transferred from the terminal to the server, which then generates a response based on this data. The data is analyzed by a natural language processing model to identify the information and actions the user is seeking. In this process, for example, generative AI models are used to perform advanced language analysis.

[0634] The server then extracts relevant information from the database and generates a response based on the sentiment analysis results. The generated response is sent to the terminal and presented to the user as audio or text. The terminal adjusts the tone and expression presented to match the user's emotions, enabling more natural communication.

[0635] Such systems have diverse applications, including customer support. For example, they can provide appropriate responses to customers who are experiencing stress.

[0636] A concrete example of a prompt message is that it can be input into a generative AI model in the form of, "Please tell me how to generate an appropriate response when the user is excited."

[0637] Thus, the present invention is an efficient system that achieves human-like interaction by analyzing the user's emotions and reflecting them in the response.

[0638] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0639] Step 1:

[0640] The user inputs information into the device. Input methods include voice or text. In the case of voice input, the device captures the voice using a microphone. Speech recognition is performed based on the input voice data, and this is converted into text data. The output of speech recognition is text data.

[0641] Step 2:

[0642] The terminal sends the converted text data to an emotion analysis unit. This unit uses natural language processing techniques to extract emotions from the text data. The input data is text data, and the output is data representing the user's emotional state. This analysis is achieved by utilizing publicly available emotion analysis libraries.

[0643] Step 3:

[0644] The device sends the analyzed sentiment data and the original text data to the server. The server analyzes the received data using a natural language processing engine to recognize the user's intent. This analysis process identifies the information necessary to provide the desired response from the input text and sentiment information.

[0645] Step 4:

[0646] The server searches the database based on the recognized intent and retrieves relevant information. In this step, the input is data about the user's intent, and the output is relevant database information.

[0647] Step 5:

[0648] The server uses acquired information and sentiment information to generate an appropriate response. It performs advanced language generation using a generative AI model. The input data consists of relevant and sentiment information, and the output is the generated response.

[0649] Step 6:

[0650] The device presents the generated response to the user. Visual output is provided through a screen display, and auditory output through speech synthesis. The tone and expression of the response are adjusted based on the user's emotional information. The user then decides on their next action.

[0651] Step 7:

[0652] The device records the feedback provided by the user and sends it to the server. This feedback is stored as data used to improve the system in the future. The server analyzes this information and uses it to improve the system's functionality.

[0653] (Application Example 2)

[0654] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0655] In current interactive systems, generating appropriate responses that take user emotions into account is difficult. Furthermore, there is a lack of adequate responses in emergencies, and the means to provide users with a sense of security are limited. As a result, users may experience inadequate communication and support. Therefore, there is a need to develop methods for generating responses based on the user's emotional state and providing rapid responses in emergencies.

[0656] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0657] In this invention, the server includes means for communicating using natural language processing, means for analyzing the user's emotional state, and means for adjusting the tone of the response based on the user's emotional state. This makes it possible to provide the user with emotionally sensitive, personalized responses, and furthermore, to provide rapid support in emergencies.

[0658] "Natural language processing" is the technology that enables computers to understand and process human language.

[0659] "Means of communication" refers to technologies for sending and receiving information, such as the internet and wireless communication.

[0660] "Means of converting user input into speech or text data" refers to technology that converts information provided by a user into a format that can be processed by a machine.

[0661] "Means of supporting multiple languages" refers to technologies that enable information processing in different languages.

[0662] "Methods for searching information from a database" refer to technologies for finding relevant information from stored data.

[0663] "Means for generating responses based on retrieved information" refers to technologies that create responses for users based on acquired data.

[0664] "Means of presenting the generated response to the user" refers to technologies that display or communicate the answer generated by the system to the user.

[0665] "Methods for analyzing a user's emotional state" refer to technologies that identify emotions from a user's tone of voice and word choice.

[0666] "Means of adjusting the tone of response" refers to techniques that change the way a response is expressed to match the user's emotions.

[0667] "Means for collecting and recording user feedback" refers to technologies that collect user opinions and reactions and store them in a format that can be used later.

[0668] "Means of providing rapid support in emergencies" refers to technologies that enable a rapid response to emergency situations and provide necessary assistance.

[0669] This invention is an interactive system that takes user emotions into account and is comprised of a combination of an emotion analysis engine and a natural language processing engine. The system processes user voice and text input and provides appropriate responses.

[0670] First, the user inputs information into a device such as a smartphone via voice or text. The device has built-in speech recognition software, which uses libraries such as a "speech recognition API" to convert the voice into text data. This converted text data is then sent to a server located in the cloud.

[0671] The server uses an "emotion analysis engine" to analyze the emotional state from the wording and nuances of the voice in the received text data. For example, an emotion analysis tool like "Affectiva" is used in this process. Once the analysis results are obtained, the data is then sent to a "natural language processing engine" where the user's intentions are analyzed in detail.

[0672] Based on this analysis and emotional information, the server searches the database and retrieves relevant information. Based on the retrieved information, it generates a response tailored to the user's emotional state and sends it to the terminal. In this process, the "Natural Language Processing API" is used to generate the text.

[0673] For example, if a user is feeling anxious during an emergency, the server can quickly generate the most appropriate support information nearby and provide reassurance to the user.

[0674] A concrete example to understand how this system works is when a user mutters "I'm scared" while walking alone at night and enters it into the application. In response to this input, the system will present the user with information about safe places to alleviate their anxiety.

[0675] Examples of prompt statements include the following:

[0676] "Users feel uneasy walking alone at night. Please consider ways to reassure users who are expressing fear."

[0677] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0678] Step 1:

[0679] The device receives voice or text input from the user. Once this input is received, if it's voice, it's converted into text data using a "speech recognition API." This transforms the raw voice input into concrete text output.

[0680] Step 2:

[0681] The terminal sends the converted text data to the server. The server prepares to process the received data and passes it to the sentiment analysis engine for the next analysis step. It receives text data from the terminal as input and formats it into a parseable data format as output.

[0682] Step 3:

[0683] The server uses an emotion analysis engine to analyze the emotional state of text data. For example, it might use "Affectiva" to extract emotional information based on the nuances and tone of the input text. The input is text data, and the output is metadata indicating the user's emotional state.

[0684] Step 4:

[0685] The server uses a natural language processing engine to analyze text data containing emotional information in detail. Here, natural language processing APIs are used to process the data and extract the user's intent. Based on the text data and emotional metadata obtained as input, the server outputs the user's specific intent.

[0686] Step 5:

[0687] The server searches the database based on the analysis results and collects information related to the user's emotions and intentions. The input is the analysis results, and the output is the relevant information to be presented to the user. A rapid search algorithm works in conjunction with this process.

[0688] Step 6:

[0689] The server generates the optimal response based on the collected information and emotional state. By using a generative AI model, flexible responses are possible. The input consists of relevant information and emotional state, while the output is a response text that matches the emotional state.

[0690] Step 7:

[0691] The terminal presents the user with responses obtained from the server. It provides visual or auditory feedback and, as a termination process, offers information in a tone appropriate to the user's emotions. The input is response data from the server, and the output is the display or audio as feedback to the user.

[0692] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0693] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0694] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0695] [Fourth Embodiment]

[0696] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0697] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0698] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0699] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0700] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0701] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0702] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0703] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0704] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0705] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0706] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0707] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0708] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0709] This invention is a user-terminal-server type interactive system that utilizes natural language processing. This system is designed to provide comprehensive and personalized services to users who speak different languages ​​or who are unfamiliar with digital technologies.

[0710] The user inputs information into the device via voice or text. The device uses speech recognition technology to convert the input into text data and sends it to the server. The server analyzes this text data through a natural language processing engine to accurately understand the user's intent and requests. It also uses translation functions as needed to facilitate communication between different languages.

[0711] Based on the analysis results, the server searches the database for relevant information. For example, if a user is looking for nearby restaurants, the server collects local restaurant information based on their location. Next, it generates a response based on the searched information and sends it to the terminal.

[0712] The terminal receives this response and communicates it to the user through visual and auditory means. This allows the user to obtain the necessary information in real time. The user can also input feedback about the information provided. The terminal records this feedback and sends it to the server. The server records this as collected feedback data in a database and uses it to improve the system's services.

[0713] In this way, the present invention is a system that responds to the diverse needs of users and supports service providers in performing their duties more efficiently. For example, the present invention can be implemented in a wide range of applications, such as setting reminders to help elderly people remember their medication schedules, or providing interpretation services to travelers from different countries.

[0714] The following describes the processing flow.

[0715] Step 1:

[0716] The user enters information into the device via voice or text. In the case of voice input, the device uses speech recognition to convert the voice into text data.

[0717] Step 2:

[0718] The terminal sends user input data to the server. This data includes information in text format.

[0719] Step 3:

[0720] The server analyzes the received text data using a natural language processing engine to understand the user's intent and question. The data is translated as needed.

[0721] Step 4:

[0722] The server searches the database for relevant information based on the analysis results. For example, it collects information about the vicinity of the proposed location.

[0723] Step 5:

[0724] The server generates a response using the retrieved information and sends it to the terminal in text format.

[0725] Step 6:

[0726] The terminal presents the response received from the server to the user through visual or audio feedback.

[0727] Step 7:

[0728] The user generates feedback about the information provided and enters it into the device.

[0729] Step 8:

[0730] The device collects user feedback and sends it to the server.

[0731] Step 9:

[0732] The server records feedback in a database to help improve the service.

[0733] (Example 1)

[0734] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0735] The challenge lies in improving the efficiency and accuracy of information provision to multilingual users and users unfamiliar with digital technology, as well as effectively collecting and utilizing user feedback to improve system services.

[0736] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0737] In this invention, the server includes means for communicating using natural language processing, means for converting speech to text using speech recognition technology, and means for generating prompt sentences from the analysis results using a generative AI model. This enables efficient information exchange between diverse languages, accurate understanding of user intent, and effective collection of feedback for system improvement.

[0738] "Natural language processing" is a technology that uses computers to understand, analyze, and generate language that humans use on a daily basis.

[0739] "Means of communication" refers to a function that enables the sending and receiving of data and the exchange of information.

[0740] "Speech recognition technology" is a technology that converts voice input into text data.

[0741] "Means of converting to text data" refers to functions that convert audio or other data formats into textual information.

[0742] "Means of supporting multiple languages" refers to a function that has the ability to understand different languages ​​and translate / convert between them.

[0743] A "data set" is a general term for databases and records that collect and manage diverse information.

[0744] A "means of searching for information" is a function that finds necessary information from a data set based on specific conditions.

[0745] "Means for generating responses" refers to functions that create appropriate answers or information based on input data and search results.

[0746] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate and analyze text and data.

[0747] A "prompt" refers to an instruction or question given to a generative AI model.

[0748] "Means of ensuring security" refer to functions that maintain data safety and prevent unauthorized access and data leaks.

[0749] "Means for collecting and recording feedback" refers to a function that collects and saves user reactions and opinions to be used for future improvements.

[0750] This invention is an interactive information provision system that handles multiple languages ​​and is based on a three-tier structure consisting of a user, terminal, and server. Specific embodiments of this system are described below.

[0751] Terminal roles and functions

[0752] The user first inputs information into the terminal via voice or text. If voice input is used, the terminal receives the information through the microphone and utilizes speech recognition technology. This speech recognition includes the ability to convert speech into text data using a common API (e.g., a speech recognition API). If text input is provided, it is sent directly to the server. The terminal also displays the response from the server and provides the information to the user as voice through speech synthesis technology.

[0753] Server roles and functions

[0754] The server receives text data sent from the terminal and parses it via a natural language processing engine. During this process, a generative AI model is used to generate prompt sentences. For example, in response to a user requesting "I'm looking for a nearby Italian restaurant," the server might generate a prompt such as "Provide information on nearby Italian restaurants."

[0755] Based on the analysis, the server retrieves relevant information from a data set. This data set is managed in a common database system (e.g., a relational database), enabling efficient searching. Furthermore, the server supports multiple languages, utilizing translation functions as needed to tailor responses to each user's language. To ensure secure communication, security protocols (e.g., SSL / TLS) are used for data transmission during this process.

[0756] Gathering feedback and improving the system

[0757] Users can provide feedback on the information provided. The device receives the user's feedback and sends it to the server. The server collects this feedback data and records it in a data set. This recorded feedback is regularly analyzed to update the generative AI model and improve the quality of responses, leading to continuous improvement of the system.

[0758] This system will enable smoother information exchange and meet diverse user needs. For example, when tourists are looking for a restaurant at their travel destination, they can instantly obtain the most suitable information, overcoming language barriers.

[0759] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0760] Step 1:

[0761] The user inputs information into the device via voice or text. If the input is voice, they might speak into the microphone, for example, "Find a nearby Italian restaurant." The input data is either voice or text data. This data forms the basis for subsequent processing.

[0762] Step 2:

[0763] The terminal uses speech recognition technology to convert voice input into text data. Specifically, it uses a speech recognition API to analyze the audio waveform and generate the corresponding text. This output text is then sent to the server.

[0764] Step 3:

[0765] The terminal sends the converted text data to the server. The data is encrypted using a security protocol and transferred securely. This ensures that the user's request reaches the server accurately and securely.

[0766] Step 4:

[0767] The server inputs text data received from the terminal into a natural language processing engine. Using a generative AI model, it analyzes the user's intent from the text. This process generates prompt sentences, such as "Provide information on nearby Italian restaurants." Based on this analysis, a specific data search is performed.

[0768] Step 5:

[0769] The server searches the data set for relevant information based on the prompt message and the analysis results. It executes a database query to retrieve, for example, a list of Italian restaurants closest to the current location. This search result forms the core of the information to be provided to the user.

[0770] Step 6:

[0771] The server generates a response based on the retrieved information. Using a multilingual module, it formats the text to match the user's language and prepares it for visual and auditory presentation. This generated response is then sent to the terminal.

[0772] Step 7:

[0773] The terminal presents the received response to the user. It displays the information on the screen and plays the content aloud using speech synthesis technology. This allows the user to easily confirm the information.

[0774] Step 8:

[0775] Users provide feedback on the information provided. They input comments via text through their device, such as "This information was helpful" or "I'd like to know about other restaurants." This feedback becomes valuable data for improving the system.

[0776] Step 9:

[0777] The device sends user feedback to the server. The transmitted feedback data is recorded and analyzed on the server side for future improvements. Based on this feedback, the server takes measures to improve the quality of the service.

[0778] (Application Example 1)

[0779] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0780] In physical stores, customers often face challenges in quickly and appropriately obtaining information due to language barriers and unfamiliarity with technology. In particular, there is a lack of readily available means to guide foreign-speaking tourists and the elderly with product information and store services. Therefore, there is a need for ways to improve the customer experience in stores.

[0781] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0782] In this invention, the server includes means for communicating using natural language processing, means for converting user input into voice or text data, and means for acquiring and providing product information in response to user questions using speech recognition in order to provide guidance within a physical store. This enables users to receive the information they need in real time, regardless of language, thereby improving the customer experience in physical stores.

[0783] "Natural language processing" is a technology that enables computers to understand, interpret, and generate human language.

[0784] "Speech recognition" is a technology that converts speech into digital data and interprets the content of that speech as textual information.

[0785] "Multilingual support" refers to the ability to process and provide information in multiple different languages.

[0786] A "database" is a structured digital collection of information designed to efficiently store, retrieve, and manage information.

[0787] "Feedback" refers to responses and evaluations from system users, and is data used to improve the service.

[0788] A "physical store" is a facility that exists in a physical location and is accessible to customers for sales or service provision.

[0789] "Product information" refers to data containing detailed descriptions and specifications about a product or service.

[0790] "Foreign language speakers" refer to people who speak a language other than their native language and receive information within that linguistic environment.

[0791] The system implementing this invention links a terminal installed in a physical store environment with the user's smart device and uses natural language processing technology to smoothly provide product information in response to the user's questions.

[0792] Applications on terminals or smart devices capture user voice input and convert it into text data using open-source or commercially available speech recognition software. Google Cloud Speech-to-Text technology is preferably used for this purpose. The converted text data is transferred to a server in real time.

[0793] The server uses IBM Watson Natural Language Understanding to analyze the content of the received text data and identify the user's intent. Based on the analysis results, it searches a database (e.g., Firebase Realtime Database) to retrieve relevant product information. The retrieved information is translated as needed via the Microsoft Translator Text API and adapted to the user's language.

[0794] Users can visually or audibly verify the information provided by their device. Multilingual support is available, allowing foreign language speakers to receive guidance in their own language. For example, if a tourist asks a question like, "I'm looking for traditional Japanese souvenirs," they will be provided with real-time information on store locations and products that match their intent.

[0795] Examples of prompt statements are as follows:

[0796] "The user has asked you for information about Japanese souvenirs in the store in English. Search the database for relevant information and answer in both English and Japanese."

[0797] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0798] Step 1:

[0799] The user inputs a question by voice into the device's microphone. This voice input is converted into text data using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is audio data, and the output is text data. The audio waveform data is sampled, features are extracted, and the audio is documented based on the results.

[0800] Step 2:

[0801] The terminal sends the obtained text data to the server. The server receives this text data and processes it using a natural language processing engine (e.g., IBM Watson Natural Language Understanding) for analysis. The input is text data, and the output is structured data that indicates the user's intent. The server analyzes the document's structure, keywords, and context to interpret the user's intent.

[0802] Step 3:

[0803] The server searches the database based on the analysis results. The database contains product and store information, and extracts information that matches the user's intent. The input is structured intent data, and the output is related information data. The server generates queries and retrieves the necessary entries from the database.

[0804] Step 4:

[0805] If necessary, the server translates the extracted information into the user's native language. A translation API (e.g., Microsoft Translator Text API) is used for translation. The input is informational data, and the output is translated informational data. The server sends the informational data to the translation engine to generate text in the new language.

[0806] Step 5:

[0807] The server sends the translated information back to the terminal. After receiving this information, the terminal provides it to the user visually or audibly. The input is the translated information data, and the output is a display or audio output to the user. The terminal displays the information on the screen or plays synthesized speech through its speaker.

[0808] Step 6:

[0809] Users provide feedback on the information presented. This feedback is sent from the terminal to the server, which records its contents in a database. The input is the feedback data, and the output is the recorded feedback data. The server receives the feedback, stores it as evaluation information, and uses it to improve future processes.

[0810] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0811] This invention is a user-terminal-server type interactive system that combines natural language processing and an emotion engine. This system is designed to understand not only the words but also the emotions behind them in response to user input, and to generate more appropriate and personalized responses.

[0812] Users input information into the device via voice or text. The device then uses speech recognition technology to convert the voice into text data, which is then sent to an emotion engine for analysis. The emotion engine extracts emotional information from the user's voice tone and text wording.

[0813] The terminal sends the converted and analyzed data to the server. The server analyzes the text data using a natural language processing engine to understand the user's intent. Based on the results, it searches the database and retrieves relevant information.

[0814] The server generates an appropriate response based on the acquired information and emotional information, and sends it to the terminal. For example, if the user is dissatisfied, it will generate a response that provides reassurance.

[0815] The terminal presents the response received from the server to the user using visual or audio feedback. In doing so, it adjusts the tone and expression of the feedback based on emotional information.

[0816] Furthermore, when users input feedback into their devices, data is collected for the purpose of improving the service. The server records this feedback in a database and uses it to continuously improve the system, including sentiment analysis.

[0817] In this way, the present invention is a system that achieves more human-like interaction by taking user emotions into consideration. For example, the present invention can be implemented in customer support, such as providing a friendly response that quickly resolves problems when it is determined that a customer is feeling stressed.

[0818] The following describes the processing flow.

[0819] Step 1:

[0820] The user enters information into the device via voice or text. This input includes user questions and requests.

[0821] Step 2:

[0822] When the device receives voice input, it uses its speech recognition function to convert the speech into text data. At the same time, it also records characteristics such as the tone and speed of the speech.

[0823] Step 3:

[0824] The device passes the converted text data to an emotion engine, which then recognizes emotions from the input. For example, it can determine whether the user is angry or happy based on the wording and voice features of the input.

[0825] Step 4:

[0826] The terminal sends the sentiment analysis results along with text data to the server. The server receives this data and prepares it for analysis.

[0827] Step 5:

[0828] The server analyzes the received data using a natural language processing engine to understand the user's intent. For example, it identifies what kind of information is needed.

[0829] Step 6:

[0830] Based on the analysis results, the server searches the database for relevant information and extracts the necessary data.

[0831] Step 7:

[0832] The server uses the retrieved information and sentiment analysis results to generate a response for the user. For example, if it determines that the user is dissatisfied, it will include reassuring language.

[0833] Step 8:

[0834] The server sends the generated response to the terminal. The terminal receives it and prepares to present it to the user.

[0835] Step 9:

[0836] The device presents responses to the user visually or audibly. During presentation, it adjusts the tone and expression based on the sentiment analysis results.

[0837] Step 10:

[0838] Users can provide feedback on the information presented. The device receives this feedback and sends it to the server.

[0839] Step 11:

[0840] The server records feedback data in a database and uses it for analysis to improve the service. This includes sentiment data.

[0841] (Example 2)

[0842] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0843] Conventional technologies only provide linguistic responses to user input, making it difficult to deeply understand the user's emotional state and intentions. Furthermore, they lacked the means to generate responses tailored to the user's emotions, resulting in insufficient means of achieving personalized interaction. Additionally, the effective collection and utilization of user feedback was challenging, posing a challenge to continuous system improvement.

[0844] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0845] In this invention, the server includes means for analyzing information using natural language processing functions, means for extracting data from information resources, and a module for generating individual responses based on the extracted information. This makes it possible to generate more human-like and personalized responses that take into account the user's emotions and intentions, thereby achieving effective interaction with the user.

[0846] "Natural language processing" refers to the technology that enables computers to understand human language and analyze its meaning.

[0847] A "processing unit for converting to audio or text data" refers to hardware and software components for converting data into different formats, such as converting user voice input to text.

[0848] An "interface for identifying multiple languages" is a function that accurately identifies different language inputs and performs corresponding processing.

[0849] "Means of extracting data from information resources" refers to a function or process for obtaining necessary information from existing databases, etc.

[0850] A "module that generates individual responses" is a component that creates responses tailored to the user's specific conditions and needs based on the information it has acquired.

[0851] A "display device" is a device that presents a generated response to the user visually or audibly.

[0852] "Means for analyzing emotional information and reflecting it in response generation" refers to the processes and tools for identifying a user's emotions and reflecting the results in the response.

[0853] A "feedback collection and storage device" is a device designed to efficiently gather user reactions and opinions so that they can be analyzed later.

[0854] "Methods for analyzing saved feedback and using it to improve the system" refers to the process of analyzing recorded feedback data to improve the functionality and usability of the system.

[0855] The present invention is a system for interacting with a user through voice or text input, understanding the user's emotions, and generating responses. Its embodiments are described in detail below.

[0856] The process begins with the user inputting information into the device via voice or text. In the case of voice input, the device's built-in microphone is used to capture the voice. Next, the device applies speech recognition software to convert the voice data into text data. This process can utilize, for example, a common speech recognition API or open-source or commercial speech recognition technology.

[0857] The input text data is sent to an engine equipped with sentiment analysis capabilities. This engine uses natural language processing techniques to extract emotional information from the user's input. In this process, the terminal, for example, utilizes publicly available natural language processing libraries to analyze the user's intentions and emotions.

[0858] Sentimental information and text data are transferred from the terminal to the server, which then generates a response based on this data. The data is analyzed by a natural language processing model to identify the information and actions the user is seeking. In this process, for example, generative AI models are used to perform advanced language analysis.

[0859] The server then extracts relevant information from the database and generates a response based on the sentiment analysis results. The generated response is sent to the terminal and presented to the user as audio or text. The terminal adjusts the tone and expression presented to match the user's emotions, enabling more natural communication.

[0860] Such systems have diverse applications, including customer support. For example, they can provide appropriate responses to customers who are experiencing stress.

[0861] A concrete example of a prompt message is that it can be input into a generative AI model in the form of, "Please tell me how to generate an appropriate response when the user is excited."

[0862] Thus, the present invention is an efficient system that achieves human-like interaction by analyzing the user's emotions and reflecting them in the response.

[0863] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0864] Step 1:

[0865] The user inputs information into the device. Input methods include voice or text. In the case of voice input, the device captures the voice using a microphone. Speech recognition is performed based on the input voice data, and this is converted into text data. The output of speech recognition is text data.

[0866] Step 2:

[0867] The terminal sends the converted text data to an emotion analysis unit. This unit uses natural language processing techniques to extract emotions from the text data. The input data is text data, and the output is data representing the user's emotional state. This analysis is achieved by utilizing publicly available emotion analysis libraries.

[0868] Step 3:

[0869] The device sends the analyzed sentiment data and the original text data to the server. The server analyzes the received data using a natural language processing engine to recognize the user's intent. This analysis process identifies the information necessary to provide the desired response from the input text and sentiment information.

[0870] Step 4:

[0871] The server searches the database based on the recognized intent and retrieves relevant information. In this step, the input is data about the user's intent, and the output is relevant database information.

[0872] Step 5:

[0873] The server uses acquired information and sentiment information to generate an appropriate response. It performs advanced language generation using a generative AI model. The input data consists of relevant and sentiment information, and the output is the generated response.

[0874] Step 6:

[0875] The device presents the generated response to the user. Visual output is provided through a screen display, and auditory output through speech synthesis. The tone and expression of the response are adjusted based on the user's emotional information. The user then decides on their next action.

[0876] Step 7:

[0877] The device records the feedback provided by the user and sends it to the server. This feedback is stored as data used to improve the system in the future. The server analyzes this information and uses it to improve the system's functionality.

[0878] (Application Example 2)

[0879] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0880] In current interactive systems, generating appropriate responses that take user emotions into account is difficult. Furthermore, there is a lack of adequate responses in emergencies, and the means to provide users with a sense of security are limited. As a result, users may experience inadequate communication and support. Therefore, there is a need to develop methods for generating responses based on the user's emotional state and providing rapid responses in emergencies.

[0881] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0882] In this invention, the server includes means for communicating using natural language processing, means for analyzing the user's emotional state, and means for adjusting the tone of the response based on the user's emotional state. This makes it possible to provide the user with emotionally sensitive, personalized responses, and furthermore, to provide rapid support in emergencies.

[0883] "Natural language processing" is the technology that enables computers to understand and process human language.

[0884] "Means of communication" refers to technologies for sending and receiving information, such as the internet and wireless communication.

[0885] "Means of converting user input into speech or text data" refers to technology that converts information provided by a user into a format that can be processed by a machine.

[0886] "Means of supporting multiple languages" refers to technologies that enable information processing in different languages.

[0887] "Methods for searching information from a database" refer to technologies for finding relevant information from stored data.

[0888] "Means for generating responses based on retrieved information" refers to technologies that create responses for users based on acquired data.

[0889] "Means of presenting the generated response to the user" refers to technologies that display or communicate the answer generated by the system to the user.

[0890] "Methods for analyzing a user's emotional state" refer to technologies that identify emotions from a user's tone of voice and word choice.

[0891] "Means of adjusting the tone of response" refers to techniques that change the way a response is expressed to match the user's emotions.

[0892] "Means for collecting and recording user feedback" refers to technologies that collect user opinions and reactions and store them in a format that can be used later.

[0893] "Means of providing rapid support in emergencies" refers to technologies that enable a rapid response to emergency situations and provide necessary assistance.

[0894] This invention is an interactive system that takes user emotions into account and is comprised of a combination of an emotion analysis engine and a natural language processing engine. The system processes user voice and text input and provides appropriate responses.

[0895] First, the user inputs information into a device such as a smartphone via voice or text. The device has built-in speech recognition software, which uses libraries such as a "speech recognition API" to convert the voice into text data. This converted text data is then sent to a server located in the cloud.

[0896] The server uses an "emotion analysis engine" to analyze the emotional state from the wording and nuances of the voice in the received text data. For example, an emotion analysis tool like "Affectiva" is used in this process. Once the analysis results are obtained, the data is then sent to a "natural language processing engine" where the user's intentions are analyzed in detail.

[0897] Based on this analysis and emotional information, the server searches the database and retrieves relevant information. Based on the retrieved information, it generates a response tailored to the user's emotional state and sends it to the terminal. In this process, the "Natural Language Processing API" is used to generate the text.

[0898] For example, if a user is feeling anxious during an emergency, the server can quickly generate the most appropriate support information nearby and provide reassurance to the user.

[0899] A concrete example to understand how this system works is when a user mutters "I'm scared" while walking alone at night and enters it into the application. In response to this input, the system will present the user with information about safe places to alleviate their anxiety.

[0900] Examples of prompt statements include the following:

[0901] "Users feel uneasy walking alone at night. Please consider ways to reassure users who are expressing fear."

[0902] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0903] Step 1:

[0904] The device receives voice or text input from the user. Once this input is received, if it's voice, it's converted into text data using a "speech recognition API." This transforms the raw voice input into concrete text output.

[0905] Step 2:

[0906] The terminal sends the converted text data to the server. The server prepares to process the received data and passes it to the sentiment analysis engine for the next analysis step. It receives text data from the terminal as input and formats it into a parseable data format as output.

[0907] Step 3:

[0908] The server uses an emotion analysis engine to analyze the emotional state of text data. For example, it might use "Affectiva" to extract emotional information based on the nuances and tone of the input text. The input is text data, and the output is metadata indicating the user's emotional state.

[0909] Step 4:

[0910] The server uses a natural language processing engine to analyze text data containing emotional information in detail. Here, natural language processing APIs are used to process the data and extract the user's intent. Based on the text data and emotional metadata obtained as input, the server outputs the user's specific intent.

[0911] Step 5:

[0912] The server searches the database based on the analysis results and collects information related to the user's emotions and intentions. The input is the analysis results, and the output is the relevant information to be presented to the user. A rapid search algorithm works in conjunction with this process.

[0913] Step 6:

[0914] The server generates the optimal response based on the collected information and emotional state. By using a generative AI model, flexible responses are possible. The input consists of relevant information and emotional state, while the output is a response text that matches the emotional state.

[0915] Step 7:

[0916] The terminal presents the user with responses obtained from the server. It provides visual or auditory feedback and, as a termination process, offers information in a tone appropriate to the user's emotions. The input is response data from the server, and the output is the display or audio as feedback to the user.

[0917] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0918] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0919] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0920] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0921] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0922] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0923] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0924] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0925] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0926] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0927] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0928] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0929] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0930] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0931] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0932] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0933] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0934] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0935] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0936] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0937] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0938] The following is further disclosed regarding the embodiments described above.

[0939] (Claim 1)

[0940] A means of communication using natural language processing,

[0941] A means of converting user input into speech or text data,

[0942] Means to support multiple languages,

[0943] A means of searching for information in a database,

[0944] Means for generating a response based on retrieved information,

[0945] A means of presenting the generated response to the user,

[0946] Means for collecting and recording user feedback,

[0947] A system that includes this.

[0948] (Claim 2)

[0949] The system according to claim 1, which uses speech recognition to convert a user's voice input into text.

[0950] (Claim 3)

[0951] The system according to claim 1, which analyzes feedback data collected from users and uses it to improve the system.

[0952] "Example 1"

[0953] (Claim 1)

[0954] A means of communication using natural language processing,

[0955] A means of converting user input into speech or text data,

[0956] Means to support multiple languages,

[0957] A means of retrieving information from a data set,

[0958] Means for generating a response based on retrieved information,

[0959] A means of presenting the generated response to the user,

[0960] Means for collecting and recording user feedback,

[0961] A means of converting speech to text using speech recognition technology,

[0962] A means for generating prompt sentences from analysis results using a generative AI model,

[0963] Means to ensure security when transmitting data,

[0964] A system that includes this.

[0965] (Claim 2)

[0966] The system according to claim 1, which uses speech recognition to convert a user's voice input into text.

[0967] (Claim 3)

[0968] The system according to claim 1, which analyzes feedback data collected from users and uses it to improve the system.

[0969] "Application Example 1"

[0970] (Claim 1)

[0971] A means of communication using natural language processing,

[0972] A means of converting user input into speech or text data,

[0973] Means to support multiple languages,

[0974] A means of searching for information in a database,

[0975] Means for generating a response based on retrieved information,

[0976] A means of presenting the generated response to the user,

[0977] Means for collecting and recording user feedback,

[0978] In order to provide guidance within a physical store, a means of obtaining product information in response to user questions using voice recognition and providing guidance,

[0979] A means of providing information that is understandable to foreign language speakers through multilingual support,

[0980] A system that includes this.

[0981] (Claim 2)

[0982] The system according to claim 1, which uses speech recognition to convert a user's voice input into text.

[0983] (Claim 3)

[0984] The system according to claim 1, which analyzes feedback data collected from users and uses it to improve the system.

[0985] "Example 2 of combining an emotion engine"

[0986] (Claim 1)

[0987] A means of analyzing information using natural language processing functions,

[0988] A processing unit that converts user input into speech or text data,

[0989] An interface for identifying multiple languages,

[0990] Means for extracting data from information resources,

[0991] A module that generates individual responses based on extracted information,

[0992] A display device that presents the generated response to the user and adjusts its tone,

[0993] A means of analyzing user emotional information and reflecting it in the generation of responses,

[0994] A device for collecting and storing user feedback,

[0995] A means of analyzing saved feedback and using it to improve the system,

[0996] A system that includes this.

[0997] (Claim 2)

[0998] The system according to claim 1, wherein a speech recognition unit is used to convert a user's voice input into text data.

[0999] (Claim 3)

[1000] The system according to claim 1, which analyzes emotional information collected from users and reflects it in generating system responses.

[1001] "Application example 2 when combining with an emotional engine"

[1002] (Claim 1)

[1003] A means of communication using natural language processing,

[1004] A means of converting user input into speech or text data,

[1005] Means to support multiple languages,

[1006] A means of searching for information in a database,

[1007] Means for generating a response based on retrieved information,

[1008] A means of presenting the generated response to the user,

[1009] A means of analyzing the emotional state of users,

[1010] A means of adjusting the tone of response based on the user's emotional state,

[1011] Means for collecting and recording user feedback,

[1012] Means to provide rapid support in emergencies,

[1013] A system that includes this.

[1014] (Claim 2)

[1015] The system according to claim 1, which uses speech recognition to convert a user's voice input into text.

[1016] (Claim 3)

[1017] The system according to claim 1, which analyzes feedback data collected from users and uses it to improve the system. [Explanation of symbols]

[1018] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of communication using natural language processing, A means of converting user input into speech or text data, Means to support multiple languages, A means of searching for information in a database, Means for generating a response based on retrieved information, A means of presenting the generated response to the user, Means for collecting and recording user feedback, A system that includes this.

2. The system according to claim 1, which converts a user's voice input into text using speech recognition.

3. The system according to claim 1, which analyzes feedback data collected from users and uses it to improve the system.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A