system
The system addresses the challenge of providing personalized and timely information by analyzing user inquiries, utilizing natural language processing and real-time data acquisition, and incorporating feedback to enhance accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-10
- Publication Date
- 2026-06-22
AI Technical Summary
Conventional local information providing systems struggle to efficiently provide highly relevant, personalized information to users due to limitations in data collection from external sources, lack of real-time data acquisition, and difficulty in understanding user interests and concerns.
A system that analyzes user inquiries using natural language processing, references user profile data, and acquires data in real-time from external sources to generate personalized information, with feedback mechanisms to improve information provision accuracy over time.
Enables the provision of accurate and timely information tailored to individual user needs, continuously improving its relevance through user feedback.
Smart Images

Figure 2026101267000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional local information providing systems, there was a problem that it was difficult for users to efficiently obtain highly relevant information. Specifically, it was impossible to provide personalized information according to the interests and concerns of each user, and access to the information required by the user was often limited. In addition, since data collection from external information sources was not performed in real time, there was a problem that it was difficult to obtain the latest information.
Means for Solving the Problems
[0005] This invention solves the above problems by providing a system that analyzes user inquiries using natural language processing and generates personalized information by referring to the user's profile data. Furthermore, by providing means for acquiring data in real time from external information sources and selecting information according to the user's intent, the system accurately provides the latest and most relevant information that the user is looking for. In addition, by collecting user feedback and improving the information provision algorithm, the accuracy of the information can be improved over time.
[0006] A "user inquiry" is a request made by a user when they ask for information or a question from a system.
[0007] "Natural language processing" is a technology that enables computers to understand, analyze, and generate natural language that humans use in everyday life.
[0008] "Profile data" refers to a collection of data that includes an individual user's interests, preferences, past behavioral history, and personal information.
[0009] "Personalized information" refers to information that has been tailored and optimized to the individual user's needs and interests.
[0010] "External information sources" refer to information sources such as databases and APIs that exist outside the system.
[0011] "Retrieving data" means accessing and collecting necessary information from a specific database or API.
[0012] A "user interface" refers to the interface that includes screens and operating tools for a user to interact with a system.
[0013] "Feedback" refers to the reactions, evaluations, and opinions that users provide after using a system.
[0014] An "information provision algorithm" is a set of computational procedures or rules used to provide effective information to a user. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the language used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention constructs a local information provision system and is implemented according to the following procedure. First, a user makes an inquiry through a terminal. Since this inquiry is entered in natural language, the terminal sends its contents to the server. The server analyzes the input inquiry via a natural language processing (NLP) engine and understands the user's intent. This clarifies what kind of information the user is seeking.
[0037] Next, the server accesses the user's profile data, which includes information related to the user's past behavior, interests, and preferences. Based on this data, personalized information useful to the user is generated. The server then accesses external information sources to obtain relevant information in real time and selects the most relevant information based on the user's profile.
[0038] The selected information is sent from the server to the terminal. The terminal displays this information on its user interface, providing it in a format that is easy for the user to understand. During this process, details of the event, such as the date and time and participation requirements, are displayed, and the user can use the information as needed.
[0039] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm. As a result, the system can provide more accurate information that better suits the user's needs over time.
[0040] Specific example
[0041] When a user enters "Are there any events I can take my child to this weekend?" into their device, the server receives the query and uses natural language processing to extract the keywords "this weekend," "children," and "event." Next, the server identifies the appropriate event type from the user's past participation history and retrieves relevant event information by referring to an external event database. The device then displays "Family Day Events Held at the Park," providing details such as the date, time, location, and how to participate. By reviewing this information and providing feedback, the system can provide more accurate information in the future.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The user inputs information into the device using natural language. The device receives this input and prepares to send it to the server as a query.
[0045] Step 2:
[0046] The terminal sends the user's query to the server. The server receives this query and calls a natural language processing (NLP) engine to analyze the input.
[0047] Step 3:
[0048] The server uses NLP to analyze the intent of the query and extract keywords. This identifies the type of information the user is seeking.
[0049] Step 4:
[0050] The server accesses the user's profile database to retrieve the user's past behavior history and interests. This information is then used to prepare for personalizing the user's information.
[0051] Step 5:
[0052] The server accesses external information sources and retrieves the latest relevant data. It extracts the relevant events and information and selects the data that matches the user's intent.
[0053] Step 6:
[0054] The server organizes the selected information and generates personalized information. The generated information is structured in a format that is easy for the user to understand.
[0055] Step 7:
[0056] The server generates information and sends it to the terminal. The terminal receives this information and displays it in the user interface. The user can then review the presented information.
[0057] Step 8:
[0058] The user provides feedback on the information provided. The device sends this feedback to the server, which is used to improve the information provision algorithm.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] Providing local information in a timely and accurate manner that meets the individual needs and interests of users is challenging. Furthermore, if the provided information does not meet user expectations, it is necessary to improve the quality of the information based on feedback, which is also not easy. In addition, there is an increasing demand for information that takes the user's current location into account, and responding to this is essential.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user profile information and generating personalized information, means for acquiring data from external information sources and selecting optimal information according to the user's intent, and means for generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences. This enables the provision of accurate and timely information tailored to the individual needs of the user, and furthermore, allows for continuous improvement of the quality of the information provided through feedback.
[0064] "Methods for analyzing user inquiries using natural language processing" refers to technologies that analyze user questions written in natural language and convert their meaning and intent into a format that the system can understand.
[0065] "Means of referencing user profile information and generating personalized information" refers to methods of creating information tailored to the individual needs of users by utilizing data on their past behavioral history and interests.
[0066] "A means of acquiring data from external information sources and selecting the most suitable information according to the user's intent" refers to the process of gathering necessary data from external information sources, analyzing it, and providing the information that best matches the user's requirements.
[0067] "A method of generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences" refers to a technique that utilizes AI technology to generate the most suitable information for the user, and the process by which the AI finds the best answer based on the prompt sentences.
[0068] "Means of visually displaying information on a user interface" refers to technologies for displaying generated information on a screen in a way that is easy for the user to understand.
[0069] "Means of optionally acquiring the user's geographical location information to improve the accuracy of information provision" refers to methods of acquiring and utilizing location information in order to provide more appropriate information based on the user's current location.
[0070] This invention is a system that provides local information tailored to user needs and is implemented based on the following embodiments.
[0071] Users can query information using natural language through their devices. The device receives this query and sends the data to the server to accurately analyze the user's intent. The server uses a natural language processing (NLP) engine to analyze the language data received from the user. Through this analysis, the server can extract keywords and intent from the user's query.
[0072] Next, the server references the user's profile information and generates personalized information based on this information. Specifically, information is generated that takes into account the user's past behavior history and interests. In addition, the server accesses external information sources and selects the most suitable information for the user based on data collected in real time. As part of this process, a generative AI model is used to find the best candidates when the prompt message "Tell me about events happening this weekend in the area I'm interested in" is entered.
[0073] The selected information is displayed in a visually easy-to-understand format on the user interface and provided to the user through the device. This allows users to easily obtain and use the information they need. For example, detailed information such as "Family Day Event held in the park" is presented along with the date, time, and location.
[0074] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to continuously improve the information provision algorithm. As a result, the system learns with each use and can provide information that better meets user expectations the next time it is used.
[0075] Through this configuration, the system can efficiently continue to provide valuable information to users.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The user inputs information in natural language on their device. For example, they might ask, "Tell me about weekend events happening in this area." This input is then sent to the server as data for the next processing step.
[0079] Step 2:
[0080] The terminal sends the received user query to the server in its original format. This process involves encoding based on the data communication protocol. This prepares the server to analyze the user query.
[0081] Step 3:
[0082] The server uses a natural language processing (NLP) engine to analyze user inquiries. Keywords such as "weekend," "region," and "event" are extracted from the input inquiry. These extracted keywords are used to understand the user's intent and guide the subsequent processing.
[0083] Step 4:
[0084] The server references the user's profile information to generate personalized data. Input includes the user's past behavior history and interests. Based on this, customized information tailored to each user's individual needs is output.
[0085] Step 5:
[0086] The server accesses external information sources and retrieves event data in real time. The collected data is then analyzed, taking into account the user's profile information and the analysis results, to select the most relevant information. During this process, a generative AI model is used to generate optimal candidates based on the prompt "Events this weekend in the region the user is interested in".
[0087] Step 6:
[0088] The selected information is sent from the server to the terminal. The server formats the data into a user-friendly format and delivers it to the terminal as an interface display.
[0089] Step 7:
[0090] The terminal displays the received information on its user interface. Here, the information is visually organized for easy understanding, and detailed event information, such as "Family Day Event held in the park," is provided.
[0091] Step 8:
[0092] Users enter feedback about the displayed information on their device. This feedback is sent to the server, which then uses it to improve the server's information provision algorithm, thereby increasing the accuracy of future information provision.
[0093] (Application Example 1)
[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0095] In modern society, users have access to vast amounts of information, but efficiently obtaining the most relevant information remains difficult. Furthermore, while there is a demand for information provision that takes into account users' interests and behavioral history, there is also a need to improve the accuracy of natural language queries and present information clearly through user interfaces. Additionally, enhancing voice input capabilities is desirable to improve user convenience.
[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0097] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user attribute information and generating personalized information, means for acquiring information from external sources and selecting information according to the user's intent, means for converting voice input into text, and means for displaying the selected information on a user interface. As a result, users can efficiently acquire information that suits their preferences, and convenience is improved through natural and simple interaction utilizing voice input.
[0098] A "user inquiry" is a question or request from a user regarding the information or services they seek from the system.
[0099] "Natural language processing" is a technology that understands the meaning and intent behind the language that humans use in everyday life and processes it as data.
[0100] "User attribute information" refers to data related to an individual, including their interests, preferences, and past behavioral history.
[0101] "Personalized information" refers to customized data that is presented in a way that is most suitable for the individual, based on their profile.
[0102] An "external information source" is a data provider that exists outside the system and provides information in real time.
[0103] "Voice input" is a method of transmitting information to a system through the voice spoken by the user.
[0104] "Means of converting to text" refers to technologies and processes that convert non-textual data, such as audio, into textual information.
[0105] A "user interface" is an interaction environment that provides screens and methods of operation for users to interact with a system.
[0106] The system for implementing this invention consists of three main components: a server, a terminal, and a user. The server uses a natural language processing engine to analyze inquiries entered by the user through the terminal. This can be done using natural language processing tools such as Google Cloud Natural Language API. The analyzed information is stored on the server as personalized data, taking into account the user's attribute information.
[0107] The server further accesses external data sources and retrieves data as needed. Open data APIs can be used for this access. From the retrieved data, it selects the information that best matches the user's intent and sends it to the terminal for display in the user interface.
[0108] The device provides a voice input function, allowing users to make natural inquiries using their voice. This voice is converted to text using speech-to-text technologies such as the Google Speech-to-Text API and sent to the server. The converted text is then analyzed on the server, and the corresponding information is prepared.
[0109] For example, if a user asks a question via voice, such as "Tell me about jazz events happening nearby," the voice is converted to text and analyzed by the server's natural language processing engine. Based on the analysis results, the system references the user's past activity history to extract and display relevant jazz events.
[0110] An example of a prompt for a generative AI model is written as follows:
[0111] "When a user asks, 'What are some recommended music events near me?', you should recommend events they might be able to attend, taking into account their interest in jazz events in particular. You should also consider events the user has attended in the past to suggest the best options."
[0112] This allows users to quickly obtain information that is of particular interest to them, and to enjoy a more convenient information retrieval experience.
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] Users enter inquiries using their device via voice or text. For voice input, the device uses the Google Speech-to-Text API to convert speech to text. This process takes a speech waveform as input and outputs the converted text data.
[0116] Step 2:
[0117] The terminal sends the obtained text to the server. The server receives this text and parses it using a natural language processing engine (e.g., Google Cloud Natural Language API). The input for this step is the user's query text, and the output is data indicating the parsed keywords and user intent.
[0118] Step 3:
[0119] The server references user attribute information based on the analysis results and generates personalized information. The inputs here are the analyzed keywords and user profile information, and the output is a list of information optimized for the user.
[0120] Step 4:
[0121] The server accesses external information sources and collects information that matches the user's intent. Open data APIs, for example, are used. The input is the analysis results and the user's interest data, and the output is the collected external information.
[0122] Step 5:
[0123] The server compares and analyzes the collected external information and selects the most suitable information to display in the user interface. In this step, the input is a list of candidate information that forms the basis of the selection, and the output is the final information presented to the user.
[0124] Step 6:
[0125] The terminal displays the selection information received from the server in the user interface. Here, the information sent from the server is the input, and the visual presentation of that information to the user is the output. The user reviews the displayed information and takes action as needed.
[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0127] This invention is a local information provision system that combines natural language processing and an emotion engine to provide users with more personalized information. First, the user queries for information through a terminal. The terminal sends the content to a server, which uses a natural language processing (NLP) engine to analyze the query. This analysis extracts the type of information the user is seeking and related keywords.
[0128] Furthermore, this invention uses an emotion engine to recognize emotions from the user's text input. The emotion engine is used to systematically understand the user's emotional state and select appropriate information according to that state. For example, if it is determined that the user is feeling stressed, it will provide relaxation events or healing information.
[0129] Next, the server examines the user's profile data. This data contains past behavioral history and user interests, and personalized information is generated based on this data. The server accesses an external database to retrieve the most relevant information selected based on the user's intentions and emotions.
[0130] Using the acquired information, the server formats it into a format suitable for display on the user interface and sends it to the terminal. The terminal receives this information and displays it visually to the user. The user can then review and utilize this information.
[0131] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm and adjust sentiment patterns. This allows the system to continuously improve its ability to meet user needs over the long term.
[0132] Specific example
[0133] For example, if a user enters a question into their device expressing stress, such as "What events can I enjoy with my kids this weekend?", the server analyzes the inquiry and extracts "stress" and "kid-friendly events" as keywords. Next, the emotion engine detects the stress and suggests relaxing family events based on the user's profile. The device might display information such as "Weekend Family Day at the Park: With Relaxation Corner." If the user provides feedback on this event, it will be incorporated into future information provision.
[0134] The following describes the processing flow.
[0135] Step 1:
[0136] The user queries the device for event information using natural language. The device retrieves this information and prepares to send it to the server along with the emotion engine.
[0137] Step 2:
[0138] The terminal sends the user's inquiry to the server. The server processes the received data and parses the text via a natural language processing engine.
[0139] Step 3:
[0140] The server uses natural language processing to extract the main keywords of the query. This identifies the type of event the user is looking for, as well as the available time and location.
[0141] Step 4:
[0142] The server uses an emotion engine to recognize the emotions contained in the user's input. Emotions such as stress, excitement, and satisfaction are evaluated.
[0143] Step 5:
[0144] The server accesses the user's profile database to investigate their past behavior and interests. This allows for the provision of information that is best suited to the user.
[0145] Step 6:
[0146] The server connects to external information sources in real time to retrieve relevant data such as event information. This data is filtered based on user requests and sentiments.
[0147] Step 7:
[0148] The server organizes the selected information and generates personalized event information according to the user's request. This information is then converted into an appropriate format and prepared for display.
[0149] Step 8:
[0150] The server sends the generated information to the terminal. The terminal receives this information and displays it in the user interface, allowing the user to view the information.
[0151] Step 9:
[0152] The user provides feedback on the information they have been given. The device sends this feedback to the server, which is then used to improve the system.
[0153] Step 10:
[0154] The server analyzes the feedback and uses it to adjust the information delivery algorithm and emotion engine. This process further improves the user's experience on subsequent visits.
[0155] (Example 2)
[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0157] Conventional local information systems typically provide information to users, making it difficult to effectively deliver personalized information tailored to the individual user's emotions and interests. As a result, users often fail to receive the information they need, highlighting the need for improved user experience.
[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0159] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing user emotions, and means for referencing user profile information and generating personalized information. This makes it possible to provide information that is tailored to the individual emotions and needs of each user.
[0160] "Natural language processing" is the technology that enables computers to understand, analyze, and generate human language.
[0161] "Emotion recognition" is a technology that extracts and analyzes emotional elements from information provided by users.
[0162] "Profile information" refers to data about individual users, including their interests, behavioral history, and personal characteristics.
[0163] "External information sources" refer to databases, APIs, and other information sources located outside the system that provide up-to-date or additional data.
[0164] "Feedback" refers to the opinions and evaluations that users provide to a system, which are used to improve the services and information provided.
[0165] "Personalized information" refers to information tailored to each user's individual characteristics and preferences, and is provided on the user interface.
[0166] "Filtering" refers to the process of selecting information data based on certain criteria and eliminating unnecessary information.
[0167] A "user interface" is a means of displaying and inputting information between a computer system and a user, and includes visual display screens and input devices.
[0168] This invention is an information provision system that combines natural language processing technology and an emotion recognition engine to provide users with personalized information. In this system, the user queries for information from a terminal, the server analyzes the content, selects appropriate information, and provides it to the user.
[0169] The user inputs information and makes inquiries using natural language via a terminal. This terminal sends data to the server using a common communication protocol, initiating information processing. The server uses a natural language processing engine implemented in a programming language such as Python to analyze the user's input and extract the intent of the inquiry and related keywords. In this process, libraries such as NLTK may be used.
[0170] In addition, the server uses an emotion recognition engine to analyze the user's emotional state from their text. This engine further improves the accuracy of the information the user is seeking by using machine learning models (such as the BERT model) powered by the Transformers library.
[0171] Next, the server checks the user's profile information. This information is stored in a database system (e.g., MySQL®) and reflects past behavior and interests. Based on this data, the server generates personalized information that is most relevant to the user.
[0172] The server then accesses external information sources to retrieve the latest relevant events and news. This process utilizes various APIs to gather a wide range of information and select the information that best suits the user's needs.
[0173] The acquired information is formatted by the server into a user-friendly format and transmitted to the terminal. The terminal displays the information in the specified format and provides it to the user. The user can then make decisions and take actions based on this information.
[0174] Furthermore, users can input feedback using their devices. This information is sent to the server and used to improve the information provision algorithm. As a result, the information provision will continuously improve over the long term and become more valuable to users.
[0175] For example, if a user enters "Where can I relax with my family on the weekend?", the server will analyze this query and suggest events related to relaxation. For instance, it might provide information such as "Weekend events at nearby parks."
[0176] Example of a prompt:
[0177] If a user asks, "Can you recommend some places to relax on the weekend?", how would you provide them with relevant information?
[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0179] Step 1:
[0180] User input of inquiry
[0181] Specific operation: The user requests specific information on the device's input screen. For example, they might type text such as, "What are some places where my family can relax on the weekend?"
[0182] Input and Output: The text entered by the user is input into the terminal. As output, the terminal sends this query data to the server.
[0183] Step 2:
[0184] Sending information from the terminal to the server
[0185] Specific operation: The terminal sends the received user query to the server via the appropriate communication protocol (e.g., HTTPS).
[0186] Input and Output: The input is the query text received from the user. The output is the data received by the server for analysis.
[0187] Step 3:
[0188] Execution of natural language processing by a server
[0189] Specific operation: The server executes a Python script and uses a natural language processing library (e.g., NLTK) to parse the query text. Specific techniques include text tokenization and keyword extraction.
[0190] Input and Output: The input is text data received from the terminal. The output generates data containing the analyzed keywords and intent.
[0191] Step 4:
[0192] Emotion recognition performed by the server
[0193] Specific operation: The server uses a machine learning model (e.g., a BERT-based model) to recognize emotions from the user's text. Emotional scoring is performed, and the emotional state of the text is quantified.
[0194] Input and Output: The input is parsed text, and the output is a sentiment score or state.
[0195] Step 5:
[0196] User profile referencing by the server
[0197] Specific operation: The server executes database queries (e.g., MySQL) to retrieve information about the user's past behavior and interests.
[0198] Input and Output: The input is identification information such as a user ID, and the output is profile information.
[0199] Step 6:
[0200] Obtaining information from external sources
[0201] Specific operation: The server uses various APIs to request relevant information from external databases. The retrieved information is prioritized to match the user's current interests and emotional state.
[0202] Input and Output: The input consists of selected keywords or conditions. The output consists of retrieved event information or news.
[0203] Step 7:
[0204] Formatting and sending information to the device
[0205] Specific operation: The server uses an HTML template to visually format the information to be presented to the user and sends it to the terminal via the HTTP protocol.
[0206] Input and Output: Input is the acquired information data, and output is the information formatted in a way that the user can visually understand.
[0207] Step 8:
[0208] Displaying information on the device
[0209] Specific operation: The terminal displays information received from the server on the screen, making it available for the user to review. Detailed information and related information can be accessed through the user interface.
[0210] Input and Output: Input is formatted information received from the server, and output is content presented visually to the user.
[0211] Step 9:
[0212] User feedback
[0213] Specific operation: The user enters feedback on the information provided via the terminal and sends it to the server by operating the submit button.
[0214] Input and Output: Input is user ratings and opinions, and output is feedback data recorded by the server.
[0215] (Application Example 2)
[0216] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0217] In modern urban life, users need information tailored to their emotional state and individual interests at any given time. However, conventional systems struggle to accurately recognize the diverse emotions of users and provide appropriate information. Furthermore, there is a lack of methods to effectively utilize user feedback and provide information that is more suitable for individual users.
[0218] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0219] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing the user's emotional state and reflecting it in information selection, and means for referencing the user's profile data and generating personalized information. This makes it possible to provide more appropriate and personalized information based on the user's emotions and profile.
[0220] "Natural language processing" is a technology that enables computers to understand, analyze, and respond to human language.
[0221] "Emotional state" refers to the psychological state inferred from the user's input, and includes emotions such as stress, relaxation, and excitement.
[0222] "Profile data" refers to a collection of individual information, including a user's past behavioral history, interests, and preferences.
[0223] "Personalized information" refers to information that is optimized based on the user's individual attributes and emotions.
[0224] "External information sources" refer to external information sources that the server can access, such as databases and web resources.
[0225] A "user interface" is a means of displaying information visually to a user, allowing the user to confirm and manipulate that information.
[0226] "Feedback" refers to the opinions and evaluations given by users, and is information used to improve the system.
[0227] "In-city experiential information" refers to information related to tourist attractions, events, transportation, etc., within a city.
[0228] This invention provides an information delivery system that performs emotion recognition and personalization for residents of smart cities. Users use their smartphones to query information to customize their individual urban experience. The server receives the user's query in natural language and analyzes it using a natural language processing engine. In this process, it extracts the type of information requested by the user and related keywords. Furthermore, it uses an emotion engine to recognize the user's emotional state from their input. The server considers the user's profile data and selects and provides personalized information that is appropriate for their emotional state.
[0229] Specifically, the server filters information on tourist attractions, events, and transportation based on user profiles and data obtained from external sources, and formats urban experience information that matches the user's emotions. This information is transmitted to the terminal through the user interface, allowing the user to visually confirm it.
[0230] For example, if a user enters the question, "Where can I relax in Tokyo on the weekend?", the server processes this question using natural language processing and extracts "relax" and "Tokyo" as keywords. Then, using an emotion engine, it determines that the user is seeking relaxation and suggests quiet parks and healing spots based on their profile data. Information such as "Tokyo Central Park: with a quiet rest area" is displayed on the device. The user can then use these suggestions to plan their weekend.
[0231] A concrete example of a prompt is, "Please tell me some relaxing places in Tokyo that I should visit this weekend." This prompt allows the user to have a suitable urban experience.
[0232] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0233] Step 1:
[0234] The user enters the query in natural language using a smartphone. The input is received in text format, and the device sends this data to the server. The output is the query data sent to the server.
[0235] Step 2:
[0236] The server passes the received query data to a natural language processing engine for analysis. The input is the user's query text, and the output is the extraction of keywords and information types. In this step, the server analyzes important words and phrases contained in the query to determine what the user is asking for.
[0237] Step 3:
[0238] The server uses an emotion engine to recognize the user's emotional state from their input. The input is the user's query text, and the output is the recognized emotional state. In this step, the server analyzes the user's emotions from the content and context of the text and evaluates how those emotions will affect the information provided.
[0239] Step 4:
[0240] The server retrieves user profile data from the database and analyzes past behavior and interests. The input is user identification information, and the output is user-specific profile data. This step prepares the foundational data for personalizing information, taking into account the user's interests and behavioral tendencies.
[0241] Step 5:
[0242] The server accesses external information sources to obtain optimal urban experience information based on analyzed keywords, emotional states, and profile data. The input is a combination of keywords and emotional data, and the output is filtered information results. This step selects information tailored to the user's needs and creates customized recommendations.
[0243] Step 6:
[0244] The server formats the filtered information results into a format suitable for the user interface and sends it to the terminal. The input is the filtered information results, and the output is in a displayable information format. This step ensures that the information is displayed in a way that is easy for the user to understand.
[0245] Step 7:
[0246] The terminal displays formatted information on a user interface, providing users with visual access. Input is in a displayable information format, and output is the display of information to the user. This allows users to easily review and utilize proposed urban experiences.
[0247] Step 8:
[0248] The user enters feedback on the provided information via a terminal. The input is the text of the feedback, and the terminal sends this feedback to the server. The output is the feedback data sent to the server.
[0249] Step 9:
[0250] The server uses the received feedback data to improve its information provision algorithm. The input is user feedback data, and the output is the improved algorithm. In this step, the feedback is analyzed and improvements are made to make future information provision more appropriate.
[0251] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0252] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0253] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0254] [Second Embodiment]
[0255] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0256] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0257] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0258] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0259] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0260] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0261] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0262] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0263] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0264] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0265] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0266] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0267] This invention constructs a local information provision system and is implemented according to the following procedure. First, a user makes an inquiry through a terminal. Since this inquiry is entered in natural language, the terminal sends its contents to the server. The server analyzes the input inquiry via a natural language processing (NLP) engine and understands the user's intent. This clarifies what kind of information the user is seeking.
[0268] Next, the server accesses the user's profile data, which includes information related to the user's past behavior, interests, and preferences. Based on this data, personalized information useful to the user is generated. The server then accesses external information sources to obtain relevant information in real time and selects the most relevant information based on the user's profile.
[0269] The selected information is sent from the server to the terminal. The terminal displays this information on its user interface, providing it in a format that is easy for the user to understand. During this process, details of the event, such as the date and time and participation requirements, are displayed, and the user can use the information as needed.
[0270] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm. As a result, the system can provide more accurate information that better suits the user's needs over time.
[0271] Specific example
[0272] When a user enters "Are there any events I can take my child to this weekend?" into their device, the server receives the query and uses natural language processing to extract the keywords "this weekend," "children," and "event." Next, the server identifies the appropriate event type from the user's past participation history and retrieves relevant event information by referring to an external event database. The device then displays "Family Day Events Held at the Park," providing details such as the date, time, location, and how to participate. By reviewing this information and providing feedback, the system can provide more accurate information in the future.
[0273] The following describes the processing flow.
[0274] Step 1:
[0275] The user inputs information in natural language to the terminal. The terminal receives this input and prepares to send it to the server as a query.
[0276] Step 2:
[0277] The terminal sends the user's query to the server. The server receives this query, calls the natural language processing (NLP) engine, and analyzes the input content.
[0278] Step 3:
[0279] The server analyzes the intent of the query using NLP and extracts keywords. This identifies the type of information the user is seeking.
[0280] Step 4:
[0281] The server accesses the user's profile database and obtains the user's past behavior history, interests, and concerns. This is used to prepare for personalizing the information.
[0282] Step 5:
[0283] The server accesses external information providers and obtains the latest relevant data. It extracts the corresponding events and information and selects the data that matches the user's intent.
[0284] Step 6:
[0285] The server organizes the selected information and generates personalized information. The generated information is structured in a format that is easy for the user to understand.
[0286] Step 7:
[0287] The server sends the generated information to the terminal. The terminal receives this information and displays it on the user interface. The user can view the presented information.
[0288] Step 8:
[0289] The user provides feedback on the information provided. The device sends this feedback to the server, which is used to improve the information provision algorithm.
[0290] (Example 1)
[0291] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0292] Providing local information in a timely and accurate manner that meets the individual needs and interests of users is challenging. Furthermore, if the provided information does not meet user expectations, it is necessary to improve the quality of the information based on feedback, which is also not easy. In addition, there is an increasing demand for information that takes the user's current location into account, and responding to this is essential.
[0293] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0294] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user profile information and generating personalized information, means for acquiring data from external information sources and selecting optimal information according to the user's intent, and means for generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences. This enables the provision of accurate and timely information tailored to the individual needs of the user, and furthermore, allows for continuous improvement of the quality of the information provided through feedback.
[0295] "Methods for analyzing user inquiries using natural language processing" refers to technologies that analyze user questions written in natural language and convert their meaning and intent into a format that the system can understand.
[0296] "Means of referencing user profile information and generating personalized information" refers to methods of creating information tailored to the individual needs of users by utilizing data on their past behavioral history and interests.
[0297] "A means of acquiring data from external information sources and selecting the most suitable information according to the user's intent" refers to the process of gathering necessary data from external information sources, analyzing it, and providing the information that best matches the user's requirements.
[0298] "A method of generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences" refers to a technique that utilizes AI technology to generate the most suitable information for the user, and the process by which the AI finds the best answer based on the prompt sentences.
[0299] "Means of visually displaying information on a user interface" refers to technologies for displaying generated information on a screen in a way that is easy for the user to understand.
[0300] "Means of optionally acquiring the user's geographical location information to improve the accuracy of information provision" refers to methods of acquiring and utilizing location information in order to provide more appropriate information based on the user's current location.
[0301] This invention is a system that provides local information tailored to user needs and is implemented based on the following embodiments.
[0302] Users can query information using natural language through their devices. The device receives this query and sends the data to the server to accurately analyze the user's intent. The server uses a natural language processing (NLP) engine to analyze the language data received from the user. Through this analysis, the server can extract keywords and intent from the user's query.
[0303] Next, the server refers to the user's profile information and generates personalized information based on this information. Specifically, information considering the user's past behavior history and interests is generated. In addition, the server accesses external information sources and selects the most suitable information for the user based on the data collected in real time. As part of this process, by utilizing a generative AI model and inputting the prompt sentence "Tell me about the events this weekend in the area the user is interested in", the optimal candidates are searched for.
[0304] The selected information is displayed on the user interface in a visually easy-to-understand format and provided to the user through the terminal. As a result, the user can easily obtain and utilize the necessary information. For example, detailed information such as "Family Day event held in the park" is presented together with the date and location.
[0305] Furthermore, the user can input feedback on the provided information via the terminal. This feedback is sent to the server and used for the continuous improvement of the information provision algorithm. As a result, the system learns every time it is used and can provide information that better meets the user's expectations next time.
[0306] Through the above form, the system can continuously and efficiently provide information valuable to the user.
[0307] The flow of the specific process in Example 1 will be described using FIG. 11.
[0308] Step 1:
[0309] The user inputs information in natural language using the terminal. For example, an inquiry such as "Tell me about the weekend events held in this area" is made. This input is sent to the server as data for the next process.
[0310] Step 2:
[0311] The terminal sends the received user query to the server in its original format. This process involves encoding based on the data communication protocol. This prepares the server to analyze the user query.
[0312] Step 3:
[0313] The server uses a natural language processing (NLP) engine to analyze user inquiries. Keywords such as "weekend," "region," and "event" are extracted from the input inquiry. These extracted keywords are used to understand the user's intent and guide the subsequent processing.
[0314] Step 4:
[0315] The server references the user's profile information to generate personalized data. Input includes the user's past behavior history and interests. Based on this, customized information tailored to each user's individual needs is output.
[0316] Step 5:
[0317] The server accesses external information sources and retrieves event data in real time. The collected data is then analyzed, taking into account the user's profile information and the analysis results, to select the most relevant information. During this process, a generative AI model is used to generate optimal candidates based on the prompt "Events this weekend in the region the user is interested in".
[0318] Step 6:
[0319] The selected information is sent from the server to the terminal. The server formats the data into a user-friendly format and delivers it to the terminal as an interface display.
[0320] Step 7:
[0321] The terminal displays the received information on its user interface. Here, the information is visually organized for easy understanding, and detailed event information, such as "Family Day Event held in the park," is provided.
[0322] Step 8:
[0323] Users enter feedback about the displayed information on their device. This feedback is sent to the server, which then uses it to improve the server's information provision algorithm, thereby increasing the accuracy of future information provision.
[0324] (Application Example 1)
[0325] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0326] In modern society, users have access to vast amounts of information, but efficiently obtaining the most relevant information remains difficult. Furthermore, while there is a demand for information provision that takes into account users' interests and behavioral history, there is also a need to improve the accuracy of natural language queries and present information clearly through user interfaces. Additionally, enhancing voice input capabilities is desirable to improve user convenience.
[0327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0328] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user attribute information and generating personalized information, means for acquiring information from external sources and selecting information according to the user's intent, means for converting voice input into text, and means for displaying the selected information on a user interface. As a result, users can efficiently acquire information that suits their preferences, and convenience is improved through natural and simple interaction utilizing voice input.
[0329] A "user inquiry" is a question or request from a user regarding the information or services they seek from the system.
[0330] "Natural language processing" is a technology that understands the meaning and intent behind the language that humans use in everyday life and processes it as data.
[0331] "User attribute information" refers to data related to an individual, including their interests, preferences, and past behavioral history.
[0332] "Personalized information" refers to customized data that is presented in a way that is most suitable for the individual, based on their profile.
[0333] An "external information source" is a data provider that exists outside the system and provides information in real time.
[0334] "Voice input" is a method of transmitting information to a system through the voice spoken by the user.
[0335] "Means of converting to text" refers to technologies and processes that convert non-textual data, such as audio, into textual information.
[0336] A "user interface" is an interaction environment that provides screens and methods of operation for users to interact with a system.
[0337] The system for implementing this invention consists of three main components: a server, a terminal, and a user. The server uses a natural language processing engine to analyze inquiries entered by the user through the terminal. This can be done using natural language processing tools such as the Google Cloud Natural Language API. The analyzed information is stored on the server as personalized data, taking into account the user's attribute information.
[0338] The server further accesses external data sources and retrieves data as needed. Open data APIs can be used for this access. From the retrieved data, it selects the information that best matches the user's intent and sends it to the terminal for display in the user interface.
[0339] The device provides a voice input function, allowing users to make natural inquiries using their voice. This voice is converted to text using speech-to-text technologies such as the Google Speech-to-Text API and sent to the server. The converted text is then analyzed on the server, and the corresponding information is prepared.
[0340] For example, if a user asks a question via voice, such as "Tell me about jazz events happening nearby," the voice is converted to text and analyzed by the server's natural language processing engine. Based on the analysis results, the system references the user's past activity history to extract and display relevant jazz events.
[0341] An example of a prompt for a generative AI model is written as follows:
[0342] "When a user asks, 'What are some recommended music events near me?', you should recommend events they might be able to attend, taking into account their interest in jazz events in particular. You should also consider events the user has attended in the past to suggest the best options."
[0343] This allows users to quickly obtain information that is of particular interest to them, and to enjoy a more convenient information retrieval experience.
[0344] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0345] Step 1:
[0346] Users enter inquiries using their device via voice or text. For voice input, the device uses the Google Speech-to-Text API to convert speech to text. This process takes a speech waveform as input and outputs the converted text data.
[0347] Step 2:
[0348] The terminal sends the obtained text to the server. The server receives this text and parses it using a natural language processing engine (e.g., Google Cloud Natural Language API). The input for this step is the user's query text, and the output is data indicating the parsed keywords and user intent.
[0349] Step 3:
[0350] The server references user attribute information based on the analysis results and generates personalized information. The inputs here are the analyzed keywords and user profile information, and the output is a list of information optimized for the user.
[0351] Step 4:
[0352] The server accesses external information sources and collects information that matches the user's intent. Open data APIs, for example, are used. The input is the analysis results and the user's interest data, and the output is the collected external information.
[0353] Step 5:
[0354] The server compares and analyzes the collected external information and selects the most suitable information to display in the user interface. In this step, the input is a list of candidate information that forms the basis of the selection, and the output is the final information presented to the user.
[0355] Step 6:
[0356] The terminal displays the selection information received from the server in the user interface. Here, the information sent from the server is the input, and the visual presentation of that information to the user is the output. The user reviews the displayed information and takes action as needed.
[0357] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0358] This invention is a local information provision system that combines natural language processing and an emotion engine to provide users with more personalized information. First, the user queries for information through a terminal. The terminal sends the content to a server, which uses a natural language processing (NLP) engine to analyze the query. This analysis extracts the type of information the user is seeking and related keywords.
[0359] Furthermore, this invention uses an emotion engine to recognize emotions from the user's text input. The emotion engine is used to systematically understand the user's emotional state and select appropriate information according to that state. For example, if it is determined that the user is feeling stressed, it will provide relaxation events or healing information.
[0360] Next, the server examines the user's profile data. This data contains past behavioral history and user interests, and personalized information is generated based on this data. The server accesses an external database to retrieve the most relevant information selected based on the user's intentions and emotions.
[0361] Using the acquired information, the server formats it into a format suitable for display on the user interface and sends it to the terminal. The terminal receives this information and displays it visually to the user. The user can then review and utilize this information.
[0362] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm and adjust sentiment patterns. This allows the system to continuously improve its ability to meet user needs over the long term.
[0363] Specific example
[0364] For example, if a user enters a question into their device expressing stress, such as "What events can I enjoy with my kids this weekend?", the server analyzes the inquiry and extracts "stress" and "kid-friendly events" as keywords. Next, the emotion engine detects the stress and suggests relaxing family events based on the user's profile. The device might display information such as "Weekend Family Day at the Park: With Relaxation Corner." If the user provides feedback on this event, it will be incorporated into future information provision.
[0365] The following describes the processing flow.
[0366] Step 1:
[0367] The user queries the device for event information using natural language. The device retrieves this information and prepares to send it to the server along with the emotion engine.
[0368] Step 2:
[0369] The terminal sends the user's inquiry to the server. The server processes the received data and parses the text via a natural language processing engine.
[0370] Step 3:
[0371] The server uses natural language processing to extract the main keywords of the query. This identifies the type of event the user is looking for, as well as the available time and location.
[0372] Step 4:
[0373] The server uses an emotion engine to recognize the emotions contained in the user's input. Emotions such as stress, excitement, and satisfaction are evaluated.
[0374] Step 5:
[0375] The server accesses the user's profile database to investigate their past behavior and interests. This allows for the provision of information that is best suited to the user.
[0376] Step 6:
[0377] The server connects to external information sources in real time to retrieve relevant data such as event information. This data is filtered based on user requests and sentiments.
[0378] Step 7:
[0379] The server organizes the selected information and generates personalized event information according to the user's request. This information is then converted into an appropriate format and prepared for display.
[0380] Step 8:
[0381] The server sends the generated information to the terminal. The terminal receives this information and displays it in the user interface, allowing the user to view the information.
[0382] Step 9:
[0383] The user provides feedback on the information they have been given. The device sends this feedback to the server, which is then used to improve the system.
[0384] Step 10:
[0385] The server analyzes the feedback and uses it to adjust the information delivery algorithm and emotion engine. This process further improves the user's experience on subsequent visits.
[0386] (Example 2)
[0387] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0388] Conventional local information systems typically provide information to users, making it difficult to effectively deliver personalized information tailored to the individual user's emotions and interests. As a result, users often fail to receive the information they need, highlighting the need for improved user experience.
[0389] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0390] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing user emotions, and means for referencing user profile information and generating personalized information. This makes it possible to provide information that is tailored to the individual emotions and needs of each user.
[0391] "Natural language processing" is the technology that enables computers to understand, analyze, and generate human language.
[0392] "Emotion recognition" is a technology that extracts and analyzes emotional elements from information provided by users.
[0393] "Profile information" refers to data about individual users, including their interests, behavioral history, and personal characteristics.
[0394] "External information sources" refer to databases, APIs, and other information sources located outside the system that provide up-to-date or additional data.
[0395] "Feedback" refers to the opinions and evaluations that users provide to a system, which are used to improve the services and information provided.
[0396] "Personalized information" refers to information tailored to each user's individual characteristics and preferences, and is provided on the user interface.
[0397] "Filtering" refers to the process of selecting information data based on certain criteria and eliminating unnecessary information.
[0398] A "user interface" is a means of displaying and inputting information between a computer system and a user, and includes visual display screens and input devices.
[0399] This invention is an information provision system that combines natural language processing technology and an emotion recognition engine to provide users with personalized information. In this system, the user queries for information from a terminal, the server analyzes the content, selects appropriate information, and provides it to the user.
[0400] The user inputs information and makes inquiries using natural language via a terminal. This terminal sends data to the server using a common communication protocol, initiating information processing. The server uses a natural language processing engine implemented in a programming language such as Python to analyze the user's input and extract the intent of the inquiry and related keywords. In this process, libraries such as NLTK may be used.
[0401] In addition, the server uses an emotion recognition engine to analyze the user's emotional state from their text. This engine further improves the accuracy of the information the user is seeking by using machine learning models (such as the BERT model) powered by the Transformers library.
[0402] Next, the server checks the user's profile information. This information is stored in a database system (e.g., MySQL) and reflects past behavior and interests. Based on this data, the server generates personalized information that is most relevant to the user.
[0403] The server then accesses external information sources to retrieve the latest relevant events and news. This process utilizes various APIs to gather a wide range of information and select the information that best suits the user's needs.
[0404] The acquired information is formatted by the server into a user-friendly format and transmitted to the terminal. The terminal displays the information in the specified format and provides it to the user. The user can then make decisions and take actions based on this information.
[0405] Furthermore, users can input feedback using their devices. This information is sent to the server and used to improve the information provision algorithm. As a result, the information provision will continuously improve over the long term and become more valuable to users.
[0406] For example, if a user enters "Where can I relax with my family on the weekend?", the server will analyze this query and suggest events related to relaxation. For instance, it might provide information such as "Weekend events at nearby parks."
[0407] Example of a prompt:
[0408] If a user asks, "Can you recommend some places to relax on the weekend?", how would you provide them with relevant information?
[0409] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0410] Step 1:
[0411] User input of inquiry
[0412] Specific operation: The user requests specific information on the device's input screen. For example, they might type text such as, "What are some places where my family can relax on the weekend?"
[0413] Input and Output: The text entered by the user is input into the terminal. As output, the terminal sends this query data to the server.
[0414] Step 2:
[0415] Sending information from the terminal to the server
[0416] Specific operation: The terminal sends the received user query to the server via the appropriate communication protocol (e.g., HTTPS).
[0417] Input and Output: The input is the query text received from the user. The output is the data received by the server for analysis.
[0418] Step 3:
[0419] Execution of natural language processing by a server
[0420] Specific operation: The server executes a Python script and uses a natural language processing library (e.g., NLTK) to parse the query text. Specific techniques include text tokenization and keyword extraction.
[0421] Input and Output: The input is text data received from the terminal. The output generates data containing the analyzed keywords and intent.
[0422] Step 4:
[0423] Emotion recognition performed by the server
[0424] Specific operation: The server uses a machine learning model (e.g., a BERT-based model) to recognize emotions from the user's text. Emotional scoring is performed, and the emotional state of the text is quantified.
[0425] Input and Output: The input is parsed text, and the output is a sentiment score or state.
[0426] Step 5:
[0427] User profile referencing by the server
[0428] Specific operation: The server executes database queries (e.g., MySQL) to retrieve information about the user's past behavior and interests.
[0429] Input and Output: The input is identification information such as a user ID, and the output is profile information.
[0430] Step 6:
[0431] Obtaining information from external sources
[0432] Specific operation: The server uses various APIs to request relevant information from external databases. The retrieved information is prioritized to match the user's current interests and emotional state.
[0433] Input and Output: The input consists of selected keywords or conditions. The output consists of retrieved event information or news.
[0434] Step 7:
[0435] Formatting and sending information to the device
[0436] Specific operation: The server uses an HTML template to visually format the information to be presented to the user and sends it to the terminal via the HTTP protocol.
[0437] Input and Output: Input is the acquired information data, and output is the information formatted in a way that the user can visually understand.
[0438] Step 8:
[0439] Displaying information on the device
[0440] Specific operation: The terminal displays information received from the server on the screen, making it available for the user to review. Detailed information and related information can be accessed through the user interface.
[0441] Input and Output: Input is formatted information received from the server, and output is content presented visually to the user.
[0442] Step 9:
[0443] User feedback
[0444] Specific operation: The user enters feedback on the information provided via the terminal and sends it to the server by operating the submit button.
[0445] Input and Output: Input is user ratings and opinions, and output is feedback data recorded by the server.
[0446] (Application Example 2)
[0447] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0448] In modern urban life, users need information tailored to their emotional state and individual interests at any given time. However, conventional systems struggle to accurately recognize the diverse emotions of users and provide appropriate information. Furthermore, there is a lack of methods to effectively utilize user feedback and provide information that is more suitable for individual users.
[0449] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0450] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing the user's emotional state and reflecting it in information selection, and means for referencing the user's profile data and generating personalized information. This makes it possible to provide more appropriate and personalized information based on the user's emotions and profile.
[0451] "Natural language processing" is a technology that enables computers to understand, analyze, and respond to human language.
[0452] "Emotional state" refers to the psychological state inferred from the user's input, and includes emotions such as stress, relaxation, and excitement.
[0453] "Profile data" refers to a collection of individual information, including a user's past behavioral history, interests, and preferences.
[0454] "Personalized information" refers to information that is optimized based on the user's individual attributes and emotions.
[0455] "External information sources" refer to external information sources that the server can access, such as databases and web resources.
[0456] A "user interface" is a means of displaying information visually to a user, allowing the user to confirm and manipulate that information.
[0457] "Feedback" refers to the opinions and evaluations given by users, and is information used to improve the system.
[0458] "In-city experiential information" refers to information related to tourist attractions, events, transportation, etc., within a city.
[0459] This invention provides an information delivery system that performs emotion recognition and personalization for residents of smart cities. Users use their smartphones to query information to customize their individual urban experience. The server receives the user's query in natural language and analyzes it using a natural language processing engine. In this process, it extracts the type of information requested by the user and related keywords. Furthermore, it uses an emotion engine to recognize the user's emotional state from their input. The server considers the user's profile data and selects and provides personalized information that is appropriate for their emotional state.
[0460] Specifically, the server filters information on tourist attractions, events, and transportation based on user profiles and data obtained from external sources, and formats urban experience information that matches the user's emotions. This information is transmitted to the terminal through the user interface, allowing the user to visually confirm it.
[0461] For example, if a user enters the question, "Where can I relax in Tokyo on the weekend?", the server processes this question using natural language processing and extracts "relax" and "Tokyo" as keywords. Then, using an emotion engine, it determines that the user is seeking relaxation and suggests quiet parks and healing spots based on their profile data. Information such as "Tokyo Central Park: with a quiet rest area" is displayed on the device. The user can then use these suggestions to plan their weekend.
[0462] A concrete example of a prompt is, "Please tell me some relaxing places in Tokyo that I should visit this weekend." This prompt allows the user to have a suitable urban experience.
[0463] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0464] Step 1:
[0465] The user enters the query in natural language using a smartphone. The input is received in text format, and the device sends this data to the server. The output is the query data sent to the server.
[0466] Step 2:
[0467] The server passes the received query data to a natural language processing engine for analysis. The input is the user's query text, and the output is the extraction of keywords and information types. In this step, the server analyzes important words and phrases contained in the query to determine what the user is asking for.
[0468] Step 3:
[0469] The server uses an emotion engine to recognize the user's emotional state from their input. The input is the user's query text, and the output is the recognized emotional state. In this step, the server analyzes the user's emotions from the content and context of the text and evaluates how those emotions will affect the information provided.
[0470] Step 4:
[0471] The server retrieves user profile data from the database and analyzes past behavior and interests. The input is user identification information, and the output is user-specific profile data. This step prepares the foundational data for personalizing information, taking into account the user's interests and behavioral tendencies.
[0472] Step 5:
[0473] The server accesses external information sources to obtain optimal urban experience information based on analyzed keywords, emotional states, and profile data. The input is a combination of keywords and emotional data, and the output is filtered information results. This step selects information tailored to the user's needs and creates customized recommendations.
[0474] Step 6:
[0475] The server formats the filtered information results into a format suitable for the user interface and sends it to the terminal. The input is the filtered information results, and the output is in a displayable information format. This step ensures that the information is displayed in a way that is easy for the user to understand.
[0476] Step 7:
[0477] The terminal displays formatted information on a user interface, providing users with visual access. Input is in a displayable information format, and output is the display of information to the user. This allows users to easily review and utilize proposed urban experiences.
[0478] Step 8:
[0479] The user enters feedback on the provided information via a terminal. The input is the text of the feedback, and the terminal sends this feedback to the server. The output is the feedback data sent to the server.
[0480] Step 9:
[0481] The server uses the received feedback data to improve its information provision algorithm. The input is user feedback data, and the output is the improved algorithm. In this step, the feedback is analyzed and improvements are made to make future information provision more appropriate.
[0482] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0483] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0484] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0485] [Third Embodiment]
[0486] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0487] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0488] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0489] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0490] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0491] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0492] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0493] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0494] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0495] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0496] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0497] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0498] This invention constructs a local information provision system and is implemented according to the following procedure. First, a user makes an inquiry through a terminal. Since this inquiry is entered in natural language, the terminal sends its contents to the server. The server analyzes the input inquiry via a natural language processing (NLP) engine and understands the user's intent. This clarifies what kind of information the user is seeking.
[0499] Next, the server accesses the user's profile data, which includes information related to the user's past behavior, interests, and preferences. Based on this data, personalized information useful to the user is generated. The server then accesses external information sources to obtain relevant information in real time and selects the most relevant information based on the user's profile.
[0500] The selected information is sent from the server to the terminal. The terminal displays this information on its user interface, providing it in a format that is easy for the user to understand. During this process, details of the event, such as the date and time and participation requirements, are displayed, and the user can use the information as needed.
[0501] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm. As a result, the system can provide more accurate information that better suits the user's needs over time.
[0502] Specific example
[0503] When a user enters "Are there any events I can take my child to this weekend?" into their device, the server receives the query and uses natural language processing to extract the keywords "this weekend," "children," and "event." Next, the server identifies the appropriate event type from the user's past participation history and retrieves relevant event information by referring to an external event database. The device then displays "Family Day Events Held at the Park," providing details such as the date, time, location, and how to participate. By reviewing this information and providing feedback, the system can provide more accurate information in the future.
[0504] The following describes the processing flow.
[0505] Step 1:
[0506] The user inputs information into the device using natural language. The device receives this input and prepares to send it to the server as a query.
[0507] Step 2:
[0508] The terminal sends the user's query to the server. The server receives this query and calls a natural language processing (NLP) engine to analyze the input.
[0509] Step 3:
[0510] The server uses NLP to analyze the intent of the query and extract keywords. This identifies the type of information the user is seeking.
[0511] Step 4:
[0512] The server accesses the user's profile database to retrieve the user's past behavior history and interests. This information is then used to prepare for personalizing the user's information.
[0513] Step 5:
[0514] The server accesses external information sources and retrieves the latest relevant data. It extracts the relevant events and information and selects the data that matches the user's intent.
[0515] Step 6:
[0516] The server organizes the selected information and generates personalized information. The generated information is structured in a format that is easy for the user to understand.
[0517] Step 7:
[0518] The server generates information and sends it to the terminal. The terminal receives this information and displays it in the user interface. The user can then review the presented information.
[0519] Step 8:
[0520] The user provides feedback on the information provided. The device sends this feedback to the server, which is used to improve the information provision algorithm.
[0521] (Example 1)
[0522] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0523] Providing local information in a timely and accurate manner that meets the individual needs and interests of users is challenging. Furthermore, if the provided information does not meet user expectations, it is necessary to improve the quality of the information based on feedback, which is also not easy. In addition, there is an increasing demand for information that takes the user's current location into account, and responding to this is essential.
[0524] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0525] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user profile information and generating personalized information, means for acquiring data from external information sources and selecting optimal information according to the user's intent, and means for generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences. This enables the provision of accurate and timely information tailored to the individual needs of the user, and furthermore, allows for continuous improvement of the quality of the information provided through feedback.
[0526] "Methods for analyzing user inquiries using natural language processing" refers to technologies that analyze user questions written in natural language and convert their meaning and intent into a format that the system can understand.
[0527] "Means of referencing user profile information and generating personalized information" refers to methods of creating information tailored to the individual needs of users by utilizing data on their past behavioral history and interests.
[0528] "A means of acquiring data from external information sources and selecting the most suitable information according to the user's intent" refers to the process of gathering necessary data from external information sources, analyzing it, and providing the information that best matches the user's requirements.
[0529] "A method of generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences" refers to a technique that utilizes AI technology to generate the most suitable information for the user, and the process by which the AI finds the best answer based on the prompt sentences.
[0530] "Means of visually displaying information on a user interface" refers to technologies for displaying generated information on a screen in a way that is easy for the user to understand.
[0531] "Means of optionally acquiring the user's geographical location information to improve the accuracy of information provision" refers to methods of acquiring and utilizing location information in order to provide more appropriate information based on the user's current location.
[0532] This invention is a system that provides local information tailored to user needs and is implemented based on the following embodiments.
[0533] Users can query information using natural language through their devices. The device receives this query and sends the data to the server to accurately analyze the user's intent. The server uses a natural language processing (NLP) engine to analyze the language data received from the user. Through this analysis, the server can extract keywords and intent from the user's query.
[0534] Next, the server references the user's profile information and generates personalized information based on this information. Specifically, information is generated that takes into account the user's past behavior history and interests. In addition, the server accesses external information sources and selects the most suitable information for the user based on data collected in real time. As part of this process, a generative AI model is used to find the best candidates when the prompt message "Tell me about events happening this weekend in the area I'm interested in" is entered.
[0535] The selected information is displayed in a visually easy-to-understand format on the user interface and provided to the user through the device. This allows users to easily obtain and use the information they need. For example, detailed information such as "Family Day Event held in the park" is presented along with the date, time, and location.
[0536] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to continuously improve the information provision algorithm. As a result, the system learns with each use and can provide information that better meets user expectations the next time it is used.
[0537] Through this configuration, the system can efficiently continue to provide valuable information to users.
[0538] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0539] Step 1:
[0540] The user inputs information in natural language on their device. For example, they might ask, "Tell me about weekend events happening in this area." This input is then sent to the server as data for the next processing step.
[0541] Step 2:
[0542] The terminal sends the received user query to the server in its original format. This process involves encoding based on the data communication protocol. This prepares the server to analyze the user query.
[0543] Step 3:
[0544] The server uses a natural language processing (NLP) engine to analyze user inquiries. Keywords such as "weekend," "region," and "event" are extracted from the input inquiry. These extracted keywords are used to understand the user's intent and guide the subsequent processing.
[0545] Step 4:
[0546] The server references the user's profile information to generate personalized data. Input includes the user's past behavior history and interests. Based on this, customized information tailored to each user's individual needs is output.
[0547] Step 5:
[0548] The server accesses external information sources and retrieves event data in real time. The collected data is then analyzed, taking into account the user's profile information and the analysis results, to select the most relevant information. During this process, a generative AI model is used to generate optimal candidates based on the prompt "Events this weekend in the region the user is interested in".
[0549] Step 6:
[0550] The selected information is sent from the server to the terminal. The server formats the data into a user-friendly format and delivers it to the terminal as an interface display.
[0551] Step 7:
[0552] The terminal displays the received information on its user interface. Here, the information is visually organized for easy understanding, and detailed event information, such as "Family Day Event held in the park," is provided.
[0553] Step 8:
[0554] Users enter feedback about the displayed information on their device. This feedback is sent to the server, which then uses it to improve the server's information provision algorithm, thereby increasing the accuracy of future information provision.
[0555] (Application Example 1)
[0556] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0557] In modern society, users have access to vast amounts of information, but efficiently obtaining the most relevant information remains difficult. Furthermore, while there is a demand for information provision that takes into account users' interests and behavioral history, there is also a need to improve the accuracy of natural language queries and present information clearly through user interfaces. Additionally, enhancing voice input capabilities is desirable to improve user convenience.
[0558] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0559] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user attribute information and generating personalized information, means for acquiring information from external sources and selecting information according to the user's intent, means for converting voice input into text, and means for displaying the selected information on a user interface. As a result, users can efficiently acquire information that suits their preferences, and convenience is improved through natural and simple interaction utilizing voice input.
[0560] A "user inquiry" is a question or request from a user regarding the information or services they seek from the system.
[0561] "Natural language processing" is a technology that understands the meaning and intent behind the language that humans use in everyday life and processes it as data.
[0562] "User attribute information" refers to data related to an individual, including their interests, preferences, and past behavioral history.
[0563] "Personalized information" refers to customized data that is presented in a way that is most suitable for the individual, based on their profile.
[0564] An "external information source" is a data provider that exists outside the system and provides information in real time.
[0565] "Voice input" is a method of transmitting information to a system through the voice spoken by the user.
[0566] "Means of converting to text" refers to technologies and processes that convert non-textual data, such as audio, into textual information.
[0567] A "user interface" is an interaction environment that provides screens and methods of operation for users to interact with a system.
[0568] The system for implementing this invention consists of three main components: a server, a terminal, and a user. The server uses a natural language processing engine to analyze inquiries entered by the user through the terminal. This can be done using natural language processing tools such as the Google Cloud Natural Language API. The analyzed information is stored on the server as personalized data, taking into account the user's attribute information.
[0569] The server further accesses external data sources and retrieves data as needed. Open data APIs can be used for this access. From the retrieved data, it selects the information that best matches the user's intent and sends it to the terminal for display in the user interface.
[0570] The device provides a voice input function, allowing users to make natural inquiries using their voice. This voice is converted to text using speech-to-text technologies such as the Google Speech-to-Text API and sent to the server. The converted text is then analyzed on the server, and the corresponding information is prepared.
[0571] For example, if a user asks a question via voice, such as "Tell me about jazz events happening nearby," the voice is converted to text and analyzed by the server's natural language processing engine. Based on the analysis results, the system references the user's past activity history to extract and display relevant jazz events.
[0572] An example of a prompt for a generative AI model is written as follows:
[0573] "When a user asks, 'What are some recommended music events near me?', you should recommend events they might be able to attend, taking into account their interest in jazz events in particular. You should also consider events the user has attended in the past to suggest the best options."
[0574] This allows users to quickly obtain information that is of particular interest to them, and to enjoy a more convenient information retrieval experience.
[0575] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0576] Step 1:
[0577] Users enter inquiries using their device via voice or text. For voice input, the device uses the Google Speech-to-Text API to convert speech to text. This process takes a speech waveform as input and outputs the converted text data.
[0578] Step 2:
[0579] The terminal sends the obtained text to the server. The server receives this text and parses it using a natural language processing engine (e.g., Google Cloud Natural Language API). The input for this step is the user's query text, and the output is data indicating the parsed keywords and user intent.
[0580] Step 3:
[0581] The server references user attribute information based on the analysis results and generates personalized information. The inputs here are the analyzed keywords and user profile information, and the output is a list of information optimized for the user.
[0582] Step 4:
[0583] The server accesses external information sources and collects information that matches the user's intent. Open data APIs, for example, are used. The input is the analysis results and the user's interest data, and the output is the collected external information.
[0584] Step 5:
[0585] The server compares and analyzes the collected external information and selects the most suitable information to display in the user interface. In this step, the input is a list of candidate information that forms the basis of the selection, and the output is the final information presented to the user.
[0586] Step 6:
[0587] The terminal displays the selection information received from the server in the user interface. Here, the information sent from the server is the input, and the visual presentation of that information to the user is the output. The user reviews the displayed information and takes action as needed.
[0588] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0589] This invention is a local information provision system that combines natural language processing and an emotion engine to provide users with more personalized information. First, the user queries for information through a terminal. The terminal sends the content to a server, which uses a natural language processing (NLP) engine to analyze the query. This analysis extracts the type of information the user is seeking and related keywords.
[0590] Furthermore, this invention uses an emotion engine to recognize emotions from the user's text input. The emotion engine is used to systematically understand the user's emotional state and select appropriate information according to that state. For example, if it is determined that the user is feeling stressed, it will provide relaxation events or healing information.
[0591] Next, the server examines the user's profile data. This data contains past behavioral history and user interests, and personalized information is generated based on this data. The server accesses an external database to retrieve the most relevant information selected based on the user's intentions and emotions.
[0592] Using the acquired information, the server formats it into a format suitable for display on the user interface and sends it to the terminal. The terminal receives this information and displays it visually to the user. The user can then review and utilize this information.
[0593] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm and adjust sentiment patterns. This allows the system to continuously improve its ability to meet user needs over the long term.
[0594] Specific example
[0595] For example, if a user enters a question into their device expressing stress, such as "What events can I enjoy with my kids this weekend?", the server analyzes the inquiry and extracts "stress" and "kid-friendly events" as keywords. Next, the emotion engine detects the stress and suggests relaxing family events based on the user's profile. The device might display information such as "Weekend Family Day at the Park: With Relaxation Corner." If the user provides feedback on this event, it will be incorporated into future information provision.
[0596] The following describes the processing flow.
[0597] Step 1:
[0598] The user queries the device for event information using natural language. The device retrieves this information and prepares to send it to the server along with the emotion engine.
[0599] Step 2:
[0600] The terminal sends the user's inquiry to the server. The server processes the received data and parses the text via a natural language processing engine.
[0601] Step 3:
[0602] The server uses natural language processing to extract the main keywords of the query. This identifies the type of event the user is looking for, as well as the available time and location.
[0603] Step 4:
[0604] The server uses an emotion engine to recognize the emotions contained in the user's input. Emotions such as stress, excitement, and satisfaction are evaluated.
[0605] Step 5:
[0606] The server accesses the user's profile database to investigate their past behavior and interests. This allows for the provision of information that is best suited to the user.
[0607] Step 6:
[0608] The server connects to external information sources in real time to retrieve relevant data such as event information. This data is filtered based on user requests and sentiments.
[0609] Step 7:
[0610] The server organizes the selected information and generates personalized event information according to the user's request. This information is then converted into an appropriate format and prepared for display.
[0611] Step 8:
[0612] The server sends the generated information to the terminal. The terminal receives this information and displays it in the user interface, allowing the user to view the information.
[0613] Step 9:
[0614] The user provides feedback on the information they have been given. The device sends this feedback to the server, which is then used to improve the system.
[0615] Step 10:
[0616] The server analyzes the feedback and uses it to adjust the information delivery algorithm and emotion engine. This process further improves the user's experience on subsequent visits.
[0617] (Example 2)
[0618] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0619] Conventional local information systems typically provide information to users, making it difficult to effectively deliver personalized information tailored to the individual user's emotions and interests. As a result, users often fail to receive the information they need, highlighting the need for improved user experience.
[0620] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0621] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing user emotions, and means for referencing user profile information and generating personalized information. This makes it possible to provide information that is tailored to the individual emotions and needs of each user.
[0622] "Natural language processing" is the technology that enables computers to understand, analyze, and generate human language.
[0623] "Emotion recognition" is a technology that extracts and analyzes emotional elements from information provided by users.
[0624] "Profile information" refers to data about individual users, including their interests, behavioral history, and personal characteristics.
[0625] "External information sources" refer to databases, APIs, and other information sources located outside the system that provide up-to-date or additional data.
[0626] "Feedback" refers to the opinions and evaluations that users provide to a system, which are used to improve the services and information provided.
[0627] "Personalized information" refers to information tailored to each user's individual characteristics and preferences, and is provided on the user interface.
[0628] "Filtering" refers to the process of selecting information data based on certain criteria and eliminating unnecessary information.
[0629] A "user interface" is a means of displaying and inputting information between a computer system and a user, and includes visual display screens and input devices.
[0630] This invention is an information provision system that combines natural language processing technology and an emotion recognition engine to provide users with personalized information. In this system, the user queries for information from a terminal, the server analyzes the content, selects appropriate information, and provides it to the user.
[0631] The user inputs information and makes inquiries using natural language via a terminal. This terminal sends data to the server using a common communication protocol, initiating information processing. The server uses a natural language processing engine implemented in a programming language such as Python to analyze the user's input and extract the intent of the inquiry and related keywords. In this process, libraries such as NLTK may be used.
[0632] In addition, the server uses an emotion recognition engine to analyze the user's emotional state from their text. This engine further improves the accuracy of the information the user is seeking by using machine learning models (such as the BERT model) powered by the Transformers library.
[0633] Next, the server checks the user's profile information. This information is stored in a database system (e.g., MySQL) and reflects past behavior and interests. Based on this data, the server generates personalized information that is most relevant to the user.
[0634] The server then accesses external information sources to retrieve the latest relevant events and news. This process utilizes various APIs to gather a wide range of information and select the information that best suits the user's needs.
[0635] The acquired information is formatted by the server into a user-friendly format and transmitted to the terminal. The terminal displays the information in the specified format and provides it to the user. The user can then make decisions and take actions based on this information.
[0636] Furthermore, users can input feedback using their devices. This information is sent to the server and used to improve the information provision algorithm. As a result, the information provision will continuously improve over the long term and become more valuable to users.
[0637] For example, if a user enters "Where can I relax with my family on the weekend?", the server will analyze this query and suggest events related to relaxation. For instance, it might provide information such as "Weekend events at nearby parks."
[0638] Example of a prompt:
[0639] If a user asks, "Can you recommend some places to relax on the weekend?", how would you provide them with relevant information?
[0640] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0641] Step 1:
[0642] User input of inquiry
[0643] Specific operation: The user requests specific information on the device's input screen. For example, they might type text such as, "What are some places where my family can relax on the weekend?"
[0644] Input and Output: The text entered by the user is input into the terminal. As output, the terminal sends this query data to the server.
[0645] Step 2:
[0646] Sending information from the terminal to the server
[0647] Specific operation: The terminal sends the received user query to the server via the appropriate communication protocol (e.g., HTTPS).
[0648] Input and Output: The input is the query text received from the user. The output is the data received by the server for analysis.
[0649] Step 3:
[0650] Execution of natural language processing by a server
[0651] Specific operation: The server executes a Python script and uses a natural language processing library (e.g., NLTK) to parse the query text. Specific techniques include text tokenization and keyword extraction.
[0652] Input and Output: The input is text data received from the terminal. The output generates data containing the analyzed keywords and intent.
[0653] Step 4:
[0654] Emotion recognition performed by the server
[0655] Specific operation: The server uses a machine learning model (e.g., a BERT-based model) to recognize emotions from the user's text. Emotional scoring is performed, and the emotional state of the text is quantified.
[0656] Input and Output: The input is parsed text, and the output is a sentiment score or state.
[0657] Step 5:
[0658] User profile referencing by the server
[0659] Specific operation: The server executes database queries (e.g., MySQL) to retrieve information about the user's past behavior and interests.
[0660] Input and Output: The input is identification information such as a user ID, and the output is profile information.
[0661] Step 6:
[0662] Obtaining information from external sources
[0663] Specific operation: The server uses various APIs to request relevant information from external databases. The retrieved information is prioritized to match the user's current interests and emotional state.
[0664] Input and Output: The input consists of selected keywords or conditions. The output consists of retrieved event information or news.
[0665] Step 7:
[0666] Formatting and sending information to the device
[0667] Specific operation: The server uses an HTML template to visually format the information to be presented to the user and sends it to the terminal via the HTTP protocol.
[0668] Input and Output: Input is the acquired information data, and output is the information formatted in a way that the user can visually understand.
[0669] Step 8:
[0670] Displaying information on the device
[0671] Specific operation: The terminal displays information received from the server on the screen, making it available for the user to review. Detailed information and related information can be accessed through the user interface.
[0672] Input and Output: Input is formatted information received from the server, and output is content presented visually to the user.
[0673] Step 9:
[0674] User feedback
[0675] Specific operation: The user enters feedback on the information provided via the terminal and sends it to the server by operating the submit button.
[0676] Input and Output: Input is user ratings and opinions, and output is feedback data recorded by the server.
[0677] (Application Example 2)
[0678] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0679] In modern urban life, users need information tailored to their emotional state and individual interests at any given time. However, conventional systems struggle to accurately recognize the diverse emotions of users and provide appropriate information. Furthermore, there is a lack of methods to effectively utilize user feedback and provide information that is more suitable for individual users.
[0680] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0681] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing the user's emotional state and reflecting it in information selection, and means for referencing the user's profile data and generating personalized information. This makes it possible to provide more appropriate and personalized information based on the user's emotions and profile.
[0682] "Natural language processing" is a technology that enables computers to understand, analyze, and respond to human language.
[0683] "Emotional state" refers to the psychological state inferred from the user's input, and includes emotions such as stress, relaxation, and excitement.
[0684] "Profile data" refers to a collection of individual information, including a user's past behavioral history, interests, and preferences.
[0685] "Personalized information" refers to information that is optimized based on the user's individual attributes and emotions.
[0686] "External information sources" refer to external information sources that the server can access, such as databases and web resources.
[0687] A "user interface" is a means of displaying information visually to a user, allowing the user to confirm and manipulate that information.
[0688] "Feedback" refers to the opinions and evaluations given by users, and is information used to improve the system.
[0689] "In-city experiential information" refers to information related to tourist attractions, events, transportation, etc., within a city.
[0690] This invention provides an information delivery system that performs emotion recognition and personalization for residents of smart cities. Users use their smartphones to query information to customize their individual urban experience. The server receives the user's query in natural language and analyzes it using a natural language processing engine. In this process, it extracts the type of information requested by the user and related keywords. Furthermore, it uses an emotion engine to recognize the user's emotional state from their input. The server considers the user's profile data and selects and provides personalized information that is appropriate for their emotional state.
[0691] Specifically, the server filters information on tourist attractions, events, and transportation based on user profiles and data obtained from external sources, and formats urban experience information that matches the user's emotions. This information is transmitted to the terminal through the user interface, allowing the user to visually confirm it.
[0692] For example, if a user enters the question, "Where can I relax in Tokyo on the weekend?", the server processes this question using natural language processing and extracts "relax" and "Tokyo" as keywords. Then, using an emotion engine, it determines that the user is seeking relaxation and suggests quiet parks and healing spots based on their profile data. Information such as "Tokyo Central Park: with a quiet rest area" is displayed on the device. The user can then use these suggestions to plan their weekend.
[0693] A concrete example of a prompt is, "Please tell me some relaxing places in Tokyo that I should visit this weekend." This prompt allows the user to have a suitable urban experience.
[0694] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0695] Step 1:
[0696] The user enters the query in natural language using a smartphone. The input is received in text format, and the device sends this data to the server. The output is the query data sent to the server.
[0697] Step 2:
[0698] The server passes the received query data to a natural language processing engine for analysis. The input is the user's query text, and the output is the extraction of keywords and information types. In this step, the server analyzes important words and phrases contained in the query to determine what the user is asking for.
[0699] Step 3:
[0700] The server uses an emotion engine to recognize the user's emotional state from their input. The input is the user's query text, and the output is the recognized emotional state. In this step, the server analyzes the user's emotions from the content and context of the text and evaluates how those emotions will affect the information provided.
[0701] Step 4:
[0702] The server retrieves user profile data from the database and analyzes past behavior and interests. The input is user identification information, and the output is user-specific profile data. This step prepares the foundational data for personalizing information, taking into account the user's interests and behavioral tendencies.
[0703] Step 5:
[0704] The server accesses external information sources to obtain optimal urban experience information based on analyzed keywords, emotional states, and profile data. The input is a combination of keywords and emotional data, and the output is filtered information results. This step selects information tailored to the user's needs and creates customized recommendations.
[0705] Step 6:
[0706] The server formats the filtered information results into a format suitable for the user interface and sends it to the terminal. The input is the filtered information results, and the output is in a displayable information format. This step ensures that the information is displayed in a way that is easy for the user to understand.
[0707] Step 7:
[0708] The terminal displays formatted information on a user interface, providing users with visual access. Input is in a displayable information format, and output is the display of information to the user. This allows users to easily review and utilize proposed urban experiences.
[0709] Step 8:
[0710] The user enters feedback on the provided information via a terminal. The input is the text of the feedback, and the terminal sends this feedback to the server. The output is the feedback data sent to the server.
[0711] Step 9:
[0712] The server uses the received feedback data to improve its information provision algorithm. The input is user feedback data, and the output is the improved algorithm. In this step, the feedback is analyzed and improvements are made to make future information provision more appropriate.
[0713] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0714] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0715] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0716] [Fourth Embodiment]
[0717] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0718] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0719] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0720] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0721] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0722] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0723] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0724] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0725] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0726] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0727] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0728] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0729] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0730] This invention constructs a local information provision system and is implemented according to the following procedure. First, a user makes an inquiry through a terminal. Since this inquiry is entered in natural language, the terminal sends its contents to the server. The server analyzes the input inquiry via a natural language processing (NLP) engine and understands the user's intent. This clarifies what kind of information the user is seeking.
[0731] Next, the server accesses the user's profile data, which includes information related to the user's past behavior, interests, and preferences. Based on this data, personalized information useful to the user is generated. The server then accesses external information sources to obtain relevant information in real time and selects the most relevant information based on the user's profile.
[0732] The selected information is sent from the server to the terminal. The terminal displays this information on its user interface, providing it in a format that is easy for the user to understand. During this process, details of the event, such as the date and time and participation requirements, are displayed, and the user can use the information as needed.
[0733] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm. As a result, the system can provide more accurate information that better suits the user's needs over time.
[0734] Specific example
[0735] When a user enters "Are there any events I can take my child to this weekend?" into their device, the server receives the query and uses natural language processing to extract the keywords "this weekend," "children," and "event." Next, the server identifies the appropriate event type from the user's past participation history and retrieves relevant event information by referring to an external event database. The device then displays "Family Day Events Held at the Park," providing details such as the date, time, location, and how to participate. By reviewing this information and providing feedback, the system can provide more accurate information in the future.
[0736] The following describes the processing flow.
[0737] Step 1:
[0738] The user inputs information into the device using natural language. The device receives this input and prepares to send it to the server as a query.
[0739] Step 2:
[0740] The terminal sends the user's query to the server. The server receives this query and calls a natural language processing (NLP) engine to analyze the input.
[0741] Step 3:
[0742] The server uses NLP to analyze the intent of the query and extract keywords. This identifies the type of information the user is seeking.
[0743] Step 4:
[0744] The server accesses the user's profile database to retrieve the user's past behavior history and interests. This information is then used to prepare for personalizing the user's information.
[0745] Step 5:
[0746] The server accesses external information sources and retrieves the latest relevant data. It extracts the relevant events and information and selects the data that matches the user's intent.
[0747] Step 6:
[0748] The server organizes the selected information and generates personalized information. The generated information is structured in a format that is easy for the user to understand.
[0749] Step 7:
[0750] The server generates information and sends it to the terminal. The terminal receives this information and displays it in the user interface. The user can then review the presented information.
[0751] Step 8:
[0752] The user provides feedback on the information provided. The device sends this feedback to the server, which is used to improve the information provision algorithm.
[0753] (Example 1)
[0754] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0755] Providing local information in a timely and accurate manner that meets the individual needs and interests of users is challenging. Furthermore, if the provided information does not meet user expectations, it is necessary to improve the quality of the information based on feedback, which is also not easy. In addition, there is an increasing demand for information that takes the user's current location into account, and responding to this is essential.
[0756] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0757] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user profile information and generating personalized information, means for acquiring data from external information sources and selecting optimal information according to the user's intent, and means for generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences. This enables the provision of accurate and timely information tailored to the individual needs of the user, and furthermore, allows for continuous improvement of the quality of the information provided through feedback.
[0758] "Methods for analyzing user inquiries using natural language processing" refers to technologies that analyze user questions written in natural language and convert their meaning and intent into a format that the system can understand.
[0759] "Means of referencing user profile information and generating personalized information" refers to methods of creating information tailored to the individual needs of users by utilizing data on their past behavioral history and interests.
[0760] "A means of acquiring data from external information sources and selecting the most suitable information according to the user's intent" refers to the process of gathering necessary data from external information sources, analyzing it, and providing the information that best matches the user's requirements.
[0761] "A method of generating information tailored to the user using a generative AI model and searching for appropriate candidates by inputting prompt sentences" refers to a technique that utilizes AI technology to generate the most suitable information for the user, and the process by which the AI finds the best answer based on the prompt sentences.
[0762] "Means of visually displaying information on a user interface" refers to technologies for displaying generated information on a screen in a way that is easy for the user to understand.
[0763] "Means of optionally acquiring the user's geographical location information to improve the accuracy of information provision" refers to methods of acquiring and utilizing location information in order to provide more appropriate information based on the user's current location.
[0764] This invention is a system that provides local information tailored to user needs and is implemented based on the following embodiments.
[0765] Users can query information using natural language through their devices. The device receives this query and sends the data to the server to accurately analyze the user's intent. The server uses a natural language processing (NLP) engine to analyze the language data received from the user. Through this analysis, the server can extract keywords and intent from the user's query.
[0766] Next, the server references the user's profile information and generates personalized information based on this information. Specifically, information is generated that takes into account the user's past behavior history and interests. In addition, the server accesses external information sources and selects the most suitable information for the user based on data collected in real time. As part of this process, a generative AI model is used to find the best candidates when the prompt message "Tell me about events happening this weekend in the area I'm interested in" is entered.
[0767] The selected information is displayed in a visually easy-to-understand format on the user interface and provided to the user through the device. This allows users to easily obtain and use the information they need. For example, detailed information such as "Family Day Event held in the park" is presented along with the date, time, and location.
[0768] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to continuously improve the information provision algorithm. As a result, the system learns with each use and can provide information that better meets user expectations the next time it is used.
[0769] Through this configuration, the system can efficiently continue to provide valuable information to users.
[0770] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0771] Step 1:
[0772] The user inputs information in natural language on their device. For example, they might ask, "Tell me about weekend events happening in this area." This input is then sent to the server as data for the next processing step.
[0773] Step 2:
[0774] The terminal sends the received user query to the server in its original format. This process involves encoding based on the data communication protocol. This prepares the server to analyze the user query.
[0775] Step 3:
[0776] The server uses a natural language processing (NLP) engine to analyze user inquiries. Keywords such as "weekend," "region," and "event" are extracted from the input inquiry. These extracted keywords are used to understand the user's intent and guide the subsequent processing.
[0777] Step 4:
[0778] The server references the user's profile information to generate personalized data. Input includes the user's past behavior history and interests. Based on this, customized information tailored to each user's individual needs is output.
[0779] Step 5:
[0780] The server accesses external information sources and retrieves event data in real time. The collected data is then analyzed, taking into account the user's profile information and the analysis results, to select the most relevant information. During this process, a generative AI model is used to generate optimal candidates based on the prompt "Events this weekend in the region the user is interested in".
[0781] Step 6:
[0782] The selected information is sent from the server to the terminal. The server formats the data into a user-friendly format and delivers it to the terminal as an interface display.
[0783] Step 7:
[0784] The terminal displays the received information on its user interface. Here, the information is visually organized for easy understanding, and detailed event information, such as "Family Day Event held in the park," is provided.
[0785] Step 8:
[0786] Users enter feedback about the displayed information on their device. This feedback is sent to the server, which then uses it to improve the server's information provision algorithm, thereby increasing the accuracy of future information provision.
[0787] (Application Example 1)
[0788] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0789] In modern society, users have access to vast amounts of information, but efficiently obtaining the most relevant information remains difficult. Furthermore, while there is a demand for information provision that takes into account users' interests and behavioral history, there is also a need to improve the accuracy of natural language queries and present information clearly through user interfaces. Additionally, enhancing voice input capabilities is desirable to improve user convenience.
[0790] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0791] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for referencing user attribute information and generating personalized information, means for acquiring information from external sources and selecting information according to the user's intent, means for converting voice input into text, and means for displaying the selected information on a user interface. As a result, users can efficiently acquire information that suits their preferences, and convenience is improved through natural and simple interaction utilizing voice input.
[0792] A "user inquiry" is a question or request from a user regarding the information or services they seek from the system.
[0793] "Natural language processing" is a technology that understands the meaning and intent behind the language that humans use in everyday life and processes it as data.
[0794] "User attribute information" refers to data related to an individual, including their interests, preferences, and past behavioral history.
[0795] "Personalized information" refers to customized data that is presented in a way that is most suitable for the individual, based on their profile.
[0796] An "external information source" is a data provider that exists outside the system and provides information in real time.
[0797] "Voice input" is a method of transmitting information to a system through the voice spoken by the user.
[0798] "Means of converting to text" refers to technologies and processes that convert non-textual data, such as audio, into textual information.
[0799] A "user interface" is an interaction environment that provides screens and methods of operation for users to interact with a system.
[0800] The system for implementing this invention consists of three main components: a server, a terminal, and a user. The server uses a natural language processing engine to analyze inquiries entered by the user through the terminal. This can be done using natural language processing tools such as the Google Cloud Natural Language API. The analyzed information is stored on the server as personalized data, taking into account the user's attribute information.
[0801] The server further accesses external data sources and retrieves data as needed. Open data APIs can be used for this access. From the retrieved data, it selects the information that best matches the user's intent and sends it to the terminal for display in the user interface.
[0802] The device provides a voice input function, allowing users to make natural inquiries using their voice. This voice is converted to text using speech-to-text technologies such as the Google Speech-to-Text API and sent to the server. The converted text is then analyzed on the server, and the corresponding information is prepared.
[0803] For example, if a user asks a question via voice, such as "Tell me about jazz events happening nearby," the voice is converted to text and analyzed by the server's natural language processing engine. Based on the analysis results, the system references the user's past activity history to extract and display relevant jazz events.
[0804] An example of a prompt for a generative AI model is written as follows:
[0805] "When a user asks, 'What are some recommended music events near me?', you should recommend events they might be able to attend, taking into account their interest in jazz events in particular. You should also consider events the user has attended in the past to suggest the best options."
[0806] This allows users to quickly obtain information that is of particular interest to them, and to enjoy a more convenient information retrieval experience.
[0807] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0808] Step 1:
[0809] Users enter inquiries using their device via voice or text. For voice input, the device uses the Google Speech-to-Text API to convert speech to text. This process takes a speech waveform as input and outputs the converted text data.
[0810] Step 2:
[0811] The terminal sends the obtained text to the server. The server receives this text and parses it using a natural language processing engine (e.g., Google Cloud Natural Language API). The input for this step is the user's query text, and the output is data indicating the parsed keywords and user intent.
[0812] Step 3:
[0813] The server references user attribute information based on the analysis results and generates personalized information. The inputs here are the analyzed keywords and user profile information, and the output is a list of information optimized for the user.
[0814] Step 4:
[0815] The server accesses external information sources and collects information that matches the user's intent. Open data APIs, for example, are used. The input is the analysis results and the user's interest data, and the output is the collected external information.
[0816] Step 5:
[0817] The server compares and analyzes the collected external information and selects the most suitable information to display in the user interface. In this step, the input is a list of candidate information that forms the basis of the selection, and the output is the final information presented to the user.
[0818] Step 6:
[0819] The terminal displays the selection information received from the server in the user interface. Here, the information sent from the server is the input, and the visual presentation of that information to the user is the output. The user reviews the displayed information and takes action as needed.
[0820] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0821] This invention is a local information provision system that combines natural language processing and an emotion engine to provide users with more personalized information. First, the user queries for information through a terminal. The terminal sends the content to a server, which uses a natural language processing (NLP) engine to analyze the query. This analysis extracts the type of information the user is seeking and related keywords.
[0822] Furthermore, this invention uses an emotion engine to recognize emotions from the user's text input. The emotion engine is used to systematically understand the user's emotional state and select appropriate information according to that state. For example, if it is determined that the user is feeling stressed, it will provide relaxation events or healing information.
[0823] Next, the server examines the user's profile data. This data contains past behavioral history and user interests, and personalized information is generated based on this data. The server accesses an external database to retrieve the most relevant information selected based on the user's intentions and emotions.
[0824] Using the acquired information, the server formats it into a format suitable for display on the user interface and sends it to the terminal. The terminal receives this information and displays it visually to the user. The user can then review and utilize this information.
[0825] Furthermore, users can input feedback on the information provided via their device. This feedback is sent to the server and used to improve the information provision algorithm and adjust sentiment patterns. This allows the system to continuously improve its ability to meet user needs over the long term.
[0826] Specific example
[0827] For example, if a user enters a question into their device expressing stress, such as "What events can I enjoy with my kids this weekend?", the server analyzes the inquiry and extracts "stress" and "kid-friendly events" as keywords. Next, the emotion engine detects the stress and suggests relaxing family events based on the user's profile. The device might display information such as "Weekend Family Day at the Park: With Relaxation Corner." If the user provides feedback on this event, it will be incorporated into future information provision.
[0828] The following describes the processing flow.
[0829] Step 1:
[0830] The user queries the device for event information using natural language. The device retrieves this information and prepares to send it to the server along with the emotion engine.
[0831] Step 2:
[0832] The terminal sends the user's inquiry to the server. The server processes the received data and parses the text via a natural language processing engine.
[0833] Step 3:
[0834] The server uses natural language processing to extract the main keywords of the query. This identifies the type of event the user is looking for, as well as the available time and location.
[0835] Step 4:
[0836] The server uses an emotion engine to recognize the emotions contained in the user's input. Emotions such as stress, excitement, and satisfaction are evaluated.
[0837] Step 5:
[0838] The server accesses the user's profile database to investigate their past behavior and interests. This allows for the provision of information that is best suited to the user.
[0839] Step 6:
[0840] The server connects to external information sources in real time to retrieve relevant data such as event information. This data is filtered based on user requests and sentiments.
[0841] Step 7:
[0842] The server organizes the selected information and generates personalized event information according to the user's request. This information is then converted into an appropriate format and prepared for display.
[0843] Step 8:
[0844] The server sends the generated information to the terminal. The terminal receives this information and displays it in the user interface, allowing the user to view the information.
[0845] Step 9:
[0846] The user provides feedback on the information they have been given. The device sends this feedback to the server, which is then used to improve the system.
[0847] Step 10:
[0848] The server analyzes the feedback and uses it to adjust the information delivery algorithm and emotion engine. This process further improves the user's experience on subsequent visits.
[0849] (Example 2)
[0850] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0851] Conventional local information systems typically provide information to users, making it difficult to effectively deliver personalized information tailored to the individual user's emotions and interests. As a result, users often fail to receive the information they need, highlighting the need for improved user experience.
[0852] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0853] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing user emotions, and means for referencing user profile information and generating personalized information. This makes it possible to provide information that is tailored to the individual emotions and needs of each user.
[0854] "Natural language processing" is the technology that enables computers to understand, analyze, and generate human language.
[0855] "Emotion recognition" is a technology that extracts and analyzes emotional elements from information provided by users.
[0856] "Profile information" refers to data about individual users, including their interests, behavioral history, and personal characteristics.
[0857] "External information sources" refer to databases, APIs, and other information sources located outside the system that provide up-to-date or additional data.
[0858] "Feedback" refers to the opinions and evaluations that users provide to a system, which are used to improve the services and information provided.
[0859] "Personalized information" refers to information tailored to each user's individual characteristics and preferences, and is provided on the user interface.
[0860] "Filtering" refers to the process of selecting information data based on certain criteria and eliminating unnecessary information.
[0861] A "user interface" is a means of displaying and inputting information between a computer system and a user, and includes visual display screens and input devices.
[0862] This invention is an information provision system that combines natural language processing technology and an emotion recognition engine to provide users with personalized information. In this system, the user queries for information from a terminal, the server analyzes the content, selects appropriate information, and provides it to the user.
[0863] The user inputs information and makes inquiries using natural language via a terminal. This terminal sends data to the server using a common communication protocol, initiating information processing. The server uses a natural language processing engine implemented in a programming language such as Python to analyze the user's input and extract the intent of the inquiry and related keywords. In this process, libraries such as NLTK may be used.
[0864] In addition, the server uses an emotion recognition engine to analyze the user's emotional state from their text. This engine further improves the accuracy of the information the user is seeking by using machine learning models (such as the BERT model) powered by the Transformers library.
[0865] Next, the server checks the user's profile information. This information is stored in a database system (e.g., MySQL) and reflects past behavior and interests. Based on this data, the server generates personalized information that is most relevant to the user.
[0866] The server then accesses external information sources to retrieve the latest relevant events and news. This process utilizes various APIs to gather a wide range of information and select the information that best suits the user's needs.
[0867] The acquired information is formatted by the server into a user-friendly format and transmitted to the terminal. The terminal displays the information in the specified format and provides it to the user. The user can then make decisions and take actions based on this information.
[0868] Furthermore, users can input feedback using their devices. This information is sent to the server and used to improve the information provision algorithm. As a result, the information provision will continuously improve over the long term and become more valuable to users.
[0869] For example, if a user enters "Where can I relax with my family on the weekend?", the server will analyze this query and suggest events related to relaxation. For instance, it might provide information such as "Weekend events at nearby parks."
[0870] Example of a prompt:
[0871] If a user asks, "Can you recommend some places to relax on the weekend?", how would you provide them with relevant information?
[0872] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0873] Step 1:
[0874] User input of inquiry
[0875] Specific operation: The user requests specific information on the device's input screen. For example, they might type text such as, "What are some places where my family can relax on the weekend?"
[0876] Input and Output: The text entered by the user is input into the terminal. As output, the terminal sends this query data to the server.
[0877] Step 2:
[0878] Sending information from the terminal to the server
[0879] Specific operation: The terminal sends the received user query to the server via the appropriate communication protocol (e.g., HTTPS).
[0880] Input and Output: The input is the query text received from the user. The output is the data received by the server for analysis.
[0881] Step 3:
[0882] Execution of natural language processing by a server
[0883] Specific operation: The server executes a Python script and uses a natural language processing library (e.g., NLTK) to parse the query text. Specific techniques include text tokenization and keyword extraction.
[0884] Input and Output: The input is text data received from the terminal. The output generates data containing the analyzed keywords and intent.
[0885] Step 4:
[0886] Emotion recognition performed by the server
[0887] Specific operation: The server uses a machine learning model (e.g., a BERT-based model) to recognize emotions from the user's text. Emotional scoring is performed, and the emotional state of the text is quantified.
[0888] Input and Output: The input is parsed text, and the output is a sentiment score or state.
[0889] Step 5:
[0890] User profile referencing by the server
[0891] Specific operation: The server executes database queries (e.g., MySQL) to retrieve information about the user's past behavior and interests.
[0892] Input and Output: The input is identification information such as a user ID, and the output is profile information.
[0893] Step 6:
[0894] Obtaining information from external sources
[0895] Specific operation: The server uses various APIs to request relevant information from external databases. The retrieved information is prioritized to match the user's current interests and emotional state.
[0896] Input and Output: The input consists of selected keywords or conditions. The output consists of retrieved event information or news.
[0897] Step 7:
[0898] Formatting and sending information to the device
[0899] Specific operation: The server uses an HTML template to visually format the information to be presented to the user and sends it to the terminal via the HTTP protocol.
[0900] Input and Output: Input is the acquired information data, and output is the information formatted in a way that the user can visually understand.
[0901] Step 8:
[0902] Displaying information on the device
[0903] Specific operation: The terminal displays information received from the server on the screen, making it available for the user to review. Detailed information and related information can be accessed through the user interface.
[0904] Input and Output: Input is formatted information received from the server, and output is content presented visually to the user.
[0905] Step 9:
[0906] User feedback
[0907] Specific operation: The user enters feedback on the information provided via the terminal and sends it to the server by operating the submit button.
[0908] Input and Output: Input is user ratings and opinions, and output is feedback data recorded by the server.
[0909] (Application Example 2)
[0910] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0911] In modern urban life, users need information tailored to their emotional state and individual interests at any given time. However, conventional systems struggle to accurately recognize the diverse emotions of users and provide appropriate information. Furthermore, there is a lack of methods to effectively utilize user feedback and provide information that is more suitable for individual users.
[0912] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0913] In this invention, the server includes means for analyzing user inquiries using natural language processing, means for recognizing the user's emotional state and reflecting it in information selection, and means for referencing the user's profile data and generating personalized information. This makes it possible to provide more appropriate and personalized information based on the user's emotions and profile.
[0914] "Natural language processing" is a technology that enables computers to understand, analyze, and respond to human language.
[0915] "Emotional state" refers to the psychological state inferred from the user's input, and includes emotions such as stress, relaxation, and excitement.
[0916] "Profile data" refers to a collection of individual information, including a user's past behavioral history, interests, and preferences.
[0917] "Personalized information" refers to information that is optimized based on the user's individual attributes and emotions.
[0918] "External information sources" refer to external information sources that the server can access, such as databases and web resources.
[0919] A "user interface" is a means of displaying information visually to a user, allowing the user to confirm and manipulate that information.
[0920] "Feedback" refers to the opinions and evaluations given by users, and is information used to improve the system.
[0921] "In-city experiential information" refers to information related to tourist attractions, events, transportation, etc., within a city.
[0922] This invention provides an information delivery system that performs emotion recognition and personalization for residents of smart cities. Users use their smartphones to query information to customize their individual urban experience. The server receives the user's query in natural language and analyzes it using a natural language processing engine. In this process, it extracts the type of information requested by the user and related keywords. Furthermore, it uses an emotion engine to recognize the user's emotional state from their input. The server considers the user's profile data and selects and provides personalized information that is appropriate for their emotional state.
[0923] Specifically, the server filters information on tourist attractions, events, and transportation based on user profiles and data obtained from external sources, and formats urban experience information that matches the user's emotions. This information is transmitted to the terminal through the user interface, allowing the user to visually confirm it.
[0924] For example, if a user enters the question, "Where can I relax in Tokyo on the weekend?", the server processes this question using natural language processing and extracts "relax" and "Tokyo" as keywords. Then, using an emotion engine, it determines that the user is seeking relaxation and suggests quiet parks and healing spots based on their profile data. Information such as "Tokyo Central Park: with a quiet rest area" is displayed on the device. The user can then use these suggestions to plan their weekend.
[0925] A concrete example of a prompt is, "Please tell me some relaxing places in Tokyo that I should visit this weekend." This prompt allows the user to have a suitable urban experience.
[0926] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0927] Step 1:
[0928] The user enters the query in natural language using a smartphone. The input is received in text format, and the device sends this data to the server. The output is the query data sent to the server.
[0929] Step 2:
[0930] The server passes the received query data to a natural language processing engine for analysis. The input is the user's query text, and the output is the extraction of keywords and information types. In this step, the server analyzes important words and phrases contained in the query to determine what the user is asking for.
[0931] Step 3:
[0932] The server uses an emotion engine to recognize the user's emotional state from their input. The input is the user's query text, and the output is the recognized emotional state. In this step, the server analyzes the user's emotions from the content and context of the text and evaluates how those emotions will affect the information provided.
[0933] Step 4:
[0934] The server retrieves user profile data from the database and analyzes past behavior and interests. The input is user identification information, and the output is user-specific profile data. This step prepares the foundational data for personalizing information, taking into account the user's interests and behavioral tendencies.
[0935] Step 5:
[0936] The server accesses external information sources to obtain optimal urban experience information based on analyzed keywords, emotional states, and profile data. The input is a combination of keywords and emotional data, and the output is filtered information results. This step selects information tailored to the user's needs and creates customized recommendations.
[0937] Step 6:
[0938] The server formats the filtered information results into a format suitable for the user interface and sends it to the terminal. The input is the filtered information results, and the output is in a displayable information format. This step ensures that the information is displayed in a way that is easy for the user to understand.
[0939] Step 7:
[0940] The terminal displays formatted information on a user interface, providing users with visual access. Input is in a displayable information format, and output is the display of information to the user. This allows users to easily review and utilize proposed urban experiences.
[0941] Step 8:
[0942] The user enters feedback on the provided information via a terminal. The input is the text of the feedback, and the terminal sends this feedback to the server. The output is the feedback data sent to the server.
[0943] Step 9:
[0944] The server uses the received feedback data to improve its information provision algorithm. The input is user feedback data, and the output is the improved algorithm. In this step, the feedback is analyzed and improvements are made to make future information provision more appropriate.
[0945] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0946] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0947] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0948] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0949] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0950] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0951] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0952] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0953] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0954] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0955] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0956] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0957] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0958] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0959] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0960] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0961] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0962] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0963] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0964] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0965] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0966] The following is further disclosed regarding the embodiments described above.
[0967] (Claim 1)
[0968] A means of analyzing user inquiries using natural language processing,
[0969] A means of referencing user profile data and generating personalized information,
[0970] A means of acquiring data from external information sources and selecting information according to the user's intent,
[0971] A means of displaying the selected information on the user interface,
[0972] A system that includes this.
[0973] (Claim 2)
[0974] The system according to claim 1, comprising means for collecting user feedback and improving the information provision algorithm based on that feedback.
[0975] (Claim 3)
[0976] The system according to claim 1, comprising means for filtering event information by taking into account the user's past behavior history.
[0977] "Example 1"
[0978] (Claim 1)
[0979] A means of analyzing user inquiries using natural language processing,
[0980] A means of referencing user profile information and generating personalized information,
[0981] A means of acquiring data from external information sources and selecting the most suitable information according to the user's intent,
[0982] A means of visually displaying the generated information on a user interface,
[0983] A means to optionally obtain the user's geographical location information and improve the accuracy of information provision,
[0984] A system that includes this.
[0985] (Claim 2)
[0986] The system according to claim 1, comprising means for collecting user feedback and improving the information provision algorithm based on that feedback.
[0987] (Claim 3)
[0988] The system according to claim 1, comprising a means for generating information tailored to the user using a generative AI model and for searching for appropriate candidates by inputting a prompt sentence.
[0989] "Application Example 1"
[0990] (Claim 1)
[0991] A means of analyzing user inquiries using natural language processing,
[0992] A means for referencing user attribute information and generating personalized information,
[0993] A means of obtaining information from external sources and selecting information according to the user's intent,
[0994] A means of displaying the selected information on the user interface,
[0995] A means of converting voice input to text,
[0996] A system that includes this.
[0997] (Claim 2)
[0998] The system according to claim 1, comprising means for collecting user feedback and improving the information provision method based on that feedback.
[0999] (Claim 3)
[1000] The system according to claim 1, comprising means for filtering activity information, taking into account the user's past behavioral records.
[1001] "Example 2 of combining an emotion engine"
[1002] (Claim 1)
[1003] A means of analyzing user inquiries using natural language processing,
[1004] Means of recognizing user emotions,
[1005] A means of referencing user profile information and generating personalized information,
[1006] A means of acquiring data from external sources and selecting information according to the user's intent and emotional state,
[1007] A means for displaying the selected information on a user display device,
[1008] A system that includes this.
[1009] (Claim 2)
[1010] The system according to claim 1, comprising means for collecting user feedback and improving the information provision algorithm based on that feedback.
[1011] (Claim 3)
[1012] The system according to claim 1, comprising means for filtering user content information, taking into account the user's past behavior history.
[1013] "Application example 2 when combining with an emotional engine"
[1014] (Claim 1)
[1015] A means of analyzing user inquiries using natural language processing,
[1016] A means of recognizing the user's emotional state and reflecting it in information selection,
[1017] A means of referencing user profile data and generating personalized information,
[1018] A means of acquiring data from external information sources and selecting information that matches the user's intent and emotions,
[1019] A means of displaying the selected information on the user interface,
[1020] A system that includes this.
[1021] (Claim 2)
[1022] The system according to claim 1, comprising means for collecting user feedback and improving the information provision algorithm based on that feedback.
[1023] (Claim 3)
[1024] The system according to claim 1, comprising means for filtering urban experience information, taking into account the user's past behavioral history. [Explanation of Symbols]
[1025] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of analyzing user inquiries using natural language processing, A means for referencing user attribute information and generating personalized information, A means of obtaining information from external sources and selecting information according to the user's intent, A means of displaying the selected information on the user interface, A means of converting voice input to text, A system that includes this.
2. The system according to claim 1, comprising means for collecting user feedback and improving the information provision method based on that feedback.
3. The system according to claim 1, comprising means for filtering activity information, taking into account the user's past behavioral records.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A