system
The system addresses the lack of tailored support for children in need by efficiently analyzing user information, generating customized measures, and matching them with suitable support organizations, ensuring rapid and accurate support provision.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Children with economic difficulties and their guardians lack sufficient opportunities for tailored support, and existing systems fail to efficiently match supporters with those in need, leading to inadequate and delayed advisory advice.
A system that efficiently receives and analyzes user information, generates customized support measures, searches for matching support organizations, and provides interactive support tailored to individual needs.
Enables rapid and accurate provision of support tailored to individual needs by efficiently matching users with appropriate support organizations.
Smart Images

Figure 2026073426000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, it is a problem that children with economic difficulties and their guardians do not have sufficient opportunities to receive appropriate support. In particular, since there is a lack of a system that can quickly and appropriately receive support tailored to individual needs, there is a problem that support for those with needs that cannot be covered by conventional support methods is not thorough. In addition, since the matching between supporters and people who need support is not efficiently carried out, advisory advice and support are often not provided quickly.
Means for Solving the Problems
[0005] This invention features a function to efficiently receive and analyze information from users, thereby enabling an accurate understanding of the user's situation. Based on the analysis results, it generates customized optimal measures and provides these measures to the user in an interactive format, enabling support tailored to individual needs. Furthermore, it includes a function to search for support organizations that match the user's needs and automatically perform matching, allowing for the rapid and accurate provision of support. In addition, by notifying the user of the matching results, it facilitates people who need support in obtaining cooperation from appropriate support organizations, providing a system that helps solve problems.
[0006] A "user" is an individual or group that uses the system and is the entity that inputs information in need of assistance.
[0007] "Means for receiving information" refers to a device or program that has the function of acquiring data transmitted by a user and preparing it in a format that can be processed within the system.
[0008] "Means of analysis" refers to technologies for understanding received information and extracting user needs and problems, and includes processes such as natural language processing and data analysis.
[0009] "Means for generating optimal measures" refers to algorithms and logic used to construct the most suitable solutions and support plans for users based on analysis results.
[0010] "Interactive delivery methods" refer to interfaces that present generated policies and information to users in a conversational format and can respond to additional information and inquiries from users.
[0011] "Means for searching for and matching support organizations" refers to a system or mechanism for quickly identifying and connecting users with appropriate support organizations from a database according to their needs.
[0012] "Notification methods" refer to methods and technologies for quickly communicating matching results and details of support measures to users. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention begins with the user inputting their situation or challenges through a terminal. The terminal used here is an electronic device such as a smartphone or personal computer, which accepts input via a dedicated application or web interface. The user can provide information in text or voice, and in the case of voice, the terminal uses speech recognition technology to convert the data into text.
[0035] The terminal then sends the input information as data packets to the server. The server converts the received information into a format for analysis by the generating AI model, verifies the integrity of the information, and filters out unnecessary data. The converted data is analyzed using a natural language processing engine to clearly identify the user's problems and requests. Based on the analyzed information, the generating AI formulates available support options and generates the optimal solution according to the user's needs.
[0036] The generated measures are then presented to the user by a conversational AI on the device. The conversational AI can answer additional questions and explain the details of the measures through interaction with the user. Users can also provide feedback on the measures, allowing the system to offer even more personalized support.
[0037] Furthermore, the server searches its relevant database for support organizations that match the user's needs and automatically matches them with the most suitable organization. This process includes querying and filtering existing support organization databases to identify the most effective support collaboration.
[0038] Finally, the server sends the matching results as a notification message to the device. The device displays this notification to the user in real time, and the user uses the provided information to request necessary support or make contact.
[0039] For example, if a single mother struggling with insufficient income uses this system, she would input specific needs such as "I need support in finding a job and increasing my income." The generating AI would then suggest options such as vocational training programs or local support facilities for childcare. Through the conversational AI, she would be provided with more details about these options, and by connecting with support organizations, she could receive this assistance quickly.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] Users input their situation and problems into the device as text or voice using a dedicated application or web interface. In the case of voice input, the device uses speech recognition to convert it into text data.
[0043] Step 2:
[0044] The terminal sends the entered data to the server. Here, the data is protected using encryption technology to ensure secure communication.
[0045] Step 3:
[0046] The server formats the received text data in preparation for analysis. This includes text preprocessing and noise reduction.
[0047] Step 4:
[0048] The server inputs the formatted data into a generation AI model and runs a natural language processing engine. This classifies and organizes the user's needs and problems.
[0049] Step 5:
[0050] The generation AI generates multiple measures to address user needs based on the analysis results. This process includes referencing and comparing the generated measures with an internal support measure database.
[0051] Step 6:
[0052] The generated strategies are sent from the server to the terminal, where a conversational AI interactively presents them to the user. Further detailed questions and feedback from the user are then accepted.
[0053] Step 7:
[0054] The server uses a matching function to search for support organizations that meet the user's needs and performs the optimal matching. This search takes into account the support organizations' resources, location, and the services they can provide.
[0055] Step 8:
[0056] The matching results are sent from the server to the terminal and notified to the user. The terminal displays this information and instructs the user on the next steps to receive the provided assistance.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] Currently, there is a need to understand user needs in various ways and provide appropriate support based on those needs. However, existing systems have limited methods for inputting and analyzing information, making it difficult to respond to diverse user needs. In addition, the matching process with support organizations is often inefficient, and users may not be able to receive appropriate support quickly. This invention aims to address these problems by accurately analyzing diverse user input information, providing optimal measures, and achieving rapid and appropriate matching with relevant support organizations.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] This invention includes a server that converts user information into an analysis format after performing integrity checks and filtering; a server that analyzes the information converted into the analysis format using a natural language processing engine and generates optimal measures using a generative AI model; and a server that interactively provides the generated measures to the user using conversational artificial intelligence and receives feedback through interaction. This enables the formulation of measures that meet the diverse needs of users and the rapid matching of users with appropriate support organizations.
[0062] "Means of converting information into an analysis format after verifying its integrity and filtering" refers to a process of converting information sent by a user into an accurate and consistent format, thereby removing unnecessary data and errors, and then converting it into a format suitable for further analysis.
[0063] "A means of analyzing using a natural language processing engine and generating optimal measures using a generative AI model" refers to a process of interpreting and analyzing text and audio information using advanced algorithms, and then using an AI model generated based on the analysis results to derive the most suitable support measures for the user's needs.
[0064] "A means of interactively providing information to users using conversational artificial intelligence and receiving feedback through interaction" refers to a process that uses artificial intelligence technology to enable natural conversations with users, provides generated policy information to users, accepts questions and comments from users, and takes further action based on them.
[0065] This invention is a system that begins with a user inputting their situation or challenges through a terminal, and then optimally processes the input information to provide support. Users can input information using electronic devices such as smartphones or personal computers through a dedicated application or web interface. In this case, input is possible not only in text format but also in voice format, and the terminal uses speech recognition technology (e.g., speech recognition API) to convert the voice information into text.
[0066] The terminal is responsible for receiving information and sending it to the server as data packets. The server first verifies the integrity of this information and filters it, then converts it into a format for analysis. This is often done using software written in scripting languages such as Python. This converted information is then analyzed by a natural language processing engine (e.g., spaCy or NLTK).
[0067] Based on the analysis results, a generative AI model (e.g., Generative AI Model GPT) is used to formulate optimal measures that meet the user's needs. The generated measures are presented to the user on their device using conversational artificial intelligence (e.g., conversational AI system). The conversational AI can explain the details of the proposal and answer questions through interaction with the user.
[0068] During implementation, users will provide prompts such as the following. Examples of specific prompts include questions like, "Please tell me how I can increase my income," or "I'm looking for local support regarding my child's education."
[0069] Furthermore, based on the information provided by the user, the server searches its database for relevant support organizations and automatically matches the user with the most suitable organization. The matching results are notified to the terminal in real time, allowing the user to quickly obtain the necessary support based on this information.
[0070] This invention provides a comprehensive system that utilizes information technology to effectively and quickly provide support that meets diverse needs.
[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0072] Step 1:
[0073] Users input their situation and challenges using a dedicated application or web interface via a device such as a smartphone or PC. Input can be in text or voice. In the case of voice input, the device uses speech recognition technology to convert the speech to text. The output of the input data is formatted text data.
[0074] Step 2:
[0075] The terminal sends formatted text data to the server as data packets. The input is text data, which is sent to the server as data packets. The output on the server side is data packets ready for analysis.
[0076] Step 3:
[0077] The server first performs a consistency check on the received data packets and filters out unnecessary data and errors before converting them into an analysis format. This process is carried out using a data cleaning technique with Python. The input is the transmitted data packet, and the output is the analysis format after consistency checks have been completed.
[0078] Step 4:
[0079] The server analyzes the analysis format using a natural language processing engine to clearly identify the user's needs and problems. Here, the input is the data in the analysis format, and the output is the analyzed information on the user's needs.
[0080] Step 5:
[0081] The server generates optimal strategies using a generative AI model based on the analyzed data. In this process, the generative AI model generates "suggestions that address the user's specific needs" as prompt messages. The input is the analyzed data, and the output is specific strategy prompt messages.
[0082] Step 6:
[0083] The generated policy prompts are presented to the user by an interactive artificial intelligence (AI) on the terminal. The AI can explain the details of the policy and answer questions through conversation with the user. User input is feedback, and output is a detailed explanation of the policy.
[0084] Step 7:
[0085] Based on user feedback, the server searches its database for relevant support organizations and automatically matches the user with the most suitable one. The input consists of user feedback and analysis results, and the output is a list of the most suitable support organizations.
[0086] Step 8:
[0087] The server sends the matching results as a notification to the terminal. The terminal displays this notification to the user in real time. The input is a list of support organizations, and the output is a notification message. Based on this information, the user can take action to quickly receive support.
[0088] (Application Example 1)
[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0090] Modern consumers have diverse needs and demand personalized recommendations when purchasing products. However, traditional systems have struggled to effectively recommend the best products based on user preferences and past purchase history. Furthermore, they fail to provide comprehensive support and information that users need, resulting in a lower level of satisfaction with the consumer experience.
[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0092] In this invention, the server includes means for receiving and analyzing information from the user, means for generating optimal measures based on the analysis results, means for interactively providing the generated measures to the user, means for recommending products based on the user's purchase history and preferences, and means for presenting the recommended product information to the user. This makes it possible to quickly and effectively recommend products suitable for the user and to support optimal purchasing decisions.
[0093] "Means for receiving and analyzing information from users" refers to devices or programs that acquire voice or text input provided by users via smart devices or computer networks and analyze its content.
[0094] "Means for generating optimal measures based on analysis results" refers to devices or programs that construct appropriate action suggestions and support measures based on acquired information, tailored to the user's needs and objectives.
[0095] "Means of interactively providing generated measures to users" refers to devices or programs that use AI to interact with users and clearly communicate the details and options of the measures.
[0096] "Means of recommending products based on a user's purchase history and preferences" refers to devices or programs that analyze past transaction data and user preferences to suggest the most suitable products and services.
[0097] "Means of presenting recommended product information to users" refers to devices or programs that effectively communicate the features and benefits of recommended products to users visually or audibly.
[0098] The system that realizes this invention consists of a user-operated terminal, a server, and a generative AI model that handles data processing. First, the user inputs information using a smartphone or computer in the form of voice or text. The terminal receives this input data and, if necessary, uses a speech recognition engine (e.g., Google® Speech-to-Text) to convert the voice into text.
[0099] The converted data is sent from the terminal to the server. The server uses a natural language processing engine (e.g., spaCy or Google NLP API) to analyze the user's input and understand their needs. Based on this analysis, a generative AI model (e.g., OpenAI® GPT-4®) formulates the most suitable strategy for the user.
[0100] Furthermore, the server selects suitable products from the database based on the user's purchase history and preferences, and recommends them to the user. This process uses machine learning algorithms to analyze trends from past data and generate a final product list. The generated product information and strategies are provided to the user through an interactive AI on the terminal, with detailed information and options explained in an easy-to-understand manner.
[0101] For example, if a user enters "I'm looking for summer casual wear," the system will recommend the best products and discount information based on past purchase history and product reviews. The generating AI model analyzes the user's input and uses the prompt "The user is looking for summer casual wear. Please suggest the best products and discount information based on past purchase history and user reviews" to select recommended products. This allows the user to efficiently find products that meet their needs.
[0102] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0103] Step 1:
[0104] Users input information via voice or text using a smartphone or computer terminal. The input information describes the user's needs and objectives. In the case of voice input, the terminal uses a speech recognition engine to convert it to text and adds it to the data stream as output.
[0105] Step 2:
[0106] The terminal sends the converted text data to the server. The server receives the input data and analyzes its content using a natural language processing engine. Through data analysis, it extracts keywords and topics related to the user's needs and outputs them as analysis results.
[0107] Step 3:
[0108] The server uses an AI model generated based on the analysis results to produce optimal strategies and product recommendations. The user communicates their needs to the AI model using prompts, and the algorithm constructs appropriate suggestions. In this step, strategy and product information is generated and output as suggestions.
[0109] Step 4:
[0110] The server optimizes the generated product recommendations based on the user's history and preferences. It scrutinizes the recommended product list using past purchase history and rating data, filtering as needed. This optimized product list is then prepared as the final output.
[0111] Step 5:
[0112] The well-organized recommended products and promotional information are presented to the user through the terminal's interactive AI. The terminal interactively explains the reasons for the recommendations and product details to the user, and answers the user's questions, thereby supporting their purchase and decision-making. This allows the user to make the best purchase choice based on the information provided.
[0113] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0114] In this embodiment of the invention, an interface incorporating an emotion engine is implemented in the user's terminal. The user can input information about their situation and challenges through a terminal such as a smartphone or personal computer, and in the process, multimedia input such as audio and video is possible.
[0115] The emotion engine installed in the device analyzes voice and facial expression data input by the user in real time to estimate the user's emotional state. This analysis is performed using voice tone changes and facial expression analysis technology, and the estimation results are sent to the server as part of the user information.
[0116] The server receives user text and sentiment data and formats it in preparation for analysis. The generative AI model uses natural language processing to extract the user's specific needs and problems, and then, taking sentiment data into consideration, proposes optimal measures adapted to the user's emotional state. This results in the generation of more personalized support measures that are tailored to the user's psychological state.
[0117] Policy information provided by the server is presented to the user via conversational AI. The conversational AI not only simply lists the generated policies, but also interacts with the user based on their emotional state, creating an environment where the user can provide feedback with confidence.
[0118] Furthermore, the server searches for support organizations that match the user's needs and emotions, and performs the optimal matching. By using emotional information, the likelihood of selecting a support organization with which the user can interact more comfortably increases.
[0119] For example, in the case of a user seeking support due to workplace stress, if tension is detected in their voice, the generating AI can prioritize suggesting relaxation-promoting measures and psychological care support. This system provides users with more appropriate and effective support measures than traditional responses based solely on text information.
[0120] The following describes the processing flow.
[0121] Step 1:
[0122] Users input their worries and challenges into their devices via text or voice input through a dedicated application or web interface that incorporates an emotion engine. In the case of voice input, the input is recorded as audio data using the device's microphone.
[0123] Step 2:
[0124] The device processes the input audio data in real time and converts it into text data using speech recognition technology. Simultaneously, an emotion engine analyzes the tone and volume of the audio data, as well as changes in facial expressions in the video input, to estimate emotions.
[0125] Step 3:
[0126] The device sends the converted text data and estimated sentiment information to the server. Encryption technology is used during transmission to maintain data confidentiality.
[0127] Step 4:
[0128] The server formats the text data and sentiment information received from the terminal into an analysis format. This performs data cleansing, removing noise that would interfere with the analysis.
[0129] Step 5:
[0130] The generative AI model runs on a server and uses formatted data to interpret user problems and needs through natural language processing. Furthermore, it incorporates emotional information into the analysis results to generate strategies adapted to those emotions.
[0131] Step 6:
[0132] The server sends the generated measures to the terminal, and the conversational AI presents them to the user. The conversational AI explains the details of the generated measures and receives feedback and additional information from the user.
[0133] Step 7:
[0134] The server searches its database for the most suitable support group based on the user's needs and feelings, and performs a matching process. Prioritizing groups that the user can comfortably interact with is a key consideration.
[0135] Step 8:
[0136] Matched support group information is sent from the server to the terminal and notified to the user in real time. The terminal displays this information to the user and provides advice on what to do next.
[0137] (Example 2)
[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0139] In modern society, there is a need to provide more accurate and personalized support quickly to address the psychological and practical problems faced by individual users. However, conventional systems have been insufficient in optimizing support while considering the user's emotional state. Furthermore, they have not been able to effectively match users with support organizations.
[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0141] In this invention, the server includes means for analyzing diverse input information from the user and estimating the emotional state, means for generating AI that generates optimal measures according to the estimated emotional state, and means for searching for and matching support organizations based on the user's needs and emotions. This enables the provision of more personalized support according to the user's psychological state and rapid matching of the user with the most suitable support organization.
[0142] A "user" refers to an entity that uses a system to provide information and receive support and policy suggestions.
[0143] "Diverse input information" refers to information in multiple formats, such as text data, audio data, and video data, provided by the user.
[0144] "Emotional state" refers to the emotional and mood state expressed by the user, and is estimated through analysis of voice and facial expressions.
[0145] "Means of estimation" refers to technical methods used to identify and analyze emotional states based on user input information.
[0146] "Generative AI methods" refer to systems that utilize artificial intelligence technology to generate optimal measures and suggestions based on the user's needs and emotional state.
[0147] "Interactive delivery methods" refer to technologies that interactively display or communicate generated measures to users.
[0148] A "support group" refers to an organization or institution that provides support for problems that users are facing.
[0149] "Matching methods" refer to technical means for selecting and introducing support organizations and resources that are suitable for the user's needs and emotional state.
[0150] In this embodiment of the invention, the user accesses the system via a computer terminal such as a smartphone or personal computer. The user can input information about their daily situation or specific issues into the terminal. In this process, the terminal can collect audio and video data using its voice recognition function and camera.
[0151] The emotion engine built into the device processes voice and video data from the user in real time, analyzing changes in voice tone and facial features. Based on this analysis, the user's emotional state is estimated and added as an emotion label.
[0152] Both estimated emotion data and user input data are sent to the server. The server receives this data and uses a generative AI model to extract the user's needs and challenges. This model uses natural language processing technology to automatically generate measures that take emotional states into account. For example, a user experiencing stress at work might be offered measures specifically focused on relaxation techniques and psychological care. This process begins with a prompt such as, "I've recently been feeling stressed due to work pressure, and I'd like to know what to do about it."
[0153] The measures generated by the server are delivered to the user through an interactive AI. This AI not only presents the measures but also facilitates interactions that reflect the user's emotional state, providing a reassuring experience. Furthermore, the server searches for and matches the user with a suitable support organization based on emotional information and notifies the user of their contact information. This allows the user to receive the support that is best suited to them.
[0154] This system, following a predetermined algorithm, leverages known emotion recognition and natural language processing technologies to provide users with innovative and effective support.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] Users access the system via devices such as smartphones and computers and input information about their situation and challenges. This input includes audio, video, and text data. For example, a user might record a statement such as "I've been feeling stressed at work lately" as audio data using the device's microphone.
[0158] Step 2:
[0159] The device processes collected audio and video data using an emotion engine, analyzing voice tone and facial expression data. This includes aspects such as voice pitch, rhythm, and facial muscle movements. Raw audio and facial expression data are used as input, and emotion labels (e.g., "feeling stressed") are generated as output. Specifically, "tension" is detected through voice tone analysis, and "fatigue" is estimated through facial expression analysis.
[0160] Step 3:
[0161] The terminal sends emotion labels and input information to the server. The server formats the received data and performs error checking. The input consists of text data related to the audio data and emotion labels, and the output is data ready for analysis, which is supplied to the generating AI model.
[0162] Step 4:
[0163] The AI model on the server extracts user needs and problems based on the configured prompt. Using natural language processing technology, it analyzes the content of the text data and sentiment labels to generate optimal solutions. For example, using the prompt "Recently, my workload has been heavy; how can I alleviate it?", it generates solutions such as "relaxation methods" and "improvements to time management."
[0164] Step 5:
[0165] The server passes the generated measures to the conversational AI, which then presents them interactively to the user. The AI engages in dialogue that takes the user's emotional state into account, providing a detailed explanation of the measures. The input is the generated measures information, and the output is voice or text feedback to the user. For example, the conversational AI might say, "Why not try yoga for stress management?"
[0166] Step 6:
[0167] The server uses user needs and emotional information to search for and match the user with the most suitable support organization. A search algorithm generates a list of support organizations that match the user's requirements. Input includes the user's location and areas of interest, and output lists information on appropriate support organizations. Specifically, it searches for "nearby organizations specializing in stress management" and notifies the user of their contact information.
[0168] (Application Example 2)
[0169] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0170] Currently, many systems process information without considering users' emotions or psychological states, providing uniform responses and thus failing to offer personalized support tailored to individual needs. Furthermore, in the security field, immediate anomaly detection based on emotional states is difficult, potentially leading to missed potential risks. To address these challenges, there is a need for technology that can specifically assess user information based on emotional states and provide appropriate measures and alerts.
[0171] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0172] In this invention, the server includes means for receiving information from the user and analyzing that information, means for generating optimal measures based on the analysis results, and means for interactively providing the generated measures to the user. This enables the provision of appropriate and personalized support measures that are tailored to the user's emotional state and needs, as well as rapid and accurate anomaly detection at the site.
[0173] "User information" refers to information provided by users, including audio, visual, and text data.
[0174] "Means of analysis" refers to a function that performs a process of evaluating the user's emotions and psychological state based on the received information.
[0175] "Means for generating optimal measures" refers to devising support measures tailored to the individual needs of users based on analysis results.
[0176] "Interactive delivery methods" refer to functions that deliver generated measures through two-way communication with users.
[0177] "Means of detecting and notifying of anomalies based on emotional state" refers to a function that analyzes the user's emotional state to determine the risk and issues an immediate alert if necessary.
[0178] "A means of searching for and matching support organizations" refers to the process of identifying and collaborating with organizations that can provide appropriate support tailored to the user's needs.
[0179] "Means of notifying users of matching results" refers to a function that communicates matching information obtained through a search to the user.
[0180] The system implementing this invention consists of a user terminal and a server. The terminal uses a device such as Google Glass® or Microsoft® HoloLens® to receive audio, visual, and text information from the user and transmit it to the server. The Google Cloud Speech-to-Text API is installed on this device for analyzing audio data, and OpenCV and Dlib are installed for analyzing visual data.
[0181] The server takes in the received data and performs sentiment analysis. This uses OpenAI GPT-3® as a generative AI model to examine the user's emotional state. Based on the results, it generates optimal support measures tailored to the user. These support measures include life support, vocational training, psychological care, and anomaly detection.
[0182] The generated support measures are presented to the user interactively. This is to facilitate user understanding and make it easier to obtain appropriate feedback by providing an approach tailored to the user's situation through dialogue. Furthermore, the server searches for support organizations that meet the required support needs and matches the user with them. Information on the selected support organization is accurately notified to the user.
[0183] As a concrete example, in commercial facilities, security guards may use smart glasses to check the emotional state of visitors and identify suspicious individuals. When an anomaly is detected, an alert can be issued immediately, allowing for a rapid response to the situation. Using prompts such as, "Detect anomalies from visual and audio data. Example: Analyze suspicious behavioral patterns and speech patterns exhibited by visitors and notify security personnel," it is possible to provide information that responds to the user's intentions in real time.
[0184] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0185] Step 1:
[0186] The device acquires audio, visual, and text information from the user. In this step, data is captured using the camera and microphone of smart glasses. The input is the user's real-time audio and video information, and the output is analyzable digital data.
[0187] Step 2:
[0188] The terminal preprocesses the acquired data, using the Google Cloud Speech-to-Text API for audio data analysis and OpenCV and Dlib for visual data analysis. The input is the raw data acquired in step 1, and the output is frame-by-frame facial expression features and transcribed audio information. This processing converts the data into the format necessary for analysis.
[0189] Step 3:
[0190] The server receives data sent from the terminal and uses the generative AI model OpenAI GPT-3 to estimate the user's emotional state. The input is the feature data and text data generated in the previous step, and the output is the user's emotional state and its confidence score. In this step, emotions are quantified by evaluating voice tone and facial expression changes.
[0191] Step 4:
[0192] The server generates optimal measures based on the emotional state. The generating AI model considers the estimated emotional state and creates measures such as relaxation recommendations or warnings. The input is the emotional state and confidence level from step 3, and the output is the suggested measures for the user.
[0193] Step 5:
[0194] The server interactively provides the generated measures to the user. During this process, it uses prompts to present the user with details of the measures. The input consists of internal rules regarding the measures and their delivery methods, while the output is the measures information visualized for the user. Specifically, notifications may appear on the user's glasses, and supplementary explanations may be provided via audio.
[0195] Step 6:
[0196] The server searches for relevant support organizations based on the user's needs and makes appropriate matches. The input is the user's needs and emotional state, and the output is the contact information of the support organizations. This step ensures that the user receives appropriate assistance.
[0197] Step 7:
[0198] The server notifies the user of the matching results and provides guidance for further assistance as needed. The input is information about support organizations obtained through the search, and the output is detailed assistance guidance displayed on the device. The user can access the relevant facilities through smart glasses.
[0199] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0200] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0201] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0202] [Second Embodiment]
[0203] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0204] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0205] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0206] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0207] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0208] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0209] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0210] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0211] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0212] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0213] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0214] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0215] This invention begins with the user inputting their situation or challenges through a terminal. The terminal used here is an electronic device such as a smartphone or personal computer, which accepts input via a dedicated application or web interface. The user can provide information in text or voice, and in the case of voice, the terminal uses speech recognition technology to convert the data into text.
[0216] The terminal then sends the input information as data packets to the server. The server converts the received information into a format for analysis by the generating AI model, verifies the integrity of the information, and filters out unnecessary data. The converted data is analyzed using a natural language processing engine to clearly identify the user's problems and requests. Based on the analyzed information, the generating AI formulates available support options and generates the optimal solution according to the user's needs.
[0217] The generated measures are then presented to the user by a conversational AI on the device. The conversational AI can answer additional questions and explain the details of the measures through interaction with the user. Users can also provide feedback on the measures, allowing the system to offer even more personalized support.
[0218] Furthermore, the server searches its relevant database for support organizations that match the user's needs and automatically matches them with the most suitable organization. This process includes querying and filtering existing support organization databases to identify the most effective support collaboration.
[0219] Finally, the server sends the matching results as a notification message to the device. The device displays this notification to the user in real time, and the user uses the provided information to request necessary support or make contact.
[0220] For example, if a single mother struggling with insufficient income uses this system, she would input specific needs such as "I need support in finding a job and increasing my income." The generating AI would then suggest options such as vocational training programs or local support facilities for childcare. Through the conversational AI, she would be provided with more details about these options, and by connecting with support organizations, she could receive this assistance quickly.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] Users input their situation and problems into the device as text or voice using a dedicated application or web interface. In the case of voice input, the device uses speech recognition to convert it into text data.
[0224] Step 2:
[0225] The terminal sends the entered data to the server. Here, the data is protected using encryption technology to ensure secure communication.
[0226] Step 3:
[0227] The server formats the received text data in preparation for analysis. This includes text preprocessing and noise reduction.
[0228] Step 4:
[0229] The server inputs the formatted data into a generation AI model and runs a natural language processing engine. This classifies and organizes the user's needs and problems.
[0230] Step 5:
[0231] The generation AI generates multiple measures to address user needs based on the analysis results. This process includes referencing and comparing the generated measures with an internal support measure database.
[0232] Step 6:
[0233] The generated strategies are sent from the server to the terminal, where a conversational AI interactively presents them to the user. Further detailed questions and feedback from the user are then accepted.
[0234] Step 7:
[0235] The server uses a matching function to search for support organizations that meet the user's needs and performs the optimal matching. This search takes into account the support organizations' resources, location, and the services they can provide.
[0236] Step 8:
[0237] The matching results are sent from the server to the terminal and notified to the user. The terminal displays this information and instructs the user on the next steps to receive the provided assistance.
[0238] (Example 1)
[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0240] Currently, there is a need to understand user needs in various ways and provide appropriate support based on those needs. However, existing systems have limited methods for inputting and analyzing information, making it difficult to respond to diverse user needs. In addition, the matching process with support organizations is often inefficient, and users may not be able to receive appropriate support quickly. This invention aims to address these problems by accurately analyzing diverse user input information, providing optimal measures, and achieving rapid and appropriate matching with relevant support organizations.
[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0242] This invention includes a server that converts user information into an analysis format after performing integrity checks and filtering; a server that analyzes the information converted into the analysis format using a natural language processing engine and generates optimal measures using a generative AI model; and a server that interactively provides the generated measures to the user using conversational artificial intelligence and receives feedback through interaction. This enables the formulation of measures that meet the diverse needs of users and the rapid matching of users with appropriate support organizations.
[0243] "Means of converting information into an analysis format after verifying its integrity and filtering" refers to a process of converting information sent by a user into an accurate and consistent format, thereby removing unnecessary data and errors, and then converting it into a format suitable for further analysis.
[0244] "A means of analyzing using a natural language processing engine and generating optimal measures using a generative AI model" refers to a process of interpreting and analyzing text and audio information using advanced algorithms, and then using an AI model generated based on the analysis results to derive the most suitable support measures for the user's needs.
[0245] "A means of interactively providing information to users using conversational artificial intelligence and receiving feedback through interaction" refers to a process that uses artificial intelligence technology to enable natural conversations with users, provides generated policy information to users, accepts questions and comments from users, and takes further action based on them.
[0246] This invention is a system that begins with a user inputting their situation or challenges through a terminal, and then optimally processes the input information to provide support. Users can input information using electronic devices such as smartphones or personal computers through a dedicated application or web interface. In this case, input is possible not only in text format but also in voice format, and the terminal uses speech recognition technology (e.g., speech recognition API) to convert the voice information into text.
[0247] The terminal is responsible for receiving information and sending it to the server as data packets. The server first verifies the integrity of this information and filters it, then converts it into a format for analysis. This is often done using software written in scripting languages such as Python. This converted information is then analyzed by a natural language processing engine (e.g., spaCy or NLTK).
[0248] Based on the analysis results, a generative AI model (e.g., Generative AI Model GPT) is used to formulate optimal measures that meet the user's needs. The generated measures are presented to the user on their device using conversational artificial intelligence (e.g., conversational AI system). The conversational AI can explain the details of the proposal and answer questions through interaction with the user.
[0249] During implementation, users will provide prompts such as the following. Examples of specific prompts include questions like, "Please tell me how I can increase my income," or "I'm looking for local support regarding my child's education."
[0250] Furthermore, based on the information provided by the user, the server searches its database for relevant support organizations and automatically matches the user with the most suitable organization. The matching results are notified to the terminal in real time, allowing the user to quickly obtain the necessary support based on this information.
[0251] This invention provides a comprehensive system that utilizes information technology to effectively and quickly provide support that meets diverse needs.
[0252] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0253] Step 1:
[0254] Users input their situation and challenges using a dedicated application or web interface via a device such as a smartphone or PC. Input can be in text or voice. In the case of voice input, the device uses speech recognition technology to convert the speech to text. The output of the input data is formatted text data.
[0255] Step 2:
[0256] The terminal sends formatted text data to the server as data packets. The input is text data, which is sent to the server as data packets. The output on the server side is data packets ready for analysis.
[0257] Step 3:
[0258] The server first performs a consistency check on the received data packets and filters out unnecessary data and errors before converting them into an analysis format. This process is carried out using a data cleaning technique with Python. The input is the transmitted data packet, and the output is the analysis format after consistency checks have been completed.
[0259] Step 4:
[0260] The server analyzes the analysis format using a natural language processing engine to clearly identify the user's needs and problems. Here, the input is the data in the analysis format, and the output is the analyzed information on the user's needs.
[0261] Step 5:
[0262] The server generates optimal strategies using a generative AI model based on the analyzed data. In this process, the generative AI model generates "suggestions that address the user's specific needs" as prompt messages. The input is the analyzed data, and the output is specific strategy prompt messages.
[0263] Step 6:
[0264] The generated policy prompts are presented to the user by an interactive artificial intelligence (AI) on the terminal. The AI can explain the details of the policy and answer questions through conversation with the user. User input is feedback, and output is a detailed explanation of the policy.
[0265] Step 7:
[0266] Based on user feedback, the server searches its database for relevant support organizations and automatically matches the user with the most suitable one. The input consists of user feedback and analysis results, and the output is a list of the most suitable support organizations.
[0267] Step 8:
[0268] The server sends the matching results as a notification to the terminal. The terminal displays this notification to the user in real time. The input is a list of support organizations, and the output is a notification message. Based on this information, the user can take action to quickly receive support.
[0269] (Application Example 1)
[0270] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0271] Modern consumers have diverse needs and demand personalized recommendations when purchasing products. However, traditional systems have struggled to effectively recommend the best products based on user preferences and past purchase history. Furthermore, they fail to provide comprehensive support and information that users need, resulting in a lower level of satisfaction with the consumer experience.
[0272] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0273] In this invention, the server includes means for receiving and analyzing information from the user, means for generating optimal measures based on the analysis results, means for interactively providing the generated measures to the user, means for recommending products based on the user's purchase history and preferences, and means for presenting the recommended product information to the user. This makes it possible to quickly and effectively recommend products suitable for the user and to support optimal purchasing decisions.
[0274] "Means for receiving and analyzing information from users" refers to devices or programs that acquire voice or text input provided by users via smart devices or computer networks and analyze its content.
[0275] "Means for generating optimal measures based on analysis results" refers to devices or programs that construct appropriate action suggestions and support measures based on acquired information, tailored to the user's needs and objectives.
[0276] "Means of interactively providing generated measures to users" refers to devices or programs that use AI to interact with users and clearly communicate the details and options of the measures.
[0277] "Means of recommending products based on a user's purchase history and preferences" refers to devices or programs that analyze past transaction data and user preferences to suggest the most suitable products and services.
[0278] "Means of presenting recommended product information to users" refers to devices or programs that effectively communicate the features and benefits of recommended products to users visually or audibly.
[0279] The system that realizes this invention consists of a user-operated terminal, a server, and a generative AI model that handles data processing. First, the user inputs information using voice or text via a smartphone or computer. The terminal receives this input data and, if necessary, uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice to text.
[0280] The converted data is sent from the terminal to the server. The server uses a natural language processing engine (e.g., spaCy or Google NLP API) to analyze the user's input and understand their needs. Based on this analysis, a generative AI model (e.g., OpenAI GPT-4) formulates the most suitable strategy for the user.
[0281] Furthermore, the server selects suitable products from the database based on the user's purchase history and preferences, and recommends them to the user. This process uses machine learning algorithms to analyze trends from past data and generate a final product list. The generated product information and strategies are provided to the user through an interactive AI on the terminal, with detailed information and options explained in an easy-to-understand manner.
[0282] As a specific example, when a user inputs that they are looking for "summer casual wear", the system recommends the optimal products and discount information based on the past purchase history and product reviews. The generative AI model analyzes the user's statement and uses a prompt sentence like "The user is looking for summer casual wear. Please provide and propose the optimal products and discount information based on the past purchase history and user reviews." to select the recommended products. This enables the user to efficiently find products that meet their needs.
[0283] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0284] Step 1:
[0285] The user uses a smartphone or computer terminal to input information either by voice or text. The input information describes the user's needs and purposes. In the case of voice input, the terminal uses a speech recognition engine to convert it into text and adds it to the data stream as output.
[0286] Step 2:
[0287] The terminal sends the converted text data to the server. The server receives the input data and analyzes the content using a natural language processing engine. By analyzing the data, keywords and topics related to the user's needs are extracted and output as the analysis result.
[0288] Step 3:
[0289] Based on the analysis result, the server utilizes the generative AI model to generate optimal measures and product recommendations. The user's needs are conveyed to the AI model using a prompt sentence, and appropriate proposals are constructed through an algorithm. In this step, measures and product information are generated and output as proposals.
[0290] Step 4:
[0291] The server optimizes the generated product recommendations based on the user's history and preferences. It scrutinizes the recommended product list using past purchase history and rating data, filtering as needed. This optimized product list is then prepared as the final output.
[0292] Step 5:
[0293] The well-organized recommended products and promotional information are presented to the user through the terminal's interactive AI. The terminal interactively explains the reasons for the recommendations and product details to the user, and answers the user's questions, thereby supporting their purchase and decision-making. This allows the user to make the best purchase choice based on the information provided.
[0294] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0295] In this embodiment of the invention, an interface incorporating an emotion engine is implemented in the user's terminal. The user can input information about their situation and challenges through a terminal such as a smartphone or personal computer, and in the process, multimedia input such as audio and video is possible.
[0296] The emotion engine installed in the device analyzes voice and facial expression data input by the user in real time to estimate the user's emotional state. This analysis is performed using voice tone changes and facial expression analysis technology, and the estimation results are sent to the server as part of the user information.
[0297] The server receives user text and sentiment data and formats it in preparation for analysis. The generative AI model uses natural language processing to extract the user's specific needs and problems, and then, taking sentiment data into consideration, proposes optimal measures adapted to the user's emotional state. This results in the generation of more personalized support measures that are tailored to the user's psychological state.
[0298] Policy information provided by the server is presented to the user via conversational AI. The conversational AI not only simply lists the generated policies, but also interacts with the user based on their emotional state, creating an environment where the user can provide feedback with confidence.
[0299] Furthermore, the server searches for support organizations that match the user's needs and emotions, and performs the optimal matching. By using emotional information, the likelihood of selecting a support organization with which the user can interact more comfortably increases.
[0300] For example, in the case of a user seeking support due to workplace stress, if tension is detected in their voice, the generating AI can prioritize suggesting relaxation-promoting measures and psychological care support. This system provides users with more appropriate and effective support measures than traditional responses based solely on text information.
[0301] The following describes the processing flow.
[0302] Step 1:
[0303] Users input their worries and challenges into their devices via text or voice input through a dedicated application or web interface that incorporates an emotion engine. In the case of voice input, the input is recorded as audio data using the device's microphone.
[0304] Step 2:
[0305] The device processes the input audio data in real time and converts it into text data using speech recognition technology. Simultaneously, an emotion engine analyzes the tone and volume of the audio data, as well as changes in facial expressions in the video input, to estimate emotions.
[0306] Step 3:
[0307] The terminal transmits the converted text data and the estimated emotion information to the server. At the time of transmission, encryption technology is used to maintain the confidentiality of the data.
[0308] Step 4:
[0309] The server formats the text data and emotion information received from the terminal into an analysis format. As a result, data cleansing is performed and noise that hinders analysis is removed.
[0310] Step 5:
[0311] The generated AI model is executed on the server, and uses the formatted data to interpret the user's problems and needs through natural language processing. Furthermore, the emotion information is reflected in the analysis results to generate measures adapted to the emotion.
[0312] Step 6:
[0313] The server transmits the generated measures to the terminal, and the dialogue AI presents them to the user. The dialogue AI explains the details of the generated measures and receives feedback and additional information from the user.
[0314] Step 7:
[0315] The server searches the relevant database for the most suitable support groups based on the user's needs and emotions, and performs matching. Here, groups that the user can comfortably contact are preferentially selected.
[0316] Step 8:
[0317] The matched support group information is transmitted from the server to the terminal and notified to the user in real time. The terminal displays this information to the user and provides advice on the next action to take.
[0318] (Example 2)
[0319] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0320] In modern society, there is a need to provide more accurate and personalized support quickly to address the psychological and practical problems faced by individual users. However, conventional systems have been insufficient in optimizing support while considering the user's emotional state. Furthermore, they have not been able to effectively match users with support organizations.
[0321] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0322] In this invention, the server includes means for analyzing diverse input information from the user and estimating the emotional state, means for generating AI that generates optimal measures according to the estimated emotional state, and means for searching for and matching support organizations based on the user's needs and emotions. This enables the provision of more personalized support according to the user's psychological state and rapid matching of the user with the most suitable support organization.
[0323] A "user" refers to an entity that uses a system to provide information and receive support and policy suggestions.
[0324] "Diverse input information" refers to information in multiple formats, such as text data, audio data, and video data, provided by the user.
[0325] "Emotional state" refers to the emotional and mood state expressed by the user, and is estimated through analysis of voice and facial expressions.
[0326] "Means of estimation" refers to technical methods used to identify and analyze emotional states based on user input information.
[0327] "Generative AI methods" refer to systems that utilize artificial intelligence technology to generate optimal measures and suggestions based on the user's needs and emotional state.
[0328] "Interactive delivery methods" refer to technologies that interactively display or communicate generated measures to users.
[0329] A "support group" refers to an organization or institution that provides support for problems that users are facing.
[0330] "Matching methods" refer to technical means for selecting and introducing support organizations and resources that are suitable for the user's needs and emotional state.
[0331] In this embodiment of the invention, the user accesses the system via a computer terminal such as a smartphone or personal computer. The user can input information about their daily situation or specific issues into the terminal. In this process, the terminal can collect audio and video data using its voice recognition function and camera.
[0332] The emotion engine built into the device processes voice and video data from the user in real time, analyzing changes in voice tone and facial features. Based on this analysis, the user's emotional state is estimated and added as an emotion label.
[0333] Both estimated emotion data and user input data are sent to the server. The server receives this data and uses a generative AI model to extract the user's needs and challenges. This model uses natural language processing technology to automatically generate measures that take emotional states into account. For example, a user experiencing stress at work might be offered measures specifically focused on relaxation techniques and psychological care. This process begins with a prompt such as, "I've recently been feeling stressed due to work pressure, and I'd like to know what to do about it."
[0334] The measures generated by the server are delivered to the user through an interactive AI. This AI not only presents the measures but also facilitates interactions that reflect the user's emotional state, providing a reassuring experience. Furthermore, the server searches for and matches the user with a suitable support organization based on emotional information and notifies the user of their contact information. This allows the user to receive the support that is best suited to them.
[0335] This system, following a predetermined algorithm, leverages known emotion recognition and natural language processing technologies to provide users with innovative and effective support.
[0336] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0337] Step 1:
[0338] Users access the system via devices such as smartphones and computers and input information about their situation and challenges. This input includes audio, video, and text data. For example, a user might record a statement such as "I've been feeling stressed at work lately" as audio data using the device's microphone.
[0339] Step 2:
[0340] The device processes collected audio and video data using an emotion engine, analyzing voice tone and facial expression data. This includes aspects such as voice pitch, rhythm, and facial muscle movements. Raw audio and facial expression data are used as input, and emotion labels (e.g., "feeling stressed") are generated as output. Specifically, "tension" is detected through voice tone analysis, and "fatigue" is estimated through facial expression analysis.
[0341] Step 3:
[0342] The terminal sends emotion labels and input information to the server. The server formats the received data and performs error checking. The input consists of text data related to the audio data and emotion labels, and the output is data ready for analysis, which is supplied to the generating AI model.
[0343] Step 4:
[0344] The AI model on the server extracts user needs and problems based on the configured prompt. Using natural language processing technology, it analyzes the content of the text data and sentiment labels to generate optimal solutions. For example, using the prompt "Recently, my workload has been heavy; how can I alleviate it?", it generates solutions such as "relaxation methods" and "improvements to time management."
[0345] Step 5:
[0346] The server passes the generated measures to the conversational AI, which then presents them interactively to the user. The AI engages in dialogue that takes the user's emotional state into account, providing a detailed explanation of the measures. The input is the generated measures information, and the output is voice or text feedback to the user. For example, the conversational AI might say, "Why not try yoga for stress management?"
[0347] Step 6:
[0348] The server uses user needs and emotional information to search for and match the user with the most suitable support organization. A search algorithm generates a list of support organizations that match the user's requirements. Input includes the user's location and areas of interest, and output lists information on appropriate support organizations. Specifically, it searches for "nearby organizations specializing in stress management" and notifies the user of their contact information.
[0349] (Application Example 2)
[0350] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0351] Currently, many systems process information without considering users' emotions or psychological states, providing uniform responses and thus failing to offer personalized support tailored to individual needs. Furthermore, in the security field, immediate anomaly detection based on emotional states is difficult, potentially leading to missed potential risks. To address these challenges, there is a need for technology that can specifically assess user information based on emotional states and provide appropriate measures and alerts.
[0352] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0353] In this invention, the server includes means for receiving information from the user and analyzing that information, means for generating optimal measures based on the analysis results, and means for interactively providing the generated measures to the user. This enables the provision of appropriate and personalized support measures that are tailored to the user's emotional state and needs, as well as rapid and accurate anomaly detection at the site.
[0354] "User information" refers to information provided by users, including audio, visual, and text data.
[0355] "Means of analysis" refers to a function that performs a process of evaluating the user's emotions and psychological state based on the received information.
[0356] "Means for generating optimal measures" refers to devising support measures tailored to the individual needs of users based on analysis results.
[0357] "Interactive delivery methods" refer to functions that deliver generated measures through two-way communication with users.
[0358] "Means of detecting and notifying of anomalies based on emotional state" refers to a function that analyzes the user's emotional state to determine the risk and issues an immediate alert if necessary.
[0359] "A means of searching for and matching support organizations" refers to the process of identifying and collaborating with organizations that can provide appropriate support tailored to the user's needs.
[0360] "Means of notifying users of matching results" refers to a function that communicates matching information obtained through a search to the user.
[0361] The system implementing this invention consists of a user terminal and a server. The terminal uses a device such as Google Glass or Microsoft HoloLens to receive audio, visual, and text information from the user and transmit it to the server. The Google Cloud Speech-to-Text API is installed on this device for analyzing audio data, and OpenCV and Dlib are installed for analyzing visual data.
[0362] The server takes in the received data and performs sentiment analysis. This uses OpenAI GPT-3 as a generative AI model to examine the user's emotional state. Based on the results, it generates personalized support measures. These measures include life support, vocational training, psychological care, and anomaly detection.
[0363] The generated support measures are presented to the user interactively. This is to facilitate user understanding and make it easier to obtain appropriate feedback by providing an approach tailored to the user's situation through dialogue. Furthermore, the server searches for support organizations that meet the required support needs and matches the user with them. Information on the selected support organization is accurately notified to the user.
[0364] As a concrete example, in commercial facilities, security guards may use smart glasses to check the emotional state of visitors and identify suspicious individuals. When an anomaly is detected, an alert can be issued immediately, allowing for a rapid response to the situation. Using prompts such as, "Detect anomalies from visual and audio data. Example: Analyze suspicious behavioral patterns and speech patterns exhibited by visitors and notify security personnel," it is possible to provide information that responds to the user's intentions in real time.
[0365] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0366] Step 1:
[0367] The device acquires audio, visual, and text information from the user. In this step, data is captured using the camera and microphone of smart glasses. The input is the user's real-time audio and video information, and the output is analyzable digital data.
[0368] Step 2:
[0369] The terminal preprocesses the acquired data, using the Google Cloud Speech-to-Text API for audio data analysis and OpenCV and Dlib for visual data analysis. The input is the raw data acquired in step 1, and the output is frame-by-frame facial expression features and transcribed audio information. This processing converts the data into the format necessary for analysis.
[0370] Step 3:
[0371] The server receives data sent from the terminal and uses the generative AI model OpenAI GPT-3 to estimate the user's emotional state. The input is the feature data and text data generated in the previous step, and the output is the user's emotional state and its confidence score. In this step, emotions are quantified by evaluating voice tone and facial expression changes.
[0372] Step 4:
[0373] The server generates optimal measures based on the emotional state. The generating AI model considers the estimated emotional state and creates measures such as relaxation recommendations or warnings. The input is the emotional state and confidence level from step 3, and the output is the suggested measures for the user.
[0374] Step 5:
[0375] The server interactively provides the generated measures to the user. During this process, it uses prompts to present the user with details of the measures. The input consists of internal rules regarding the measures and their delivery methods, while the output is the measures information visualized for the user. Specifically, notifications may appear on the user's glasses, and supplementary explanations may be provided via audio.
[0376] Step 6:
[0377] The server searches for relevant support organizations based on the user's needs and makes appropriate matches. The input is the user's needs and emotional state, and the output is the contact information of the support organizations. This step ensures that the user receives appropriate assistance.
[0378] Step 7:
[0379] The server notifies the user of the matching results and provides guidance for further assistance as needed. The input is information about support organizations obtained through the search, and the output is detailed assistance guidance displayed on the device. The user can access the relevant facilities through smart glasses.
[0380] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0381] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0382] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0383] [Third Embodiment]
[0384] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0385] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0386] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0387] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0388] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0389] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0390] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0391] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0392] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0393] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0394] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0395] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0396] This invention begins with the user inputting their situation or challenges through a terminal. The terminal used here is an electronic device such as a smartphone or personal computer, which accepts input via a dedicated application or web interface. The user can provide information in text or voice, and in the case of voice, the terminal uses speech recognition technology to convert the data into text.
[0397] The terminal then sends the input information as data packets to the server. The server converts the received information into a format for analysis by the generating AI model, verifies the integrity of the information, and filters out unnecessary data. The converted data is analyzed using a natural language processing engine to clearly identify the user's problems and requests. Based on the analyzed information, the generating AI formulates available support options and generates the optimal solution according to the user's needs.
[0398] The generated measures are then presented to the user by a conversational AI on the device. The conversational AI can answer additional questions and explain the details of the measures through interaction with the user. Users can also provide feedback on the measures, allowing the system to offer even more personalized support.
[0399] Furthermore, the server searches its relevant database for support organizations that match the user's needs and automatically matches them with the most suitable organization. This process includes querying and filtering existing support organization databases to identify the most effective support collaboration.
[0400] Finally, the server sends the matching results as a notification message to the device. The device displays this notification to the user in real time, and the user uses the provided information to request necessary support or make contact.
[0401] For example, if a single mother struggling with insufficient income uses this system, she would input specific needs such as "I need support in finding a job and increasing my income." The generating AI would then suggest options such as vocational training programs or local support facilities for childcare. Through the conversational AI, she would be provided with more details about these options, and by connecting with support organizations, she could receive this assistance quickly.
[0402] The following describes the processing flow.
[0403] Step 1:
[0404] Users input their situation and problems into the device as text or voice using a dedicated application or web interface. In the case of voice input, the device uses speech recognition to convert it into text data.
[0405] Step 2:
[0406] The terminal sends the entered data to the server. Here, the data is protected using encryption technology to ensure secure communication.
[0407] Step 3:
[0408] The server formats the received text data in preparation for analysis. This includes text preprocessing and noise reduction.
[0409] Step 4:
[0410] The server inputs the formatted data into a generation AI model and runs a natural language processing engine. This classifies and organizes the user's needs and problems.
[0411] Step 5:
[0412] The generation AI generates multiple measures to address user needs based on the analysis results. This process includes referencing and comparing the generated measures with an internal support measure database.
[0413] Step 6:
[0414] The generated strategies are sent from the server to the terminal, where a conversational AI interactively presents them to the user. Further detailed questions and feedback from the user are then accepted.
[0415] Step 7:
[0416] The server uses a matching function to search for support organizations that meet the user's needs and performs the optimal matching. This search takes into account the support organizations' resources, location, and the services they can provide.
[0417] Step 8:
[0418] The matching results are sent from the server to the terminal and notified to the user. The terminal displays this information and instructs the user on the next steps to receive the provided assistance.
[0419] (Example 1)
[0420] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0421] Currently, there is a need to understand user needs in various ways and provide appropriate support based on those needs. However, existing systems have limited methods for inputting and analyzing information, making it difficult to respond to diverse user needs. In addition, the matching process with support organizations is often inefficient, and users may not be able to receive appropriate support quickly. This invention aims to address these problems by accurately analyzing diverse user input information, providing optimal measures, and achieving rapid and appropriate matching with relevant support organizations.
[0422] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0423] This invention includes a server that converts user information into an analysis format after performing integrity checks and filtering; a server that analyzes the information converted into the analysis format using a natural language processing engine and generates optimal measures using a generative AI model; and a server that interactively provides the generated measures to the user using conversational artificial intelligence and receives feedback through interaction. This enables the formulation of measures that meet the diverse needs of users and the rapid matching of users with appropriate support organizations.
[0424] "Means of converting information into an analysis format after verifying its integrity and filtering" refers to a process of converting information sent by a user into an accurate and consistent format, thereby removing unnecessary data and errors, and then converting it into a format suitable for further analysis.
[0425] "A means of analyzing using a natural language processing engine and generating optimal measures using a generative AI model" refers to a process of interpreting and analyzing text and audio information using advanced algorithms, and then using an AI model generated based on the analysis results to derive the most suitable support measures for the user's needs.
[0426] "A means of interactively providing information to users using conversational artificial intelligence and receiving feedback through interaction" refers to a process that uses artificial intelligence technology to enable natural conversations with users, provides generated policy information to users, accepts questions and comments from users, and takes further action based on them.
[0427] This invention is a system that begins with a user inputting their situation or challenges through a terminal, and then optimally processes the input information to provide support. Users can input information using electronic devices such as smartphones or personal computers through a dedicated application or web interface. In this case, input is possible not only in text format but also in voice format, and the terminal uses speech recognition technology (e.g., speech recognition API) to convert the voice information into text.
[0428] The terminal is responsible for receiving information and sending it to the server as data packets. The server first verifies the integrity of this information and filters it, then converts it into a format for analysis. This is often done using software written in scripting languages such as Python. This converted information is then analyzed by a natural language processing engine (e.g., spaCy or NLTK).
[0429] Based on the analysis results, a generative AI model (e.g., Generative AI Model GPT) is used to formulate optimal measures that meet the user's needs. The generated measures are presented to the user on their device using conversational artificial intelligence (e.g., conversational AI system). The conversational AI can explain the details of the proposal and answer questions through interaction with the user.
[0430] During implementation, users will provide prompts such as the following. Examples of specific prompts include questions like, "Please tell me how I can increase my income," or "I'm looking for local support regarding my child's education."
[0431] Furthermore, based on the information provided by the user, the server searches its database for relevant support organizations and automatically matches the user with the most suitable organization. The matching results are notified to the terminal in real time, allowing the user to quickly obtain the necessary support based on this information.
[0432] This invention provides a comprehensive system that utilizes information technology to effectively and quickly provide support that meets diverse needs.
[0433] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0434] Step 1:
[0435] Users input their situation and challenges using a dedicated application or web interface via a device such as a smartphone or PC. Input can be in text or voice. In the case of voice input, the device uses speech recognition technology to convert the speech to text. The output of the input data is formatted text data.
[0436] Step 2:
[0437] The terminal sends formatted text data to the server as data packets. The input is text data, which is sent to the server as data packets. The output on the server side is data packets ready for analysis.
[0438] Step 3:
[0439] The server first performs a consistency check on the received data packets and filters out unnecessary data and errors before converting them into an analysis format. This process is carried out using a data cleaning technique with Python. The input is the transmitted data packet, and the output is the analysis format after consistency checks have been completed.
[0440] Step 4:
[0441] The server analyzes the analysis format using a natural language processing engine to clearly identify the user's needs and problems. Here, the input is the data in the analysis format, and the output is the analyzed information on the user's needs.
[0442] Step 5:
[0443] The server generates optimal strategies using a generative AI model based on the analyzed data. In this process, the generative AI model generates "suggestions that address the user's specific needs" as prompt messages. The input is the analyzed data, and the output is specific strategy prompt messages.
[0444] Step 6:
[0445] The generated policy prompts are presented to the user by an interactive artificial intelligence (AI) on the terminal. The AI can explain the details of the policy and answer questions through conversation with the user. User input is feedback, and output is a detailed explanation of the policy.
[0446] Step 7:
[0447] Based on user feedback, the server searches its database for relevant support organizations and automatically matches the user with the most suitable one. The input consists of user feedback and analysis results, and the output is a list of the most suitable support organizations.
[0448] Step 8:
[0449] The server sends the matching results as a notification to the terminal. The terminal displays this notification to the user in real time. The input is a list of support organizations, and the output is a notification message. Based on this information, the user can take action to quickly receive support.
[0450] (Application Example 1)
[0451] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0452] Modern consumers have diverse needs and demand personalized recommendations when purchasing products. However, traditional systems have struggled to effectively recommend the best products based on user preferences and past purchase history. Furthermore, they fail to provide comprehensive support and information that users need, resulting in a lower level of satisfaction with the consumer experience.
[0453] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0454] In this invention, the server includes means for receiving and analyzing information from the user, means for generating optimal measures based on the analysis results, means for interactively providing the generated measures to the user, means for recommending products based on the user's purchase history and preferences, and means for presenting the recommended product information to the user. This makes it possible to quickly and effectively recommend products suitable for the user and to support optimal purchasing decisions.
[0455] "Means for receiving and analyzing information from users" refers to devices or programs that acquire voice or text input provided by users via smart devices or computer networks and analyze its content.
[0456] "Means for generating optimal measures based on analysis results" refers to devices or programs that construct appropriate action suggestions and support measures based on acquired information, tailored to the user's needs and objectives.
[0457] "Means of interactively providing generated measures to users" refers to devices or programs that use AI to interact with users and clearly communicate the details and options of the measures.
[0458] "Means of recommending products based on a user's purchase history and preferences" refers to devices or programs that analyze past transaction data and user preferences to suggest the most suitable products and services.
[0459] "Means of presenting recommended product information to users" refers to devices or programs that effectively communicate the features and benefits of recommended products to users visually or audibly.
[0460] The system that realizes this invention consists of a user-operated terminal, a server, and a generative AI model that handles data processing. First, the user inputs information using voice or text via a smartphone or computer. The terminal receives this input data and, if necessary, uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice to text.
[0461] The converted data is sent from the terminal to the server. The server uses a natural language processing engine (e.g., spaCy or Google NLP API) to analyze the user's input and understand their needs. Based on this analysis, a generative AI model (e.g., OpenAI GPT-4) formulates the most suitable strategy for the user.
[0462] Furthermore, the server selects suitable products from the database based on the user's purchase history and preferences, and recommends them to the user. This process uses machine learning algorithms to analyze trends from past data and generate a final product list. The generated product information and strategies are provided to the user through an interactive AI on the terminal, with detailed information and options explained in an easy-to-understand manner.
[0463] For example, if a user enters "I'm looking for summer casual wear," the system will recommend the best products and discount information based on past purchase history and product reviews. The generating AI model analyzes the user's input and uses the prompt "The user is looking for summer casual wear. Please suggest the best products and discount information based on past purchase history and user reviews" to select recommended products. This allows the user to efficiently find products that meet their needs.
[0464] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0465] Step 1:
[0466] Users input information via voice or text using a smartphone or computer terminal. The input information describes the user's needs and objectives. In the case of voice input, the terminal uses a speech recognition engine to convert it to text and adds it to the data stream as output.
[0467] Step 2:
[0468] The terminal sends the converted text data to the server. The server receives the input data and analyzes its content using a natural language processing engine. Through data analysis, it extracts keywords and topics related to the user's needs and outputs them as analysis results.
[0469] Step 3:
[0470] The server uses an AI model generated based on the analysis results to produce optimal strategies and product recommendations. The user communicates their needs to the AI model using prompts, and the algorithm constructs appropriate suggestions. In this step, strategy and product information is generated and output as suggestions.
[0471] Step 4:
[0472] The server optimizes the generated product recommendations based on the user's history and preferences. It scrutinizes the recommended product list using past purchase history and rating data, filtering as needed. This optimized product list is then prepared as the final output.
[0473] Step 5:
[0474] The well-organized recommended products and promotional information are presented to the user through the terminal's interactive AI. The terminal interactively explains the reasons for the recommendations and product details to the user, and answers the user's questions, thereby supporting their purchase and decision-making. This allows the user to make the best purchase choice based on the information provided.
[0475] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0476] In this embodiment of the invention, an interface incorporating an emotion engine is implemented in the user's terminal. The user can input information about their situation and challenges through a terminal such as a smartphone or personal computer, and in the process, multimedia input such as audio and video is possible.
[0477] The emotion engine installed in the device analyzes voice and facial expression data input by the user in real time to estimate the user's emotional state. This analysis is performed using voice tone changes and facial expression analysis technology, and the estimation results are sent to the server as part of the user information.
[0478] The server receives user text and sentiment data and formats it in preparation for analysis. The generative AI model uses natural language processing to extract the user's specific needs and problems, and then, taking sentiment data into consideration, proposes optimal measures adapted to the user's emotional state. This results in the generation of more personalized support measures that are tailored to the user's psychological state.
[0479] Policy information provided by the server is presented to the user via conversational AI. The conversational AI not only simply lists the generated policies, but also interacts with the user based on their emotional state, creating an environment where the user can provide feedback with confidence.
[0480] Furthermore, the server searches for support organizations that match the user's needs and emotions, and performs the optimal matching. By using emotional information, the likelihood of selecting a support organization with which the user can interact more comfortably increases.
[0481] For example, in the case of a user seeking support due to workplace stress, if tension is detected in their voice, the generating AI can prioritize suggesting relaxation-promoting measures and psychological care support. This system provides users with more appropriate and effective support measures than traditional responses based solely on text information.
[0482] The following describes the processing flow.
[0483] Step 1:
[0484] Users input their worries and challenges into their devices via text or voice input through a dedicated application or web interface that incorporates an emotion engine. In the case of voice input, the input is recorded as audio data using the device's microphone.
[0485] Step 2:
[0486] The device processes the input audio data in real time and converts it into text data using speech recognition technology. Simultaneously, an emotion engine analyzes the tone and volume of the audio data, as well as changes in facial expressions in the video input, to estimate emotions.
[0487] Step 3:
[0488] The device sends the converted text data and estimated sentiment information to the server. Encryption technology is used during transmission to maintain data confidentiality.
[0489] Step 4:
[0490] The server formats the text data and sentiment information received from the terminal into an analysis format. This performs data cleansing, removing noise that would interfere with the analysis.
[0491] Step 5:
[0492] The generative AI model runs on a server and uses formatted data to interpret user problems and needs through natural language processing. Furthermore, it incorporates emotional information into the analysis results to generate strategies adapted to those emotions.
[0493] Step 6:
[0494] The server sends the generated measures to the terminal, and the conversational AI presents them to the user. The conversational AI explains the details of the generated measures and receives feedback and additional information from the user.
[0495] Step 7:
[0496] The server searches its database for the most suitable support group based on the user's needs and feelings, and performs a matching process. Prioritizing groups that the user can comfortably interact with is a key consideration.
[0497] Step 8:
[0498] Matched support group information is sent from the server to the terminal and notified to the user in real time. The terminal displays this information to the user and provides advice on what to do next.
[0499] (Example 2)
[0500] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0501] In modern society, there is a need to provide more accurate and personalized support quickly to address the psychological and practical problems faced by individual users. However, conventional systems have been insufficient in optimizing support while considering the user's emotional state. Furthermore, they have not been able to effectively match users with support organizations.
[0502] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0503] In this invention, the server includes means for analyzing diverse input information from the user and estimating the emotional state, means for generating AI that generates optimal measures according to the estimated emotional state, and means for searching for and matching support organizations based on the user's needs and emotions. This enables the provision of more personalized support according to the user's psychological state and rapid matching of the user with the most suitable support organization.
[0504] A "user" refers to an entity that uses a system to provide information and receive support and policy suggestions.
[0505] "Diverse input information" refers to information in multiple formats, such as text data, audio data, and video data, provided by the user.
[0506] "Emotional state" refers to the emotional and mood state expressed by the user, and is estimated through analysis of voice and facial expressions.
[0507] "Means of estimation" refers to technical methods used to identify and analyze emotional states based on user input information.
[0508] "Generative AI methods" refer to systems that utilize artificial intelligence technology to generate optimal measures and suggestions based on the user's needs and emotional state.
[0509] "Interactive delivery methods" refer to technologies that interactively display or communicate generated measures to users.
[0510] A "support group" refers to an organization or institution that provides support for problems that users are facing.
[0511] "Matching methods" refer to technical means for selecting and introducing support organizations and resources that are suitable for the user's needs and emotional state.
[0512] In this embodiment of the invention, the user accesses the system via a computer terminal such as a smartphone or personal computer. The user can input information about their daily situation or specific issues into the terminal. In this process, the terminal can collect audio and video data using its voice recognition function and camera.
[0513] The emotion engine built into the device processes voice and video data from the user in real time, analyzing changes in voice tone and facial features. Based on this analysis, the user's emotional state is estimated and added as an emotion label.
[0514] Both estimated emotion data and user input data are sent to the server. The server receives this data and uses a generative AI model to extract the user's needs and challenges. This model uses natural language processing technology to automatically generate measures that take emotional states into account. For example, a user experiencing stress at work might be offered measures specifically focused on relaxation techniques and psychological care. This process begins with a prompt such as, "I've recently been feeling stressed due to work pressure, and I'd like to know what to do about it."
[0515] The measures generated by the server are delivered to the user through an interactive AI. This AI not only presents the measures but also facilitates interactions that reflect the user's emotional state, providing a reassuring experience. Furthermore, the server searches for and matches the user with a suitable support organization based on emotional information and notifies the user of their contact information. This allows the user to receive the support that is best suited to them.
[0516] This system, following a predetermined algorithm, leverages known emotion recognition and natural language processing technologies to provide users with innovative and effective support.
[0517] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0518] Step 1:
[0519] Users access the system via devices such as smartphones and computers and input information about their situation and challenges. This input includes audio, video, and text data. For example, a user might record a statement such as "I've been feeling stressed at work lately" as audio data using the device's microphone.
[0520] Step 2:
[0521] The device processes collected audio and video data using an emotion engine, analyzing voice tone and facial expression data. This includes aspects such as voice pitch, rhythm, and facial muscle movements. Raw audio and facial expression data are used as input, and emotion labels (e.g., "feeling stressed") are generated as output. Specifically, "tension" is detected through voice tone analysis, and "fatigue" is estimated through facial expression analysis.
[0522] Step 3:
[0523] The terminal sends emotion labels and input information to the server. The server formats the received data and performs error checking. The input consists of text data related to the audio data and emotion labels, and the output is data ready for analysis, which is supplied to the generating AI model.
[0524] Step 4:
[0525] The AI model on the server extracts user needs and problems based on the configured prompt. Using natural language processing technology, it analyzes the content of the text data and sentiment labels to generate optimal solutions. For example, using the prompt "Recently, my workload has been heavy; how can I alleviate it?", it generates solutions such as "relaxation methods" and "improvements to time management."
[0526] Step 5:
[0527] The server passes the generated measures to the conversational AI, which then presents them interactively to the user. The AI engages in dialogue that takes the user's emotional state into account, providing a detailed explanation of the measures. The input is the generated measures information, and the output is voice or text feedback to the user. For example, the conversational AI might say, "Why not try yoga for stress management?"
[0528] Step 6:
[0529] The server uses user needs and emotional information to search for and match the user with the most suitable support organization. A search algorithm generates a list of support organizations that match the user's requirements. Input includes the user's location and areas of interest, and output lists information on appropriate support organizations. Specifically, it searches for "nearby organizations specializing in stress management" and notifies the user of their contact information.
[0530] (Application Example 2)
[0531] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0532] Currently, many systems process information without considering users' emotions or psychological states, providing uniform responses and thus failing to offer personalized support tailored to individual needs. Furthermore, in the security field, immediate anomaly detection based on emotional states is difficult, potentially leading to missed potential risks. To address these challenges, there is a need for technology that can specifically assess user information based on emotional states and provide appropriate measures and alerts.
[0533] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0534] In this invention, the server includes means for receiving information from the user and analyzing that information, means for generating optimal measures based on the analysis results, and means for interactively providing the generated measures to the user. This enables the provision of appropriate and personalized support measures that are tailored to the user's emotional state and needs, as well as rapid and accurate anomaly detection at the site.
[0535] "User information" refers to information provided by users, including audio, visual, and text data.
[0536] "Means of analysis" refers to a function that performs a process of evaluating the user's emotions and psychological state based on the received information.
[0537] "Means for generating optimal measures" refers to devising support measures tailored to the individual needs of users based on analysis results.
[0538] "Interactive delivery methods" refer to functions that deliver generated measures through two-way communication with users.
[0539] "Means of detecting and notifying of anomalies based on emotional state" refers to a function that analyzes the user's emotional state to determine the risk and issues an immediate alert if necessary.
[0540] "A means of searching for and matching support organizations" refers to the process of identifying and collaborating with organizations that can provide appropriate support tailored to the user's needs.
[0541] "Means of notifying users of matching results" refers to a function that communicates matching information obtained through a search to the user.
[0542] The system implementing this invention consists of a user terminal and a server. The terminal uses a device such as Google Glass or Microsoft HoloLens to receive audio, visual, and text information from the user and transmit it to the server. The Google Cloud Speech-to-Text API is installed on this device for analyzing audio data, and OpenCV and Dlib are installed for analyzing visual data.
[0543] The server takes in the received data and performs sentiment analysis. This uses OpenAI GPT-3 as a generative AI model to examine the user's emotional state. Based on the results, it generates personalized support measures. These measures include life support, vocational training, psychological care, and anomaly detection.
[0544] The generated support measures are presented to the user interactively. This is to facilitate user understanding and make it easier to obtain appropriate feedback by providing an approach tailored to the user's situation through dialogue. Furthermore, the server searches for support organizations that meet the required support needs and matches the user with them. Information on the selected support organization is accurately notified to the user.
[0545] As a concrete example, in commercial facilities, security guards may use smart glasses to check the emotional state of visitors and identify suspicious individuals. When an anomaly is detected, an alert can be issued immediately, allowing for a rapid response to the situation. Using prompts such as, "Detect anomalies from visual and audio data. Example: Analyze suspicious behavioral patterns and speech patterns exhibited by visitors and notify security personnel," it is possible to provide information that responds to the user's intentions in real time.
[0546] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0547] Step 1:
[0548] The device acquires audio, visual, and text information from the user. In this step, data is captured using the camera and microphone of smart glasses. The input is the user's real-time audio and video information, and the output is analyzable digital data.
[0549] Step 2:
[0550] The terminal preprocesses the acquired data, using the Google Cloud Speech-to-Text API for audio data analysis and OpenCV and Dlib for visual data analysis. The input is the raw data acquired in step 1, and the output is frame-by-frame facial expression features and transcribed audio information. This processing converts the data into the format necessary for analysis.
[0551] Step 3:
[0552] The server receives data sent from the terminal and uses the generative AI model OpenAI GPT-3 to estimate the user's emotional state. The input is the feature data and text data generated in the previous step, and the output is the user's emotional state and its confidence score. In this step, emotions are quantified by evaluating voice tone and facial expression changes.
[0553] Step 4:
[0554] The server generates optimal measures based on the emotional state. The generating AI model considers the estimated emotional state and creates measures such as relaxation recommendations or warnings. The input is the emotional state and confidence level from step 3, and the output is the suggested measures for the user.
[0555] Step 5:
[0556] The server interactively provides the generated measures to the user. During this process, it uses prompts to present the user with details of the measures. The input consists of internal rules regarding the measures and their delivery methods, while the output is the measures information visualized for the user. Specifically, notifications may appear on the user's glasses, and supplementary explanations may be provided via audio.
[0557] Step 6:
[0558] The server searches for relevant support organizations based on the user's needs and makes appropriate matches. The input is the user's needs and emotional state, and the output is the contact information of the support organizations. This step ensures that the user receives appropriate assistance.
[0559] Step 7:
[0560] The server notifies the user of the matching results and provides guidance for further assistance as needed. The input is information about support organizations obtained through the search, and the output is detailed assistance guidance displayed on the device. The user can access the relevant facilities through smart glasses.
[0561] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0562] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0563] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0564] [Fourth Embodiment]
[0565] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0566] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0567] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0568] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0569] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0570] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0571] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0572] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0573] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0574] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0575] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0576] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0577] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0578] This invention begins with the user inputting their situation or challenges through a terminal. The terminal used here is an electronic device such as a smartphone or personal computer, which accepts input via a dedicated application or web interface. The user can provide information in text or voice, and in the case of voice, the terminal uses speech recognition technology to convert the data into text.
[0579] The terminal then sends the input information as data packets to the server. The server converts the received information into a format for analysis by the generating AI model, verifies the integrity of the information, and filters out unnecessary data. The converted data is analyzed using a natural language processing engine to clearly identify the user's problems and requests. Based on the analyzed information, the generating AI formulates available support options and generates the optimal solution according to the user's needs.
[0580] The generated measures are then presented to the user by a conversational AI on the device. The conversational AI can answer additional questions and explain the details of the measures through interaction with the user. Users can also provide feedback on the measures, allowing the system to offer even more personalized support.
[0581] Furthermore, the server searches its relevant database for support organizations that match the user's needs and automatically matches them with the most suitable organization. This process includes querying and filtering existing support organization databases to identify the most effective support collaboration.
[0582] Finally, the server sends the matching results as a notification message to the device. The device displays this notification to the user in real time, and the user uses the provided information to request necessary support or make contact.
[0583] For example, if a single mother struggling with insufficient income uses this system, she would input specific needs such as "I need support in finding a job and increasing my income." The generating AI would then suggest options such as vocational training programs or local support facilities for childcare. Through the conversational AI, she would be provided with more details about these options, and by connecting with support organizations, she could receive this assistance quickly.
[0584] The following describes the processing flow.
[0585] Step 1:
[0586] Users input their situation and problems into the device as text or voice using a dedicated application or web interface. In the case of voice input, the device uses speech recognition to convert it into text data.
[0587] Step 2:
[0588] The terminal sends the entered data to the server. Here, the data is protected using encryption technology to ensure secure communication.
[0589] Step 3:
[0590] The server formats the received text data in preparation for analysis. This includes text preprocessing and noise reduction.
[0591] Step 4:
[0592] The server inputs the formatted data into a generation AI model and runs a natural language processing engine. This classifies and organizes the user's needs and problems.
[0593] Step 5:
[0594] The generation AI generates multiple measures to address user needs based on the analysis results. This process includes referencing and comparing the generated measures with an internal support measure database.
[0595] Step 6:
[0596] The generated strategies are sent from the server to the terminal, where a conversational AI interactively presents them to the user. Further detailed questions and feedback from the user are then accepted.
[0597] Step 7:
[0598] The server uses a matching function to search for support organizations that meet the user's needs and performs the optimal matching. This search takes into account the support organizations' resources, location, and the services they can provide.
[0599] Step 8:
[0600] The matching results are sent from the server to the terminal and notified to the user. The terminal displays this information and instructs the user on the next steps to receive the provided assistance.
[0601] (Example 1)
[0602] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0603] Currently, there is a need to understand user needs in various ways and provide appropriate support based on those needs. However, existing systems have limited methods for inputting and analyzing information, making it difficult to respond to diverse user needs. In addition, the matching process with support organizations is often inefficient, and users may not be able to receive appropriate support quickly. This invention aims to address these problems by accurately analyzing diverse user input information, providing optimal measures, and achieving rapid and appropriate matching with relevant support organizations.
[0604] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0605] This invention includes a server that converts user information into an analysis format after performing integrity checks and filtering; a server that analyzes the information converted into the analysis format using a natural language processing engine and generates optimal measures using a generative AI model; and a server that interactively provides the generated measures to the user using conversational artificial intelligence and receives feedback through interaction. This enables the formulation of measures that meet the diverse needs of users and the rapid matching of users with appropriate support organizations.
[0606] "Means of converting information into an analysis format after verifying its integrity and filtering" refers to a process of converting information sent by a user into an accurate and consistent format, thereby removing unnecessary data and errors, and then converting it into a format suitable for further analysis.
[0607] "A means of analyzing using a natural language processing engine and generating optimal measures using a generative AI model" refers to a process of interpreting and analyzing text and audio information using advanced algorithms, and then using an AI model generated based on the analysis results to derive the most suitable support measures for the user's needs.
[0608] "A means of interactively providing information to users using conversational artificial intelligence and receiving feedback through interaction" refers to a process that uses artificial intelligence technology to enable natural conversations with users, provides generated policy information to users, accepts questions and comments from users, and takes further action based on them.
[0609] This invention is a system that begins with a user inputting their situation or challenges through a terminal, and then optimally processes the input information to provide support. Users can input information using electronic devices such as smartphones or personal computers through a dedicated application or web interface. In this case, input is possible not only in text format but also in voice format, and the terminal uses speech recognition technology (e.g., speech recognition API) to convert the voice information into text.
[0610] The terminal is responsible for receiving information and sending it to the server as data packets. The server first verifies the integrity of this information and filters it, then converts it into a format for analysis. This is often done using software written in scripting languages such as Python. This converted information is then analyzed by a natural language processing engine (e.g., spaCy or NLTK).
[0611] Based on the analysis results, a generative AI model (e.g., Generative AI Model GPT) is used to formulate optimal measures that meet the user's needs. The generated measures are presented to the user on their device using conversational artificial intelligence (e.g., conversational AI system). The conversational AI can explain the details of the proposal and answer questions through interaction with the user.
[0612] During implementation, users will provide prompts such as the following. Examples of specific prompts include questions like, "Please tell me how I can increase my income," or "I'm looking for local support regarding my child's education."
[0613] Furthermore, based on the information provided by the user, the server searches its database for relevant support organizations and automatically matches the user with the most suitable organization. The matching results are notified to the terminal in real time, allowing the user to quickly obtain the necessary support based on this information.
[0614] This invention provides a comprehensive system that utilizes information technology to effectively and quickly provide support that meets diverse needs.
[0615] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0616] Step 1:
[0617] Users input their situation and challenges using a dedicated application or web interface via a device such as a smartphone or PC. Input can be in text or voice. In the case of voice input, the device uses speech recognition technology to convert the speech to text. The output of the input data is formatted text data.
[0618] Step 2:
[0619] The terminal sends formatted text data to the server as data packets. The input is text data, which is sent to the server as data packets. The output on the server side is data packets ready for analysis.
[0620] Step 3:
[0621] The server first performs a consistency check on the received data packets and filters out unnecessary data and errors before converting them into an analysis format. This process is carried out using a data cleaning technique with Python. The input is the transmitted data packet, and the output is the analysis format after consistency checks have been completed.
[0622] Step 4:
[0623] The server analyzes the analysis format using a natural language processing engine to clearly identify the user's needs and problems. Here, the input is the data in the analysis format, and the output is the analyzed information on the user's needs.
[0624] Step 5:
[0625] The server generates optimal strategies using a generative AI model based on the analyzed data. In this process, the generative AI model generates "suggestions that address the user's specific needs" as prompt messages. The input is the analyzed data, and the output is specific strategy prompt messages.
[0626] Step 6:
[0627] The generated policy prompts are presented to the user by an interactive artificial intelligence (AI) on the terminal. The AI can explain the details of the policy and answer questions through conversation with the user. User input is feedback, and output is a detailed explanation of the policy.
[0628] Step 7:
[0629] Based on user feedback, the server searches its database for relevant support organizations and automatically matches the user with the most suitable one. The input consists of user feedback and analysis results, and the output is a list of the most suitable support organizations.
[0630] Step 8:
[0631] The server sends the matching results as a notification to the terminal. The terminal displays this notification to the user in real time. The input is a list of support organizations, and the output is a notification message. Based on this information, the user can take action to quickly receive support.
[0632] (Application Example 1)
[0633] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0634] Modern consumers have diverse needs and demand personalized recommendations when purchasing products. However, traditional systems have struggled to effectively recommend the best products based on user preferences and past purchase history. Furthermore, they fail to provide comprehensive support and information that users need, resulting in a lower level of satisfaction with the consumer experience.
[0635] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0636] In this invention, the server includes means for receiving and analyzing information from the user, means for generating optimal measures based on the analysis results, means for interactively providing the generated measures to the user, means for recommending products based on the user's purchase history and preferences, and means for presenting the recommended product information to the user. This makes it possible to quickly and effectively recommend products suitable for the user and to support optimal purchasing decisions.
[0637] "Means for receiving and analyzing information from users" refers to devices or programs that acquire voice or text input provided by users via smart devices or computer networks and analyze its content.
[0638] "Means for generating optimal measures based on analysis results" refers to devices or programs that construct appropriate action suggestions and support measures based on acquired information, tailored to the user's needs and objectives.
[0639] "Means of interactively providing generated measures to users" refers to devices or programs that use AI to interact with users and clearly communicate the details and options of the measures.
[0640] "Means of recommending products based on a user's purchase history and preferences" refers to devices or programs that analyze past transaction data and user preferences to suggest the most suitable products and services.
[0641] "Means of presenting recommended product information to users" refers to devices or programs that effectively communicate the features and benefits of recommended products to users visually or audibly.
[0642] The system that realizes this invention consists of a user-operated terminal, a server, and a generative AI model that handles data processing. First, the user inputs information using voice or text via a smartphone or computer. The terminal receives this input data and, if necessary, uses a speech recognition engine (e.g., Google Speech-to-Text) to convert the voice to text.
[0643] The converted data is sent from the terminal to the server. The server uses a natural language processing engine (e.g., spaCy or Google NLP API) to analyze the user's input and understand their needs. Based on this analysis, a generative AI model (e.g., OpenAI GPT-4) formulates the most suitable strategy for the user.
[0644] Furthermore, the server selects suitable products from the database based on the user's purchase history and preferences, and recommends them to the user. This process uses machine learning algorithms to analyze trends from past data and generate a final product list. The generated product information and strategies are provided to the user through an interactive AI on the terminal, with detailed information and options explained in an easy-to-understand manner.
[0645] For example, if a user enters "I'm looking for summer casual wear," the system will recommend the best products and discount information based on past purchase history and product reviews. The generating AI model analyzes the user's input and uses the prompt "The user is looking for summer casual wear. Please suggest the best products and discount information based on past purchase history and user reviews" to select recommended products. This allows the user to efficiently find products that meet their needs.
[0646] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0647] Step 1:
[0648] Users input information via voice or text using a smartphone or computer terminal. The input information describes the user's needs and objectives. In the case of voice input, the terminal uses a speech recognition engine to convert it to text and adds it to the data stream as output.
[0649] Step 2:
[0650] The terminal sends the converted text data to the server. The server receives the input data and analyzes its content using a natural language processing engine. Through data analysis, it extracts keywords and topics related to the user's needs and outputs them as analysis results.
[0651] Step 3:
[0652] The server uses an AI model generated based on the analysis results to produce optimal strategies and product recommendations. The user communicates their needs to the AI model using prompts, and the algorithm constructs appropriate suggestions. In this step, strategy and product information is generated and output as suggestions.
[0653] Step 4:
[0654] The server optimizes the generated product recommendations based on the user's history and preferences. It scrutinizes the recommended product list using past purchase history and rating data, filtering as needed. This optimized product list is then prepared as the final output.
[0655] Step 5:
[0656] The well-organized recommended products and promotional information are presented to the user through the terminal's interactive AI. The terminal interactively explains the reasons for the recommendations and product details to the user, and answers the user's questions, thereby supporting their purchase and decision-making. This allows the user to make the best purchase choice based on the information provided.
[0657] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0658] In this embodiment of the invention, an interface incorporating an emotion engine is implemented in the user's terminal. The user can input information about their situation and challenges through a terminal such as a smartphone or personal computer, and in the process, multimedia input such as audio and video is possible.
[0659] The emotion engine installed in the device analyzes voice and facial expression data input by the user in real time to estimate the user's emotional state. This analysis is performed using voice tone changes and facial expression analysis technology, and the estimation results are sent to the server as part of the user information.
[0660] The server receives user text and sentiment data and formats it in preparation for analysis. The generative AI model uses natural language processing to extract the user's specific needs and problems, and then, taking sentiment data into consideration, proposes optimal measures adapted to the user's emotional state. This results in the generation of more personalized support measures that are tailored to the user's psychological state.
[0661] Policy information provided by the server is presented to the user via conversational AI. The conversational AI not only simply lists the generated policies, but also interacts with the user based on their emotional state, creating an environment where the user can provide feedback with confidence.
[0662] Furthermore, the server searches for support organizations that match the user's needs and emotions, and performs the optimal matching. By using emotional information, the likelihood of selecting a support organization with which the user can interact more comfortably increases.
[0663] For example, in the case of a user seeking support due to workplace stress, if tension is detected in their voice, the generating AI can prioritize suggesting relaxation-promoting measures and psychological care support. This system provides users with more appropriate and effective support measures than traditional responses based solely on text information.
[0664] The following describes the processing flow.
[0665] Step 1:
[0666] Users input their worries and challenges into their devices via text or voice input through a dedicated application or web interface that incorporates an emotion engine. In the case of voice input, the input is recorded as audio data using the device's microphone.
[0667] Step 2:
[0668] The device processes the input audio data in real time and converts it into text data using speech recognition technology. Simultaneously, an emotion engine analyzes the tone and volume of the audio data, as well as changes in facial expressions in the video input, to estimate emotions.
[0669] Step 3:
[0670] The device sends the converted text data and estimated sentiment information to the server. Encryption technology is used during transmission to maintain data confidentiality.
[0671] Step 4:
[0672] The server formats the text data and sentiment information received from the terminal into an analysis format. This performs data cleansing, removing noise that would interfere with the analysis.
[0673] Step 5:
[0674] The generative AI model runs on a server and uses formatted data to interpret user problems and needs through natural language processing. Furthermore, it incorporates emotional information into the analysis results to generate strategies adapted to those emotions.
[0675] Step 6:
[0676] The server sends the generated measures to the terminal, and the conversational AI presents them to the user. The conversational AI explains the details of the generated measures and receives feedback and additional information from the user.
[0677] Step 7:
[0678] The server searches its database for the most suitable support group based on the user's needs and feelings, and performs a matching process. Prioritizing groups that the user can comfortably interact with is a key consideration.
[0679] Step 8:
[0680] Matched support group information is sent from the server to the terminal and notified to the user in real time. The terminal displays this information to the user and provides advice on what to do next.
[0681] (Example 2)
[0682] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0683] In modern society, there is a need to provide more accurate and personalized support quickly to address the psychological and practical problems faced by individual users. However, conventional systems have been insufficient in optimizing support while considering the user's emotional state. Furthermore, they have not been able to effectively match users with support organizations.
[0684] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0685] In this invention, the server includes means for analyzing diverse input information from the user and estimating the emotional state, means for generating AI that generates optimal measures according to the estimated emotional state, and means for searching for and matching support organizations based on the user's needs and emotions. This enables the provision of more personalized support according to the user's psychological state and rapid matching of the user with the most suitable support organization.
[0686] A "user" refers to an entity that uses a system to provide information and receive support and policy suggestions.
[0687] "Diverse input information" refers to information in multiple formats, such as text data, audio data, and video data, provided by the user.
[0688] "Emotional state" refers to the emotional and mood state expressed by the user, and is estimated through analysis of voice and facial expressions.
[0689] "Means of estimation" refers to technical methods used to identify and analyze emotional states based on user input information.
[0690] "Generative AI methods" refer to systems that utilize artificial intelligence technology to generate optimal measures and suggestions based on the user's needs and emotional state.
[0691] "Interactive delivery methods" refer to technologies that interactively display or communicate generated measures to users.
[0692] A "support group" refers to an organization or institution that provides support for problems that users are facing.
[0693] "Matching methods" refer to technical means for selecting and introducing support organizations and resources that are suitable for the user's needs and emotional state.
[0694] In this embodiment of the invention, the user accesses the system via a computer terminal such as a smartphone or personal computer. The user can input information about their daily situation or specific issues into the terminal. In this process, the terminal can collect audio and video data using its voice recognition function and camera.
[0695] The emotion engine built into the device processes voice and video data from the user in real time, analyzing changes in voice tone and facial features. Based on this analysis, the user's emotional state is estimated and added as an emotion label.
[0696] Both estimated emotion data and user input data are sent to the server. The server receives this data and uses a generative AI model to extract the user's needs and challenges. This model uses natural language processing technology to automatically generate measures that take emotional states into account. For example, a user experiencing stress at work might be offered measures specifically focused on relaxation techniques and psychological care. This process begins with a prompt such as, "I've recently been feeling stressed due to work pressure, and I'd like to know what to do about it."
[0697] The measures generated by the server are delivered to the user through an interactive AI. This AI not only presents the measures but also facilitates interactions that reflect the user's emotional state, providing a reassuring experience. Furthermore, the server searches for and matches the user with a suitable support organization based on emotional information and notifies the user of their contact information. This allows the user to receive the support that is best suited to them.
[0698] This system, following a predetermined algorithm, leverages known emotion recognition and natural language processing technologies to provide users with innovative and effective support.
[0699] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0700] Step 1:
[0701] Users access the system via devices such as smartphones and computers and input information about their situation and challenges. This input includes audio, video, and text data. For example, a user might record a statement such as "I've been feeling stressed at work lately" as audio data using the device's microphone.
[0702] Step 2:
[0703] The device processes collected audio and video data using an emotion engine, analyzing voice tone and facial expression data. This includes aspects such as voice pitch, rhythm, and facial muscle movements. Raw audio and facial expression data are used as input, and emotion labels (e.g., "feeling stressed") are generated as output. Specifically, "tension" is detected through voice tone analysis, and "fatigue" is estimated through facial expression analysis.
[0704] Step 3:
[0705] The terminal sends emotion labels and input information to the server. The server formats the received data and performs error checking. The input consists of text data related to the audio data and emotion labels, and the output is data ready for analysis, which is supplied to the generating AI model.
[0706] Step 4:
[0707] The AI model on the server extracts user needs and problems based on the configured prompt. Using natural language processing technology, it analyzes the content of the text data and sentiment labels to generate optimal solutions. For example, using the prompt "Recently, my workload has been heavy; how can I alleviate it?", it generates solutions such as "relaxation methods" and "improvements to time management."
[0708] Step 5:
[0709] The server passes the generated measures to the conversational AI, which then presents them interactively to the user. The AI engages in dialogue that takes the user's emotional state into account, providing a detailed explanation of the measures. The input is the generated measures information, and the output is voice or text feedback to the user. For example, the conversational AI might say, "Why not try yoga for stress management?"
[0710] Step 6:
[0711] The server uses user needs and emotional information to search for and match the user with the most suitable support organization. A search algorithm generates a list of support organizations that match the user's requirements. Input includes the user's location and areas of interest, and output lists information on appropriate support organizations. Specifically, it searches for "nearby organizations specializing in stress management" and notifies the user of their contact information.
[0712] (Application Example 2)
[0713] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0714] Currently, many systems process information without considering users' emotions or psychological states, providing uniform responses and thus failing to offer personalized support tailored to individual needs. Furthermore, in the security field, immediate anomaly detection based on emotional states is difficult, potentially leading to missed potential risks. To address these challenges, there is a need for technology that can specifically assess user information based on emotional states and provide appropriate measures and alerts.
[0715] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0716] In this invention, the server includes means for receiving information from the user and analyzing that information, means for generating optimal measures based on the analysis results, and means for interactively providing the generated measures to the user. This enables the provision of appropriate and personalized support measures that are tailored to the user's emotional state and needs, as well as rapid and accurate anomaly detection at the site.
[0717] "User information" refers to information provided by users, including audio, visual, and text data.
[0718] "Means of analysis" refers to a function that performs a process of evaluating the user's emotions and psychological state based on the received information.
[0719] "Means for generating optimal measures" refers to devising support measures tailored to the individual needs of users based on analysis results.
[0720] "Interactive delivery methods" refer to functions that deliver generated measures through two-way communication with users.
[0721] "Means of detecting and notifying of anomalies based on emotional state" refers to a function that analyzes the user's emotional state to determine the risk and issues an immediate alert if necessary.
[0722] "A means of searching for and matching support organizations" refers to the process of identifying and collaborating with organizations that can provide appropriate support tailored to the user's needs.
[0723] "Means of notifying users of matching results" refers to a function that communicates matching information obtained through a search to the user.
[0724] The system implementing this invention consists of a user terminal and a server. The terminal uses a device such as Google Glass or Microsoft HoloLens to receive audio, visual, and text information from the user and transmit it to the server. The Google Cloud Speech-to-Text API is installed on this device for analyzing audio data, and OpenCV and Dlib are installed for analyzing visual data.
[0725] The server takes in the received data and performs sentiment analysis. This uses OpenAI GPT-3 as a generative AI model to examine the user's emotional state. Based on the results, it generates personalized support measures. These measures include life support, vocational training, psychological care, and anomaly detection.
[0726] The generated support measures are presented to the user interactively. This is to facilitate user understanding and make it easier to obtain appropriate feedback by providing an approach tailored to the user's situation through dialogue. Furthermore, the server searches for support organizations that meet the required support needs and matches the user with them. Information on the selected support organization is accurately notified to the user.
[0727] As a concrete example, in commercial facilities, security guards may use smart glasses to check the emotional state of visitors and identify suspicious individuals. When an anomaly is detected, an alert can be issued immediately, allowing for a rapid response to the situation. Using prompts such as, "Detect anomalies from visual and audio data. Example: Analyze suspicious behavioral patterns and speech patterns exhibited by visitors and notify security personnel," it is possible to provide information that responds to the user's intentions in real time.
[0728] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0729] Step 1:
[0730] The device acquires audio, visual, and text information from the user. In this step, data is captured using the camera and microphone of smart glasses. The input is the user's real-time audio and video information, and the output is analyzable digital data.
[0731] Step 2:
[0732] The terminal preprocesses the acquired data, using the Google Cloud Speech-to-Text API for audio data analysis and OpenCV and Dlib for visual data analysis. The input is the raw data acquired in step 1, and the output is frame-by-frame facial expression features and transcribed audio information. This processing converts the data into the format necessary for analysis.
[0733] Step 3:
[0734] The server receives data sent from the terminal and uses the generative AI model OpenAI GPT-3 to estimate the user's emotional state. The input is the feature data and text data generated in the previous step, and the output is the user's emotional state and its confidence score. In this step, emotions are quantified by evaluating voice tone and facial expression changes.
[0735] Step 4:
[0736] The server generates optimal measures based on the emotional state. The generating AI model considers the estimated emotional state and creates measures such as relaxation recommendations or warnings. The input is the emotional state and confidence level from step 3, and the output is the suggested measures for the user.
[0737] Step 5:
[0738] The server interactively provides the generated measures to the user. During this process, it uses prompts to present the user with details of the measures. The input consists of internal rules regarding the measures and their delivery methods, while the output is the measures information visualized for the user. Specifically, notifications may appear on the user's glasses, and supplementary explanations may be provided via audio.
[0739] Step 6:
[0740] The server searches for relevant support organizations based on the user's needs and makes appropriate matches. The input is the user's needs and emotional state, and the output is the contact information of the support organizations. This step ensures that the user receives appropriate assistance.
[0741] Step 7:
[0742] The server notifies the user of the matching results and provides guidance for further assistance as needed. The input is information about support organizations obtained through the search, and the output is detailed assistance guidance displayed on the device. The user can access the relevant facilities through smart glasses.
[0743] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0744] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0745] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0746] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0747] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0748] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0749] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0750] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0751] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0752] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0753] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0754] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0755] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0756] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0757] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0758] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0759] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0760] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0761] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0762] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0763] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0764] The following is further disclosed regarding the embodiments described above.
[0765] (Claim 1)
[0766] A means of receiving information from users and analyzing that information,
[0767] A means of generating optimal measures based on analysis results,
[0768] A means of interactively providing the generated measures to the user,
[0769] A means of searching for and matching support organizations that meet the user's needs,
[0770] A means of notifying users of the matching results,
[0771] A system that includes this.
[0772] (Claim 2)
[0773] The system according to claim 1, wherein the generated measures include one of the following: life support, vocational training, or psychological care.
[0774] (Claim 3)
[0775] The system according to claim 1, wherein user input information includes voice and text.
[0776] "Example 1"
[0777] (Claim 1)
[0778] A means for receiving information from users, verifying its integrity and filtering it, and then converting it into an analysis format,
[0779] A means of analyzing information converted into an analysis format using a natural language processing engine, and generating optimal measures using a generated AI model based on the analysis results,
[0780] The generated measures are provided to the user interactively using conversational artificial intelligence, and feedback is received through further interaction.
[0781] A means of automatically matching users with support organizations that meet their needs by searching relevant databases,
[0782] A means of notifying users of the matching results from support organizations,
[0783] A system that includes this.
[0784] (Claim 2)
[0785] The system according to claim 1, wherein the generated measures include any of the following: social support, professional skills development, or mental health care.
[0786] (Claim 3)
[0787] The system according to claim 1, wherein user input information includes voice and text formats.
[0788] "Application Example 1"
[0789] (Claim 1)
[0790] A means of receiving information from users and analyzing that information,
[0791] A means of generating optimal measures based on analysis results,
[0792] A means of interactively providing the generated measures to the user,
[0793] A means of recommending products based on the user's purchase history and preferences,
[0794] A means of presenting recommended product information to the user,
[0795] A means of searching for and matching support organizations that meet the user's needs,
[0796] A means of notifying users of the matching results,
[0797] A system that includes this.
[0798] (Claim 2)
[0799] The system according to claim 1, wherein the generated measures include any of the following: life support, vocational training, psychological care, or product purchase support.
[0800] (Claim 3)
[0801] The system according to claim 1, wherein user input information includes voice and text.
[0802] "Example 2 of combining an emotion engine"
[0803] (Claim 1)
[0804] A means of analyzing diverse input information from users to estimate their emotional state,
[0805] A generative AI means that generates the optimal measures according to the estimated emotional state,
[0806] A means of interactively providing the generated measures to the user,
[0807] A means of searching for and matching support organizations based on user needs and emotions,
[0808] A means of notifying users of the matching results,
[0809] A system that includes this.
[0810] (Claim 2)
[0811] The system according to claim 1, wherein the generated measures include health management, educational support, or emotional care.
[0812] (Claim 3)
[0813] The system according to claim 1, wherein user input information includes multimedia data and text.
[0814] "Application example 2 when combining with an emotional engine"
[0815] (Claim 1)
[0816] A means of receiving information from users and analyzing that information,
[0817] A means of generating optimal measures based on analysis results,
[0818] A means of interactively providing the generated measures to the user,
[0819] A means of determining and notifying of an anomaly based on the user's emotional state,
[0820] A means of searching for and matching support organizations that meet the user's needs,
[0821] A means of notifying users of the matching results,
[0822] A system that includes this.
[0823] (Claim 2)
[0824] The system according to claim 1, wherein the generated measures include any of the following: life support, vocational training, psychological care, or anomaly detection.
[0825] (Claim 3)
[0826] The system according to claim 1, wherein user input information includes voice, visual, and text. [Explanation of Symbols]
[0827] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving information from users and analyzing that information, A means of generating optimal measures based on analysis results, A means of interactively providing the generated measures to the user, A means of searching for and matching support organizations that meet the user's needs, A means of notifying users of the matching results, A system that includes this.
2. The system according to claim 1, wherein the generated measures include one of the following: life support, vocational training, or psychological care.
3. The system according to claim 1, wherein user input information includes voice and text.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A