system

A system using natural language processing and virtual reality to analyze user inputs and provide feedback enhances the understanding and practice of dementia prevention behaviors, addressing the challenge of complexity in existing information systems.

JP2026068455APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Individuals find it difficult to understand and implement effective dementia prevention actions due to the complexity and diversity of available information, necessitating intuitive and effective means for understanding and practicing such behaviors.

Method used

A system that utilizes natural language processing to analyze user inputs, provides relevant information, generates virtual reality experiences, and analyzes user behavior data to offer feedback, thereby facilitating intuitive understanding and practice of dementia prevention behaviors.

Benefits of technology

Enables users to concretely understand and incorporate dementia prevention behaviors into their daily lives, supporting healthy lifestyle habits through continuous data analysis and personalized feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068455000001_ABST
    Figure 2026068455000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A natural language processing means that receives voice or text input data, analyzes the input data to understand the user's intent, Information provision means that retrieves information based on analyzed intent from a database and generates an appropriate response, A virtual reality generation means for displaying selected content in order to provide a virtual reality experience to the user, A data analysis means that records user behavior data, analyzes said behavior data, and provides feedback, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, dementia has become a major health problem associated with aging, and its prevention is an urgent task. However, it is not easy for individual users to understand actions effective for dementia prevention and incorporate them into their daily lives. Since information is diverse and its acquisition and implementation methods may be difficult for users to understand, there is a need for means that can be understood and practiced more intuitively and effectively.

Means for Solving the Problems

[0005] This invention provides a natural language processing means that receives voice or text input data and understands the user's intent through analysis. Based on the analysis results, it appropriately provides the user with acquired information and offers a virtual reality experience related to specific dementia prevention behaviors. Furthermore, it constructs a system equipped with a data analysis means that records user behavior data and analyzes that data to provide feedback for improving healthy habits, thereby realizing an environment in which users can intuitively understand and practice dementia prevention behaviors.

[0006] "Voice or text input data" refers to information that a user sends to the system using voice or text.

[0007] "Natural language processing means" refers to a device or system that analyzes input data in the form of speech or text and performs technical processing to understand the user's intent.

[0008] "Information provision means" refers to the means of generating and presenting appropriate responses to users based on information obtained from servers and databases.

[0009] "Virtual reality generation means" refers to technologies and devices for providing users with a virtual space visually and aurally.

[0010] "Data analysis means" refers to a technology or method for collecting and analyzing user behavior data and providing feedback based on that data. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] This invention relates to a support system that enables users to intuitively understand and practice dementia prevention behaviors. The main components consist of a terminal operated by the user, a server that processes data, and a system that generates a virtual reality experience.

[0033] First, the user inputs a question using their device via voice or text. For example, they might ask, "What exercises are effective for preventing dementia?" The device then sends this input to the server.

[0034] The server performs natural language processing to parse the input it receives. This is the process of understanding the user's question and identifying their intent and requests. The server then refers to databases and external knowledge bases to generate answers to the user's question. For example, it might provide information such as, "Walking and swimming are effective."

[0035] The generated response is sent from the server to the terminal, which then displays it to the user. The user can then decide on further actions based on this information.

[0036] Next, if a user wishes to experience a virtual reality (VR) activity related to a specific dementia prevention behavior, the request is received from the device. The device receives instructions from the server and generates VR content based on the selected experience. This content features 360-degree video and sound effects, allowing the user to fully enjoy the atmosphere of walking.

[0037] During the experience, the device collects user behavior data. This data is sent to the server as the user's usage history. The server analyzes the collected data and provides feedback with suggestions for improvement to help the user more effectively prevent dementia.

[0038] This invention provides an innovative means to help users understand dementia prevention behaviors more concretely and incorporate them into their daily lives. Furthermore, it supports the realization of healthy lifestyle habits based on continuous data analysis.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The user uses the device to input questions or requests via voice or text. The device receives this data and prepares to send it to the server.

[0042] Step 2:

[0043] The terminal sends the entered data to the server. The server receives the data and begins preparing to analyze its contents.

[0044] Step 3:

[0045] The server uses natural language processing to analyze the input data. Through this analysis, it clearly understands the intent of what the user wants to ask and identifies the necessary information.

[0046] Step 4:

[0047] Based on the identified intent, the server searches for relevant information from databases and external knowledge bases and generates an appropriate response for the user. This response includes specific suggestions and information regarding the user's question.

[0048] Step 5:

[0049] The server sends the generated response data to the terminal. The terminal receives this data and displays or provides audio guidance in a format that is easy for the user to understand.

[0050] Step 6:

[0051] Based on the information provided, users can request a virtual reality experience. This request is sent to the server via their device.

[0052] Step 7:

[0053] The device generates relevant virtual reality content based on the user's chosen virtual reality experience. It creates 360-degree video and audio to provide an immersive experience for the user.

[0054] Step 8:

[0055] Users experience virtual reality on their devices and virtually engage in relevant dementia prevention behaviors. During this time, the device records data on usage and behavior.

[0056] Step 9:

[0057] The device sends the collected behavioral data to the server. The server receives this data and analyzes the user's behavior patterns and usage frequency.

[0058] Step 10:

[0059] Based on the analyzed data, the server generates feedback suggesting areas for improvement and further healthy behaviors for the user. This feedback is then delivered to the user via their device.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In an aging society, a key challenge in dementia prevention is ensuring that users understand how to effectively implement preventive activities and sustain their effects. This challenge goes beyond mere information provision; interactive support is essential to encourage users to take action and promote sustainable health maintenance. In particular, it is crucial that users not only receive information but also perceive it through real-life experiences and connect it to their daily lives. However, conventional systems are insufficient in this regard and are not intuitive for users, which is why they have not been widely adopted.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes a natural language understanding means that receives information in the form of voice or text and analyzes that information to understand the user's intent; an information response means that retrieves information based on the analyzed intent from a storage device and generates an appropriate response; and a virtual environment generation means that displays selected content in order to provide the user with a virtual environment experience. As a result, the user not only receives information about dementia prevention, but also deepens their understanding through real-world experiences based on that information, enabling them to effectively and sustainably practice preventive activities.

[0065] "Natural language understanding methods" are techniques for analyzing information in the form of speech or text to understand the user's intentions and requests.

[0066] An "information response means" is a method for obtaining relevant information from a storage device based on an analyzed intent and generating an appropriate response.

[0067] A "virtual environment generation method" is a system that provides a virtual reality experience to a user based on their selected content.

[0068] "Information analysis methods" refer to techniques for recording user behavior information, analyzing that information, and providing feedback.

[0069] The "response generation means" is a function that creates and provides suggestions for improving healthy habits based on the user's behavioral information.

[0070] This invention is a comprehensive support system for users to understand and practice dementia prevention behaviors. The system primarily includes a user-operated terminal, a data processing server, and means for generating virtual reality experiences.

[0071] The device receives voice or text input from the user. This input is made through a dedicated application, allowing the user to enter prompts such as, "Please tell me about exercises that are effective for preventing dementia."

[0072] The server uses natural language understanding (NLP) tools to analyze the input it receives. This process utilizes generative AI models such as BERT and GPT to determine the intent of the user's question and generate appropriate information. Specifically, it accesses databases and external knowledge bases to extract and generate information such as "Walking and swimming are effective."

[0073] This information is sent from the server to the terminal, which then displays it to the user. The terminal's display is used for this display, enabling immediate visual feedback.

[0074] Furthermore, if a user wishes to experience virtual reality, the device sends a request to the server. Based on the user's selection, the server generates a virtual environment utilizing 360-degree video and sound, allowing the user to experience real-world activities. For example, a walking course within the virtual reality environment is provided, which the user can explore.

[0075] During the trial, user behavior data is recorded on the device and sent to the server. The server analyzes this data and generates feedback to help users improve their healthy habits, offering suggestions for further improvement.

[0076] An example of a prompt is, "Please give specific suggestions for a healthy diet," and the response to this input would be specific information such as, "A diet rich in fish is recommended."

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The user enters questions via voice or text using a terminal. This involves opening a dedicated application and entering a prompt. For example, "Please tell me about exercises that are effective for preventing dementia" might be entered. This input data is formatted and sent to the server. The output here is structured data sent to the server.

[0080] Step 2:

[0081] The server analyzes the received input data using natural language understanding (NLP) tools. Specifically, it uses a generative AI model to understand the intent of the input prompt. This analysis identifies the user's question and its purpose. The output after analysis provides specific information about the user's intent.

[0082] Step 3:

[0083] Based on the analyzed intent, the server retrieves relevant information from databases and external knowledge bases. This process generates specific responses, such as "Walking and swimming are effective." The output obtained using the information response means is appropriate answer information to the user's question.

[0084] Step 4:

[0085] The generated response information is sent from the server to the terminal. The terminal receives this information and displays it so that the user can visually confirm it. Specifically, the information is presented on the display screen as text or audio. The output of this step appears in a state where the user can see or hear the information.

[0086] Step 5:

[0087] When a user wishes to experience virtual reality, a request is sent from the device to the server. The request includes the desired virtual reality experience and detailed settings. The output in response to this input is the server's preparation for the virtual environment experience.

[0088] Step 6:

[0089] The server generates a virtual reality experience based on the requested content. This involves generating 360-degree video and audio, preparing to provide the user with an immersive experience. The output is content data ready for playback on the device.

[0090] Step 7:

[0091] The terminal provides the user with generated virtual reality content. Here, the virtual environment that the user experiences is recreated. During the experience, user behavior data is recorded. This record includes the user's eye movements and selected options. The output of this step is the behavior data sent to the server.

[0092] Step 8:

[0093] The server analyzes collected behavioral data and generates feedback for improving health habits. This analysis includes evaluating trends and effects based on the behavioral data. The feedback includes specific improvement suggestions and next steps, which are provided to the user. The output of this step is improvement suggestions presented to the user.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] Activities aimed at preventing dementia and maintaining health among the elderly face challenges in terms of being difficult to understand and implement. As a result, effective preventive actions are not taken, and continuous health management is difficult. Furthermore, existing health services do not adequately adapt to individual needs and cannot provide effective approaches for each person.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes natural language processing means for receiving voice or text input data and analyzing the input data to understand the user's intent; information providing means for obtaining information based on the analyzed intent from a storage medium and generating an appropriate response; and virtual reality generating means for displaying selected content to provide the user with a virtual reality experience. This enables the user to intuitively understand and practice health maintenance behaviors.

[0099] "Natural language processing means" are technical means for analyzing input data in the form of speech or text to understand the user's intent.

[0100] "Information provision means" refers to a technical means that, based on the analyzed user's intent, retrieves appropriate information from a storage medium and generates a response.

[0101] A "virtual reality generation means" is a technical means for displaying selected virtual reality content to a user and providing them with an experience.

[0102] "Data analysis means" refers to technical means that record and analyze user behavior data to provide feedback.

[0103] A "visual information device" is a device used to allow users to experience and facilitate health-related activities in real time.

[0104] "Interface generation means" refers to technical means for promoting health-related activities and providing operability to users through a visual information device.

[0105] The system for carrying out this invention consists mainly of an information terminal operated by the user, a server for data processing, and a visual information device for generating virtual reality. The server has the capability to process user questions in natural language, either by voice or text. This can be achieved, for example, using a Python-based natural language processing library (e.g., spaCy). User input is sent to the server via the information terminal, where the server analyzes the input and extracts the user's intent. Based on the analysis results, it retrieves appropriate information from a database as a storage medium and sends a response generated by an AI model to the terminal. In this case, the use of a generative AI model can be considered.

[0106] Furthermore, the virtual reality generation system utilizes a virtual reality simulation engine such as Unity to display user-selected content on a visual information device in real time. This allows users to easily experience health-related activities. The visual information device is envisioned to be smart glasses or similar devices and functions as an interface generation means to facilitate user behavior. The system also records user behavior data as accumulated data and provides feedback based on its patterns. This data analysis aims to present users with concrete action plans to improve their lifestyle habits.

[0107] As a concrete example, when a user enters a specific area in a supermarket, a visual information device displays health food information for that area. Based on a prompt such as "Please tell me about exercises that are effective in preventing dementia," relevant exercise information and walking routes are displayed on the user's glasses to support them in exercising.

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The user inputs voice or text data via an information terminal. The input data is sent to a server to identify the information the user is requesting. Voice data, as input, is converted to text using speech recognition software on the terminal.

[0111] Step 2:

[0112] The server analyzes the received input data using natural language processing techniques. Specifically, it uses a generative AI model to understand the user's intent through syntactic and semantic analysis of the input data. The analyzed intent is output and passed on to the next processing step.

[0113] Step 3:

[0114] Based on the analyzed intent, the server retrieves appropriate information from the database on the storage medium. For example, if the user wants to know about exercises related to dementia prevention, information such as walking and swimming will be retrieved from the database. This search result is output as the appropriate response.

[0115] Step 4:

[0116] The server sends an appropriate response generated through a generative AI model to the information terminal. The terminal then displays the response to the user in natural language. Based on this information, the user can choose actions to prevent dementia.

[0117] Step 5:

[0118] When a user requests a virtual reality experience, a corresponding request is sent from the terminal to the server. The server generates appropriate virtual reality content via a virtual reality generation system. The generated content is sent to a visual information device such as smart glasses and displayed to the user in real time.

[0119] Step 6:

[0120] A visual information device collects behavioral data during the user's activities. This data is transmitted to a server, where data analysis tools analyze the user's habits and patterns, and process it into feedback for health management. The analysis results are used to suggest future activities for the user.

[0121] Step 7:

[0122] The server generates feedback for the user based on collected behavioral data. Insights gained from the data are used to generate specific suggestions for more effective health maintenance methods, which are then delivered to the user via their device.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention relates to a system that incorporates an emotion engine to help users understand information related to dementia prevention and support their actions. The system consists of a terminal that processes user input, a server that handles data processing, an emotion engine that performs emotion recognition, and components that generate a virtual reality experience.

[0125] Users input questions and requests via voice or text through their device. For example, they might ask, "What kind of exercise is good for preventing dementia?" The device then prepares to send this input to the server.

[0126] The server performs natural language processing on the received input data to analyze the user's intent. Based on this analysis, the server retrieves necessary information from the database. Subsequently, the emotion engine recognizes the user's emotional state based on the user's input, voice tone, past behavior patterns, and other factors.

[0127] Based on the emotions recognized by the emotion engine, the server customizes the information it provides. For example, for a user experiencing stress, it can suggest exercises that promote relaxation. The optimized information is then sent to the device and presented to the user.

[0128] Furthermore, if a user chooses a virtual reality experience to practice healthy habits, the content of that experience will also be adjusted based on the results of the emotion recognition. For example, a user experiencing anxiety will be provided with a virtual environment featuring calming scenery and music.

[0129] During the virtual reality experience, the device records other behavioral data. This data is sent to a server, where an emotion engine performs further cumulative analysis. The server uses the results of this analysis to provide the user with suggestions for improvement and feedback. The feedback is presented as personalized advice based on the user's emotional state and behavioral history.

[0130] Thus, according to the present invention, users can not only receive information but also gain an experience that promotes more effective dementia prevention behaviors tailored to their individual emotional state. This system aims to improve the user's quality of life and effectively delay or prevent the onset of dementia.

[0131] The following describes the processing flow.

[0132] Step 1:

[0133] Users input questions and requests via voice or text through the device. For example, they might ask, "What exercises are good for preventing dementia?" The device collects this input data.

[0134] Step 2:

[0135] The terminal sends the collected input data to the server. The server receives this data and uses natural language processing technology to analyze the user's intent and the content of the question.

[0136] Step 3:

[0137] Based on the analysis results, the server searches for relevant information from databases and external sources and generates an appropriate response. This information includes specific answers to the user's questions.

[0138] Step 4:

[0139] The server simultaneously uses an emotion engine to recognize the user's emotional state from input data and past behavioral data. For example, it can determine if the user is stressed based on their tone of voice and word choice.

[0140] Step 5:

[0141] Based on the user's emotional state, the server adjusts the information and suggestions it generates. For example, a user who needs to relax might be suggested gentle exercise or simple workouts.

[0142] Step 6:

[0143] The adjusted information is sent from the server to the terminal, which then guides the user through the information on the display or via voice. The user then uses this information to decide on specific actions.

[0144] Step 7:

[0145] If a user wishes to experience virtual reality, they select this request on their device, and the virtual reality experience begins. Based on the results of the emotion engine's recognition, the server customizes the content of the virtual reality experience, for example, by providing relaxing images and music.

[0146] Step 8:

[0147] The device provides users with a virtual reality experience and records behavioral data collected during that time. This includes usage details such as the duration of the experience and the content selected.

[0148] Step 9:

[0149] The device sends recorded behavioral data to the server. The server analyzes this data to understand the user's behavioral trends and emotions, and prepares feedback for the next time.

[0150] Step 10:

[0151] Based on the analysis results, the server generates new suggestions and feedback for the user. This feedback is customized based on the user's emotional state and behavioral history and is delivered to the user's device.

[0152] (Example 2)

[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0154] Conventional dementia prevention support systems lack mechanisms to efficiently recognize users' emotional states and individual behavioral patterns, and to provide appropriate responses and feedback. Furthermore, they are unable to go beyond simply providing information and deliver content optimized according to the user's emotional state.

[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0156] In this invention, the server includes natural language processing means that receive input information in the form of voice or text and analyze the input information to understand the user's purpose; emotion recognition means that analyze the input content and past behavioral patterns in order to recognize the user's emotional state; and data analysis means that record the user's behavioral data and analyze the behavioral data to provide feedback. This makes it possible to customize information and feedback according to the user's emotional state.

[0157] "Voice or text input information" refers to audio data spoken by the user to the system, or text data entered using a keyboard or other means.

[0158] "Natural language processing means" refers to technologies that analyze input language data to understand its context and meaning, and functions to identify the requests and questions that the user intends to ask the system.

[0159] "Information provision means" refers to methods and functions for collecting appropriate information based on the analyzed user's objectives and communicating that information to the user.

[0160] "Emotion recognition methods" are technologies that analyze user input data and behavioral history to identify the emotions and psychological states that a user exhibits.

[0161] "Information optimization means" refers to methods or functions that adjust the information provided to the user based on the results of emotion recognition, according to the user's emotional state, and present it in a more appropriate form.

[0162] "Virtual reality generation means" refers to technologies and functions that generate and display content in order to provide users with a visual and experiential virtual environment.

[0163] "Data analysis means" refers to technologies that collect and analyze user behavior data, extract useful information from that data, and provide feedback.

[0164] A "feedback generation method" is a function that takes into account the user's behavioral history and emotional state to generate and provide personalized improvement suggestions and advice to the user.

[0165] This invention is a system for supporting dementia prevention that recognizes the user's emotional state based on input information and provides information and virtual reality experiences tailored to that state.

[0166] Users input questions and requests into the system via voice input devices or text input interfaces. For example, they might input a question such as, "What kind of exercise is helpful in preventing dementia?" The terminal sends the input voice or text data to the server, which is then prepared for analysis.

[0167] The server uses natural language processing tools to analyze user input and identify the underlying intent. Specifically, it uses software libraries such as spaCy and NLTK to analyze the input text. Based on the detected intent, the server retrieves the necessary information from the database.

[0168] Subsequently, the emotion engine analyzes the user's input and past behavior history to recognize the user's emotional state. This analysis utilizes voice tone analysis and logs of previous interactions, among other things. The user's emotions are evaluated using tools such as IBM Watson® Tone Analyzer.

[0169] The server optimizes and customizes the information provided to the user based on the emotion recognition results. For example, if the user is feeling stressed, it can suggest exercise methods that help with relaxation. The optimized information is sent to the terminal and presented to the user.

[0170] Furthermore, if the user chooses a virtual reality experience, the device will provide an appropriate experience through VR equipment (e.g., HTC Vive, Oculus Rift). Depending on the user's emotional state, it will visualize calming natural scenery in the virtual environment to promote a relaxation experience.

[0171] As a concrete example, if you input the prompt sentence "Please tell me a good way to fall asleep when I'm feeling anxious" into the AI ​​model, the system will generate a response that explains appropriate relaxation techniques.

[0172] The device records user behavior data during the virtual experience and sends it back to the server. The server aggregates this data and, through further analysis by an emotion engine, provides the user with personalized feedback and improvement suggestions.

[0173] The system based on this invention aims to enhance the effectiveness of dementia prevention through the provision of appropriate information and experiences tailored to the user's emotional state.

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] Users input questions and requests into the system via voice input devices or text input interfaces. This input is received by the terminal as audio data captured by the microphone in the case of voice input, or as text data in the case of text input. Specifically, this involves converting the voice to text using speech-to-text conversion software. The input data might be something like, "What kind of exercise is good for preventing dementia?", and this text data is sent to the server as output.

[0177] Step 2:

[0178] The terminal packets the text data received from the user and sends it to the server. The input is text data, and encryption technology is used to securely transmit it to the server. The output is encrypted text data received on the server side. Specific operations include data transfer via the HTTPS protocol.

[0179] Step 3:

[0180] The server performs natural language processing on the received text data to analyze the topic the user is interested in. The input is decrypted text data. Specifically, it uses spaCy or NLTK to tokenize the data, recognize parts of speech, and identify the user's intent. As output, a data object representing the user's intent is generated.

[0181] Step 4:

[0182] The server retrieves relevant information from the database based on the analyzed user intent. The input is a data object representing the user's intent. Database queries are used to retrieve corresponding health information and exercise methods. The output is information data tailored to the user's intent. Specifically, SQL queries are executed.

[0183] Step 5:

[0184] The emotion engine analyzes user input, voice tone, and past behavioral history to recognize the user's emotional state. Input consists of user text data and behavioral history sent from the server. This data is analyzed using an AI algorithm, and an evaluation result indicating the emotional state is generated as output. Specific operations include frequency analysis of voice tone and text sentiment analysis.

[0185] Step 6:

[0186] The server customizes the information it provides based on the emotional evaluation results generated by the emotion engine. Inputs include the emotional state evaluation results and previously acquired informational data. Filtering and prioritization are performed to best suit the user. Outputs are customized informational data. Specific operations include ranking and selecting information.

[0187] Step 7:

[0188] The terminal presents customized information sent from the server to the user. The input is customized information data. The terminal outputs this data using a user interface to display it visually or audibly. Specific operations include rendering the information to a graphical user interface.

[0189] Step 8:

[0190] If the user chooses a virtual reality experience, the device delivers this experience through VR equipment. Inputs are data indicating the virtual reality content and the user's state. The content is loaded into the hardware device, and feedback is provided to the user. Outputs are the visual and auditory virtual reality experience. Specific actions include launching the VR simulator and adjusting the scene based on the user's emotional state.

[0191] Step 9:

[0192] During the virtual experience, the device continuously records user behavior data and sends it to the server. The input is the user's movement data during the VR experience, including eye tracking and reaction time. This transmitted data is used in the feedback generation process. The output is the accumulated behavior data.

[0193] Step 10:

[0194] The server generates feedback and improvement suggestions based on behavioral data. The input is the result of emotion and behavior analysis. The feedback generation mechanism creates personalized recommendations based on the collected data. The output is the provision of customized feedback to the user. Specific actions include generating and personalizing feedback templates.

[0195] (Application Example 2)

[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0197] In modern society, there is a demand for personalized information based on individual emotional states. In particular, there is a need for systems that reduce user stress and anxiety and provide a more comfortable experience. However, current systems struggle to quickly and accurately analyze a user's emotional state and provide corresponding information. This invention provides a system to solve these problems.

[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0199] In this invention, the server includes: natural language processing means for receiving information in the form of voice or text and analyzing the information to understand the user's intent; information providing means for obtaining information from information sources based on the analyzed intent and generating an appropriate response; virtual reality generating means for displaying selected virtual content in order to provide an optimized virtual reality environment based on the user's emotional state; and information analysis means for recording the user's behavioral information, analyzing the behavioral information, and providing feedback. This makes it possible to provide more appropriate and effective information and experiences that take into account the user's emotional state.

[0200] "Natural language processing means" refers to technologies that analyze spoken or written information and have the function of understanding the user's intent and meaning.

[0201] An "information provision tool" is a technology that has the function of obtaining necessary information from information sources based on analyzed intent and generating an appropriate response for the user.

[0202] "Virtual reality generation means" refers to a technology that displays selected content to provide an optimized virtual reality environment, taking into account the user's emotional state.

[0203] "Information analysis means" refers to technology that has the function of recording user behavior information and providing feedback by analyzing it.

[0204] "Emotional state" refers to the user's psychological and emotional condition, and is an element that enables the provision of appropriate information and experiences based on that state.

[0205] "Virtual content" refers to information and visual elements displayed within a virtual reality environment, which are selected according to the user's emotions and intentions.

[0206] This invention comprises a terminal with a user interface, a server for data processing, a natural language processing function for analyzing voice and text information, an emotion engine for recognizing the user's emotional state, and components for generating a virtual reality experience.

[0207] A terminal is a device that receives audio and visual information using smart glasses or a head-mounted display. The terminal provides an interface for receiving user input and transmitting it to a server.

[0208] The server runs Python-based programs. It uses Python natural language processing libraries (e.g., spaCy, NLTK) to analyze received audio or text information. The librosa library is used for speech tone analysis. This allows not only to understand the user's intent but also to identify their emotional state.

[0209] The emotion engine recognizes the user's emotions using collected data. Based on this, the server selects response information appropriate to the user's state and generates an optimal virtual reality environment. The virtual reality experience is built using Unity or Unreal Engine, providing the user with a personalized shopping or relaxation experience.

[0210] As a concrete example, when a user is feeling stressed, the system will display music and videos designed to promote relaxation based on their emotional state. Information on products that are effective in relieving stress will also be presented within the virtual environment.

[0211] An example of a prompt using a generative AI model is, "Generate a list of recommended products to offer to a user who is feeling stressed and looking for relaxation items." Based on this, the AI ​​can generate content tailored to the user's needs and provide a personalized virtual reality experience.

[0212] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0213] Step 1:

[0214] The user inputs information via voice or text through a device. The device converts this input into text data and sends it to the server. The input data includes the user's questions and requests.

[0215] Step 2:

[0216] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK) to understand the user's intent. Here, the input intent is recognized as a command, and indicators are generated to retrieve relevant information from the database. The analysis results are output as intent information.

[0217] Step 3:

[0218] The server uses a speech tone analysis library (e.g., librosa) to estimate the user's emotional state from the input speech data. The analysis generates emotional state data, which can then be used to determine, for example, whether the user is experiencing stress.

[0219] Step 4:

[0220] Based on emotional state data, the server uses virtual reality generation methods to create an optimized virtual environment. The content of the virtual environment is built using Unity or Unreal Engine and prepared as visual and auditory content tailored to the user's emotions. The prepared virtual content is then generated.

[0221] Step 5:

[0222] Based on acquired intent information and emotional state data, the server generates user-specific response information using information delivery tools. This includes product information and experiential content designed to alleviate stress. The response information is output in a customized format for display within the virtual reality environment.

[0223] Step 6:

[0224] When a user takes an action within the virtual reality experience, the device records that action data and sends it to a server. This allows behavioral patterns to be collected as data. This behavioral data is then used to generate feedback.

[0225] Step 7:

[0226] The server analyzes the collected behavioral data using information analysis tools and generates feedback for the user. This feedback is customized to help improve healthy habits and emotional state, and is intended to be useful for the next virtual reality experience. The generated feedback is saved for future use.

[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0230] [Second Embodiment]

[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0243] This invention relates to a support system that enables users to intuitively understand and practice dementia prevention behaviors. The main components consist of a terminal operated by the user, a server that processes data, and a system that generates a virtual reality experience.

[0244] First, the user inputs a question using their device via voice or text. For example, they might ask, "What exercises are effective for preventing dementia?" The device then sends this input to the server.

[0245] The server performs natural language processing to parse the input it receives. This is the process of understanding the user's question and identifying their intent and requests. The server then refers to databases and external knowledge bases to generate answers to the user's question. For example, it might provide information such as, "Walking and swimming are effective."

[0246] The generated response is sent from the server to the terminal, which then displays it to the user. The user can then decide on further actions based on this information.

[0247] Next, if a user wishes to experience a virtual reality (VR) activity related to a specific dementia prevention behavior, the request is received from the device. The device receives instructions from the server and generates VR content based on the selected experience. This content features 360-degree video and sound effects, allowing the user to fully enjoy the atmosphere of walking.

[0248] During the experience, the device collects user behavior data. This data is sent to the server as the user's usage history. The server analyzes the collected data and provides feedback with suggestions for improvement to help the user more effectively prevent dementia.

[0249] This invention provides an innovative means to help users understand dementia prevention behaviors more concretely and incorporate them into their daily lives. Furthermore, it supports the realization of healthy lifestyle habits based on continuous data analysis.

[0250] The following describes the processing flow.

[0251] Step 1:

[0252] The user uses the device to input questions or requests via voice or text. The device receives this data and prepares to send it to the server.

[0253] Step 2:

[0254] The terminal sends the entered data to the server. The server receives the data and begins preparing to analyze its contents.

[0255] Step 3:

[0256] The server uses natural language processing to analyze the input data. Through this analysis, it clearly understands the intent of what the user wants to ask and identifies the necessary information.

[0257] Step 4:

[0258] Based on the identified intent, the server searches for relevant information from databases and external knowledge bases and generates an appropriate response for the user. This response includes specific suggestions and information regarding the user's question.

[0259] Step 5:

[0260] The server sends the generated response data to the terminal. The terminal receives this data and displays or provides audio guidance in a format that is easy for the user to understand.

[0261] Step 6:

[0262] Based on the information provided, users can request a virtual reality experience. This request is sent to the server via their device.

[0263] Step 7:

[0264] The device generates relevant virtual reality content based on the user's chosen virtual reality experience. It creates 360-degree video and audio to provide an immersive experience for the user.

[0265] Step 8:

[0266] Users experience virtual reality on their devices and virtually engage in relevant dementia prevention behaviors. During this time, the device records data on usage and behavior.

[0267] Step 9:

[0268] The device sends the collected behavioral data to the server. The server receives this data and analyzes the user's behavior patterns and usage frequency.

[0269] Step 10:

[0270] Based on the analyzed data, the server generates feedback suggesting areas for improvement and further healthy behaviors for the user. This feedback is then delivered to the user via their device.

[0271] (Example 1)

[0272] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0273] In an aging society, a key challenge in dementia prevention is ensuring that users understand how to effectively implement preventive activities and sustain their effects. This challenge goes beyond mere information provision; interactive support is essential to encourage users to take action and promote sustainable health maintenance. In particular, it is crucial that users not only receive information but also perceive it through real-life experiences and connect it to their daily lives. However, conventional systems are insufficient in this regard and are not intuitive for users, which is why they have not been widely adopted.

[0274] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0275] In this invention, the server includes a natural language understanding means that receives information in the form of voice or text and analyzes that information to understand the user's intent; an information response means that retrieves information based on the analyzed intent from a storage device and generates an appropriate response; and a virtual environment generation means that displays selected content in order to provide the user with a virtual environment experience. As a result, the user not only receives information about dementia prevention, but also deepens their understanding through real-world experiences based on that information, enabling them to effectively and sustainably practice preventive activities.

[0276] "Natural language understanding methods" are techniques for analyzing information in the form of speech or text to understand the user's intentions and requests.

[0277] An "information response means" is a method for obtaining relevant information from a storage device based on an analyzed intent and generating an appropriate response.

[0278] A "virtual environment generation method" is a system that provides a virtual reality experience to a user based on their selected content.

[0279] "Information analysis methods" refer to techniques for recording user behavior information, analyzing that information, and providing feedback.

[0280] The "response generation means" is a function that creates and provides suggestions for improving healthy habits based on the user's behavioral information.

[0281] This invention is a comprehensive support system for users to understand and practice dementia prevention behaviors. The system primarily includes a user-operated terminal, a data processing server, and means for generating virtual reality experiences.

[0282] The device receives voice or text input from the user. This input is made through a dedicated application, allowing the user to enter prompts such as, "Please tell me about exercises that are effective for preventing dementia."

[0283] The server uses natural language understanding means to analyze the received input. In this process, generative AI models such as BERT and GPT are utilized to determine the intent of the user's question and generate appropriate information. Specifically, it accesses a database or an external knowledge base, and information such as "Walking and swimming are effective" is extracted and generated.

[0284] This information is sent from the server to the terminal, and the terminal displays it to the user. For display, by using the terminal's display, immediate visual feedback is enabled.

[0285] Furthermore, when the user desires a virtual reality experience, the terminal sends that request to the server. The server generates a virtual environment that makes full use of 360-degree video and audio based on the user's selection, enabling the user to experience actual actions. As a specific example, a walking course within the virtual reality is provided, and the user can explore it.

[0286] The action data of the user during the experience is recorded by the terminal and sent to the server. The server analyzes this data and provides further improvement suggestions to the user by generating feedback for improving healthy habits.

[0287] As an example of a prompt sentence, "Please give specific suggestions for leading a healthy diet" is given, and as a response to this input, specific information such as "A diet rich in fish is recommended" is provided.

[0288] The flow of the specific process in Example 1 will be described using FIG. 11.

[0289] Step 1:

[0290] The user enters questions via voice or text using a terminal. This involves opening a dedicated application and entering a prompt. For example, "Please tell me about exercises that are effective for preventing dementia" might be entered. This input data is formatted and sent to the server. The output here is structured data sent to the server.

[0291] Step 2:

[0292] The server analyzes the received input data using natural language understanding (NLP) tools. Specifically, it uses a generative AI model to understand the intent of the input prompt. This analysis identifies the user's question and its purpose. The output after analysis provides specific information about the user's intent.

[0293] Step 3:

[0294] Based on the analyzed intent, the server retrieves relevant information from databases and external knowledge bases. This process generates specific responses, such as "Walking and swimming are effective." The output obtained using the information response means is appropriate answer information to the user's question.

[0295] Step 4:

[0296] The generated response information is sent from the server to the terminal. The terminal receives this information and displays it so that the user can visually confirm it. Specifically, the information is presented on the display screen as text or audio. The output of this step appears in a state where the user can see or hear the information.

[0297] Step 5:

[0298] When a user wishes to experience virtual reality, a request is sent from the device to the server. The request includes the desired virtual reality experience and detailed settings. The output in response to this input is the server's preparation for the virtual environment experience.

[0299] Step 6:

[0300] The server generates a virtual reality experience based on the requested content. At this time, it generates 360-degree video and sound and prepares to provide the user with an immersive experience. As output, content data for playback on the terminal is generated.

[0301] Step 7:

[0302] The terminal provides the user with the generated virtual reality content. Here, the virtual environment experienced by the user is played back. During the experience, the user's behavior data is recorded. This record includes movements of the user's perspective and selected options, etc. The output of this step is the behavior data sent to the server.

[0303] Step 8:

[0304] The server analyzes the collected behavior data and generates feedback for improving health habits. This analysis includes evaluations of trends and effects based on the behavior data. The feedback content includes specific improvement plans and next steps and is provided to the user. The output of this step is the improvement proposal presented to the user.

[0305] (Application Example 1)

[0306] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0307] Activities for preventing dementia and maintaining health for the elderly have the problem that they are difficult to understand and practice. As a result, effective preventive actions are not taken, and it is difficult to continuously manage health. In addition, existing health services do not sufficiently adapt to individual needs and cannot provide an effective approach for each person. <-

[0308] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0309] In this invention, the server includes natural language processing means for receiving voice or text input data and analyzing the input data to understand the user's intent; information providing means for obtaining information based on the analyzed intent from a storage medium and generating an appropriate response; and virtual reality generating means for displaying selected content to provide the user with a virtual reality experience. This enables the user to intuitively understand and practice health maintenance behaviors.

[0310] "Natural language processing means" are technical means for analyzing input data in the form of speech or text to understand the user's intent.

[0311] "Information provision means" refers to a technical means that, based on the analyzed user's intent, retrieves appropriate information from a storage medium and generates a response.

[0312] A "virtual reality generation means" is a technical means for displaying selected virtual reality content to a user and providing them with an experience.

[0313] "Data analysis means" refers to technical means that record and analyze user behavior data to provide feedback.

[0314] A "visual information device" is a device used to allow users to experience and facilitate health-related activities in real time.

[0315] "Interface generation means" refers to technical means for promoting health-related activities and providing operability to users through a visual information device.

[0316] The system for carrying out this invention consists mainly of an information terminal operated by the user, a server for data processing, and a visual information device for generating virtual reality. The server has the capability to process user questions in natural language, either by voice or text. This can be achieved, for example, using a Python-based natural language processing library (e.g., spaCy). User input is sent to the server via the information terminal, where the server analyzes the input and extracts the user's intent. Based on the analysis results, it retrieves appropriate information from a database as a storage medium and sends a response generated by an AI model to the terminal. In this case, the use of a generative AI model can be considered.

[0317] Furthermore, the virtual reality generation system utilizes a virtual reality simulation engine such as Unity to display user-selected content on a visual information device in real time. This allows users to easily experience health-related activities. The visual information device is envisioned to be smart glasses or similar devices and functions as an interface generation means to facilitate user behavior. The system also records user behavior data as accumulated data and provides feedback based on its patterns. This data analysis aims to present users with concrete action plans to improve their lifestyle habits.

[0318] As a concrete example, when a user enters a specific area in a supermarket, a visual information device displays health food information for that area. Based on a prompt such as "Please tell me about exercises that are effective in preventing dementia," relevant exercise information and walking routes are displayed on the user's glasses to support them in exercising.

[0319] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0320] Step 1:

[0321] The user inputs voice or text data via an information terminal. The input data is sent to a server to identify the information the user is requesting. Voice data, as input, is converted to text using speech recognition software on the terminal.

[0322] Step 2:

[0323] The server analyzes the received input data using natural language processing techniques. Specifically, it uses a generative AI model to understand the user's intent through syntactic and semantic analysis of the input data. The analyzed intent is output and passed on to the next processing step.

[0324] Step 3:

[0325] Based on the analyzed intent, the server retrieves appropriate information from the database on the storage medium. For example, if the user wants to know about exercises related to dementia prevention, information such as walking and swimming will be retrieved from the database. This search result is output as the appropriate response.

[0326] Step 4:

[0327] The server sends an appropriate response generated through a generative AI model to the information terminal. The terminal then displays the response to the user in natural language. Based on this information, the user can choose actions to prevent dementia.

[0328] Step 5:

[0329] When a user requests a virtual reality experience, a corresponding request is sent from the terminal to the server. The server generates appropriate virtual reality content via a virtual reality generation system. The generated content is sent to a visual information device such as smart glasses and displayed to the user in real time.

[0330] Step 6:

[0331] A visual information device collects behavioral data during the user's activities. This data is transmitted to a server, where data analysis tools analyze the user's habits and patterns, and process it into feedback for health management. The analysis results are used to suggest future activities for the user.

[0332] Step 7:

[0333] The server generates feedback for the user based on collected behavioral data. Insights gained from the data are used to generate specific suggestions for more effective health maintenance methods, which are then delivered to the user via their device.

[0334] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0335] This invention relates to a system that incorporates an emotion engine to help users understand information related to dementia prevention and support their actions. The system consists of a terminal that processes user input, a server that handles data processing, an emotion engine that performs emotion recognition, and components that generate a virtual reality experience.

[0336] Users input questions and requests via voice or text through their device. For example, they might ask, "What kind of exercise is good for preventing dementia?" The device then prepares to send this input to the server.

[0337] The server performs natural language processing on the received input data to analyze the user's intent. Based on this analysis, the server retrieves necessary information from the database. Subsequently, the emotion engine recognizes the user's emotional state based on the user's input, voice tone, past behavior patterns, and other factors.

[0338] Based on the emotions recognized by the emotion engine, the server customizes the information it provides. For example, for a user experiencing stress, it can suggest exercises that promote relaxation. The optimized information is then sent to the device and presented to the user.

[0339] Furthermore, if a user chooses a virtual reality experience to practice healthy habits, the content of that experience will also be adjusted based on the results of the emotion recognition. For example, a user experiencing anxiety will be provided with a virtual environment featuring calming scenery and music.

[0340] During the virtual reality experience, the device records other behavioral data. This data is sent to a server, where an emotion engine performs further cumulative analysis. The server uses the results of this analysis to provide the user with suggestions for improvement and feedback. The feedback is presented as personalized advice based on the user's emotional state and behavioral history.

[0341] Thus, according to the present invention, users can not only receive information but also gain an experience that promotes more effective dementia prevention behaviors tailored to their individual emotional state. This system aims to improve the user's quality of life and effectively delay or prevent the onset of dementia.

[0342] The following describes the processing flow.

[0343] Step 1:

[0344] Users input questions and requests via voice or text through the device. For example, they might ask, "What exercises are good for preventing dementia?" The device collects this input data.

[0345] Step 2:

[0346] The terminal sends the collected input data to the server. The server receives this data and uses natural language processing technology to analyze the user's intent and the content of the question.

[0347] Step 3:

[0348] Based on the analysis results, the server searches for relevant information from databases and external sources and generates an appropriate response. This information includes specific answers to the user's questions.

[0349] Step 4:

[0350] The server simultaneously uses an emotion engine to recognize the user's emotional state from input data and past behavioral data. For example, it can determine if the user is stressed based on their tone of voice and word choice.

[0351] Step 5:

[0352] Based on the user's emotional state, the server adjusts the information and suggestions it generates. For example, a user who needs to relax might be suggested gentle exercise or simple workouts.

[0353] Step 6:

[0354] The adjusted information is sent from the server to the terminal, which then guides the user through the information on the display or via voice. The user then uses this information to decide on specific actions.

[0355] Step 7:

[0356] If a user wishes to experience virtual reality, they select this request on their device, and the virtual reality experience begins. Based on the results of the emotion engine's recognition, the server customizes the content of the virtual reality experience, for example, by providing relaxing images and music.

[0357] Step 8:

[0358] The device provides users with a virtual reality experience and records behavioral data collected during that time. This includes usage details such as the duration of the experience and the content selected.

[0359] Step 9:

[0360] The device sends recorded behavioral data to the server. The server analyzes this data to understand the user's behavioral trends and emotions, and prepares feedback for the next time.

[0361] Step 10:

[0362] Based on the analysis results, the server generates new suggestions and feedback for the user. This feedback is customized based on the user's emotional state and behavioral history and is delivered to the user's device.

[0363] (Example 2)

[0364] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0365] Conventional dementia prevention support systems lack mechanisms to efficiently recognize users' emotional states and individual behavioral patterns, and to provide appropriate responses and feedback. Furthermore, they are unable to go beyond simply providing information and deliver content optimized according to the user's emotional state.

[0366] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0367] In this invention, the server includes natural language processing means that receive input information in the form of voice or text and analyze the input information to understand the user's purpose; emotion recognition means that analyze the input content and past behavioral patterns in order to recognize the user's emotional state; and data analysis means that record the user's behavioral data and analyze the behavioral data to provide feedback. This makes it possible to customize information and feedback according to the user's emotional state.

[0368] "Voice or text input information" refers to audio data spoken by the user to the system, or text data entered using a keyboard or other means.

[0369] "Natural language processing means" refers to technologies that analyze input language data to understand its context and meaning, and functions to identify the requests and questions that the user intends to ask the system.

[0370] "Information provision means" refers to methods and functions for collecting appropriate information based on the analyzed user's objectives and communicating that information to the user.

[0371] "Emotion recognition methods" are technologies that analyze user input data and behavioral history to identify the emotions and psychological states that a user exhibits.

[0372] "Information optimization means" refers to methods or functions that adjust the information provided to the user based on the results of emotion recognition, according to the user's emotional state, and present it in a more appropriate form.

[0373] "Virtual reality generation means" refers to technologies and functions that generate and display content in order to provide users with a visual and experiential virtual environment.

[0374] "Data analysis means" refers to technologies that collect and analyze user behavior data, extract useful information from that data, and provide feedback.

[0375] A "feedback generation method" is a function that takes into account the user's behavioral history and emotional state to generate and provide personalized improvement suggestions and advice to the user.

[0376] This invention is a system for supporting dementia prevention that recognizes the user's emotional state based on input information and provides information and virtual reality experiences tailored to that state.

[0377] Users input questions and requests into the system via voice input devices or text input interfaces. For example, they might input a question such as, "What kind of exercise is helpful in preventing dementia?" The terminal sends the input voice or text data to the server, which is then prepared for analysis.

[0378] The server uses natural language processing tools to analyze user input and identify the underlying intent. Specifically, it uses software libraries such as spaCy and NLTK to analyze the input text. Based on the detected intent, the server retrieves the necessary information from the database.

[0379] Subsequently, the emotion engine analyzes the user's input and past behavior history to recognize the user's emotional state. This analysis utilizes voice tone analysis and logs of previous interactions, among other things. The user's emotions are then evaluated using tools such as IBM Watson Tone Analyzer.

[0380] The server optimizes and customizes the information provided to the user based on the emotion recognition results. For example, if the user is feeling stressed, it can suggest exercise methods that help with relaxation. The optimized information is sent to the terminal and presented to the user.

[0381] Furthermore, if the user chooses a virtual reality experience, the device will provide an appropriate experience through VR equipment (e.g., HTC Vive, Oculus Rift). Depending on the user's emotional state, it will visualize calming natural scenery in the virtual environment to promote a relaxation experience.

[0382] As a concrete example, if you input the prompt sentence "Please tell me a good way to fall asleep when I'm feeling anxious" into the AI ​​model, the system will generate a response that explains appropriate relaxation techniques.

[0383] The device records user behavior data during the virtual experience and sends it back to the server. The server aggregates this data and, through further analysis by an emotion engine, provides the user with personalized feedback and improvement suggestions.

[0384] The system based on this invention aims to enhance the effectiveness of dementia prevention through the provision of appropriate information and experiences tailored to the user's emotional state.

[0385] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0386] Step 1:

[0387] Users input questions and requests into the system via voice input devices or text input interfaces. This input is received by the terminal as audio data captured by the microphone in the case of voice input, or as text data in the case of text input. Specifically, this involves converting the voice to text using speech-to-text conversion software. The input data might be something like, "What kind of exercise is good for preventing dementia?", and this text data is sent to the server as output.

[0388] Step 2:

[0389] The terminal packets the text data received from the user and sends it to the server. The input is text data, and encryption technology is used to securely transmit it to the server. The output is encrypted text data received on the server side. Specific operations include data transfer via the HTTPS protocol.

[0390] Step 3:

[0391] The server performs natural language processing on the received text data to analyze the topic the user is interested in. The input is decrypted text data. Specifically, it uses spaCy or NLTK to tokenize the data, recognize parts of speech, and identify the user's intent. As output, a data object representing the user's intent is generated.

[0392] Step 4:

[0393] The server retrieves relevant information from the database based on the analyzed user intent. The input is a data object representing the user's intent. Database queries are used to retrieve corresponding health information and exercise methods. The output is information data tailored to the user's intent. Specifically, SQL queries are executed.

[0394] Step 5:

[0395] The emotion engine analyzes user input, voice tone, and past behavioral history to recognize the user's emotional state. Input consists of user text data and behavioral history sent from the server. This data is analyzed using an AI algorithm, and an evaluation result indicating the emotional state is generated as output. Specific operations include frequency analysis of voice tone and text sentiment analysis.

[0396] Step 6:

[0397] The server customizes the information it provides based on the emotional evaluation results generated by the emotion engine. Inputs include the emotional state evaluation results and previously acquired informational data. Filtering and prioritization are performed to best suit the user. Outputs are customized informational data. Specific operations include ranking and selecting information.

[0398] Step 7:

[0399] The terminal presents customized information sent from the server to the user. The input is customized information data. The terminal outputs this data using a user interface to display it visually or audibly. Specific operations include rendering the information to a graphical user interface.

[0400] Step 8:

[0401] If the user chooses a virtual reality experience, the device delivers this experience through VR equipment. Inputs are data indicating the virtual reality content and the user's state. The content is loaded into the hardware device, and feedback is provided to the user. Outputs are the visual and auditory virtual reality experience. Specific actions include launching the VR simulator and adjusting the scene based on the user's emotional state.

[0402] Step 9:

[0403] During the virtual experience, the device continuously records user behavior data and sends it to the server. The input is the user's movement data during the VR experience, including eye tracking and reaction time. This transmitted data is used in the feedback generation process. The output is the accumulated behavior data.

[0404] Step 10:

[0405] The server generates feedback and improvement suggestions based on behavioral data. The input is the result of emotion and behavior analysis. The feedback generation mechanism creates personalized recommendations based on the collected data. The output is the provision of customized feedback to the user. Specific actions include generating and personalizing feedback templates.

[0406] (Application Example 2)

[0407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0408] In modern society, there is a demand for personalized information based on individual emotional states. In particular, there is a need for systems that reduce user stress and anxiety and provide a more comfortable experience. However, current systems struggle to quickly and accurately analyze a user's emotional state and provide corresponding information. This invention provides a system to solve these problems.

[0409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0410] In this invention, the server includes: natural language processing means for receiving information in the form of voice or text and analyzing the information to understand the user's intent; information providing means for obtaining information from information sources based on the analyzed intent and generating an appropriate response; virtual reality generating means for displaying selected virtual content in order to provide an optimized virtual reality environment based on the user's emotional state; and information analysis means for recording the user's behavioral information, analyzing the behavioral information, and providing feedback. This makes it possible to provide more appropriate and effective information and experiences that take into account the user's emotional state.

[0411] "Natural language processing means" refers to technologies that analyze spoken or written information and have the function of understanding the user's intent and meaning.

[0412] An "information provision tool" is a technology that has the function of obtaining necessary information from information sources based on analyzed intent and generating an appropriate response for the user.

[0413] "Virtual reality generation means" refers to a technology that displays selected content to provide an optimized virtual reality environment, taking into account the user's emotional state.

[0414] "Information analysis means" refers to technology that has the function of recording user behavior information and providing feedback by analyzing it.

[0415] "Emotional state" refers to the user's psychological and emotional condition, and is an element that enables the provision of appropriate information and experiences based on that state.

[0416] "Virtual content" refers to information and visual elements displayed within a virtual reality environment, which are selected according to the user's emotions and intentions.

[0417] This invention comprises a terminal with a user interface, a server for data processing, a natural language processing function for analyzing voice and text information, an emotion engine for recognizing the user's emotional state, and components for generating a virtual reality experience.

[0418] A terminal is a device that receives audio and visual information using smart glasses or a head-mounted display. The terminal provides an interface for receiving user input and transmitting it to a server.

[0419] The server runs Python-based programs. It uses Python natural language processing libraries (e.g., spaCy, NLTK) to analyze received audio or text information. The librosa library is used for speech tone analysis. This allows not only to understand the user's intent but also to identify their emotional state.

[0420] The emotion engine recognizes the user's emotions using collected data. Based on this, the server selects response information appropriate to the user's state and generates an optimal virtual reality environment. The virtual reality experience is built using Unity or Unreal Engine, providing the user with a personalized shopping or relaxation experience.

[0421] As a concrete example, when a user is feeling stressed, the system will display music and videos designed to promote relaxation based on their emotional state. Information on products that are effective in relieving stress will also be presented within the virtual environment.

[0422] An example of a prompt using a generative AI model is, "Generate a list of recommended products to offer to a user who is feeling stressed and looking for relaxation items." Based on this, the AI ​​can generate content tailored to the user's needs and provide a personalized virtual reality experience.

[0423] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0424] Step 1:

[0425] The user inputs information via voice or text through a device. The device converts this input into text data and sends it to the server. The input data includes the user's questions and requests.

[0426] Step 2:

[0427] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK) to understand the user's intent. Here, the input intent is recognized as a command, and indicators are generated to retrieve relevant information from the database. The analysis results are output as intent information.

[0428] Step 3:

[0429] The server uses a speech tone analysis library (e.g., librosa) to estimate the user's emotional state from the input speech data. The analysis generates emotional state data, which can then be used to determine, for example, whether the user is experiencing stress.

[0430] Step 4:

[0431] Based on emotional state data, the server uses virtual reality generation methods to create an optimized virtual environment. The content of the virtual environment is built using Unity or Unreal Engine and prepared as visual and auditory content tailored to the user's emotions. The prepared virtual content is then generated.

[0432] Step 5:

[0433] Based on acquired intent information and emotional state data, the server generates user-specific response information using information delivery tools. This includes product information and experiential content designed to alleviate stress. The response information is output in a customized format for display within the virtual reality environment.

[0434] Step 6:

[0435] When a user takes an action within the virtual reality experience, the device records that action data and sends it to a server. This allows behavioral patterns to be collected as data. This behavioral data is then used to generate feedback.

[0436] Step 7:

[0437] The server analyzes the collected behavioral data using information analysis tools and generates feedback for the user. This feedback is customized to help improve healthy habits and emotional state, and is intended to be useful for the next virtual reality experience. The generated feedback is saved for future use.

[0438] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0439] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0440] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0441] [Third Embodiment]

[0442] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0443] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0444] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0445] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0446] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0447] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0448] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0449] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0450] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0451] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0452] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0453] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0454] This invention relates to a support system that enables users to intuitively understand and practice dementia prevention behaviors. The main components consist of a terminal operated by the user, a server that processes data, and a system that generates a virtual reality experience.

[0455] First, the user inputs a question using their device via voice or text. For example, they might ask, "What exercises are effective for preventing dementia?" The device then sends this input to the server.

[0456] The server performs natural language processing to parse the input it receives. This is the process of understanding the user's question and identifying their intent and requests. The server then refers to databases and external knowledge bases to generate answers to the user's question. For example, it might provide information such as, "Walking and swimming are effective."

[0457] The generated response is sent from the server to the terminal, which then displays it to the user. The user can then decide on further actions based on this information.

[0458] Next, if a user wishes to experience a virtual reality (VR) activity related to a specific dementia prevention behavior, the request is received from the device. The device receives instructions from the server and generates VR content based on the selected experience. This content features 360-degree video and sound effects, allowing the user to fully enjoy the atmosphere of walking.

[0459] During the experience, the device collects user behavior data. This data is sent to the server as the user's usage history. The server analyzes the collected data and provides feedback with suggestions for improvement to help the user more effectively prevent dementia.

[0460] This invention provides an innovative means to help users understand dementia prevention behaviors more concretely and incorporate them into their daily lives. Furthermore, it supports the realization of healthy lifestyle habits based on continuous data analysis.

[0461] The following describes the processing flow.

[0462] Step 1:

[0463] The user uses the device to input questions or requests via voice or text. The device receives this data and prepares to send it to the server.

[0464] Step 2:

[0465] The terminal sends the entered data to the server. The server receives the data and begins preparing to analyze its contents.

[0466] Step 3:

[0467] The server uses natural language processing to analyze the input data. Through this analysis, it clearly understands the intent of what the user wants to ask and identifies the necessary information.

[0468] Step 4:

[0469] Based on the identified intent, the server searches for relevant information from databases and external knowledge bases and generates an appropriate response for the user. This response includes specific suggestions and information regarding the user's question.

[0470] Step 5:

[0471] The server sends the generated response data to the terminal. The terminal receives this data and displays or provides audio guidance in a format that is easy for the user to understand.

[0472] Step 6:

[0473] Based on the information provided, users can request a virtual reality experience. This request is sent to the server via their device.

[0474] Step 7:

[0475] The device generates relevant virtual reality content based on the user's chosen virtual reality experience. It creates 360-degree video and audio to provide an immersive experience for the user.

[0476] Step 8:

[0477] Users experience virtual reality on their devices and virtually engage in relevant dementia prevention behaviors. During this time, the device records data on usage and behavior.

[0478] Step 9:

[0479] The device sends the collected behavioral data to the server. The server receives this data and analyzes the user's behavior patterns and usage frequency.

[0480] Step 10:

[0481] Based on the analyzed data, the server generates feedback suggesting areas for improvement and further healthy behaviors for the user. This feedback is then delivered to the user via their device.

[0482] (Example 1)

[0483] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0484] In an aging society, a key challenge in dementia prevention is ensuring that users understand how to effectively implement preventive activities and sustain their effects. This challenge goes beyond mere information provision; interactive support is essential to encourage users to take action and promote sustainable health maintenance. In particular, it is crucial that users not only receive information but also perceive it through real-life experiences and connect it to their daily lives. However, conventional systems are insufficient in this regard and are not intuitive for users, which is why they have not been widely adopted.

[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0486] In this invention, the server includes a natural language understanding means that receives information in the form of voice or text and analyzes that information to understand the user's intent; an information response means that retrieves information based on the analyzed intent from a storage device and generates an appropriate response; and a virtual environment generation means that displays selected content in order to provide the user with a virtual environment experience. As a result, the user not only receives information about dementia prevention, but also deepens their understanding through real-world experiences based on that information, enabling them to effectively and sustainably practice preventive activities.

[0487] "Natural language understanding methods" are techniques for analyzing information in the form of speech or text to understand the user's intentions and requests.

[0488] An "information response means" is a method for obtaining relevant information from a storage device based on an analyzed intent and generating an appropriate response.

[0489] A "virtual environment generation method" is a system that provides a virtual reality experience to a user based on their selected content.

[0490] "Information analysis methods" refer to techniques for recording user behavior information, analyzing that information, and providing feedback.

[0491] The "response generation means" is a function that creates and provides suggestions for improving healthy habits based on the user's behavioral information.

[0492] This invention is a comprehensive support system for users to understand and practice dementia prevention behaviors. The system primarily includes a user-operated terminal, a data processing server, and means for generating virtual reality experiences.

[0493] The device receives voice or text input from the user. This input is made through a dedicated application, allowing the user to enter prompts such as, "Please tell me about exercises that are effective for preventing dementia."

[0494] The server uses natural language understanding (NLP) tools to analyze the input it receives. This process utilizes generative AI models such as BERT and GPT to determine the intent of the user's question and generate appropriate information. Specifically, it accesses databases and external knowledge bases to extract and generate information such as "Walking and swimming are effective."

[0495] This information is sent from the server to the terminal, which then displays it to the user. The terminal's display is used for this display, enabling immediate visual feedback.

[0496] Furthermore, if a user wishes to experience virtual reality, the device sends a request to the server. Based on the user's selection, the server generates a virtual environment utilizing 360-degree video and sound, allowing the user to experience real-world activities. For example, a walking course within the virtual reality environment is provided, which the user can explore.

[0497] During the trial, user behavior data is recorded on the device and sent to the server. The server analyzes this data and generates feedback to help users improve their healthy habits, offering suggestions for further improvement.

[0498] An example of a prompt is, "Please give specific suggestions for a healthy diet," and the response to this input would be specific information such as, "A diet rich in fish is recommended."

[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0500] Step 1:

[0501] The user enters questions via voice or text using a terminal. This involves opening a dedicated application and entering a prompt. For example, "Please tell me about exercises that are effective for preventing dementia" might be entered. This input data is formatted and sent to the server. The output here is structured data sent to the server.

[0502] Step 2:

[0503] The server analyzes the received input data using natural language understanding (NLP) tools. Specifically, it uses a generative AI model to understand the intent of the input prompt. This analysis identifies the user's question and its purpose. The output after analysis provides specific information about the user's intent.

[0504] Step 3:

[0505] Based on the analyzed intent, the server retrieves relevant information from databases and external knowledge bases. This process generates specific responses, such as "Walking and swimming are effective." The output obtained using the information response means is appropriate answer information to the user's question.

[0506] Step 4:

[0507] The generated response information is sent from the server to the terminal. The terminal receives this information and displays it so that the user can visually confirm it. Specifically, the information is presented on the display screen as text or audio. The output of this step appears in a state where the user can see or hear the information.

[0508] Step 5:

[0509] When a user wishes to experience virtual reality, a request is sent from the device to the server. The request includes the desired virtual reality experience and detailed settings. The output in response to this input is the server's preparation for the virtual environment experience.

[0510] Step 6:

[0511] The server generates a virtual reality experience based on the requested content. This involves generating 360-degree video and audio, preparing to provide the user with an immersive experience. The output is content data ready for playback on the device.

[0512] Step 7:

[0513] The terminal provides the user with generated virtual reality content. Here, the virtual environment that the user experiences is recreated. During the experience, user behavior data is recorded. This record includes the user's eye movements and selected options. The output of this step is the behavior data sent to the server.

[0514] Step 8:

[0515] The server analyzes collected behavioral data and generates feedback for improving health habits. This analysis includes evaluating trends and effects based on the behavioral data. The feedback includes specific improvement suggestions and next steps, which are provided to the user. The output of this step is improvement suggestions presented to the user.

[0516] (Application Example 1)

[0517] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0518] Activities aimed at preventing dementia and maintaining health among the elderly face challenges in terms of being difficult to understand and implement. As a result, effective preventive actions are not taken, and continuous health management is difficult. Furthermore, existing health services do not adequately adapt to individual needs and cannot provide effective approaches for each person.

[0519] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0520] In this invention, the server includes natural language processing means for receiving voice or text input data and analyzing the input data to understand the user's intent; information providing means for obtaining information based on the analyzed intent from a storage medium and generating an appropriate response; and virtual reality generating means for displaying selected content to provide the user with a virtual reality experience. This enables the user to intuitively understand and practice health maintenance behaviors.

[0521] "Natural language processing means" are technical means for analyzing input data in the form of speech or text to understand the user's intent.

[0522] "Information provision means" refers to a technical means that, based on the analyzed user's intent, retrieves appropriate information from a storage medium and generates a response.

[0523] A "virtual reality generation means" is a technical means for displaying selected virtual reality content to a user and providing them with an experience.

[0524] "Data analysis means" refers to technical means that record and analyze user behavior data to provide feedback.

[0525] A "visual information device" is a device used to allow users to experience and facilitate health-related activities in real time.

[0526] "Interface generation means" refers to technical means for promoting health-related activities and providing operability to users through a visual information device.

[0527] The system for carrying out this invention consists mainly of an information terminal operated by the user, a server for data processing, and a visual information device for generating virtual reality. The server has the capability to process user questions in natural language, either by voice or text. This can be achieved, for example, using a Python-based natural language processing library (e.g., spaCy). User input is sent to the server via the information terminal, where the server analyzes the input and extracts the user's intent. Based on the analysis results, it retrieves appropriate information from a database as a storage medium and sends a response generated by an AI model to the terminal. In this case, the use of a generative AI model can be considered.

[0528] Furthermore, the virtual reality generation system utilizes a virtual reality simulation engine such as Unity to display user-selected content on a visual information device in real time. This allows users to easily experience health-related activities. The visual information device is envisioned to be smart glasses or similar devices and functions as an interface generation means to facilitate user behavior. The system also records user behavior data as accumulated data and provides feedback based on its patterns. This data analysis aims to present users with concrete action plans to improve their lifestyle habits.

[0529] As a concrete example, when a user enters a specific area in a supermarket, a visual information device displays health food information for that area. Based on a prompt such as "Please tell me about exercises that are effective in preventing dementia," relevant exercise information and walking routes are displayed on the user's glasses to support them in exercising.

[0530] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0531] Step 1:

[0532] The user inputs voice or text data via an information terminal. The input data is sent to a server to identify the information the user is requesting. Voice data, as input, is converted to text using speech recognition software on the terminal.

[0533] Step 2:

[0534] The server analyzes the received input data using natural language processing techniques. Specifically, it uses a generative AI model to understand the user's intent through syntactic and semantic analysis of the input data. The analyzed intent is output and passed on to the next processing step.

[0535] Step 3:

[0536] Based on the analyzed intent, the server retrieves appropriate information from the database on the storage medium. For example, if the user wants to know about exercises related to dementia prevention, information such as walking and swimming will be retrieved from the database. This search result is output as the appropriate response.

[0537] Step 4:

[0538] The server sends an appropriate response generated through a generative AI model to the information terminal. The terminal then displays the response to the user in natural language. Based on this information, the user can choose actions to prevent dementia.

[0539] Step 5:

[0540] When a user requests a virtual reality experience, a corresponding request is sent from the terminal to the server. The server generates appropriate virtual reality content via a virtual reality generation system. The generated content is sent to a visual information device such as smart glasses and displayed to the user in real time.

[0541] Step 6:

[0542] A visual information device collects behavioral data during the user's activities. This data is transmitted to a server, where data analysis tools analyze the user's habits and patterns, and process it into feedback for health management. The analysis results are used to suggest future activities for the user.

[0543] Step 7:

[0544] The server generates feedback for the user based on collected behavioral data. Insights gained from the data are used to generate specific suggestions for more effective health maintenance methods, which are then delivered to the user via their device.

[0545] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0546] This invention relates to a system that incorporates an emotion engine to help users understand information related to dementia prevention and support their actions. The system consists of a terminal that processes user input, a server that handles data processing, an emotion engine that performs emotion recognition, and components that generate a virtual reality experience.

[0547] Users input questions and requests via voice or text through their device. For example, they might ask, "What kind of exercise is good for preventing dementia?" The device then prepares to send this input to the server.

[0548] The server performs natural language processing on the received input data to analyze the user's intent. Based on this analysis, the server retrieves necessary information from the database. Subsequently, the emotion engine recognizes the user's emotional state based on the user's input, voice tone, past behavior patterns, and other factors.

[0549] Based on the emotions recognized by the emotion engine, the server customizes the information it provides. For example, for a user experiencing stress, it can suggest exercises that promote relaxation. The optimized information is then sent to the device and presented to the user.

[0550] Furthermore, if a user chooses a virtual reality experience to practice healthy habits, the content of that experience will also be adjusted based on the results of the emotion recognition. For example, a user experiencing anxiety will be provided with a virtual environment featuring calming scenery and music.

[0551] During the virtual reality experience, the device records other behavioral data. This data is sent to a server, where an emotion engine performs further cumulative analysis. The server uses the results of this analysis to provide the user with suggestions for improvement and feedback. The feedback is presented as personalized advice based on the user's emotional state and behavioral history.

[0552] Thus, according to the present invention, users can not only receive information but also gain an experience that promotes more effective dementia prevention behaviors tailored to their individual emotional state. This system aims to improve the user's quality of life and effectively delay or prevent the onset of dementia.

[0553] The following describes the processing flow.

[0554] Step 1:

[0555] Users input questions and requests via voice or text through the device. For example, they might ask, "What exercises are good for preventing dementia?" The device collects this input data.

[0556] Step 2:

[0557] The terminal sends the collected input data to the server. The server receives this data and uses natural language processing technology to analyze the user's intent and the content of the question.

[0558] Step 3:

[0559] Based on the analysis results, the server searches for relevant information from databases and external sources and generates an appropriate response. This information includes specific answers to the user's questions.

[0560] Step 4:

[0561] The server simultaneously uses an emotion engine to recognize the user's emotional state from input data and past behavioral data. For example, it can determine if the user is stressed based on their tone of voice and word choice.

[0562] Step 5:

[0563] Based on the user's emotional state, the server adjusts the information and suggestions it generates. For example, a user who needs to relax might be suggested gentle exercise or simple workouts.

[0564] Step 6:

[0565] The adjusted information is sent from the server to the terminal, which then guides the user through the information on the display or via voice. The user then uses this information to decide on specific actions.

[0566] Step 7:

[0567] If a user wishes to experience virtual reality, they select this request on their device, and the virtual reality experience begins. Based on the results of the emotion engine's recognition, the server customizes the content of the virtual reality experience, for example, by providing relaxing images and music.

[0568] Step 8:

[0569] The device provides users with a virtual reality experience and records behavioral data collected during that time. This includes usage details such as the duration of the experience and the content selected.

[0570] Step 9:

[0571] The device sends recorded behavioral data to the server. The server analyzes this data to understand the user's behavioral trends and emotions, and prepares feedback for the next time.

[0572] Step 10:

[0573] Based on the analysis results, the server generates new suggestions and feedback for the user. This feedback is customized based on the user's emotional state and behavioral history and is delivered to the user's device.

[0574] (Example 2)

[0575] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0576] Conventional dementia prevention support systems lack mechanisms to efficiently recognize users' emotional states and individual behavioral patterns, and to provide appropriate responses and feedback. Furthermore, they are unable to go beyond simply providing information and deliver content optimized according to the user's emotional state.

[0577] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0578] In this invention, the server includes natural language processing means that receive input information in the form of voice or text and analyze the input information to understand the user's purpose; emotion recognition means that analyze the input content and past behavioral patterns in order to recognize the user's emotional state; and data analysis means that record the user's behavioral data and analyze the behavioral data to provide feedback. This makes it possible to customize information and feedback according to the user's emotional state.

[0579] "Voice or text input information" refers to audio data spoken by the user to the system, or text data entered using a keyboard or other means.

[0580] "Natural language processing means" refers to technologies that analyze input language data to understand its context and meaning, and functions to identify the requests and questions that the user intends to ask the system.

[0581] "Information provision means" refers to methods and functions for collecting appropriate information based on the analyzed user's objectives and communicating that information to the user.

[0582] "Emotion recognition methods" are technologies that analyze user input data and behavioral history to identify the emotions and psychological states that a user exhibits.

[0583] "Information optimization means" refers to methods or functions that adjust the information provided to the user based on the results of emotion recognition, according to the user's emotional state, and present it in a more appropriate form.

[0584] "Virtual reality generation means" refers to technologies and functions that generate and display content in order to provide users with a visual and experiential virtual environment.

[0585] "Data analysis means" refers to technologies that collect and analyze user behavior data, extract useful information from that data, and provide feedback.

[0586] A "feedback generation method" is a function that takes into account the user's behavioral history and emotional state to generate and provide personalized improvement suggestions and advice to the user.

[0587] This invention is a system for supporting dementia prevention that recognizes the user's emotional state based on input information and provides information and virtual reality experiences tailored to that state.

[0588] Users input questions and requests into the system via voice input devices or text input interfaces. For example, they might input a question such as, "What kind of exercise is helpful in preventing dementia?" The terminal sends the input voice or text data to the server, which is then prepared for analysis.

[0589] The server uses natural language processing tools to analyze user input and identify the underlying intent. Specifically, it uses software libraries such as spaCy and NLTK to analyze the input text. Based on the detected intent, the server retrieves the necessary information from the database.

[0590] Subsequently, the emotion engine analyzes the user's input and past behavior history to recognize the user's emotional state. This analysis utilizes voice tone analysis and logs of previous interactions, among other things. The user's emotions are then evaluated using tools such as IBM Watson Tone Analyzer.

[0591] The server optimizes and customizes the information provided to the user based on the emotion recognition results. For example, if the user is feeling stressed, it can suggest exercise methods that help with relaxation. The optimized information is sent to the terminal and presented to the user.

[0592] Furthermore, if the user chooses a virtual reality experience, the device will provide an appropriate experience through VR equipment (e.g., HTC Vive, Oculus Rift). Depending on the user's emotional state, it will visualize calming natural scenery in the virtual environment to promote a relaxation experience.

[0593] As a concrete example, if you input the prompt sentence "Please tell me a good way to fall asleep when I'm feeling anxious" into the AI ​​model, the system will generate a response that explains appropriate relaxation techniques.

[0594] The device records user behavior data during the virtual experience and sends it back to the server. The server aggregates this data and, through further analysis by an emotion engine, provides the user with personalized feedback and improvement suggestions.

[0595] The system based on this invention aims to enhance the effectiveness of dementia prevention through the provision of appropriate information and experiences tailored to the user's emotional state.

[0596] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0597] Step 1:

[0598] Users input questions and requests into the system via voice input devices or text input interfaces. This input is received by the terminal as audio data captured by the microphone in the case of voice input, or as text data in the case of text input. Specifically, this involves converting the voice to text using speech-to-text conversion software. The input data might be something like, "What kind of exercise is good for preventing dementia?", and this text data is sent to the server as output.

[0599] Step 2:

[0600] The terminal packets the text data received from the user and sends it to the server. The input is text data, and encryption technology is used to securely transmit it to the server. The output is encrypted text data received on the server side. Specific operations include data transfer via the HTTPS protocol.

[0601] Step 3:

[0602] The server performs natural language processing on the received text data to analyze the topic the user is interested in. The input is decrypted text data. Specifically, it uses spaCy or NLTK to tokenize the data, recognize parts of speech, and identify the user's intent. As output, a data object representing the user's intent is generated.

[0603] Step 4:

[0604] The server retrieves relevant information from the database based on the analyzed user intent. The input is a data object representing the user's intent. Database queries are used to retrieve corresponding health information and exercise methods. The output is information data tailored to the user's intent. Specifically, SQL queries are executed.

[0605] Step 5:

[0606] The emotion engine analyzes user input, voice tone, and past behavioral history to recognize the user's emotional state. Input consists of user text data and behavioral history sent from the server. This data is analyzed using an AI algorithm, and an evaluation result indicating the emotional state is generated as output. Specific operations include frequency analysis of voice tone and text sentiment analysis.

[0607] Step 6:

[0608] The server customizes the information it provides based on the emotional evaluation results generated by the emotion engine. Inputs include the emotional state evaluation results and previously acquired informational data. Filtering and prioritization are performed to best suit the user. Outputs are customized informational data. Specific operations include ranking and selecting information.

[0609] Step 7:

[0610] The terminal presents customized information sent from the server to the user. The input is customized information data. The terminal outputs this data using a user interface to display it visually or audibly. Specific operations include rendering the information to a graphical user interface.

[0611] Step 8:

[0612] If the user chooses a virtual reality experience, the device delivers this experience through VR equipment. Inputs are data indicating the virtual reality content and the user's state. The content is loaded into the hardware device, and feedback is provided to the user. Outputs are the visual and auditory virtual reality experience. Specific actions include launching the VR simulator and adjusting the scene based on the user's emotional state.

[0613] Step 9:

[0614] During the virtual experience, the device continuously records user behavior data and sends it to the server. The input is the user's movement data during the VR experience, including eye tracking and reaction time. This transmitted data is used in the feedback generation process. The output is the accumulated behavior data.

[0615] Step 10:

[0616] The server generates feedback and improvement suggestions based on behavioral data. The input is the result of emotion and behavior analysis. The feedback generation mechanism creates personalized recommendations based on the collected data. The output is the provision of customized feedback to the user. Specific actions include generating and personalizing feedback templates.

[0617] (Application Example 2)

[0618] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0619] In modern society, there is a demand for personalized information based on individual emotional states. In particular, there is a need for systems that reduce user stress and anxiety and provide a more comfortable experience. However, current systems struggle to quickly and accurately analyze a user's emotional state and provide corresponding information. This invention provides a system to solve these problems.

[0620] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0621] In this invention, the server includes: natural language processing means for receiving information in the form of voice or text and analyzing the information to understand the user's intent; information providing means for obtaining information from information sources based on the analyzed intent and generating an appropriate response; virtual reality generating means for displaying selected virtual content in order to provide an optimized virtual reality environment based on the user's emotional state; and information analysis means for recording the user's behavioral information, analyzing the behavioral information, and providing feedback. This makes it possible to provide more appropriate and effective information and experiences that take into account the user's emotional state.

[0622] "Natural language processing means" refers to technologies that analyze spoken or written information and have the function of understanding the user's intent and meaning.

[0623] An "information provision tool" is a technology that has the function of obtaining necessary information from information sources based on analyzed intent and generating an appropriate response for the user.

[0624] "Virtual reality generation means" refers to a technology that displays selected content to provide an optimized virtual reality environment, taking into account the user's emotional state.

[0625] "Information analysis means" refers to technology that has the function of recording user behavior information and providing feedback by analyzing it.

[0626] "Emotional state" refers to the user's psychological and emotional condition, and is an element that enables the provision of appropriate information and experiences based on that state.

[0627] "Virtual content" refers to information and visual elements displayed within a virtual reality environment, which are selected according to the user's emotions and intentions.

[0628] This invention comprises a terminal with a user interface, a server for data processing, a natural language processing function for analyzing voice and text information, an emotion engine for recognizing the user's emotional state, and components for generating a virtual reality experience.

[0629] A terminal is a device that receives audio and visual information using smart glasses or a head-mounted display. The terminal provides an interface for receiving user input and transmitting it to a server.

[0630] The server runs Python-based programs. It uses Python natural language processing libraries (e.g., spaCy, NLTK) to analyze received audio or text information. The librosa library is used for speech tone analysis. This allows not only to understand the user's intent but also to identify their emotional state.

[0631] The emotion engine recognizes the user's emotions using collected data. Based on this, the server selects response information appropriate to the user's state and generates an optimal virtual reality environment. The virtual reality experience is built using Unity or Unreal Engine, providing the user with a personalized shopping or relaxation experience.

[0632] As a concrete example, when a user is feeling stressed, the system will display music and videos designed to promote relaxation based on their emotional state. Information on products that are effective in relieving stress will also be presented within the virtual environment.

[0633] An example of a prompt using a generative AI model is, "Generate a list of recommended products to offer to a user who is feeling stressed and looking for relaxation items." Based on this, the AI ​​can generate content tailored to the user's needs and provide a personalized virtual reality experience.

[0634] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0635] Step 1:

[0636] The user inputs information via voice or text through a device. The device converts this input into text data and sends it to the server. The input data includes the user's questions and requests.

[0637] Step 2:

[0638] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK) to understand the user's intent. Here, the input intent is recognized as a command, and indicators are generated to retrieve relevant information from the database. The analysis results are output as intent information.

[0639] Step 3:

[0640] The server uses a speech tone analysis library (e.g., librosa) to estimate the user's emotional state from the input speech data. The analysis generates emotional state data, which can then be used to determine, for example, whether the user is experiencing stress.

[0641] Step 4:

[0642] Based on emotional state data, the server uses virtual reality generation methods to create an optimized virtual environment. The content of the virtual environment is built using Unity or Unreal Engine and prepared as visual and auditory content tailored to the user's emotions. The prepared virtual content is then generated.

[0643] Step 5:

[0644] Based on acquired intent information and emotional state data, the server generates user-specific response information using information delivery tools. This includes product information and experiential content designed to alleviate stress. The response information is output in a customized format for display within the virtual reality environment.

[0645] Step 6:

[0646] When a user takes an action within the virtual reality experience, the device records that action data and sends it to a server. This allows behavioral patterns to be collected as data. This behavioral data is then used to generate feedback.

[0647] Step 7:

[0648] The server analyzes the collected behavioral data using information analysis tools and generates feedback for the user. This feedback is customized to help improve healthy habits and emotional state, and is intended to be useful for the next virtual reality experience. The generated feedback is saved for future use.

[0649] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0650] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0651] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0652] [Fourth Embodiment]

[0653] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0654] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0655] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0656] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0657] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0658] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0659] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0660] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0661] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0662] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0663] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0664] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0665] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0666] This invention relates to a support system that enables users to intuitively understand and practice dementia prevention behaviors. The main components consist of a terminal operated by the user, a server that processes data, and a system that generates a virtual reality experience.

[0667] First, the user inputs a question using their device via voice or text. For example, they might ask, "What exercises are effective for preventing dementia?" The device then sends this input to the server.

[0668] The server performs natural language processing to parse the input it receives. This is the process of understanding the user's question and identifying their intent and requests. The server then refers to databases and external knowledge bases to generate answers to the user's question. For example, it might provide information such as, "Walking and swimming are effective."

[0669] The generated response is sent from the server to the terminal, which then displays it to the user. The user can then decide on further actions based on this information.

[0670] Next, if a user wishes to experience a virtual reality (VR) activity related to a specific dementia prevention behavior, the request is received from the device. The device receives instructions from the server and generates VR content based on the selected experience. This content features 360-degree video and sound effects, allowing the user to fully enjoy the atmosphere of walking.

[0671] During the experience, the device collects user behavior data. This data is sent to the server as the user's usage history. The server analyzes the collected data and provides feedback with suggestions for improvement to help the user more effectively prevent dementia.

[0672] This invention provides an innovative means to help users understand dementia prevention behaviors more concretely and incorporate them into their daily lives. Furthermore, it supports the realization of healthy lifestyle habits based on continuous data analysis.

[0673] The following describes the processing flow.

[0674] Step 1:

[0675] The user uses the device to input questions or requests via voice or text. The device receives this data and prepares to send it to the server.

[0676] Step 2:

[0677] The terminal sends the entered data to the server. The server receives the data and begins preparing to analyze its contents.

[0678] Step 3:

[0679] The server uses natural language processing to analyze the input data. Through this analysis, it clearly understands the intent of what the user wants to ask and identifies the necessary information.

[0680] Step 4:

[0681] Based on the identified intent, the server searches for relevant information from databases and external knowledge bases and generates an appropriate response for the user. This response includes specific suggestions and information regarding the user's question.

[0682] Step 5:

[0683] The server sends the generated response data to the terminal. The terminal receives this data and displays or provides audio guidance in a format that is easy for the user to understand.

[0684] Step 6:

[0685] Based on the information provided, users can request a virtual reality experience. This request is sent to the server via their device.

[0686] Step 7:

[0687] The device generates relevant virtual reality content based on the user's chosen virtual reality experience. It creates 360-degree video and audio to provide an immersive experience for the user.

[0688] Step 8:

[0689] Users experience virtual reality on their devices and virtually engage in relevant dementia prevention behaviors. During this time, the device records data on usage and behavior.

[0690] Step 9:

[0691] The device sends the collected behavioral data to the server. The server receives this data and analyzes the user's behavior patterns and usage frequency.

[0692] Step 10:

[0693] Based on the analyzed data, the server generates feedback suggesting areas for improvement and further healthy behaviors for the user. This feedback is then delivered to the user via their device.

[0694] (Example 1)

[0695] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0696] In an aging society, a key challenge in dementia prevention is ensuring that users understand how to effectively implement preventive activities and sustain their effects. This challenge goes beyond mere information provision; interactive support is essential to encourage users to take action and promote sustainable health maintenance. In particular, it is crucial that users not only receive information but also perceive it through real-life experiences and connect it to their daily lives. However, conventional systems are insufficient in this regard and are not intuitive for users, which is why they have not been widely adopted.

[0697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0698] In this invention, the server includes a natural language understanding means that receives information in the form of voice or text and analyzes that information to understand the user's intent; an information response means that retrieves information based on the analyzed intent from a storage device and generates an appropriate response; and a virtual environment generation means that displays selected content in order to provide the user with a virtual environment experience. As a result, the user not only receives information about dementia prevention, but also deepens their understanding through real-world experiences based on that information, enabling them to effectively and sustainably practice preventive activities.

[0699] "Natural language understanding methods" are techniques for analyzing information in the form of speech or text to understand the user's intentions and requests.

[0700] An "information response means" is a method for obtaining relevant information from a storage device based on an analyzed intent and generating an appropriate response.

[0701] A "virtual environment generation method" is a system that provides a virtual reality experience to a user based on their selected content.

[0702] "Information analysis methods" refer to techniques for recording user behavior information, analyzing that information, and providing feedback.

[0703] The "response generation means" is a function that creates and provides suggestions for improving healthy habits based on the user's behavioral information.

[0704] This invention is a comprehensive support system for users to understand and practice dementia prevention behaviors. The system primarily includes a user-operated terminal, a data processing server, and means for generating virtual reality experiences.

[0705] The device receives voice or text input from the user. This input is made through a dedicated application, allowing the user to enter prompts such as, "Please tell me about exercises that are effective for preventing dementia."

[0706] The server uses natural language understanding (NLP) tools to analyze the input it receives. This process utilizes generative AI models such as BERT and GPT to determine the intent of the user's question and generate appropriate information. Specifically, it accesses databases and external knowledge bases to extract and generate information such as "Walking and swimming are effective."

[0707] This information is sent from the server to the terminal, which then displays it to the user. The terminal's display is used for this display, enabling immediate visual feedback.

[0708] Furthermore, if a user wishes to experience virtual reality, the device sends a request to the server. Based on the user's selection, the server generates a virtual environment utilizing 360-degree video and sound, allowing the user to experience real-world activities. For example, a walking course within the virtual reality environment is provided, which the user can explore.

[0709] During the trial, user behavior data is recorded on the device and sent to the server. The server analyzes this data and generates feedback to help users improve their healthy habits, offering suggestions for further improvement.

[0710] An example of a prompt is, "Please give specific suggestions for a healthy diet," and the response to this input would be specific information such as, "A diet rich in fish is recommended."

[0711] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0712] Step 1:

[0713] The user enters questions via voice or text using a terminal. This involves opening a dedicated application and entering a prompt. For example, "Please tell me about exercises that are effective for preventing dementia" might be entered. This input data is formatted and sent to the server. The output here is structured data sent to the server.

[0714] Step 2:

[0715] The server analyzes the received input data using natural language understanding (NLP) tools. Specifically, it uses a generative AI model to understand the intent of the input prompt. This analysis identifies the user's question and its purpose. The output after analysis provides specific information about the user's intent.

[0716] Step 3:

[0717] Based on the analyzed intent, the server retrieves relevant information from databases and external knowledge bases. This process generates specific responses, such as "Walking and swimming are effective." The output obtained using the information response means is appropriate answer information to the user's question.

[0718] Step 4:

[0719] The generated response information is sent from the server to the terminal. The terminal receives this information and displays it so that the user can visually confirm it. Specifically, the information is presented on the display screen as text or audio. The output of this step appears in a state where the user can see or hear the information.

[0720] Step 5:

[0721] When a user wishes to experience virtual reality, a request is sent from the device to the server. The request includes the desired virtual reality experience and detailed settings. The output in response to this input is the server's preparation for the virtual environment experience.

[0722] Step 6:

[0723] The server generates a virtual reality experience based on the requested content. This involves generating 360-degree video and audio, preparing to provide the user with an immersive experience. The output is content data ready for playback on the device.

[0724] Step 7:

[0725] The terminal provides the user with generated virtual reality content. Here, the virtual environment that the user experiences is recreated. During the experience, user behavior data is recorded. This record includes the user's eye movements and selected options. The output of this step is the behavior data sent to the server.

[0726] Step 8:

[0727] The server analyzes collected behavioral data and generates feedback for improving health habits. This analysis includes evaluating trends and effects based on the behavioral data. The feedback includes specific improvement suggestions and next steps, which are provided to the user. The output of this step is improvement suggestions presented to the user.

[0728] (Application Example 1)

[0729] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0730] Activities aimed at preventing dementia and maintaining health among the elderly face challenges in terms of being difficult to understand and implement. As a result, effective preventive actions are not taken, and continuous health management is difficult. Furthermore, existing health services do not adequately adapt to individual needs and cannot provide effective approaches for each person.

[0731] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0732] In this invention, the server includes natural language processing means for receiving voice or text input data and analyzing the input data to understand the user's intent; information providing means for obtaining information based on the analyzed intent from a storage medium and generating an appropriate response; and virtual reality generating means for displaying selected content to provide the user with a virtual reality experience. This enables the user to intuitively understand and practice health maintenance behaviors.

[0733] "Natural language processing means" are technical means for analyzing input data in the form of speech or text to understand the user's intent.

[0734] "Information provision means" refers to a technical means that, based on the analyzed user's intent, retrieves appropriate information from a storage medium and generates a response.

[0735] A "virtual reality generation means" is a technical means for displaying selected virtual reality content to a user and providing them with an experience.

[0736] "Data analysis means" refers to technical means that record and analyze user behavior data to provide feedback.

[0737] A "visual information device" is a device used to allow users to experience and facilitate health-related activities in real time.

[0738] "Interface generation means" refers to technical means for promoting health-related activities and providing operability to users through a visual information device.

[0739] The system for carrying out this invention consists mainly of an information terminal operated by the user, a server for data processing, and a visual information device for generating virtual reality. The server has the capability to process user questions in natural language, either by voice or text. This can be achieved, for example, using a Python-based natural language processing library (e.g., spaCy). User input is sent to the server via the information terminal, where the server analyzes the input and extracts the user's intent. Based on the analysis results, it retrieves appropriate information from a database as a storage medium and sends a response generated by an AI model to the terminal. In this case, the use of a generative AI model can be considered.

[0740] Furthermore, the virtual reality generation system utilizes a virtual reality simulation engine such as Unity to display user-selected content on a visual information device in real time. This allows users to easily experience health-related activities. The visual information device is envisioned to be smart glasses or similar devices and functions as an interface generation means to facilitate user behavior. The system also records user behavior data as accumulated data and provides feedback based on its patterns. This data analysis aims to present users with concrete action plans to improve their lifestyle habits.

[0741] As a concrete example, when a user enters a specific area in a supermarket, a visual information device displays health food information for that area. Based on a prompt such as "Please tell me about exercises that are effective in preventing dementia," relevant exercise information and walking routes are displayed on the user's glasses to support them in exercising.

[0742] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0743] Step 1:

[0744] The user inputs voice or text data via an information terminal. The input data is sent to a server to identify the information the user is requesting. Voice data, as input, is converted to text using speech recognition software on the terminal.

[0745] Step 2:

[0746] The server analyzes the received input data using natural language processing techniques. Specifically, it uses a generative AI model to understand the user's intent through syntactic and semantic analysis of the input data. The analyzed intent is output and passed on to the next processing step.

[0747] Step 3:

[0748] Based on the analyzed intent, the server retrieves appropriate information from the database on the storage medium. For example, if the user wants to know about exercises related to dementia prevention, information such as walking and swimming will be retrieved from the database. This search result is output as the appropriate response.

[0749] Step 4:

[0750] The server sends an appropriate response generated through a generative AI model to the information terminal. The terminal then displays the response to the user in natural language. Based on this information, the user can choose actions to prevent dementia.

[0751] Step 5:

[0752] When a user requests a virtual reality experience, a corresponding request is sent from the terminal to the server. The server generates appropriate virtual reality content via a virtual reality generation system. The generated content is sent to a visual information device such as smart glasses and displayed to the user in real time.

[0753] Step 6:

[0754] A visual information device collects behavioral data during the user's activities. This data is transmitted to a server, where data analysis tools analyze the user's habits and patterns, and process it into feedback for health management. The analysis results are used to suggest future activities for the user.

[0755] Step 7:

[0756] The server generates feedback for the user based on collected behavioral data. Insights gained from the data are used to generate specific suggestions for more effective health maintenance methods, which are then delivered to the user via their device.

[0757] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0758] This invention relates to a system that incorporates an emotion engine to help users understand information related to dementia prevention and support their actions. The system consists of a terminal that processes user input, a server that handles data processing, an emotion engine that performs emotion recognition, and components that generate a virtual reality experience.

[0759] Users input questions and requests via voice or text through their device. For example, they might ask, "What kind of exercise is good for preventing dementia?" The device then prepares to send this input to the server.

[0760] The server performs natural language processing on the received input data to analyze the user's intent. Based on this analysis, the server retrieves necessary information from the database. Subsequently, the emotion engine recognizes the user's emotional state based on the user's input, voice tone, past behavior patterns, and other factors.

[0761] Based on the emotions recognized by the emotion engine, the server customizes the information it provides. For example, for a user experiencing stress, it can suggest exercises that promote relaxation. The optimized information is then sent to the device and presented to the user.

[0762] Furthermore, if a user chooses a virtual reality experience to practice healthy habits, the content of that experience will also be adjusted based on the results of the emotion recognition. For example, a user experiencing anxiety will be provided with a virtual environment featuring calming scenery and music.

[0763] During the virtual reality experience, the device records other behavioral data. This data is sent to a server, where an emotion engine performs further cumulative analysis. The server uses the results of this analysis to provide the user with suggestions for improvement and feedback. The feedback is presented as personalized advice based on the user's emotional state and behavioral history.

[0764] Thus, according to the present invention, users can not only receive information but also gain an experience that promotes more effective dementia prevention behaviors tailored to their individual emotional state. This system aims to improve the user's quality of life and effectively delay or prevent the onset of dementia.

[0765] The following describes the processing flow.

[0766] Step 1:

[0767] Users input questions and requests via voice or text through the device. For example, they might ask, "What exercises are good for preventing dementia?" The device collects this input data.

[0768] Step 2:

[0769] The terminal sends the collected input data to the server. The server receives this data and uses natural language processing technology to analyze the user's intent and the content of the question.

[0770] Step 3:

[0771] Based on the analysis results, the server searches for relevant information from databases and external sources and generates an appropriate response. This information includes specific answers to the user's questions.

[0772] Step 4:

[0773] The server simultaneously uses an emotion engine to recognize the user's emotional state from input data and past behavioral data. For example, it can determine if the user is stressed based on their tone of voice and word choice.

[0774] Step 5:

[0775] Based on the user's emotional state, the server adjusts the information and suggestions it generates. For example, a user who needs to relax might be suggested gentle exercise or simple workouts.

[0776] Step 6:

[0777] The adjusted information is sent from the server to the terminal, which then guides the user through the information on the display or via voice. The user then uses this information to decide on specific actions.

[0778] Step 7:

[0779] If a user wishes to experience virtual reality, they select this request on their device, and the virtual reality experience begins. Based on the results of the emotion engine's recognition, the server customizes the content of the virtual reality experience, for example, by providing relaxing images and music.

[0780] Step 8:

[0781] The device provides users with a virtual reality experience and records behavioral data collected during that time. This includes usage details such as the duration of the experience and the content selected.

[0782] Step 9:

[0783] The device sends recorded behavioral data to the server. The server analyzes this data to understand the user's behavioral trends and emotions, and prepares feedback for the next time.

[0784] Step 10:

[0785] Based on the analysis results, the server generates new suggestions and feedback for the user. This feedback is customized based on the user's emotional state and behavioral history and is delivered to the user's device.

[0786] (Example 2)

[0787] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0788] Conventional dementia prevention support systems lack mechanisms to efficiently recognize users' emotional states and individual behavioral patterns, and to provide appropriate responses and feedback. Furthermore, they are unable to go beyond simply providing information and deliver content optimized according to the user's emotional state.

[0789] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0790] In this invention, the server includes natural language processing means that receive input information in the form of voice or text and analyze the input information to understand the user's purpose; emotion recognition means that analyze the input content and past behavioral patterns in order to recognize the user's emotional state; and data analysis means that record the user's behavioral data and analyze the behavioral data to provide feedback. This makes it possible to customize information and feedback according to the user's emotional state.

[0791] "Voice or text input information" refers to audio data spoken by the user to the system, or text data entered using a keyboard or other means.

[0792] "Natural language processing means" refers to technologies that analyze input language data to understand its context and meaning, and functions to identify the requests and questions that the user intends to ask the system.

[0793] "Information provision means" refers to methods and functions for collecting appropriate information based on the analyzed user's objectives and communicating that information to the user.

[0794] "Emotion recognition methods" are technologies that analyze user input data and behavioral history to identify the emotions and psychological states that a user exhibits.

[0795] "Information optimization means" refers to methods or functions that adjust the information provided to the user based on the results of emotion recognition, according to the user's emotional state, and present it in a more appropriate form.

[0796] "Virtual reality generation means" refers to technologies and functions that generate and display content in order to provide users with a visual and experiential virtual environment.

[0797] "Data analysis means" refers to technologies that collect and analyze user behavior data, extract useful information from that data, and provide feedback.

[0798] A "feedback generation method" is a function that takes into account the user's behavioral history and emotional state to generate and provide personalized improvement suggestions and advice to the user.

[0799] This invention is a system for supporting dementia prevention that recognizes the user's emotional state based on input information and provides information and virtual reality experiences tailored to that state.

[0800] Users input questions and requests into the system via voice input devices or text input interfaces. For example, they might input a question such as, "What kind of exercise is helpful in preventing dementia?" The terminal sends the input voice or text data to the server, which is then prepared for analysis.

[0801] The server uses natural language processing tools to analyze user input and identify the underlying intent. Specifically, it uses software libraries such as spaCy and NLTK to analyze the input text. Based on the detected intent, the server retrieves the necessary information from the database.

[0802] Subsequently, the emotion engine analyzes the user's input and past behavior history to recognize the user's emotional state. This analysis utilizes voice tone analysis and logs of previous interactions, among other things. The user's emotions are then evaluated using tools such as IBM Watson Tone Analyzer.

[0803] The server optimizes and customizes the information provided to the user based on the emotion recognition results. For example, if the user is feeling stressed, it can suggest exercise methods that help with relaxation. The optimized information is sent to the terminal and presented to the user.

[0804] Furthermore, if the user chooses a virtual reality experience, the device will provide an appropriate experience through VR equipment (e.g., HTC Vive, Oculus Rift). Depending on the user's emotional state, it will visualize calming natural scenery in the virtual environment to promote a relaxation experience.

[0805] As a concrete example, if you input the prompt sentence "Please tell me a good way to fall asleep when I'm feeling anxious" into the AI ​​model, the system will generate a response that explains appropriate relaxation techniques.

[0806] The device records user behavior data during the virtual experience and sends it back to the server. The server aggregates this data and, through further analysis by an emotion engine, provides the user with personalized feedback and improvement suggestions.

[0807] The system based on this invention aims to enhance the effectiveness of dementia prevention through the provision of appropriate information and experiences tailored to the user's emotional state.

[0808] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0809] Step 1:

[0810] Users input questions and requests into the system via voice input devices or text input interfaces. This input is received by the terminal as audio data captured by the microphone in the case of voice input, or as text data in the case of text input. Specifically, this involves converting the voice to text using speech-to-text conversion software. The input data might be something like, "What kind of exercise is good for preventing dementia?", and this text data is sent to the server as output.

[0811] Step 2:

[0812] The terminal packets the text data received from the user and sends it to the server. The input is text data, and encryption technology is used to securely transmit it to the server. The output is encrypted text data received on the server side. Specific operations include data transfer via the HTTPS protocol.

[0813] Step 3:

[0814] The server performs natural language processing on the received text data to analyze the topic the user is interested in. The input is decrypted text data. Specifically, it uses spaCy or NLTK to tokenize the data, recognize parts of speech, and identify the user's intent. As output, a data object representing the user's intent is generated.

[0815] Step 4:

[0816] The server retrieves relevant information from the database based on the analyzed user intent. The input is a data object representing the user's intent. Database queries are used to retrieve corresponding health information and exercise methods. The output is information data tailored to the user's intent. Specifically, SQL queries are executed.

[0817] Step 5:

[0818] The emotion engine analyzes user input, voice tone, and past behavioral history to recognize the user's emotional state. Input consists of user text data and behavioral history sent from the server. This data is analyzed using an AI algorithm, and an evaluation result indicating the emotional state is generated as output. Specific operations include frequency analysis of voice tone and text sentiment analysis.

[0819] Step 6:

[0820] The server customizes the information it provides based on the emotional evaluation results generated by the emotion engine. Inputs include the emotional state evaluation results and previously acquired informational data. Filtering and prioritization are performed to best suit the user. Outputs are customized informational data. Specific operations include ranking and selecting information.

[0821] Step 7:

[0822] The terminal presents customized information sent from the server to the user. The input is customized information data. The terminal outputs this data using a user interface to display it visually or audibly. Specific operations include rendering the information to a graphical user interface.

[0823] Step 8:

[0824] If the user chooses a virtual reality experience, the device delivers this experience through VR equipment. Inputs are data indicating the virtual reality content and the user's state. The content is loaded into the hardware device, and feedback is provided to the user. Outputs are the visual and auditory virtual reality experience. Specific actions include launching the VR simulator and adjusting the scene based on the user's emotional state.

[0825] Step 9:

[0826] During the virtual experience, the device continuously records user behavior data and sends it to the server. The input is the user's movement data during the VR experience, including eye tracking and reaction time. This transmitted data is used in the feedback generation process. The output is the accumulated behavior data.

[0827] Step 10:

[0828] The server generates feedback and improvement suggestions based on behavioral data. The input is the result of emotion and behavior analysis. The feedback generation mechanism creates personalized recommendations based on the collected data. The output is the provision of customized feedback to the user. Specific actions include generating and personalizing feedback templates.

[0829] (Application Example 2)

[0830] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0831] In modern society, there is a demand for personalized information based on individual emotional states. In particular, there is a need for systems that reduce user stress and anxiety and provide a more comfortable experience. However, current systems struggle to quickly and accurately analyze a user's emotional state and provide corresponding information. This invention provides a system to solve these problems.

[0832] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0833] In this invention, the server includes: natural language processing means for receiving information in the form of voice or text and analyzing the information to understand the user's intent; information providing means for obtaining information from information sources based on the analyzed intent and generating an appropriate response; virtual reality generating means for displaying selected virtual content in order to provide an optimized virtual reality environment based on the user's emotional state; and information analysis means for recording the user's behavioral information, analyzing the behavioral information, and providing feedback. This makes it possible to provide more appropriate and effective information and experiences that take into account the user's emotional state.

[0834] "Natural language processing means" refers to technologies that analyze spoken or written information and have the function of understanding the user's intent and meaning.

[0835] An "information provision tool" is a technology that has the function of obtaining necessary information from information sources based on analyzed intent and generating an appropriate response for the user.

[0836] "Virtual reality generation means" refers to a technology that displays selected content to provide an optimized virtual reality environment, taking into account the user's emotional state.

[0837] "Information analysis means" refers to technology that has the function of recording user behavior information and providing feedback by analyzing it.

[0838] "Emotional state" refers to the user's psychological and emotional condition, and is an element that enables the provision of appropriate information and experiences based on that state.

[0839] "Virtual content" refers to information and visual elements displayed within a virtual reality environment, which are selected according to the user's emotions and intentions.

[0840] This invention comprises a terminal with a user interface, a server for data processing, a natural language processing function for analyzing voice and text information, an emotion engine for recognizing the user's emotional state, and components for generating a virtual reality experience.

[0841] A terminal is a device that receives audio and visual information using smart glasses or a head-mounted display. The terminal provides an interface for receiving user input and transmitting it to a server.

[0842] The server runs Python-based programs. It uses Python natural language processing libraries (e.g., spaCy, NLTK) to analyze received audio or text information. The librosa library is used for speech tone analysis. This allows not only to understand the user's intent but also to identify their emotional state.

[0843] The emotion engine recognizes the user's emotions using collected data. Based on this, the server selects response information appropriate to the user's state and generates an optimal virtual reality environment. The virtual reality experience is built using Unity or Unreal Engine, providing the user with a personalized shopping or relaxation experience.

[0844] As a concrete example, when a user is feeling stressed, the system will display music and videos designed to promote relaxation based on their emotional state. Information on products that are effective in relieving stress will also be presented within the virtual environment.

[0845] An example of a prompt using a generative AI model is, "Generate a list of recommended products to offer to a user who is feeling stressed and looking for relaxation items." Based on this, the AI ​​can generate content tailored to the user's needs and provide a personalized virtual reality experience.

[0846] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0847] Step 1:

[0848] The user inputs information via voice or text through a device. The device converts this input into text data and sends it to the server. The input data includes the user's questions and requests.

[0849] Step 2:

[0850] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK) to understand the user's intent. Here, the input intent is recognized as a command, and indicators are generated to retrieve relevant information from the database. The analysis results are output as intent information.

[0851] Step 3:

[0852] The server uses a speech tone analysis library (e.g., librosa) to estimate the user's emotional state from the input speech data. The analysis generates emotional state data, which can then be used to determine, for example, whether the user is experiencing stress.

[0853] Step 4:

[0854] Based on emotional state data, the server uses virtual reality generation methods to create an optimized virtual environment. The content of the virtual environment is built using Unity or Unreal Engine and prepared as visual and auditory content tailored to the user's emotions. The prepared virtual content is then generated.

[0855] Step 5:

[0856] Based on acquired intent information and emotional state data, the server generates user-specific response information using information delivery tools. This includes product information and experiential content designed to alleviate stress. The response information is output in a customized format for display within the virtual reality environment.

[0857] Step 6:

[0858] When a user takes an action within the virtual reality experience, the device records that action data and sends it to a server. This allows behavioral patterns to be collected as data. This behavioral data is then used to generate feedback.

[0859] Step 7:

[0860] The server analyzes the collected behavioral data using information analysis tools and generates feedback for the user. This feedback is customized to help improve healthy habits and emotional state, and is intended to be useful for the next virtual reality experience. The generated feedback is saved for future use.

[0861] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0862] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0863] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0864] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0865] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0866] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0867] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0868] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0869] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0870] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0871] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0872] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0873] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0874] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0875] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0876] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0877] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0878] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0879] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0880] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0881] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0882] The following is further disclosed regarding the embodiments described above.

[0883] (Claim 1)

[0884] A natural language processing means that receives voice or text input data, analyzes the input data to understand the user's intent,

[0885] Information provision means that retrieves information based on analyzed intent from a database and generates an appropriate response,

[0886] A virtual reality generation means for displaying selected content in order to provide a virtual reality experience to the user,

[0887] A data analysis means that records user behavior data, analyzes said behavior data, and provides feedback,

[0888] A system that includes this.

[0889] (Claim 2)

[0890] The system according to claim 1, which provides virtual reality content related to specific dementia prevention behaviors based on user input data.

[0891] (Claim 3)

[0892] The system according to claim 1, comprising a feedback generation means that uses user behavior data to provide suggestions for improving healthy habits.

[0893] "Example 1"

[0894] (Claim 1)

[0895] A natural language understanding means that receives information in the form of speech or text, analyzes that information, and understands the user's intent.

[0896] Information response means that acquires information based on the analyzed intent from a storage device and generates an appropriate response,

[0897] A means for generating a virtual environment that displays selected content in order to provide users with a virtual environment experience,

[0898] An information analysis means that records user behavior information, analyzes that behavior information, and provides feedback,

[0899] A means of generating and presenting directional and effect-driven video and audio based on the virtual environment experience desired by the user,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, which provides virtual environment content related to specific dementia prevention behaviors based on user input information.

[0903] (Claim 3)

[0904] The system according to claim 1, comprising a response generation means that uses user behavior information to make suggestions for improving healthy habits.

[0905] "Application Example 1"

[0906] (Claim 1)

[0907] A natural language processing means that receives voice or text input data, analyzes the input data to understand the user's intent,

[0908] Information provision means that acquires information based on the analyzed intent from a storage medium and generates an appropriate response,

[0909] A virtual reality generation means for displaying selected content in order to provide a virtual reality experience to the user,

[0910] A data analysis means that records user behavior data, analyzes said behavior data, and provides feedback,

[0911] An interface generation means that uses a visual information device to promote health-related activities to the user in real time,

[0912] A system that includes this.

[0913] (Claim 2)

[0914] The system according to claim 1, which provides virtual reality content related to specific health maintenance behaviors based on user input data.

[0915] (Claim 3)

[0916] The system according to claim 1, comprising a feedback generation means that uses user behavior data to provide suggestions for improving lifestyle habits.

[0917] "Example 2 of combining an emotion engine"

[0918] (Claim 1)

[0919] A natural language processing means that receives voice or text input information, analyzes the input information, and understands the user's purpose.

[0920] Information provision means that acquires data based on the analyzed purpose from a storage device and generates an appropriate response,

[0921] To recognize the user's emotional state, an emotion recognition means analyzes input content and past behavioral patterns,

[0922] Information optimization means that personalizes information based on emotion recognition results,

[0923] A virtual reality generation means for displaying selected content in order to provide a virtual reality experience to the user,

[0924] A data analysis means that records user behavior data, analyzes said behavior data, and provides feedback,

[0925] A feedback generation means that personalizes the feedback received by the user based on their behavioral history and emotional state,

[0926] A system that includes this.

[0927] (Claim 2)

[0928] The system according to claim 1, which provides virtual reality content related to specific cognitive function decline prevention behaviors based on user input information.

[0929] (Claim 3)

[0930] The system according to claim 1, comprising a feedback generation means that uses user behavior data to provide suggestions for improving healthy habits.

[0931] "Application example 2 of combining emotional engines"

[0932] (Claim 1)

[0933] A natural language processing means that receives audio or text information, analyzes the information, and understands the user's intent.

[0934] Information provision means that obtains information based on analyzed intent from information sources and generates an appropriate response,

[0935] A virtual reality generation means that displays selected virtual content in order to provide an optimized virtual reality environment based on the user's emotional state,

[0936] Information analysis means for recording user behavior information, analyzing said behavior information, and providing feedback,

[0937] A system that includes this.

[0938] (Claim 2)

[0939] The system according to claim 1, which provides a virtual environment corresponding to a specific emotion based on the user's emotional state.

[0940] (Claim 3)

[0941] The system according to claim 1, comprising a feedback generation means that uses the user's behavioral information and emotional state to provide suggestions for improving healthy habits. [Explanation of Symbols]

[0942] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A natural language processing means that receives voice or text input data, analyzes the input data to understand the user's intent, Information provision means that retrieves information based on analyzed intent from a database and generates an appropriate response, A virtual reality generation means for displaying selected content in order to provide a virtual reality experience to the user, A data analysis means that records user behavior data, analyzes said behavior data, and provides feedback, A system that includes this.

2. The system according to claim 1, which provides virtual reality content related to specific dementia prevention behaviors based on user input data.

3. The system according to claim 1, further comprising a feedback generation means that uses user behavior data to provide suggestions for improving healthy habits.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A