System

A system using a generative model for natural dialogue and exercise support addresses the challenge of cognitive decline in Alzheimer's patients by facilitating natural interaction and immediate feedback, enhancing cognitive function through conversation and exercise.

JP2026036166APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138681
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Alzheimer's disease patients, early-onset Alzheimer's disease patients, and individuals at risk of dementia, including those with mild cognitive impairment (MCI), face challenges in maintaining and improving their cognitive function through conversation and exercise due to isolation and lack of effective, natural interaction systems.

Method used

A system utilizing a generative model for natural dialogue, speech recognition, speech synthesis, exercise program generation, cognitive testing, and feedback mechanisms to facilitate communication, exercise support, and cognitive evaluation, enabling users to interact naturally with AI and receive immediate feedback.

Benefits of technology

Enables Alzheimer's patients and those at risk of dementia to maintain and improve cognitive function through natural conversation and exercise, providing immediate feedback and tailored support at home or in facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036166000001_ABST
    Figure 2026036166000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a user interface means for performing natural interaction using a generation model, a voice recognition means for converting voice input from a user into text data, a server means for transmitting the text data to a server and generating a reply based on the generation model, and a voice synthesis means for converting the generated reply into voice data and reproducing it to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention aims to solve the problem of providing a means for Alzheimer's disease patients, early-onset Alzheimer's disease patients, and people at risk of dementia, including those with mild cognitive impairment (MCI), to communicate naturally at home or in a facility without relying on pharmaceutical treatment. Specifically, the aim is to solve the problem that patients who tend to be isolated find it difficult to maintain and improve their cognitive function through conversation with others, exercise, and cognitive training. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means.

[0006] The system includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data to a server and generating a response based on the generative model, and a speech synthesis means for converting the generated response into speech data and playing it back to the user (Claim 1).

[0007] In addition, the present invention further includes an exercise program generation means that collects exercise data from the user and generates an exercise program provided by an expert, a display means that displays the exercise program to the user and provides instructions, and a feedback means that transmits the user's exercise data to a server and provides the analysis results as feedback (Claim 2).

[0008] Furthermore, the present invention includes a cognitive test providing means for providing multiple types of cognitive tests, a data collecting means for collecting results of users' cognitive tests, a report generating means for analyzing the test results and generating an evaluation report, and a display means for displaying the evaluation report to the user (Claim 3).

[0009] A "generative model" is an artificial intelligence technology that generates natural responses based on input data from the user.

[0010] "User interface means" refers to an interface device for inputting and outputting information between a user and a system.

[0011] "Speech recognition means" refers to a device or software that converts the user's voice into text data.

[0012] A "server means" is a central computing device that processes data sent by a client and generates the necessary responses or information.

[0013] The "voice synthesis means" is a device or software that converts text data into voice data and outputs the voice to the user.

[0014] The "exercise program generating means" is a device or software that generates an exercise program suited to a user based on data provided by an expert.

[0015] A "display means" is a device or software that provides system-generated information or instructions to a user visually or audibly.

[0016] The "feedback means" is a device or software that analyzes the exercise data performed by the user and provides the user with appropriate advice and evaluation.

[0017] The "cognitive test providing means" is a device or software that provides a user with tests for assessing multiple types of cognitive functions.

[0018] A "data collection means" is a device or software that collects the results of a cognitive test administered by a user.

[0019] A "report generator" is a device or software that analyzes collected cognitive test results and generates an assessment report. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0042] Natural conversation with AI

[0043] User voice input

[0044] The user speaks to the terminal, for example, asking a question such as, "What would you like to talk about today?"

[0045] Speech-to-text

[0046] The terminal receives the user's voice through a microphone and converts it into text data using a voice recognition means, and the converted text is sent to the server.

[0047] Response generation using AI models

[0048] The server generates an appropriate response to the received text data based on the generative model, converts the generated response text into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[0049] Playing back replies to the user

[0050] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0051] Exercise support and advice features

[0052] Providing exercise programs

[0053] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0054] Exercise instructions

[0055] The device displays the received exercise program to the user visually or audibly and provides instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0056] Exercise data collection and analysis

[0057] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0058] Cognitive Test Function

[0059] Cognitive testing provided

[0060] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0061] Cognitive testing

[0062] The device presents the received cognitive test to the user and allows them to perform it. For example, the device starts the test by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[0063] Collecting and evaluating test results

[0064] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0065] Through these functions, this system aims to enable Alzheimer's disease patients and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and to maintain and improve their cognitive function through exercise and cognitive training. Furthermore, by utilizing a generative model, users can enjoy such a natural conversational experience that they hardly feel they are talking to an AI.

[0066] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[0067] The processing flow will be explained below.

[0068] Natural conversation with AI

[0069] Step 1:

[0070] User: Talks to the device, for example, asking, "What would you like to talk about today?"

[0071] Step 2:

[0072] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[0073] Step 3:

[0074] Terminal: Converts voice data into text data using a voice recognition means.

[0075] Step 4:

[0076] Terminal: Sends the converted text data to the server.

[0077] Step 5:

[0078] Server: Inputs the received text data into the generative model and generates an appropriate response.

[0079] Step 6:

[0080] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[0081] Step 7:

[0082] Terminal: Plays back the audio data and tells the user a response. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[0083] Exercise support and advice features

[0084] Step 1:

[0085] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[0086] Step 2:

[0087] Server: Sends the generated exercise program to the terminal.

[0088] Step 3:

[0089] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[0090] Step 4:

[0091] User: Begins exercise as instructed, for example, performing upper body stretches.

[0092] Step 5:

[0093] Device: Collects user movement data in real time from sensors.

[0094] Step 6:

[0095] Terminal: Sends collected exercise data to the server.

[0096] Step 7:

[0097] Server: Analyzes the received motion data and generates feedback.

[0098] Step 8:

[0099] Server: Sends the generated feedback to the device.

[0100] Step 9:

[0101] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[0102] Cognitive Test Function

[0103] Step 1:

[0104] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[0105] Step 2:

[0106] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[0107] Step 3:

[0108] User: Follow the instructions to complete the cognitive test.

[0109] Step 4:

[0110] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[0111] Step 5:

[0112] Server: Analyzes the received test results and generates an evaluation report.

[0113] Step 6:

[0114] Server: Sends the generated evaluation report to the terminal.

[0115] Step 7:

[0116] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[0117] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo rehabilitation safely and effectively at home or in a facility.

[0118] Example 1

[0119] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0120] As our society ages, there is an increasing demand for systems to prevent and improve cognitive and motor decline. However, current systems often have unnatural dialogue with users and have difficulty providing appropriate exercise programs and cognitive tests. Furthermore, the feedback these systems provide is not immediate, making it difficult to motivate users or provide optimal support. Therefore, there is a need for systems that enable natural dialogue, provide individually tailored exercise programs and cognitive tests, and evaluate and provide feedback on the results in real time.

[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0122] In this invention, the server includes a speech recognition means for converting speech input into text data, a server means for generating appropriate responses based on a generative model, a speech synthesis means for converting the generated responses into speech and playing them back to the user, an exercise program generation means for collecting exercise data from the user and generating an exercise program based on expert data, a cognitive test provision means for providing multiple types of cognitive tests, and a display means for generating an evaluation report and displaying it to the user. This enables natural dialogue and immediate feedback, increasing the user's motivation while providing effective exercise support and cognitive function evaluation.

[0123] A "voice recognition means" is a device or technique for converting voice input from a user into text data.

[0124] The "server means" is a device or system that utilizes a generative model to generate an appropriate response based on received text data.

[0125] "Speech synthesis means" refers to a device or technology that converts the generated text data into speech data and plays it back to the user.

[0126] The "user interface means" is a device or technology that cooperates with the server to provide operability and information to the user.

[0127] The "exercise program generating means" is a device or system that generates an individual exercise program based on collected exercise data and expert data.

[0128] The "display means" refers to a device or technology for visually presenting the generated exercise program and various information to the user.

[0129] A "feedback means" is a device or technology for analyzing collected data and providing appropriate feedback to the user.

[0130] The "cognitive test providing means" is a device or system that selects from multiple types of cognitive tests and provides them to the user.

[0131] A "data collection tool" is a device or technology that collects the results of a cognitive test or exercise performed by a user.

[0132] A "report generation means" is a device or system for analyzing collected data and generating an evaluation report.

[0133] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0134] Natural conversation with AI

[0135] Voice input and conversion

[0136] The user speaks into the microphone of the device. For example, they may ask, "What would you like to talk about today?" The device receives the user's voice through the microphone and converts the voice data into text data using a speech recognition method (e.g., Google (registered trademark) Cloud Speech-to-Text). The converted text is then sent to the server.

[0137] Response generation and playback

[0138] The server uses a generative AI model (e.g., OpenAI (registered trademark) GPT-4 (registered trademark)) to generate an appropriate response to the received text data. This response text is converted into voice data using a speech synthesis means (e.g., Amazon Polly) and sent to the device. The device then plays the voice data received from the server and conveys the response to the user. For example, a natural conversation follows, such as, "The weather is nice today. Have you gone out anywhere?"

[0139] Exercise support and advice features

[0140] Creation and provision of exercise programs

[0141] The server uses an exercise program generation means to generate an individual exercise program based on the expert's data. This exercise program is then sent to the terminal. The terminal then displays the received exercise program to the user visually or audibly, and provides specific instructions. For example, the terminal may instruct the user on specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0142] Exercise data collection and analysis

[0143] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit). This data is sent to a server, which analyzes the collected data and generates appropriate feedback. The feedback includes specific suggestions for improvement, such as "It would be more effective if you angled your arms a little higher."

[0144] Cognitive Test Function

[0145] Providing and administering cognitive tests

[0146] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal. The terminal then presents the received cognitive test to the user and has them perform it. For example, the test can be started by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[0147] Collecting and evaluating test results

[0148] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0149] Specific examples of operation

[0150] Natural conversation examples

[0151] User voice input: The user speaks into the device, "What would you like to talk about today?"

[0152] AI response generation and playback: The response from the server is sent to the device, and a voice message saying, "The weather is nice today. Have you gone out anywhere?" is played.

[0153] Exercise support examples

[0154] Exercise instructions: The device will give you visual or audio instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[0155] Feedback: Based on the data collected during exercise, feedback is sent from the server to the device, and specific advice such as "It would be more effective if you raised the angle of your arms a little more" is displayed.

[0156] Cognitive test examples

[0157] Test instructions: The device will give instructions such as "The next test is a memory test. Please remember the words that appear on the screen," and then carry out the test.

[0158] Presentation of evaluation report: After the user completes the test, the evaluation report analyzed by the server is sent to the terminal and displayed to the user.

[0159] Example prompts for generative AI models

[0160] "When a user asks, 'What would you like to talk about today?' the AI ​​responds naturally with everyday topics."

[0161] "It tells users, 'Today, stretch your upper body. Extend your arms in front of you,' and provides specific exercise advice."

[0162] "During a memory test, you are instructed to 'remember the following words,' and the test results are collected, evaluated, and a report is generated."

[0163] The above is a mode for carrying out the present invention, and shows that natural dialogue, exercise support, and cognitive testing can be effectively provided through specific processing steps.

[0164] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0165] Step 1: Receiving voice input

[0166] The user speaks to the terminal, asking questions such as, "What would you like to talk about today?"

[0167] Input: User's voice data

[0168] Output: Audio data collected by the device's microphone

[0169] Specific operation: The user speaks into the device's microphone, and voice data is input into the device.

[0170] Step 2: Speech to text

[0171] The terminal converts the user's voice data into text data using a voice recognition means (e.g., Google Cloud Speech-to-Text).

[0172] Input: User's voice data

[0173] Output: Text data

[0174] Specific operation: The terminal's voice recognition means analyzes the voice data and converts it into corresponding text.

[0175] Step 3: Send text data

[0176] The terminal transmits the converted text data to the server.

[0177] Input: Text data

[0178] Output: Text data sent to the server

[0179] Specific operation: The device sends text data to a server via the Internet.

[0180] Step 4: Generate response text

[0181] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data.

[0182] Input: Text data

[0183] Output: Response text data

[0184] How it works: The server's AI model analyzes the text data and generates an appropriate response, such as "The weather is nice today. Have you gone out anywhere?"

[0185] Step 5: Convert response text to speech

[0186] The server converts the generated response text into voice data using a speech synthesis means (e.g., Amazon Polly) and sends it to the terminal.

[0187] Input: Response text data

[0188] Output: Audio data

[0189] Specific operation: The server's speech synthesis means converts the response text into speech and transmits the speech data to the terminal.

[0190] Step 6: Play the reply audio

[0191] The terminal reproduces the voice data received from the server and conveys the response to the user.

[0192] Input: Audio data

[0193] Output: Audio played from the device

[0194] Specific operation: The device speaker plays the audio data and conveys the response to the user.

[0195] Step 7: Generate the exercise program

[0196] The server uses an exercise program generating means to generate an individual exercise program based on the expert data.

[0197] Input: Expert data and user feedback data

[0198] Output: personalized exercise program

[0199] Specific operation: The server uses the exercise program generation means to design an individual exercise program based on the collected data.

[0200] Step 8: Providing an exercise program

[0201] The terminal displays the received exercise program to the user visually or audibly, and provides specific instructions.

[0202] Input: Exercise program data

[0203] Output: An exercise program provided to the user

[0204] Specific actions: The device displays an exercise program on the screen or gives instructions to the user by voice. For example, it may say, "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0205] Step 9: Collect exercise data

[0206] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit).

[0207] Input: User's exercise data

[0208] Output: Collected exercise data

[0209] Specific operation: The sensor collects the user's movement data in real time and transmits it to the device.

[0210] Step 10: Analyze the movement data

[0211] The server analyzes the collected movement data and generates appropriate feedback.

[0212] Input: Collected exercise data

[0213] Output: Feedback data

[0214] Specific movements: The server analyzes the movement data and generates specific feedback, such as advice like "It would be more effective if you angled your arms a little higher."

[0215] Step 11: Offer cognitive testing

[0216] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0217] Input: User context data

[0218] Output: Cognitive test data

[0219] Specific operation: The server selects the appropriate cognitive test from multiple options and sends it to the device.

[0220] Step 12: Conduct cognitive testing

[0221] The terminal presents the received cognitive test to the user and allows the user to perform the test.

[0222] Input: Cognitive test data

[0223] Output: Test result data performed by the user

[0224] Specific operation: The device displays a cognitive test on the screen and asks the user to complete it. For example, it instructs the user, "The next test is a memory test. Please remember the words displayed on the screen."

[0225] Step 13: Collecting test results

[0226] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means.

[0227] Input: User test result data

[0228] Output: Test result data sent to the server

[0229] Specific operation: The device collects the test results and sends them to the server.

[0230] Step 14: Analyze test results and generate evaluation reports

[0231] The server analyzes the results and generates an evaluation report, which includes the user's current cognitive function and areas for improvement.

[0232] Input: Test result data

[0233] Output: Evaluation report

[0234] Specific operation: The server analyzes the test result data and generates an evaluation report.

[0235] Step 15: Provide an evaluation report

[0236] The terminal displays the evaluation report to the user.

[0237] Input: Evaluation report data

[0238] Output: Evaluation report provided to the user

[0239] Specific operation: The device displays the evaluation report on the screen and shows it to the user.

[0240] (Application example 1)

[0241] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0242] Conventional online shopping systems are often cumbersome, requiring complex operations such as product searches, ordering, and return procedures, making them particularly difficult for elderly users and those with low IT literacy. Another issue is the lack of product suggestions based on individual user preferences and body types, making it time-consuming to find the perfect product. Furthermore, there is a lack of systems that can routinely conduct cognitive tests on users with concerns about cognitive decline.

[0243] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0244] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into speech data and playing it back to the user, a means for providing functions that allow the user to search for, order, and process returns of products by voice, and a means for suggesting recommended items based on the user's body type and preferences. This allows the user to easily operate the shop by voice, makes it possible to suggest products that suit the user's individual needs, and also maintains and improves cognitive function.

[0245] A "generative model" is an artificial intelligence model that generates new text or information based on given text or data.

[0246] "User interface means" are technical means that provide an interface through which a user can directly interact with a system.

[0247] "Speech recognition means" refers to technical means for converting speech input into text data.

[0248] "Server means" refers to technical means that provides the functionality of a server that receives text data and generates a response based on a generative model.

[0249] "Speech synthesis means" refers to technical means for converting text data into voice data and playing it back to the user.

[0250] "Means for providing functions for searching for, ordering, and processing returns of products" refers to technical means that allow users to search for, order, and process returns of products using voice commands.

[0251] The "means for proposing recommended items" refers to a technical means for proposing appropriate products according to the user's body type and preferences.

[0252] This invention is a virtual assistant system that integrates natural dialogue using generative models with online shopping functions. The system includes a server means that uses speech recognition, speech synthesis, generative models, and a user interface means.

[0253] Hardware and Software Configuration

[0254] Hardware:

[0255] Smartphone (iOS, ANDROID (registered trademark) compatible)

[0256] software:

[0257] Speech Recognition API (Google Cloud Speech-to-Text)

[0258] Generative AI model (GPT-4)

[0259] Speech synthesis API (Google Cloud Text-to-Speech)

[0260] System Operation

[0261] Voice input and recognition:

[0262] When a user speaks into a smartphone, the built-in microphone picks up the voice, and this voice data is converted into text data in real time via a speech recognition API.

[0263] Sending data to the server and generating a model response:

[0264] The converted text data is sent to a server, where a generative AI model (GPT-4) generates natural-sounding responses based on the user's speech. This response text is then converted back into voice data using a speech synthesis API and sent back to the smartphone.

[0265] Response playback in the user interface:

[0266] The smartphone plays back the returned voice data and provides natural and appropriate responses to the user. For example, if the user says, "I'm looking for a black dress," the speech recognition API converts this into text, and the generative AI model responds, "I'll show you the search results for black dresses."

[0267] Product search, suggestions, and ordering:

[0268] The server also allows users to search for products, place orders, and process returns using voice commands, and it also provides recommendations based on the user's body type and preferences, helping users find the perfect product.

[0269] Examples and prompts

[0270] For example, if a user says, "I'm looking for a black dress," a voice-synthesized message will be played in response, such as, "Here are the search results for black dresses. There are several options." This is followed by, "Today's recommended items are also displayed," allowing the user to enjoy shopping more intuitively through voice.

[0271] Example prompt sentence:

[0272] "A virtual shopping assistant, recommending suitable products for users looking for black dresses."

[0273] This system makes online shopping more natural and user-friendly, and is suitable for elderly people and those with low IT literacy. Furthermore, for users who are concerned about declining cognitive function, it also provides a daily cognitive test function, helping them manage their health.

[0274] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0275] Step 1:

[0276] The user speaks into the smartphone. The smartphone's built-in microphone receives the voice and collects this voice data. The input is the user's voice, and the output is voice data.

[0277] Step 2:

[0278] The device sends the collected voice data to a voice recognition API (Google Cloud Speech-to-Text) in real time, which converts the voice data into text data. The input is voice data, and the output is text data. This process recognizes the voice and converts it into text.

[0279] Step 3:

[0280] Text data is sent from the terminal to the server. When the server receives the text data, the input is text data and the output is the text data sent to the server. This process prepares the data to be used in the next step.

[0281] Step 4:

[0282] The server feeds the received text data into a generative AI model (GPT-4) to generate an appropriate response. The input is text data, and the output is the generated response text. This data processing enables natural dialogue.

[0283] Step 5:

[0284] The generated response text is converted into audio data on the server side using a speech synthesis API (Google Cloud Text-to-Speech). The input is the response text and the output is audio data. This process provides the response in audio format.

[0285] Step 6:

[0286] The generated voice data is sent from the server to the terminal. The user can continue the conversation by receiving and playing the voice data. The input is the voice data from the server, and the output is the played voice.

[0287] Step 7:

[0288] When a user wants to search for a product, the server will suggest recommended items based on the user's body type and preferences. The input is the user's detailed information and search criteria, and the output is a list of recommended products. The server generates this and sends it to the device.

[0289] Step 8:

[0290] The terminal displays the received product list to the user, who then uses voice commands to order products or perform other operations. The input is the user's voice commands, and the output is the results of operations such as ordering. Specifically, if the user says, "I'm looking for a black dress," the terminal responds, "I'll show you the search results for black dresses."

[0291] This series of steps allows users to easily shop through voice, with data processing and calculations at each step providing a highly interactive and personalized shopping experience.

[0292] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0293] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0294] Natural conversation with AI

[0295] User voice input

[0296] The user speaks to the terminal, for example, "What would you like to talk about today?"

[0297] Speech-to-text

[0298] The terminal receives the user's voice through a microphone and acquires it as voice data.

[0299] Speech and Emotion Recognition

[0300] The terminal converts the voice data into text data using a voice recognition means, and simultaneously recognizes the user's emotion using an emotion engine. The converted text and the recognized emotion data are sent to the server.

[0301] Response generation using AI models

[0302] The server generates an appropriate response to the received text data based on the generative model. It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[0303] Playing back replies to the user

[0304] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0305] Exercise support and advice features

[0306] Providing exercise programs

[0307] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0308] Exercise instructions

[0309] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0310] Exercise data collection and analysis

[0311] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0312] Cognitive Test Function

[0313] Cognitive testing provided

[0314] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0315] Cognitive testing

[0316] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[0317] Collecting and evaluating test results

[0318] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0319] Through these functions, this system aims to maintain and improve cognitive function by enabling people with Alzheimer's disease and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and by providing exercise and cognitive training. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable conversational experience for the user.

[0320] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[0321] The processing flow will be explained below.

[0322] Natural conversation with AI

[0323] Step 1:

[0324] User: Speaks into the device. For example, "What would you like to talk about today?"

[0325] Step 2:

[0326] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[0327] Step 3:

[0328] Terminal: Converts voice data into text data using a voice recognition means.

[0329] Step 4:

[0330] Terminal: Uses an emotion engine to recognize emotions from the user's voice.

[0331] Step 5:

[0332] Terminal: Transmits the converted text data and recognized emotion data to the server.

[0333] Step 6:

[0334] Server: Inputs the received text data into a generative model to generate an appropriate response. Based on the emotional data, the tone and content of the response are adjusted.

[0335] Step 7:

[0336] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[0337] Step 8:

[0338] Terminal: Plays back the audio data and responds to the user. For example, if the user is sad, the response will be something comforting like, "What happened today? Tell me."

[0339] Exercise support and advice features

[0340] Step 1:

[0341] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[0342] Step 2:

[0343] Server: Sends the generated exercise program to the terminal.

[0344] Step 3:

[0345] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[0346] Step 4:

[0347] User: Starts exercising according to instructions. For example, the user is instructed to do specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0348] Step 5:

[0349] Device: Collects user movement data in real time from sensors.

[0350] Step 6:

[0351] Terminal: Sends collected exercise data to the server.

[0352] Step 7:

[0353] Server: Analyzes the received motion data and generates feedback.

[0354] Step 8:

[0355] Server: Sends the generated feedback to the device.

[0356] Step 9:

[0357] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[0358] Cognitive Test Function

[0359] Step 1:

[0360] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[0361] Step 2:

[0362] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[0363] Step 3:

[0364] User: Follow the instructions to complete the cognitive test.

[0365] Step 4:

[0366] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[0367] Step 5:

[0368] Server: Analyzes the received test results and generates an evaluation report.

[0369] Step 6:

[0370] Server: Sends the generated evaluation report to the terminal.

[0371] Step 7:

[0372] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[0373] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo safe and effective rehabilitation at home or in a facility. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable dialogue experience for the user.

[0374] Example 2

[0375] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0376] The objective of this invention is to provide a system that integrates natural user interaction, exercise support, cognitive testing, and emotion recognition. Conventional systems often provide these functions separately, resulting in a fragmented user experience. Furthermore, emotion recognition accuracy is low, making it difficult to respond to the user's emotions. As a result, users experience low satisfaction and effectiveness.

[0377] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a voice recognition means for converting a user's voice input into text data, an emotion engine means for recognizing the user's emotions, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, and a voice synthesis means for converting the generated response into voice data and playing it back to the user. This not only enables the user to enjoy a natural dialogue experience, but also enables emotionally sensitive communication. The server also includes an exercise program generation means for collecting exercise data from the user and generating an exercise program provided by an expert, a display means for displaying the exercise program to the user and providing instructions, and a feedback means for transmitting the user's exercise data to the server and providing analysis results as feedback. The server also includes a cognitive test provision means for providing multiple types of cognitive tests, a data collection means for collecting results of the cognitive tests performed by the user, a report generation means for analyzing the test results and generating an evaluation report, and a display means for displaying the evaluation report to the user. This enables comprehensive health management, thereby increasing user satisfaction and effectiveness.

[0378] A "generative model" is an algorithm that uses artificial intelligence to generate natural-looking dialogue and responses based on given input data.

[0379] "User interface means" refers to a device or software that allows a user to interact with a system, including voice input and a display.

[0380] "Speech recognition means" refers to a technology or device that analyzes voice input from a user as digital data and converts it into character data.

[0381] "Emotion engine means" refers to a technology or device that identifies emotions from the user's voice or text and outputs the emotional state as data.

[0382] "Server Means" refers to a centralized computing resource that processes data for the entire system and generates responses based on the Generative Model.

[0383] "Speech synthesis means" refers to a technique or device that converts the generated text data into voice data and plays it back to the user.

[0384] The term "exercise program generation means" refers to a technology or device that creates an individually customized exercise program based on the user's exercise data and expert guidance.

[0385] "Display means" refers to technology or devices for visually presenting exercise instructions, cognitive test results, reports, etc. to the user.

[0386] "Feedback means" refers to technology or devices that analyze the user's exercise data and cognitive test results and provide advice on areas for improvement and next steps.

[0387] A "cognitive test delivery means" refers to a technology or device that delivers tests to a user to assess various cognitive functions.

[0388] "Data collection means" refers to sensors and software that collect user behavior and test results.

[0389] A "report generation means" is a technique or device that analyzes user data and compiles the results into a report.

[0390] MODE FOR CARRYING OUT THE INVENTION

[0391] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0392] Natural conversation with AI

[0393] User voice input

[0394] The user speaks to the terminal, for example, "What would you like to talk about today?"

[0395] Speech-to-text

[0396] The device receives the user's voice through a microphone and acquires it as voice data. The Google Cloud Speech-to-Text API is used for voice recognition.

[0397] Speech and Emotion Recognition

[0398] The device converts voice data into text data using a voice recognition means. At the same time, it recognizes the user's emotions using an emotion engine means. The emotion recognition uses the Microsoft® Azure® emotion analysis API. The converted text data and the recognized emotion data are sent to the server.

[0399] Response generation using AI models

[0400] The server generates an appropriate response to the received text data based on a generative AI model (e.g., ChatGPT (registered trademark)). It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[0401] Playing back replies to the user

[0402] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0403] Exercise support and advice features

[0404] Providing exercise programs

[0405] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0406] Exercise instructions

[0407] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0408] Exercise data collection and analysis

[0409] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0410] Cognitive Test Function

[0411] Cognitive testing provided

[0412] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0413] Cognitive testing

[0414] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[0415] Collecting and evaluating test results

[0416] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0417] This not only allows users to enjoy a natural conversational experience, but also enables emotionally sensitive communication and comprehensive health management through exercise and cognitive training.

[0418] Prompt Sentence Examples

[0419] "Create prompts that are sensitive to the user's emotions and start the conversation around the weather. For example, 'The weather is nice today. Have you been out anywhere?'"

[0420] The above is an embodiment of the invention.

[0421] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0422] Step 1:

[0423] User voice input

[0424] Input: User speech

[0425] How it works: The user speaks to the device, asking something like, "What would you like to talk about today?"

[0426] Output: User's voice data

[0427] Step 2:

[0428] Speech-to-text

[0429] Input: User's voice data

[0430] How it works: The device captures the user's voice with its built-in microphone and converts the voice data into text using the Google Cloud Speech-to-Text API.

[0431] Output: Text data

[0432] Step 3:

[0433] emotion recognition

[0434] Input: Text data

[0435] How it works: The device uses Microsoft Azure's sentiment analysis API to recognize user emotions from text data.

[0436] Output: Emotion data

[0437] Step 4:

[0438] Sending data to the server

[0439] Input: Text data, emotion data

[0440] Operation: The device sends the converted text data and the recognized emotion data to the server.

[0441] Output: Data sent to the server

[0442] Step 5:

[0443] Response generation using AI models

[0444] Input: Text data and emotion data sent to the server

[0445] How it works: The server uses a generative AI model (e.g., ChatGPT) to generate appropriate responses based on the data it receives, and also adjusts the content and tone of the responses to match the user's emotions.

[0446] Output: Response text data

[0447] Step 6:

[0448] Speech synthesis and reply sending

[0449] Input: Response text data

[0450] Operation: The server converts the response text data into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[0451] Output: Audio data

[0452] Step 7:

[0453] Playing back replies to the user

[0454] Input: Audio data

[0455] Operation: The device receives the voice data from the server, plays it back through the speaker, and responds to the user. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[0456] Output: The reply speech the user hears

[0457] Step 8:

[0458] Providing exercise programs

[0459] Input: Exercise data provided by experts

[0460] Operation: The server uses the exercise program generation means to generate an individual exercise program based on the data provided by the expert, and transmits it to the terminal.

[0461] Output: Exercise program

[0462] Step 9:

[0463] Exercise instructions

[0464] Input: Exercise program

[0465] Actions: The device displays the received exercise program to the user visually or audibly, and provides specific exercise instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[0466] Output: Exercise instruction content

[0467] Step 10:

[0468] Exercise data collection and analysis

[0469] Input: User's exercise data

[0470] How it works: When a user exercises, the device collects exercise data through sensors attached to the device and transmits the data to a server, which analyzes the data and generates appropriate feedback.

[0471] Output: Analysis results and feedback

[0472] Step 11:

[0473] Cognitive testing provided

[0474] Input: Cognitive test selection instructions

[0475] Operation: The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0476] Output: Cognitive test content

[0477] Step 12:

[0478] Cognitive testing

[0479] Input: Cognitive test content

[0480] Operation: The device presents the received cognitive test to the user and has them complete it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[0481] Output: User's cognitive test results

[0482] Step 13:

[0483] Collecting and evaluating test results

[0484] Input: User's cognitive test results

[0485] Operation: The device sends the results of the cognitive tests taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the user's current cognitive function and areas for improvement, and is displayed to the user via the device.

[0486] Output: Evaluation report

[0487] (Application example 2)

[0488] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0489] There is a need for an effective method to support the dialogue between customers and store clerks in brick-and-mortar stores. In particular, there is a need for a system that can recognize customers' emotions and respond flexibly and naturally. Another challenge is to improve customer satisfaction by suggesting products that correspond to the customer's interests and emotions.

[0490] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0491] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into voice data and playing it back to the user, and a suggestion means for making product suggestions in the store. This allows for natural dialogue with customers and enables flexible and appropriate responses based on the customer's emotions. Furthermore, by making product suggestions based on the customer's interests, customer satisfaction can be improved.

[0492] A "generative model" is an artificial intelligence algorithm that generates appropriate responses or results based on input data from a user.

[0493] A "user interface means" is a means that provides an interface for a user to interact with the system.

[0494] The "voice recognition means" is a means having a function of converting voice input from a user into text data.

[0495] "Emotional data" is data collected to describe a user's emotional state.

[0496] The "server means" is a means for receiving text data and emotion data and generating an appropriate response based on the received data.

[0497] The "voice synthesis means" is a means having the function of converting the generated response into voice data and playing it back to the user.

[0498] The "proposal means" is a means having a function for proposing products and services to customers in a store.

[0499] The "exercise program generating means" is a means having a function of generating an appropriate exercise program based on data provided by an expert.

[0500] The "display means" is a means having a function for visually presenting the generated exercise program and other information to the user.

[0501] The "feedback means" is a means having a function for analyzing the user's exercise data and providing appropriate feedback.

[0502] The "cognitive test providing means" is a means having a function for providing multiple types of cognitive tests to the user.

[0503] The "data collection means" is a means having a function for collecting the results of the cognitive tests performed by the user.

[0504] The "report generation means" is a means having a function for analyzing the results of the cognitive test and generating an evaluation report.

[0505] The system that realizes this application example provides natural dialogue, exercise support, and cognitive testing using a generative model, and incorporates an emotion engine. A specific embodiment of this system will be described below.

[0506] Basic system configuration

[0507] The system consists of the following means:

[0508] User Interface Means

[0509] Voice recognition means

[0510] emotion recognition means

[0511] Server Means

[0512] Voice synthesis means

[0513] Proposal means

[0514] Exercise program generation means

[0515] Display means

[0516] Feedback Methods

[0517] Cognitive test delivery methods

[0518] Data collection methods

[0519] Report Generation Method

[0520] Natural conversational features

[0521] User voice input

[0522] The user speaks to a terminal connected to the system. For example, they might say, "Hello, do you have any product recommendations?" The terminal receives the voice through a microphone and captures it as voice data.

[0523] Speech-to-text and emotion recognition

[0524] The device converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text), and simultaneously recognizes the user's emotions using an emotion recognition engine (e.g., Microsoft Azure Emotion API). This data is then sent to the server.

[0525] Response generation using AI models

[0526] The server generates an appropriate response to the received text data based on a generative AI model (e.g., OpenAI GPT-3 (registered trademark)). It also adjusts the content and tone of the response based on emotional data. For example, if the user shows interest, it generates a friendly response such as, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?"

[0527] Playing back replies to the user

[0528] The generated response text is converted into voice data by a speech synthesis engine, and the device plays this voice data over a speaker to convey the response to the user.

[0529] Exercise support and advice features

[0530] Providing exercise programs

[0531] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0532] Exercise instructions and feedback

[0533] The device displays the received exercise program to the user visually or audibly and provides exercise instructions. During exercise, exercise data collected through sensors is sent to a server for analysis. Based on the analysis results, the device provides feedback such as, "It would be more effective if you raised the angle of your arms a little more."

[0534] Cognitive Test Function

[0535] Providing and administering cognitive tests

[0536] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal, for example, by informing the user that "The next test is a memory test. Please remember the words displayed on the screen."

[0537] Analysis and evaluation of test results

[0538] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0539] Specific examples

[0540] As examples of use in physical stores, the following specific dialogue scenarios are envisioned.

[0541] Usage

[0542] 1. User (Customer): "Hi, do you have any product recommendations?"

[0543] 2. Device (application running on smartphone):

[0544] Acquisition speech-to-text: "Hi, do you have any product recommendations?"

[0545] Emotion Recognition: "Curious" (Emotion Engine)

[0546] Response generation: "Hello, we have the latest smartwatch that's just for you. Would you like to take a look?" (Generative AI model)

[0547] Speech synthesis: Generated responses are spoken

[0548] Speaker Output: "Hello, we have the latest smartwatch just for you. Would you like to take a look?"

[0549] Prompt Sentence Examples

[0550] Generate relevant product suggestions when customers express interest: "Create an engaging product suggestion based on the following sentence: 'Hi, can you recommend something?' Keep the tone of your suggestion friendly but professional."

[0551] This system enables natural dialogue with customers in physical stores, and flexible and appropriate responses based on customer emotions. In addition, it can improve customer satisfaction by suggesting products based on the customer's interests.

[0552] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0553] Step 1:

[0554] User voice input

[0555] The user speaks to the device, for example, "Hello, do you have any product recommendations?" This speech is picked up by the device's microphone and saved as audio data.

[0556] Input: User's voice

[0557] Output: Audio data

[0558] Step 2:

[0559] Speech-to-text

[0560] The device uses a speech recognition engine to convert the acquired voice data into text data. For this purpose, it uses a speech recognition API (for example, Google Speech-to-Text).

[0561] Input: Audio data

[0562] Output: Text data

[0563] Step 3:

[0564] emotion recognition

[0565] The device sends the text data to an emotion recognition engine to recognize the user's emotions. This uses an emotion recognition API (for example, Microsoft Azure Emotion API). Emotion data reflects the user's psychological state.

[0566] Input: Text data

[0567] Output: Emotion data

[0568] Step 4:

[0569] Response generation based on generative models

[0570] The server receives the text data and emotion data and uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The tone and content of the response are adjusted based on the emotion data. For example, if the user shows interest, the response will be more friendly and engaging.

[0571] Input: Text data and emotion data

[0572] Output: Response text

[0573] Step 5:

[0574] Response text transcription

[0575] The server converts the response text into voice data using a speech synthesis engine. For this process, it uses a speech synthesis API (for example, Amazon Polly).

[0576] Input: Reply text

[0577] Output: Reply audio data

[0578] Step 6:

[0579] Playback of response audio data

[0580] The terminal plays the received response voice data over the speaker, thereby conveying the dialogue response to the user.

[0581] Input: Response audio data

[0582] Output: The audio played to the user

[0583] Step 7:

[0584] Product proposals using proposal methods

[0585] The server then suggests appropriate products based on the user's interests and emotions. For example, it generates a response like, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?" This product suggestion is made using a generative AI model that takes into account the user's emotional data.

[0586] Input: User interest and emotion data

[0587] Output: Product suggestion text

[0588] Through these steps, the system can realize natural dialogue with the user, responding flexibly and appropriately according to their emotions. Furthermore, it can improve customer satisfaction by suggesting products based on the user's interests.

[0589] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0590] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0591] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0592] [Second embodiment]

[0593] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0594] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0595] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0596] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0597] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0598] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0599] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0600] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0601] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0602] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0603] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0604] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0605] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0606] Natural conversation with AI

[0607] User voice input

[0608] The user speaks to the terminal, for example, asking a question such as, "What would you like to talk about today?"

[0609] Speech-to-text

[0610] The terminal receives the user's voice through a microphone and converts it into text data using a voice recognition means, and the converted text is sent to the server.

[0611] Response generation using AI models

[0612] The server generates an appropriate response to the received text data based on the generative model, converts the generated response text into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[0613] Playing back replies to the user

[0614] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0615] Exercise support and advice features

[0616] Providing exercise programs

[0617] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0618] Exercise instructions

[0619] The device displays the received exercise program to the user visually or audibly and provides instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0620] Exercise data collection and analysis

[0621] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0622] Cognitive Test Function

[0623] Cognitive testing provided

[0624] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0625] Cognitive testing

[0626] The device presents the received cognitive test to the user and allows them to perform it. For example, the device starts the test by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[0627] Collecting and evaluating test results

[0628] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0629] Through these functions, this system aims to enable Alzheimer's disease patients and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and to maintain and improve their cognitive function through exercise and cognitive training. Furthermore, by utilizing a generative model, users can enjoy such a natural conversational experience that they hardly feel they are talking to an AI.

[0630] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[0631] The processing flow will be explained below.

[0632] Natural conversation with AI

[0633] Step 1:

[0634] User: Talks to the device, for example, asking, "What would you like to talk about today?"

[0635] Step 2:

[0636] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[0637] Step 3:

[0638] Terminal: Converts voice data into text data using a voice recognition means.

[0639] Step 4:

[0640] Terminal: Sends the converted text data to the server.

[0641] Step 5:

[0642] Server: Inputs the received text data into the generative model and generates an appropriate response.

[0643] Step 6:

[0644] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[0645] Step 7:

[0646] Terminal: Plays back the audio data and tells the user a response. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[0647] Exercise support and advice features

[0648] Step 1:

[0649] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[0650] Step 2:

[0651] Server: Sends the generated exercise program to the terminal.

[0652] Step 3:

[0653] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[0654] Step 4:

[0655] User: Begins exercise as instructed, for example, performing upper body stretches.

[0656] Step 5:

[0657] Device: Collects user movement data in real time from sensors.

[0658] Step 6:

[0659] Terminal: Sends collected exercise data to the server.

[0660] Step 7:

[0661] Server: Analyzes the received motion data and generates feedback.

[0662] Step 8:

[0663] Server: Sends the generated feedback to the device.

[0664] Step 9:

[0665] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[0666] Cognitive Test Function

[0667] Step 1:

[0668] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[0669] Step 2:

[0670] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[0671] Step 3:

[0672] User: Follow the instructions to complete the cognitive test.

[0673] Step 4:

[0674] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[0675] Step 5:

[0676] Server: Analyzes the received test results and generates an evaluation report.

[0677] Step 6:

[0678] Server: Sends the generated evaluation report to the terminal.

[0679] Step 7:

[0680] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[0681] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo rehabilitation safely and effectively at home or in a facility.

[0682] Example 1

[0683] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0684] As our society ages, there is an increasing demand for systems to prevent and improve cognitive and motor decline. However, current systems often have unnatural dialogue with users and have difficulty providing appropriate exercise programs and cognitive tests. Furthermore, the feedback these systems provide is not immediate, making it difficult to motivate users or provide optimal support. Therefore, there is a need for systems that enable natural dialogue, provide individually tailored exercise programs and cognitive tests, and evaluate and provide feedback on the results in real time.

[0685] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0686] In this invention, the server includes a speech recognition means for converting speech input into text data, a server means for generating appropriate responses based on a generative model, a speech synthesis means for converting the generated responses into speech and playing them back to the user, an exercise program generation means for collecting exercise data from the user and generating an exercise program based on expert data, a cognitive test provision means for providing multiple types of cognitive tests, and a display means for generating an evaluation report and displaying it to the user. This enables natural dialogue and immediate feedback, increasing the user's motivation while providing effective exercise support and cognitive function evaluation.

[0687] A "voice recognition means" is a device or technique for converting voice input from a user into text data.

[0688] The "server means" is a device or system that utilizes a generative model to generate an appropriate response based on received text data.

[0689] "Speech synthesis means" refers to a device or technology that converts the generated text data into speech data and plays it back to the user.

[0690] The "user interface means" is a device or technology that cooperates with the server to provide operability and information to the user.

[0691] The "exercise program generating means" is a device or system that generates an individual exercise program based on collected exercise data and expert data.

[0692] The "display means" refers to a device or technology for visually presenting the generated exercise program and various information to the user.

[0693] A "feedback means" is a device or technology for analyzing collected data and providing appropriate feedback to the user.

[0694] The "cognitive test providing means" is a device or system that selects from multiple types of cognitive tests and provides them to the user.

[0695] A "data collection tool" is a device or technology that collects the results of a cognitive test or exercise performed by a user.

[0696] A "report generation means" is a device or system for analyzing collected data and generating an evaluation report.

[0697] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0698] Natural conversation with AI

[0699] Voice input and conversion

[0700] The user speaks into the device's microphone. For example, they might ask, "What would you like to talk about today?" The device receives the user's voice through the microphone and converts the voice data into text data using a speech recognition method (e.g., Google Cloud Speech-to-Text). The converted text is then sent to the server.

[0701] Response generation and playback

[0702] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data. This response text is converted into voice data using a speech synthesis method (e.g., Amazon Polly) and sent to the device. The device then plays the voice data received from the server and conveys the response to the user. For example, a natural conversation follows, such as, "The weather is nice today. Have you gone out anywhere?"

[0703] Exercise support and advice features

[0704] Creation and provision of exercise programs

[0705] The server uses an exercise program generation means to generate an individual exercise program based on the expert's data. This exercise program is then sent to the terminal. The terminal then displays the received exercise program to the user visually or audibly, and provides specific instructions. For example, the terminal may instruct the user on specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0706] Exercise data collection and analysis

[0707] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit). This data is sent to a server, which analyzes the collected data and generates appropriate feedback. The feedback includes specific suggestions for improvement, such as "It would be more effective if you angled your arms a little higher."

[0708] Cognitive Test Function

[0709] Providing and administering cognitive tests

[0710] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal. The terminal then presents the received cognitive test to the user and has them perform it. For example, the test can be started by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[0711] Collecting and evaluating test results

[0712] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0713] Specific examples of operation

[0714] Natural conversation examples

[0715] User voice input: The user speaks into the device, "What would you like to talk about today?"

[0716] AI response generation and playback: The response from the server is sent to the device, and a voice message saying, "The weather is nice today. Have you gone out anywhere?" is played.

[0717] Exercise support examples

[0718] Exercise instructions: The device will give you visual or audio instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[0719] Feedback: Based on the data collected during exercise, feedback is sent from the server to the device, and specific advice such as "It would be more effective if you raised the angle of your arms a little more" is displayed.

[0720] Cognitive test examples

[0721] Test instructions: The device will give instructions such as "The next test is a memory test. Please remember the words that appear on the screen," and then carry out the test.

[0722] Presentation of evaluation report: After the user completes the test, the evaluation report analyzed by the server is sent to the terminal and displayed to the user.

[0723] Example prompts for generative AI models

[0724] "When a user asks, 'What would you like to talk about today?' the AI ​​responds naturally with everyday topics."

[0725] "It tells users, 'Today, stretch your upper body. Extend your arms in front of you,' and provides specific exercise advice."

[0726] "During a memory test, you are instructed to 'remember the following words,' and the test results are collected, evaluated, and a report is generated."

[0727] The above is a mode for carrying out the present invention, and shows that natural dialogue, exercise support, and cognitive testing can be effectively provided through specific processing steps.

[0728] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0729] Step 1: Receiving voice input

[0730] The user speaks to the terminal, asking questions such as, "What would you like to talk about today?"

[0731] Input: User's voice data

[0732] Output: Audio data collected by the device's microphone

[0733] Specific operation: The user speaks into the device's microphone, and voice data is input into the device.

[0734] Step 2: Speech to text

[0735] The terminal converts the user's voice data into text data using a voice recognition means (e.g., Google Cloud Speech-to-Text).

[0736] Input: User's voice data

[0737] Output: Text data

[0738] Specific operation: The terminal's voice recognition means analyzes the voice data and converts it into corresponding text.

[0739] Step 3: Send text data

[0740] The terminal transmits the converted text data to the server.

[0741] Input: Text data

[0742] Output: Text data sent to the server

[0743] Specific operation: The device sends text data to a server via the Internet.

[0744] Step 4: Generate response text

[0745] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data.

[0746] Input: Text data

[0747] Output: Response text data

[0748] How it works: The server's AI model analyzes the text data and generates an appropriate response, such as "The weather is nice today. Have you gone out anywhere?"

[0749] Step 5: Convert response text to speech

[0750] The server converts the generated response text into voice data using a speech synthesis means (e.g., Amazon Polly) and sends it to the terminal.

[0751] Input: Response text data

[0752] Output: Audio data

[0753] Specific operation: The server's speech synthesis means converts the response text into speech and transmits the speech data to the terminal.

[0754] Step 6: Play the reply audio

[0755] The terminal reproduces the voice data received from the server and conveys the response to the user.

[0756] Input: Audio data

[0757] Output: Audio played from the device

[0758] Specific operation: The device speaker plays the audio data and conveys the response to the user.

[0759] Step 7: Generate the exercise program

[0760] The server uses an exercise program generating means to generate an individual exercise program based on the expert data.

[0761] Input: Expert data and user feedback data

[0762] Output: personalized exercise program

[0763] Specific operation: The server uses the exercise program generation means to design an individual exercise program based on the collected data.

[0764] Step 8: Providing an exercise program

[0765] The terminal displays the received exercise program to the user visually or audibly, and provides specific instructions.

[0766] Input: Exercise program data

[0767] Output: An exercise program provided to the user

[0768] Specific actions: The device displays an exercise program on the screen or gives instructions to the user by voice. For example, it may say, "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0769] Step 9: Collect exercise data

[0770] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit).

[0771] Input: User's exercise data

[0772] Output: Collected exercise data

[0773] Specific operation: The sensor collects the user's movement data in real time and transmits it to the device.

[0774] Step 10: Analyze the movement data

[0775] The server analyzes the collected movement data and generates appropriate feedback.

[0776] Input: Collected exercise data

[0777] Output: Feedback data

[0778] Specific movements: The server analyzes the movement data and generates specific feedback, such as advice like "It would be more effective if you angled your arms a little higher."

[0779] Step 11: Offer cognitive testing

[0780] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0781] Input: User context data

[0782] Output: Cognitive test data

[0783] Specific operation: The server selects the appropriate cognitive test from multiple options and sends it to the device.

[0784] Step 12: Conduct cognitive testing

[0785] The terminal presents the received cognitive test to the user and allows the user to perform the test.

[0786] Input: Cognitive test data

[0787] Output: Test result data performed by the user

[0788] Specific operation: The device displays a cognitive test on the screen and asks the user to complete it. For example, it instructs the user, "The next test is a memory test. Please remember the words displayed on the screen."

[0789] Step 13: Collecting test results

[0790] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means.

[0791] Input: User test result data

[0792] Output: Test result data sent to the server

[0793] Specific operation: The device collects the test results and sends them to the server.

[0794] Step 14: Analyze test results and generate evaluation reports

[0795] The server analyzes the results and generates an evaluation report, which includes the user's current cognitive function and areas for improvement.

[0796] Input: Test result data

[0797] Output: Evaluation report

[0798] Specific operation: The server analyzes the test result data and generates an evaluation report.

[0799] Step 15: Provide an evaluation report

[0800] The terminal displays the evaluation report to the user.

[0801] Input: Evaluation report data

[0802] Output: Evaluation report provided to the user

[0803] Specific operation: The device displays the evaluation report on the screen and shows it to the user.

[0804] (Application example 1)

[0805] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0806] Conventional online shopping systems are often cumbersome, requiring complex operations such as product searches, ordering, and return procedures, making them particularly difficult for elderly users and those with low IT literacy. Another issue is the lack of product suggestions based on individual user preferences and body types, making it time-consuming to find the perfect product. Furthermore, there is a lack of systems that can routinely conduct cognitive tests on users with concerns about cognitive decline.

[0807] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0808] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into speech data and playing it back to the user, a means for providing functions that allow the user to search for, order, and process returns of products by voice, and a means for suggesting recommended items based on the user's body type and preferences. This allows the user to easily operate the shop by voice, makes it possible to suggest products that suit the user's individual needs, and also maintains and improves cognitive function.

[0809] A "generative model" is an artificial intelligence model that generates new text or information based on given text or data.

[0810] "User interface means" are technical means that provide an interface through which a user can directly interact with a system.

[0811] "Speech recognition means" refers to technical means for converting speech input into text data.

[0812] "Server means" refers to technical means that provides the functionality of a server that receives text data and generates a response based on a generative model.

[0813] "Speech synthesis means" refers to technical means for converting text data into voice data and playing it back to the user.

[0814] "Means for providing functions for searching for, ordering, and processing returns of products" refers to technical means that allow users to search for, order, and process returns of products using voice commands.

[0815] The "means for proposing recommended items" refers to a technical means for proposing appropriate products according to the user's body type and preferences.

[0816] This invention is a virtual assistant system that integrates natural dialogue using generative models with online shopping functions. The system includes a server means that uses speech recognition, speech synthesis, generative models, and a user interface means.

[0817] Hardware and Software Configuration

[0818] Hardware:

[0819] Smartphone (iOS, Android compatible)

[0820] software:

[0821] Speech Recognition API (Google Cloud Speech-to-Text)

[0822] Generative AI model (GPT-4)

[0823] Speech synthesis API (Google Cloud Text-to-Speech)

[0824] System Operation

[0825] Voice input and recognition:

[0826] When a user speaks into a smartphone, the built-in microphone picks up the voice, and this voice data is converted into text data in real time via a speech recognition API.

[0827] Sending data to the server and generating a model response:

[0828] The converted text data is sent to a server, where a generative AI model (GPT-4) generates natural-sounding responses based on the user's speech. This response text is then converted back into voice data using a speech synthesis API and sent back to the smartphone.

[0829] Response playback in the user interface:

[0830] The smartphone plays back the returned voice data and provides natural and appropriate responses to the user. For example, if the user says, "I'm looking for a black dress," the speech recognition API converts this into text, and the generative AI model responds, "I'll show you the search results for black dresses."

[0831] Product search, suggestions, and ordering:

[0832] The server also allows users to search for products, place orders, and process returns using voice commands, and it also provides recommendations based on the user's body type and preferences, helping users find the perfect product.

[0833] Examples and prompts

[0834] For example, if a user says, "I'm looking for a black dress," a voice-synthesized message will be played in response, such as, "Here are the search results for black dresses. There are several options." This is followed by, "Today's recommended items are also displayed," allowing the user to enjoy shopping more intuitively through voice.

[0835] Example prompt sentence:

[0836] "A virtual shopping assistant, recommending suitable products for users looking for black dresses."

[0837] This system makes online shopping more natural and user-friendly, and is suitable for elderly people and those with low IT literacy. Furthermore, for users who are concerned about declining cognitive function, it also provides a daily cognitive test function, helping them manage their health.

[0838] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0839] Step 1:

[0840] The user speaks into the smartphone. The smartphone's built-in microphone receives the voice and collects this voice data. The input is the user's voice, and the output is voice data.

[0841] Step 2:

[0842] The device sends the collected voice data to a voice recognition API (Google Cloud Speech-to-Text) in real time, which converts the voice data into text data. The input is voice data, and the output is text data. This process recognizes the voice and converts it into text.

[0843] Step 3:

[0844] Text data is sent from the terminal to the server. When the server receives the text data, the input is text data and the output is the text data sent to the server. This process prepares the data to be used in the next step.

[0845] Step 4:

[0846] The server feeds the received text data into a generative AI model (GPT-4) to generate an appropriate response. The input is text data, and the output is the generated response text. This data processing enables natural dialogue.

[0847] Step 5:

[0848] The generated response text is converted into audio data on the server side using a speech synthesis API (Google Cloud Text-to-Speech). The input is the response text and the output is audio data. This process provides the response in audio format.

[0849] Step 6:

[0850] The generated voice data is sent from the server to the terminal. The user can continue the conversation by receiving and playing the voice data. The input is the voice data from the server, and the output is the played voice.

[0851] Step 7:

[0852] When a user wants to search for a product, the server will suggest recommended items based on the user's body type and preferences. The input is the user's detailed information and search criteria, and the output is a list of recommended products. The server generates this and sends it to the device.

[0853] Step 8:

[0854] The terminal displays the received product list to the user, who then uses voice commands to order products or perform other operations. The input is the user's voice commands, and the output is the results of operations such as ordering. Specifically, if the user says, "I'm looking for a black dress," the terminal responds, "I'll show you the search results for black dresses."

[0855] This series of steps allows users to easily shop through voice, with data processing and calculations at each step providing a highly interactive and personalized shopping experience.

[0856] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0857] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0858] Natural conversation with AI

[0859] User voice input

[0860] The user speaks to the terminal, for example, "What would you like to talk about today?"

[0861] Speech-to-text

[0862] The terminal receives the user's voice through a microphone and acquires it as voice data.

[0863] Speech and Emotion Recognition

[0864] The terminal converts the voice data into text data using a voice recognition means, and simultaneously recognizes the user's emotion using an emotion engine. The converted text and the recognized emotion data are sent to the server.

[0865] Response generation using AI models

[0866] The server generates an appropriate response to the received text data based on the generative model. It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[0867] Playing back replies to the user

[0868] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0869] Exercise support and advice features

[0870] Providing exercise programs

[0871] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0872] Exercise instructions

[0873] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0874] Exercise data collection and analysis

[0875] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0876] Cognitive Test Function

[0877] Cognitive testing provided

[0878] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0879] Cognitive testing

[0880] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[0881] Collecting and evaluating test results

[0882] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0883] Through these functions, this system aims to maintain and improve cognitive function by enabling people with Alzheimer's disease and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and by providing exercise and cognitive training. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable conversational experience for the user.

[0884] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[0885] The processing flow will be explained below.

[0886] Natural conversation with AI

[0887] Step 1:

[0888] User: Speaks into the device. For example, "What would you like to talk about today?"

[0889] Step 2:

[0890] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[0891] Step 3:

[0892] Terminal: Converts voice data into text data using a voice recognition means.

[0893] Step 4:

[0894] Terminal: Uses an emotion engine to recognize emotions from the user's voice.

[0895] Step 5:

[0896] Terminal: Transmits the converted text data and recognized emotion data to the server.

[0897] Step 6:

[0898] Server: Inputs the received text data into a generative model to generate an appropriate response. Based on the emotional data, the tone and content of the response are adjusted.

[0899] Step 7:

[0900] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[0901] Step 8:

[0902] Terminal: Plays back the audio data and responds to the user. For example, if the user is sad, the response will be something comforting like, "What happened today? Tell me."

[0903] Exercise support and advice features

[0904] Step 1:

[0905] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[0906] Step 2:

[0907] Server: Sends the generated exercise program to the terminal.

[0908] Step 3:

[0909] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[0910] Step 4:

[0911] User: Starts exercising according to instructions. For example, the user is instructed to do specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[0912] Step 5:

[0913] Device: Collects user movement data in real time from sensors.

[0914] Step 6:

[0915] Terminal: Sends collected exercise data to the server.

[0916] Step 7:

[0917] Server: Analyzes the received motion data and generates feedback.

[0918] Step 8:

[0919] Server: Sends the generated feedback to the device.

[0920] Step 9:

[0921] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[0922] Cognitive Test Function

[0923] Step 1:

[0924] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[0925] Step 2:

[0926] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[0927] Step 3:

[0928] User: Follow the instructions to complete the cognitive test.

[0929] Step 4:

[0930] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[0931] Step 5:

[0932] Server: Analyzes the received test results and generates an evaluation report.

[0933] Step 6:

[0934] Server: Sends the generated evaluation report to the terminal.

[0935] Step 7:

[0936] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[0937] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo safe and effective rehabilitation at home or in a facility. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable dialogue experience for the user.

[0938] Example 2

[0939] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0940] The objective of this invention is to provide a system that integrates natural user interaction, exercise support, cognitive testing, and emotion recognition. Conventional systems often provide these functions separately, resulting in a fragmented user experience. Furthermore, emotion recognition accuracy is low, making it difficult to respond to the user's emotions. As a result, users experience low satisfaction and effectiveness.

[0941] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a voice recognition means for converting a user's voice input into text data, an emotion engine means for recognizing the user's emotions, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, and a voice synthesis means for converting the generated response into voice data and playing it back to the user. This not only enables the user to enjoy a natural dialogue experience, but also enables emotionally sensitive communication. The server also includes an exercise program generation means for collecting exercise data from the user and generating an exercise program provided by an expert, a display means for displaying the exercise program to the user and providing instructions, and a feedback means for transmitting the user's exercise data to the server and providing analysis results as feedback. The server also includes a cognitive test provision means for providing multiple types of cognitive tests, a data collection means for collecting results of the cognitive tests performed by the user, a report generation means for analyzing the test results and generating an evaluation report, and a display means for displaying the evaluation report to the user. This enables comprehensive health management, thereby increasing user satisfaction and effectiveness.

[0942] A "generative model" is an algorithm that uses artificial intelligence to generate natural-looking dialogue and responses based on given input data.

[0943] "User interface means" refers to a device or software that allows a user to interact with a system, including voice input and a display.

[0944] "Speech recognition means" refers to a technology or device that analyzes voice input from a user as digital data and converts it into character data.

[0945] "Emotion engine means" refers to a technology or device that identifies emotions from the user's voice or text and outputs the emotional state as data.

[0946] "Server Means" refers to a centralized computing resource that processes data for the entire system and generates responses based on the Generative Model.

[0947] "Speech synthesis means" refers to a technique or device that converts the generated text data into voice data and plays it back to the user.

[0948] The term "exercise program generation means" refers to a technology or device that creates an individually customized exercise program based on the user's exercise data and expert guidance.

[0949] "Display means" refers to technology or devices for visually presenting exercise instructions, cognitive test results, reports, etc. to the user.

[0950] "Feedback means" refers to technology or devices that analyze the user's exercise data and cognitive test results and provide advice on areas for improvement and next steps.

[0951] A "cognitive test delivery means" refers to a technology or device that delivers tests to a user to assess various cognitive functions.

[0952] "Data collection means" refers to sensors and software that collect user behavior and test results.

[0953] A "report generation means" is a technique or device that analyzes user data and compiles the results into a report.

[0954] MODE FOR CARRYING OUT THE INVENTION

[0955] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[0956] Natural conversation with AI

[0957] User voice input

[0958] The user speaks to the terminal, for example, "What would you like to talk about today?"

[0959] Speech-to-text

[0960] The device receives the user's voice through a microphone and acquires it as voice data. The Google Cloud Speech-to-Text API is used for voice recognition.

[0961] Speech and Emotion Recognition

[0962] The device converts voice data into text data using a voice recognition means. At the same time, it recognizes the user's emotions using an emotion engine means. The emotion recognition uses Microsoft Azure's emotion analysis API. The converted text data and the recognized emotion data are then sent to the server.

[0963] Response generation using AI models

[0964] The server generates an appropriate response to the received text data based on a generative AI model (e.g., ChatGPT). It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[0965] Playing back replies to the user

[0966] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[0967] Exercise support and advice features

[0968] Providing exercise programs

[0969] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[0970] Exercise instructions

[0971] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[0972] Exercise data collection and analysis

[0973] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[0974] Cognitive Test Function

[0975] Cognitive testing provided

[0976] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[0977] Cognitive testing

[0978] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[0979] Collecting and evaluating test results

[0980] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[0981] This not only allows users to enjoy a natural conversational experience, but also enables emotionally sensitive communication and comprehensive health management through exercise and cognitive training.

[0982] Prompt Sentence Examples

[0983] "Create prompts that are sensitive to the user's emotions and start the conversation around the weather. For example, 'The weather is nice today. Have you been out anywhere?'"

[0984] The above is an embodiment of the invention.

[0985] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0986] Step 1:

[0987] User voice input

[0988] Input: User speech

[0989] How it works: The user speaks to the device, asking something like, "What would you like to talk about today?"

[0990] Output: User's voice data

[0991] Step 2:

[0992] Speech-to-text

[0993] Input: User's voice data

[0994] How it works: The device captures the user's voice with its built-in microphone and converts the voice data into text using the Google Cloud Speech-to-Text API.

[0995] Output: Text data

[0996] Step 3:

[0997] emotion recognition

[0998] Input: Text data

[0999] How it works: The device uses Microsoft Azure's sentiment analysis API to recognize user emotions from text data.

[1000] Output: Emotion data

[1001] Step 4:

[1002] Sending data to the server

[1003] Input: Text data, emotion data

[1004] Operation: The device sends the converted text data and the recognized emotion data to the server.

[1005] Output: Data sent to the server

[1006] Step 5:

[1007] Response generation using AI models

[1008] Input: Text data and emotion data sent to the server

[1009] How it works: The server uses a generative AI model (e.g., ChatGPT) to generate appropriate responses based on the data it receives, and also adjusts the content and tone of the responses to match the user's emotions.

[1010] Output: Response text data

[1011] Step 6:

[1012] Speech synthesis and reply sending

[1013] Input: Response text data

[1014] Operation: The server converts the response text data into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[1015] Output: Audio data

[1016] Step 7:

[1017] Playing back replies to the user

[1018] Input: Audio data

[1019] Operation: The device receives the voice data from the server, plays it back through the speaker, and responds to the user. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[1020] Output: The reply speech the user hears

[1021] Step 8:

[1022] Providing exercise programs

[1023] Input: Exercise data provided by experts

[1024] Operation: The server uses the exercise program generation means to generate an individual exercise program based on the data provided by the expert, and transmits it to the terminal.

[1025] Output: Exercise program

[1026] Step 9:

[1027] Exercise instructions

[1028] Input: Exercise program

[1029] Actions: The device displays the received exercise program to the user visually or audibly, and provides specific exercise instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[1030] Output: Exercise instruction content

[1031] Step 10:

[1032] Exercise data collection and analysis

[1033] Input: User's exercise data

[1034] How it works: When a user exercises, the device collects exercise data through sensors attached to the device and transmits the data to a server, which analyzes the data and generates appropriate feedback.

[1035] Output: Analysis results and feedback

[1036] Step 11:

[1037] Cognitive testing provided

[1038] Input: Cognitive test selection instructions

[1039] Operation: The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1040] Output: Cognitive test content

[1041] Step 12:

[1042] Cognitive testing

[1043] Input: Cognitive test content

[1044] Operation: The device presents the received cognitive test to the user and has them complete it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[1045] Output: User's cognitive test results

[1046] Step 13:

[1047] Collecting and evaluating test results

[1048] Input: User's cognitive test results

[1049] Operation: The device sends the results of the cognitive tests taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the user's current cognitive function and areas for improvement, and is displayed to the user via the device.

[1050] Output: Evaluation report

[1051] (Application example 2)

[1052] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1053] There is a need for an effective method to support the dialogue between customers and store clerks in brick-and-mortar stores. In particular, there is a need for a system that can recognize customers' emotions and respond flexibly and naturally. Another challenge is to improve customer satisfaction by suggesting products that correspond to the customer's interests and emotions.

[1054] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1055] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into voice data and playing it back to the user, and a suggestion means for making product suggestions in the store. This allows for natural dialogue with customers and enables flexible and appropriate responses based on the customer's emotions. Furthermore, by making product suggestions based on the customer's interests, customer satisfaction can be improved.

[1056] A "generative model" is an artificial intelligence algorithm that generates appropriate responses or results based on input data from a user.

[1057] A "user interface means" is a means that provides an interface for a user to interact with the system.

[1058] The "voice recognition means" is a means having a function of converting voice input from a user into text data.

[1059] "Emotional data" is data collected to describe a user's emotional state.

[1060] The "server means" is a means for receiving text data and emotion data and generating an appropriate response based on the received data.

[1061] The "voice synthesis means" is a means having the function of converting the generated response into voice data and playing it back to the user.

[1062] The "proposal means" is a means having a function for proposing products and services to customers in a store.

[1063] The "exercise program generating means" is a means having a function of generating an appropriate exercise program based on data provided by an expert.

[1064] The "display means" is a means having a function for visually presenting the generated exercise program and other information to the user.

[1065] The "feedback means" is a means having a function for analyzing the user's exercise data and providing appropriate feedback.

[1066] The "cognitive test providing means" is a means having a function for providing multiple types of cognitive tests to the user.

[1067] The "data collection means" is a means having a function for collecting the results of the cognitive tests performed by the user.

[1068] The "report generation means" is a means having a function for analyzing the results of the cognitive test and generating an evaluation report.

[1069] The system that realizes this application example provides natural dialogue, exercise support, and cognitive testing using a generative model, and incorporates an emotion engine. A specific embodiment of this system will be described below.

[1070] Basic system configuration

[1071] The system consists of the following means:

[1072] User Interface Means

[1073] Voice recognition means

[1074] emotion recognition means

[1075] Server Means

[1076] Voice synthesis means

[1077] Proposal means

[1078] Exercise program generation means

[1079] Display means

[1080] Feedback Methods

[1081] Cognitive test delivery methods

[1082] Data collection methods

[1083] Report Generation Method

[1084] Natural conversational features

[1085] User voice input

[1086] The user speaks to a terminal connected to the system. For example, they might say, "Hello, do you have any product recommendations?" The terminal receives the voice through a microphone and captures it as voice data.

[1087] Speech-to-text and emotion recognition

[1088] The device converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text), and simultaneously recognizes the user's emotions using an emotion recognition engine (e.g., Microsoft Azure Emotion API). This data is then sent to the server.

[1089] Response generation using AI models

[1090] The server generates an appropriate response to the received text data based on a generative AI model (e.g., OpenAI GPT-3). It also adjusts the content and tone of the response based on emotional data. For example, if the user shows interest, it generates a friendly response such as, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?"

[1091] Playing back replies to the user

[1092] The generated response text is converted into voice data by a speech synthesis engine, and the device plays this voice data over a speaker to convey the response to the user.

[1093] Exercise support and advice features

[1094] Providing exercise programs

[1095] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1096] Exercise instructions and feedback

[1097] The device displays the received exercise program to the user visually or audibly and provides exercise instructions. During exercise, exercise data collected through sensors is sent to a server for analysis. Based on the analysis results, the device provides feedback such as, "It would be more effective if you raised the angle of your arms a little more."

[1098] Cognitive Test Function

[1099] Providing and administering cognitive tests

[1100] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal, for example, by informing the user that "The next test is a memory test. Please remember the words displayed on the screen."

[1101] Analysis and evaluation of test results

[1102] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1103] Specific examples

[1104] As examples of use in physical stores, the following specific dialogue scenarios are envisioned.

[1105] Usage

[1106] 1. User (Customer): "Hi, do you have any product recommendations?"

[1107] 2. Device (application running on smartphone):

[1108] Acquisition speech-to-text: "Hi, do you have any product recommendations?"

[1109] Emotion Recognition: "Curious" (Emotion Engine)

[1110] Response generation: "Hello, we have the latest smartwatch that's just for you. Would you like to take a look?" (Generative AI model)

[1111] Speech synthesis: Generated responses are spoken

[1112] Speaker Output: "Hello, we have the latest smartwatch just for you. Would you like to take a look?"

[1113] Prompt Sentence Examples

[1114] Generate relevant product suggestions when customers express interest: "Create an engaging product suggestion based on the following sentence: 'Hi, can you recommend something?' Keep the tone of your suggestion friendly but professional."

[1115] This system enables natural dialogue with customers in physical stores, and flexible and appropriate responses based on customer emotions. In addition, it can improve customer satisfaction by suggesting products based on the customer's interests.

[1116] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1117] Step 1:

[1118] User voice input

[1119] The user speaks to the device, for example, "Hello, do you have any product recommendations?" This speech is picked up by the device's microphone and saved as audio data.

[1120] Input: User's voice

[1121] Output: Audio data

[1122] Step 2:

[1123] Speech-to-text

[1124] The device uses a speech recognition engine to convert the acquired voice data into text data. For this purpose, it uses a speech recognition API (for example, Google Speech-to-Text).

[1125] Input: Audio data

[1126] Output: Text data

[1127] Step 3:

[1128] emotion recognition

[1129] The device sends the text data to an emotion recognition engine to recognize the user's emotions. This uses an emotion recognition API (for example, Microsoft Azure Emotion API). Emotion data reflects the user's psychological state.

[1130] Input: Text data

[1131] Output: Emotion data

[1132] Step 4:

[1133] Response generation based on generative models

[1134] The server receives the text data and emotion data and uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The tone and content of the response are adjusted based on the emotion data. For example, if the user shows interest, the response will be more friendly and engaging.

[1135] Input: Text data and emotion data

[1136] Output: Response text

[1137] Step 5:

[1138] Response text transcription

[1139] The server converts the response text into voice data using a speech synthesis engine. For this process, it uses a speech synthesis API (for example, Amazon Polly).

[1140] Input: Reply text

[1141] Output: Reply audio data

[1142] Step 6:

[1143] Playback of response audio data

[1144] The terminal plays the received response voice data over the speaker, thereby conveying the dialogue response to the user.

[1145] Input: Response audio data

[1146] Output: The audio played to the user

[1147] Step 7:

[1148] Product proposals using proposal methods

[1149] The server then suggests appropriate products based on the user's interests and emotions. For example, it generates a response like, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?" This product suggestion is made using a generative AI model that takes into account the user's emotional data.

[1150] Input: User interest and emotion data

[1151] Output: Product suggestion text

[1152] Through these steps, the system can realize natural dialogue with the user, responding flexibly and appropriately according to their emotions. Furthermore, it can improve customer satisfaction by suggesting products based on the user's interests.

[1153] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1154] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1155] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1156] [Third embodiment]

[1157] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1158] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1159] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1160] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1161] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1162] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1163] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1164] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1165] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1166] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1167] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1168] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1169] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1170] Natural conversation with AI

[1171] User voice input

[1172] The user speaks to the terminal, for example, asking a question such as, "What would you like to talk about today?"

[1173] Speech-to-text

[1174] The terminal receives the user's voice through a microphone and converts it into text data using a voice recognition means, and the converted text is sent to the server.

[1175] Response generation using AI models

[1176] The server generates an appropriate response to the received text data based on the generative model, converts the generated response text into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[1177] Playing back replies to the user

[1178] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[1179] Exercise support and advice features

[1180] Providing exercise programs

[1181] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1182] Exercise instructions

[1183] The device displays the received exercise program to the user visually or audibly and provides instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[1184] Exercise data collection and analysis

[1185] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[1186] Cognitive Test Function

[1187] Cognitive testing provided

[1188] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1189] Cognitive testing

[1190] The device presents the received cognitive test to the user and allows them to perform it. For example, the device starts the test by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[1191] Collecting and evaluating test results

[1192] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1193] Through these functions, this system aims to enable Alzheimer's disease patients and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and to maintain and improve their cognitive function through exercise and cognitive training. Furthermore, by utilizing a generative model, users can enjoy such a natural conversational experience that they hardly feel they are talking to an AI.

[1194] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[1195] The processing flow will be explained below.

[1196] Natural conversation with AI

[1197] Step 1:

[1198] User: Talks to the device, for example, asking, "What would you like to talk about today?"

[1199] Step 2:

[1200] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[1201] Step 3:

[1202] Terminal: Converts voice data into text data using a voice recognition means.

[1203] Step 4:

[1204] Terminal: Sends the converted text data to the server.

[1205] Step 5:

[1206] Server: Inputs the received text data into the generative model and generates an appropriate response.

[1207] Step 6:

[1208] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[1209] Step 7:

[1210] Terminal: Plays back the audio data and tells the user a response. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[1211] Exercise support and advice features

[1212] Step 1:

[1213] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[1214] Step 2:

[1215] Server: Sends the generated exercise program to the terminal.

[1216] Step 3:

[1217] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[1218] Step 4:

[1219] User: Begins exercise as instructed, for example, performing upper body stretches.

[1220] Step 5:

[1221] Device: Collects user movement data in real time from sensors.

[1222] Step 6:

[1223] Terminal: Sends collected exercise data to the server.

[1224] Step 7:

[1225] Server: Analyzes the received motion data and generates feedback.

[1226] Step 8:

[1227] Server: Sends the generated feedback to the device.

[1228] Step 9:

[1229] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[1230] Cognitive Test Function

[1231] Step 1:

[1232] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[1233] Step 2:

[1234] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[1235] Step 3:

[1236] User: Follow the instructions to complete the cognitive test.

[1237] Step 4:

[1238] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[1239] Step 5:

[1240] Server: Analyzes the received test results and generates an evaluation report.

[1241] Step 6:

[1242] Server: Sends the generated evaluation report to the terminal.

[1243] Step 7:

[1244] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[1245] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo rehabilitation safely and effectively at home or in a facility.

[1246] Example 1

[1247] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1248] As our society ages, there is an increasing demand for systems to prevent and improve cognitive and motor decline. However, current systems often have unnatural dialogue with users and have difficulty providing appropriate exercise programs and cognitive tests. Furthermore, the feedback these systems provide is not immediate, making it difficult to motivate users or provide optimal support. Therefore, there is a need for systems that enable natural dialogue, provide individually tailored exercise programs and cognitive tests, and evaluate and provide feedback on the results in real time.

[1249] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1250] In this invention, the server includes a speech recognition means for converting speech input into text data, a server means for generating appropriate responses based on a generative model, a speech synthesis means for converting the generated responses into speech and playing them back to the user, an exercise program generation means for collecting exercise data from the user and generating an exercise program based on expert data, a cognitive test provision means for providing multiple types of cognitive tests, and a display means for generating an evaluation report and displaying it to the user. This enables natural dialogue and immediate feedback, increasing the user's motivation while providing effective exercise support and cognitive function evaluation.

[1251] A "voice recognition means" is a device or technique for converting voice input from a user into text data.

[1252] The "server means" is a device or system that utilizes a generative model to generate an appropriate response based on received text data.

[1253] "Speech synthesis means" refers to a device or technology that converts the generated text data into speech data and plays it back to the user.

[1254] The "user interface means" is a device or technology that cooperates with the server to provide operability and information to the user.

[1255] The "exercise program generating means" is a device or system that generates an individual exercise program based on collected exercise data and expert data.

[1256] The "display means" refers to a device or technology for visually presenting the generated exercise program and various information to the user.

[1257] A "feedback means" is a device or technology for analyzing collected data and providing appropriate feedback to the user.

[1258] The "cognitive test providing means" is a device or system that selects from multiple types of cognitive tests and provides them to the user.

[1259] A "data collection tool" is a device or technology that collects the results of a cognitive test or exercise performed by a user.

[1260] A "report generation means" is a device or system for analyzing collected data and generating an evaluation report.

[1261] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1262] Natural conversation with AI

[1263] Voice input and conversion

[1264] The user speaks into the device's microphone. For example, they might ask, "What would you like to talk about today?" The device receives the user's voice through the microphone and converts the voice data into text data using a speech recognition method (e.g., Google Cloud Speech-to-Text). The converted text is then sent to the server.

[1265] Response generation and playback

[1266] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data. This response text is converted into voice data using a speech synthesis method (e.g., Amazon Polly) and sent to the device. The device then plays the voice data received from the server and conveys the response to the user. For example, a natural conversation follows, such as, "The weather is nice today. Have you gone out anywhere?"

[1267] Exercise support and advice features

[1268] Creation and provision of exercise programs

[1269] The server uses an exercise program generation means to generate an individual exercise program based on the expert's data. This exercise program is then sent to the terminal. The terminal then displays the received exercise program to the user visually or audibly, and provides specific instructions. For example, the terminal may instruct the user on specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[1270] Exercise data collection and analysis

[1271] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit). This data is sent to a server, which analyzes the collected data and generates appropriate feedback. The feedback includes specific suggestions for improvement, such as "It would be more effective if you angled your arms a little higher."

[1272] Cognitive Test Function

[1273] Providing and administering cognitive tests

[1274] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal. The terminal then presents the received cognitive test to the user and has them perform it. For example, the test can be started by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[1275] Collecting and evaluating test results

[1276] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1277] Specific examples of operation

[1278] Natural conversation examples

[1279] User voice input: The user speaks into the device, "What would you like to talk about today?"

[1280] AI response generation and playback: The response from the server is sent to the device, and a voice message saying, "The weather is nice today. Have you gone out anywhere?" is played.

[1281] Exercise support examples

[1282] Exercise instructions: The device will give you visual or audio instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[1283] Feedback: Based on the data collected during exercise, feedback is sent from the server to the device, and specific advice such as "It would be more effective if you raised the angle of your arms a little more" is displayed.

[1284] Cognitive test examples

[1285] Test instructions: The device will give instructions such as "The next test is a memory test. Please remember the words that appear on the screen," and then carry out the test.

[1286] Presentation of evaluation report: After the user completes the test, the evaluation report analyzed by the server is sent to the terminal and displayed to the user.

[1287] Example prompts for generative AI models

[1288] "When a user asks, 'What would you like to talk about today?' the AI ​​responds naturally with everyday topics."

[1289] "It tells users, 'Today, stretch your upper body. Extend your arms in front of you,' and provides specific exercise advice."

[1290] "During a memory test, you are instructed to 'remember the following words,' and the test results are collected, evaluated, and a report is generated."

[1291] The above is a mode for carrying out the present invention, and shows that natural dialogue, exercise support, and cognitive testing can be effectively provided through specific processing steps.

[1292] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1293] Step 1: Receiving voice input

[1294] The user speaks to the terminal, asking questions such as, "What would you like to talk about today?"

[1295] Input: User's voice data

[1296] Output: Audio data collected by the device's microphone

[1297] Specific operation: The user speaks into the device's microphone, and voice data is input into the device.

[1298] Step 2: Speech to text

[1299] The terminal converts the user's voice data into text data using a voice recognition means (e.g., Google Cloud Speech-to-Text).

[1300] Input: User's voice data

[1301] Output: Text data

[1302] Specific operation: The terminal's voice recognition means analyzes the voice data and converts it into corresponding text.

[1303] Step 3: Send text data

[1304] The terminal transmits the converted text data to the server.

[1305] Input: Text data

[1306] Output: Text data sent to the server

[1307] Specific operation: The device sends text data to a server via the Internet.

[1308] Step 4: Generate response text

[1309] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data.

[1310] Input: Text data

[1311] Output: Response text data

[1312] How it works: The server's AI model analyzes the text data and generates an appropriate response, such as "The weather is nice today. Have you gone out anywhere?"

[1313] Step 5: Convert response text to speech

[1314] The server converts the generated response text into voice data using a speech synthesis means (e.g., Amazon Polly) and sends it to the terminal.

[1315] Input: Response text data

[1316] Output: Audio data

[1317] Specific operation: The server's speech synthesis means converts the response text into speech and transmits the speech data to the terminal.

[1318] Step 6: Play the reply audio

[1319] The terminal reproduces the voice data received from the server and conveys the response to the user.

[1320] Input: Audio data

[1321] Output: Audio played from the device

[1322] Specific operation: The device speaker plays the audio data and conveys the response to the user.

[1323] Step 7: Generate the exercise program

[1324] The server uses an exercise program generating means to generate an individual exercise program based on the expert data.

[1325] Input: Expert data and user feedback data

[1326] Output: personalized exercise program

[1327] Specific operation: The server uses the exercise program generation means to design an individual exercise program based on the collected data.

[1328] Step 8: Providing an exercise program

[1329] The terminal displays the received exercise program to the user visually or audibly, and provides specific instructions.

[1330] Input: Exercise program data

[1331] Output: An exercise program provided to the user

[1332] Specific actions: The device displays an exercise program on the screen or gives instructions to the user by voice. For example, it may say, "Today, let's stretch your upper body. Stretch your arms out in front of you."

[1333] Step 9: Collect exercise data

[1334] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit).

[1335] Input: User's exercise data

[1336] Output: Collected exercise data

[1337] Specific operation: The sensor collects the user's movement data in real time and transmits it to the device.

[1338] Step 10: Analyze the movement data

[1339] The server analyzes the collected movement data and generates appropriate feedback.

[1340] Input: Collected exercise data

[1341] Output: Feedback data

[1342] Specific movements: The server analyzes the movement data and generates specific feedback, such as advice like "It would be more effective if you angled your arms a little higher."

[1343] Step 11: Offer cognitive testing

[1344] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1345] Input: User context data

[1346] Output: Cognitive test data

[1347] Specific operation: The server selects the appropriate cognitive test from multiple options and sends it to the device.

[1348] Step 12: Conduct cognitive testing

[1349] The terminal presents the received cognitive test to the user and allows the user to perform the test.

[1350] Input: Cognitive test data

[1351] Output: Test result data performed by the user

[1352] Specific operation: The device displays a cognitive test on the screen and asks the user to complete it. For example, it instructs the user, "The next test is a memory test. Please remember the words displayed on the screen."

[1353] Step 13: Collecting test results

[1354] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means.

[1355] Input: User test result data

[1356] Output: Test result data sent to the server

[1357] Specific operation: The device collects the test results and sends them to the server.

[1358] Step 14: Analyze test results and generate evaluation reports

[1359] The server analyzes the results and generates an evaluation report, which includes the user's current cognitive function and areas for improvement.

[1360] Input: Test result data

[1361] Output: Evaluation report

[1362] Specific operation: The server analyzes the test result data and generates an evaluation report.

[1363] Step 15: Provide an evaluation report

[1364] The terminal displays the evaluation report to the user.

[1365] Input: Evaluation report data

[1366] Output: Evaluation report provided to the user

[1367] Specific operation: The device displays the evaluation report on the screen and shows it to the user.

[1368] (Application example 1)

[1369] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1370] Conventional online shopping systems are often cumbersome, requiring complex operations such as product searches, ordering, and return procedures, making them particularly difficult for elderly users and those with low IT literacy. Another issue is the lack of product suggestions based on individual user preferences and body types, making it time-consuming to find the perfect product. Furthermore, there is a lack of systems that can routinely conduct cognitive tests on users with concerns about cognitive decline.

[1371] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1372] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into speech data and playing it back to the user, a means for providing functions that allow the user to search for, order, and process returns of products by voice, and a means for suggesting recommended items based on the user's body type and preferences. This allows the user to easily operate the shop by voice, makes it possible to suggest products that suit the user's individual needs, and also maintains and improves cognitive function.

[1373] A "generative model" is an artificial intelligence model that generates new text or information based on given text or data.

[1374] "User interface means" are technical means that provide an interface through which a user can directly interact with a system.

[1375] "Speech recognition means" refers to technical means for converting speech input into text data.

[1376] "Server means" refers to technical means that provides the functionality of a server that receives text data and generates a response based on a generative model.

[1377] "Speech synthesis means" refers to technical means for converting text data into voice data and playing it back to the user.

[1378] "Means for providing functions for searching for, ordering, and processing returns of products" refers to technical means that allow users to search for, order, and process returns of products using voice commands.

[1379] The "means for proposing recommended items" refers to a technical means for proposing appropriate products according to the user's body type and preferences.

[1380] This invention is a virtual assistant system that integrates natural dialogue using generative models with online shopping functions. The system includes a server means that uses speech recognition, speech synthesis, generative models, and a user interface means.

[1381] Hardware and Software Configuration

[1382] Hardware:

[1383] Smartphone (iOS, Android compatible)

[1384] software:

[1385] Speech Recognition API (Google Cloud Speech-to-Text)

[1386] Generative AI model (GPT-4)

[1387] Speech synthesis API (Google Cloud Text-to-Speech)

[1388] System Operation

[1389] Voice input and recognition:

[1390] When a user speaks into a smartphone, the built-in microphone picks up the voice, and this voice data is converted into text data in real time via a speech recognition API.

[1391] Sending data to the server and generating a model response:

[1392] The converted text data is sent to a server, where a generative AI model (GPT-4) generates natural-sounding responses based on the user's speech. This response text is then converted back into voice data using a speech synthesis API and sent back to the smartphone.

[1393] Response playback in the user interface:

[1394] The smartphone plays back the returned voice data and provides natural and appropriate responses to the user. For example, if the user says, "I'm looking for a black dress," the speech recognition API converts this into text, and the generative AI model responds, "I'll show you the search results for black dresses."

[1395] Product search, suggestions, and ordering:

[1396] The server also allows users to search for products, place orders, and process returns using voice commands, and it also provides recommendations based on the user's body type and preferences, helping users find the perfect product.

[1397] Examples and prompts

[1398] For example, if a user says, "I'm looking for a black dress," a voice-synthesized message will be played in response, such as, "Here are the search results for black dresses. There are several options." This is followed by, "Today's recommended items are also displayed," allowing the user to enjoy shopping more intuitively through voice.

[1399] Example prompt sentence:

[1400] "A virtual shopping assistant, recommending suitable products for users looking for black dresses."

[1401] This system makes online shopping more natural and user-friendly, and is suitable for elderly people and those with low IT literacy. Furthermore, for users who are concerned about declining cognitive function, it also provides a daily cognitive test function, helping them manage their health.

[1402] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1403] Step 1:

[1404] The user speaks into the smartphone. The smartphone's built-in microphone receives the voice and collects this voice data. The input is the user's voice, and the output is voice data.

[1405] Step 2:

[1406] The device sends the collected voice data to a voice recognition API (Google Cloud Speech-to-Text) in real time, which converts the voice data into text data. The input is voice data, and the output is text data. This process recognizes the voice and converts it into text.

[1407] Step 3:

[1408] Text data is sent from the terminal to the server. When the server receives the text data, the input is text data and the output is the text data sent to the server. This process prepares the data to be used in the next step.

[1409] Step 4:

[1410] The server feeds the received text data into a generative AI model (GPT-4) to generate an appropriate response. The input is text data, and the output is the generated response text. This data processing enables natural dialogue.

[1411] Step 5:

[1412] The generated response text is converted into audio data on the server side using a speech synthesis API (Google Cloud Text-to-Speech). The input is the response text and the output is audio data. This process provides the response in audio format.

[1413] Step 6:

[1414] The generated voice data is sent from the server to the terminal. The user can continue the conversation by receiving and playing the voice data. The input is the voice data from the server, and the output is the played voice.

[1415] Step 7:

[1416] When a user wants to search for a product, the server will suggest recommended items based on the user's body type and preferences. The input is the user's detailed information and search criteria, and the output is a list of recommended products. The server generates this and sends it to the device.

[1417] Step 8:

[1418] The terminal displays the received product list to the user, who then uses voice commands to order products or perform other operations. The input is the user's voice commands, and the output is the results of operations such as ordering. Specifically, if the user says, "I'm looking for a black dress," the terminal responds, "I'll show you the search results for black dresses."

[1419] This series of steps allows users to easily shop through voice, with data processing and calculations at each step providing a highly interactive and personalized shopping experience.

[1420] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1421] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1422] Natural conversation with AI

[1423] User voice input

[1424] The user speaks to the terminal, for example, "What would you like to talk about today?"

[1425] Speech-to-text

[1426] The terminal receives the user's voice through a microphone and acquires it as voice data.

[1427] Speech and Emotion Recognition

[1428] The terminal converts the voice data into text data using a voice recognition means, and simultaneously recognizes the user's emotion using an emotion engine. The converted text and the recognized emotion data are sent to the server.

[1429] Response generation using AI models

[1430] The server generates an appropriate response to the received text data based on the generative model. It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[1431] Playing back replies to the user

[1432] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[1433] Exercise support and advice features

[1434] Providing exercise programs

[1435] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1436] Exercise instructions

[1437] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[1438] Exercise data collection and analysis

[1439] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[1440] Cognitive Test Function

[1441] Cognitive testing provided

[1442] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1443] Cognitive testing

[1444] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[1445] Collecting and evaluating test results

[1446] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1447] Through these functions, this system aims to maintain and improve cognitive function by enabling people with Alzheimer's disease and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and by providing exercise and cognitive training. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable conversational experience for the user.

[1448] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[1449] The processing flow will be explained below.

[1450] Natural conversation with AI

[1451] Step 1:

[1452] User: Speaks into the device. For example, "What would you like to talk about today?"

[1453] Step 2:

[1454] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[1455] Step 3:

[1456] Terminal: Converts voice data into text data using a voice recognition means.

[1457] Step 4:

[1458] Terminal: Uses an emotion engine to recognize emotions from the user's voice.

[1459] Step 5:

[1460] Terminal: Transmits the converted text data and recognized emotion data to the server.

[1461] Step 6:

[1462] Server: Inputs the received text data into a generative model to generate an appropriate response. Based on the emotional data, the tone and content of the response are adjusted.

[1463] Step 7:

[1464] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[1465] Step 8:

[1466] Terminal: Plays back the audio data and responds to the user. For example, if the user is sad, the response will be something comforting like, "What happened today? Tell me."

[1467] Exercise support and advice features

[1468] Step 1:

[1469] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[1470] Step 2:

[1471] Server: Sends the generated exercise program to the terminal.

[1472] Step 3:

[1473] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[1474] Step 4:

[1475] User: Starts exercising according to instructions. For example, the user is instructed to do specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[1476] Step 5:

[1477] Device: Collects user movement data in real time from sensors.

[1478] Step 6:

[1479] Terminal: Sends collected exercise data to the server.

[1480] Step 7:

[1481] Server: Analyzes the received motion data and generates feedback.

[1482] Step 8:

[1483] Server: Sends the generated feedback to the device.

[1484] Step 9:

[1485] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[1486] Cognitive Test Function

[1487] Step 1:

[1488] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[1489] Step 2:

[1490] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[1491] Step 3:

[1492] User: Follow the instructions to complete the cognitive test.

[1493] Step 4:

[1494] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[1495] Step 5:

[1496] Server: Analyzes the received test results and generates an evaluation report.

[1497] Step 6:

[1498] Server: Sends the generated evaluation report to the terminal.

[1499] Step 7:

[1500] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[1501] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo safe and effective rehabilitation at home or in a facility. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable dialogue experience for the user.

[1502] Example 2

[1503] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1504] The objective of this invention is to provide a system that integrates natural user interaction, exercise support, cognitive testing, and emotion recognition. Conventional systems often provide these functions separately, resulting in a fragmented user experience. Furthermore, emotion recognition accuracy is low, making it difficult to respond to the user's emotions. As a result, users experience low satisfaction and effectiveness.

[1505] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a voice recognition means for converting a user's voice input into text data, an emotion engine means for recognizing the user's emotions, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, and a voice synthesis means for converting the generated response into voice data and playing it back to the user. This not only enables the user to enjoy a natural dialogue experience, but also enables emotionally sensitive communication. The server also includes an exercise program generation means for collecting exercise data from the user and generating an exercise program provided by an expert, a display means for displaying the exercise program to the user and providing instructions, and a feedback means for transmitting the user's exercise data to the server and providing analysis results as feedback. The server also includes a cognitive test provision means for providing multiple types of cognitive tests, a data collection means for collecting results of the cognitive tests performed by the user, a report generation means for analyzing the test results and generating an evaluation report, and a display means for displaying the evaluation report to the user. This enables comprehensive health management, thereby increasing user satisfaction and effectiveness.

[1506] A "generative model" is an algorithm that uses artificial intelligence to generate natural-looking dialogue and responses based on given input data.

[1507] "User interface means" refers to a device or software that allows a user to interact with a system, including voice input and a display.

[1508] "Speech recognition means" refers to a technology or device that analyzes voice input from a user as digital data and converts it into character data.

[1509] "Emotion engine means" refers to a technology or device that identifies emotions from the user's voice or text and outputs the emotional state as data.

[1510] "Server Means" refers to a centralized computing resource that processes data for the entire system and generates responses based on the Generative Model.

[1511] "Speech synthesis means" refers to a technique or device that converts the generated text data into voice data and plays it back to the user.

[1512] The term "exercise program generation means" refers to a technology or device that creates an individually customized exercise program based on the user's exercise data and expert guidance.

[1513] "Display means" refers to technology or devices for visually presenting exercise instructions, cognitive test results, reports, etc. to the user.

[1514] "Feedback means" refers to technology or devices that analyze the user's exercise data and cognitive test results and provide advice on areas for improvement and next steps.

[1515] A "cognitive test delivery means" refers to a technology or device that delivers tests to a user to assess various cognitive functions.

[1516] "Data collection means" refers to sensors and software that collect user behavior and test results.

[1517] A "report generation means" is a technique or device that analyzes user data and compiles the results into a report.

[1518] MODE FOR CARRYING OUT THE INVENTION

[1519] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1520] Natural conversation with AI

[1521] User voice input

[1522] The user speaks to the terminal, for example, "What would you like to talk about today?"

[1523] Speech-to-text

[1524] The device receives the user's voice through a microphone and acquires it as voice data. The Google Cloud Speech-to-Text API is used for voice recognition.

[1525] Speech and Emotion Recognition

[1526] The device converts voice data into text data using a voice recognition means. At the same time, it recognizes the user's emotions using an emotion engine means. The emotion recognition uses Microsoft Azure's emotion analysis API. The converted text data and the recognized emotion data are then sent to the server.

[1527] Response generation using AI models

[1528] The server generates an appropriate response to the received text data based on a generative AI model (e.g., ChatGPT). It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[1529] Playing back replies to the user

[1530] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[1531] Exercise support and advice features

[1532] Providing exercise programs

[1533] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1534] Exercise instructions

[1535] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[1536] Exercise data collection and analysis

[1537] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[1538] Cognitive Test Function

[1539] Cognitive testing provided

[1540] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1541] Cognitive testing

[1542] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[1543] Collecting and evaluating test results

[1544] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1545] This not only allows users to enjoy a natural conversational experience, but also enables emotionally sensitive communication and comprehensive health management through exercise and cognitive training.

[1546] Prompt Sentence Examples

[1547] "Create prompts that are sensitive to the user's emotions and start the conversation around the weather. For example, 'The weather is nice today. Have you been out anywhere?'"

[1548] The above is an embodiment of the invention.

[1549] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1550] Step 1:

[1551] User voice input

[1552] Input: User speech

[1553] How it works: The user speaks to the device, asking something like, "What would you like to talk about today?"

[1554] Output: User's voice data

[1555] Step 2:

[1556] Speech-to-text

[1557] Input: User's voice data

[1558] How it works: The device captures the user's voice with its built-in microphone and converts the voice data into text using the Google Cloud Speech-to-Text API.

[1559] Output: Text data

[1560] Step 3:

[1561] emotion recognition

[1562] Input: Text data

[1563] How it works: The device uses Microsoft Azure's sentiment analysis API to recognize user emotions from text data.

[1564] Output: Emotion data

[1565] Step 4:

[1566] Sending data to the server

[1567] Input: Text data, emotion data

[1568] Operation: The device sends the converted text data and the recognized emotion data to the server.

[1569] Output: Data sent to the server

[1570] Step 5:

[1571] Response generation using AI models

[1572] Input: Text data and emotion data sent to the server

[1573] How it works: The server uses a generative AI model (e.g., ChatGPT) to generate appropriate responses based on the data it receives, and also adjusts the content and tone of the responses to match the user's emotions.

[1574] Output: Response text data

[1575] Step 6:

[1576] Speech synthesis and reply sending

[1577] Input: Response text data

[1578] Operation: The server converts the response text data into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[1579] Output: Audio data

[1580] Step 7:

[1581] Playing back replies to the user

[1582] Input: Audio data

[1583] Operation: The device receives the voice data from the server, plays it back through the speaker, and responds to the user. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[1584] Output: The reply speech the user hears

[1585] Step 8:

[1586] Providing exercise programs

[1587] Input: Exercise data provided by experts

[1588] Operation: The server uses the exercise program generation means to generate an individual exercise program based on the data provided by the expert, and transmits it to the terminal.

[1589] Output: Exercise program

[1590] Step 9:

[1591] Exercise instructions

[1592] Input: Exercise program

[1593] Actions: The device displays the received exercise program to the user visually or audibly, and provides specific exercise instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[1594] Output: Exercise instruction content

[1595] Step 10:

[1596] Exercise data collection and analysis

[1597] Input: User's exercise data

[1598] How it works: When a user exercises, the device collects exercise data through sensors attached to the device and transmits the data to a server, which analyzes the data and generates appropriate feedback.

[1599] Output: Analysis results and feedback

[1600] Step 11:

[1601] Cognitive testing provided

[1602] Input: Cognitive test selection instructions

[1603] Operation: The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1604] Output: Cognitive test content

[1605] Step 12:

[1606] Cognitive testing

[1607] Input: Cognitive test content

[1608] Operation: The device presents the received cognitive test to the user and has them complete it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[1609] Output: User's cognitive test results

[1610] Step 13:

[1611] Collecting and evaluating test results

[1612] Input: User's cognitive test results

[1613] Operation: The device sends the results of the cognitive tests taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the user's current cognitive function and areas for improvement, and is displayed to the user via the device.

[1614] Output: Evaluation report

[1615] (Application example 2)

[1616] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1617] There is a need for an effective method to support the dialogue between customers and store clerks in brick-and-mortar stores. In particular, there is a need for a system that can recognize customers' emotions and respond flexibly and naturally. Another challenge is to improve customer satisfaction by suggesting products that correspond to the customer's interests and emotions.

[1618] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1619] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into voice data and playing it back to the user, and a suggestion means for making product suggestions in the store. This allows for natural dialogue with customers and enables flexible and appropriate responses based on the customer's emotions. Furthermore, by making product suggestions based on the customer's interests, customer satisfaction can be improved.

[1620] A "generative model" is an artificial intelligence algorithm that generates appropriate responses or results based on input data from a user.

[1621] A "user interface means" is a means that provides an interface for a user to interact with the system.

[1622] The "voice recognition means" is a means having a function of converting voice input from a user into text data.

[1623] "Emotional data" is data collected to describe a user's emotional state.

[1624] The "server means" is a means for receiving text data and emotion data and generating an appropriate response based on the received data.

[1625] The "voice synthesis means" is a means having the function of converting the generated response into voice data and playing it back to the user.

[1626] The "proposal means" is a means having a function for proposing products and services to customers in a store.

[1627] The "exercise program generating means" is a means having a function of generating an appropriate exercise program based on data provided by an expert.

[1628] The "display means" is a means having a function for visually presenting the generated exercise program and other information to the user.

[1629] The "feedback means" is a means having a function for analyzing the user's exercise data and providing appropriate feedback.

[1630] The "cognitive test providing means" is a means having a function for providing multiple types of cognitive tests to the user.

[1631] The "data collection means" is a means having a function for collecting the results of the cognitive tests performed by the user.

[1632] The "report generation means" is a means having a function for analyzing the results of the cognitive test and generating an evaluation report.

[1633] The system that realizes this application example provides natural dialogue, exercise support, and cognitive testing using a generative model, and incorporates an emotion engine. A specific embodiment of this system will be described below.

[1634] Basic system configuration

[1635] The system consists of the following means:

[1636] User Interface Means

[1637] Voice recognition means

[1638] emotion recognition means

[1639] Server Means

[1640] Voice synthesis means

[1641] Proposal means

[1642] Exercise program generation means

[1643] Display means

[1644] Feedback Methods

[1645] Cognitive test delivery methods

[1646] Data collection methods

[1647] Report Generation Method

[1648] Natural conversational features

[1649] User voice input

[1650] The user speaks to a terminal connected to the system. For example, they might say, "Hello, do you have any product recommendations?" The terminal receives the voice through a microphone and captures it as voice data.

[1651] Speech-to-text and emotion recognition

[1652] The device converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text), and simultaneously recognizes the user's emotions using an emotion recognition engine (e.g., Microsoft Azure Emotion API). This data is then sent to the server.

[1653] Response generation using AI models

[1654] The server generates an appropriate response to the received text data based on a generative AI model (e.g., OpenAI GPT-3). It also adjusts the content and tone of the response based on emotional data. For example, if the user shows interest, it generates a friendly response such as, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?"

[1655] Playing back replies to the user

[1656] The generated response text is converted into voice data by a speech synthesis engine, and the device plays this voice data over a speaker to convey the response to the user.

[1657] Exercise support and advice features

[1658] Providing exercise programs

[1659] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1660] Exercise instructions and feedback

[1661] The device displays the received exercise program to the user visually or audibly and provides exercise instructions. During exercise, exercise data collected through sensors is sent to a server for analysis. Based on the analysis results, the device provides feedback such as, "It would be more effective if you raised the angle of your arms a little more."

[1662] Cognitive Test Function

[1663] Providing and administering cognitive tests

[1664] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal, for example, by informing the user that "The next test is a memory test. Please remember the words displayed on the screen."

[1665] Analysis and evaluation of test results

[1666] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1667] Specific examples

[1668] As examples of use in physical stores, the following specific dialogue scenarios are envisioned.

[1669] Usage

[1670] 1. User (Customer): "Hi, do you have any product recommendations?"

[1671] 2. Device (application running on smartphone):

[1672] Acquisition speech-to-text: "Hi, do you have any product recommendations?"

[1673] Emotion Recognition: "Curious" (Emotion Engine)

[1674] Response generation: "Hello, we have the latest smartwatch that's just for you. Would you like to take a look?" (Generative AI model)

[1675] Speech synthesis: Generated responses are spoken

[1676] Speaker Output: "Hello, we have the latest smartwatch just for you. Would you like to take a look?"

[1677] Prompt Sentence Examples

[1678] Generate relevant product suggestions when customers express interest: "Create an engaging product suggestion based on the following sentence: 'Hi, can you recommend something?' Keep the tone of your suggestion friendly but professional."

[1679] This system enables natural dialogue with customers in physical stores, and flexible and appropriate responses based on customer emotions. In addition, it can improve customer satisfaction by suggesting products based on the customer's interests.

[1680] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1681] Step 1:

[1682] User voice input

[1683] The user speaks to the device, for example, "Hello, do you have any product recommendations?" This speech is picked up by the device's microphone and saved as audio data.

[1684] Input: User's voice

[1685] Output: Audio data

[1686] Step 2:

[1687] Speech-to-text

[1688] The device uses a speech recognition engine to convert the acquired voice data into text data. For this purpose, it uses a speech recognition API (for example, Google Speech-to-Text).

[1689] Input: Audio data

[1690] Output: Text data

[1691] Step 3:

[1692] emotion recognition

[1693] The device sends the text data to an emotion recognition engine to recognize the user's emotions. This uses an emotion recognition API (for example, Microsoft Azure Emotion API). Emotion data reflects the user's psychological state.

[1694] Input: Text data

[1695] Output: Emotion data

[1696] Step 4:

[1697] Response generation based on generative models

[1698] The server receives the text data and emotion data and uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The tone and content of the response are adjusted based on the emotion data. For example, if the user shows interest, the response will be more friendly and engaging.

[1699] Input: Text data and emotion data

[1700] Output: Response text

[1701] Step 5:

[1702] Response text transcription

[1703] The server converts the response text into voice data using a speech synthesis engine. For this process, it uses a speech synthesis API (for example, Amazon Polly).

[1704] Input: Reply text

[1705] Output: Reply audio data

[1706] Step 6:

[1707] Playback of response audio data

[1708] The terminal plays the received response voice data over the speaker, thereby conveying the dialogue response to the user.

[1709] Input: Response audio data

[1710] Output: The audio played to the user

[1711] Step 7:

[1712] Product proposals using proposal methods

[1713] The server then suggests appropriate products based on the user's interests and emotions. For example, it generates a response like, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?" This product suggestion is made using a generative AI model that takes into account the user's emotional data.

[1714] Input: User interest and emotion data

[1715] Output: Product suggestion text

[1716] Through these steps, the system can realize natural dialogue with the user, responding flexibly and appropriately according to their emotions. Furthermore, it can improve customer satisfaction by suggesting products based on the user's interests.

[1717] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1718] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1719] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1720] [Fourth embodiment]

[1721] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1722] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1723] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1724] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1725] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1726] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1727] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1728] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1729] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1730] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1731] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1732] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1733] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1734] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1735] Natural conversation with AI

[1736] User voice input

[1737] The user speaks to the terminal, for example, asking a question such as, "What would you like to talk about today?"

[1738] Speech-to-text

[1739] The terminal receives the user's voice through a microphone and converts it into text data using a voice recognition means, and the converted text is sent to the server.

[1740] Response generation using AI models

[1741] The server generates an appropriate response to the received text data based on the generative model, converts the generated response text into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[1742] Playing back replies to the user

[1743] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[1744] Exercise support and advice features

[1745] Providing exercise programs

[1746] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[1747] Exercise instructions

[1748] The device displays the received exercise program to the user visually or audibly and provides instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[1749] Exercise data collection and analysis

[1750] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[1751] Cognitive Test Function

[1752] Cognitive testing provided

[1753] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1754] Cognitive testing

[1755] The device presents the received cognitive test to the user and allows them to perform it. For example, the device starts the test by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[1756] Collecting and evaluating test results

[1757] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1758] Through these functions, this system aims to enable Alzheimer's disease patients and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and to maintain and improve their cognitive function through exercise and cognitive training. Furthermore, by utilizing a generative model, users can enjoy such a natural conversational experience that they hardly feel they are talking to an AI.

[1759] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[1760] The processing flow will be explained below.

[1761] Natural conversation with AI

[1762] Step 1:

[1763] User: Talks to the device, for example, asking, "What would you like to talk about today?"

[1764] Step 2:

[1765] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[1766] Step 3:

[1767] Terminal: Converts voice data into text data using a voice recognition means.

[1768] Step 4:

[1769] Terminal: Sends the converted text data to the server.

[1770] Step 5:

[1771] Server: Inputs the received text data into the generative model and generates an appropriate response.

[1772] Step 6:

[1773] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[1774] Step 7:

[1775] Terminal: Plays back the audio data and tells the user a response. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[1776] Exercise support and advice features

[1777] Step 1:

[1778] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[1779] Step 2:

[1780] Server: Sends the generated exercise program to the terminal.

[1781] Step 3:

[1782] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[1783] Step 4:

[1784] User: Begins exercise as instructed, for example, performing upper body stretches.

[1785] Step 5:

[1786] Device: Collects user movement data in real time from sensors.

[1787] Step 6:

[1788] Terminal: Sends collected exercise data to the server.

[1789] Step 7:

[1790] Server: Analyzes the received motion data and generates feedback.

[1791] Step 8:

[1792] Server: Sends the generated feedback to the device.

[1793] Step 9:

[1794] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[1795] Cognitive Test Function

[1796] Step 1:

[1797] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[1798] Step 2:

[1799] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[1800] Step 3:

[1801] User: Follow the instructions to complete the cognitive test.

[1802] Step 4:

[1803] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[1804] Step 5:

[1805] Server: Analyzes the received test results and generates an evaluation report.

[1806] Step 6:

[1807] Server: Sends the generated evaluation report to the terminal.

[1808] Step 7:

[1809] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[1810] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo rehabilitation safely and effectively at home or in a facility.

[1811] Example 1

[1812] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1813] As our society ages, there is an increasing demand for systems to prevent and improve cognitive and motor decline. However, current systems often have unnatural dialogue with users and have difficulty providing appropriate exercise programs and cognitive tests. Furthermore, the feedback these systems provide is not immediate, making it difficult to motivate users or provide optimal support. Therefore, there is a need for systems that enable natural dialogue, provide individually tailored exercise programs and cognitive tests, and evaluate and provide feedback on the results in real time.

[1814] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1815] In this invention, the server includes a speech recognition means for converting speech input into text data, a server means for generating appropriate responses based on a generative model, a speech synthesis means for converting the generated responses into speech and playing them back to the user, an exercise program generation means for collecting exercise data from the user and generating an exercise program based on expert data, a cognitive test provision means for providing multiple types of cognitive tests, and a display means for generating an evaluation report and displaying it to the user. This enables natural dialogue and immediate feedback, increasing the user's motivation while providing effective exercise support and cognitive function evaluation.

[1816] A "voice recognition means" is a device or technique for converting voice input from a user into text data.

[1817] The "server means" is a device or system that utilizes a generative model to generate an appropriate response based on received text data.

[1818] "Speech synthesis means" refers to a device or technology that converts the generated text data into speech data and plays it back to the user.

[1819] The "user interface means" is a device or technology that cooperates with the server to provide operability and information to the user.

[1820] The "exercise program generating means" is a device or system that generates an individual exercise program based on collected exercise data and expert data.

[1821] The "display means" refers to a device or technology for visually presenting the generated exercise program and various information to the user.

[1822] A "feedback means" is a device or technology for analyzing collected data and providing appropriate feedback to the user.

[1823] The "cognitive test providing means" is a device or system that selects from multiple types of cognitive tests and provides them to the user.

[1824] A "data collection tool" is a device or technology that collects the results of a cognitive test or exercise performed by a user.

[1825] A "report generation means" is a device or system for analyzing collected data and generating an evaluation report.

[1826] The present invention provides a system for performing natural interaction, exercise support, and cognitive testing using a generative model, which includes a user interface, a speech recognition unit, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1827] Natural conversation with AI

[1828] Voice input and conversion

[1829] The user speaks into the device's microphone. For example, they might ask, "What would you like to talk about today?" The device receives the user's voice through the microphone and converts the voice data into text data using a speech recognition method (e.g., Google Cloud Speech-to-Text). The converted text is then sent to the server.

[1830] Response generation and playback

[1831] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data. This response text is converted into voice data using a speech synthesis method (e.g., Amazon Polly) and sent to the device. The device then plays the voice data received from the server and conveys the response to the user. For example, a natural conversation follows, such as, "The weather is nice today. Have you gone out anywhere?"

[1832] Exercise support and advice features

[1833] Creation and provision of exercise programs

[1834] The server uses an exercise program generation means to generate an individual exercise program based on the expert's data. This exercise program is then sent to the terminal. The terminal then displays the received exercise program to the user visually or audibly, and provides specific instructions. For example, the terminal may instruct the user on specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[1835] Exercise data collection and analysis

[1836] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit). This data is sent to a server, which analyzes the collected data and generates appropriate feedback. The feedback includes specific suggestions for improvement, such as "It would be more effective if you angled your arms a little higher."

[1837] Cognitive Test Function

[1838] Providing and administering cognitive tests

[1839] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal. The terminal then presents the received cognitive test to the user and has them perform it. For example, the test can be started by saying, "The next test is a memory test. Please remember the words displayed on the screen."

[1840] Collecting and evaluating test results

[1841] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[1842] Specific examples of operation

[1843] Natural conversation examples

[1844] User voice input: The user speaks into the device, "What would you like to talk about today?"

[1845] AI response generation and playback: The response from the server is sent to the device, and a voice message saying, "The weather is nice today. Have you gone out anywhere?" is played.

[1846] Exercise support examples

[1847] Exercise instructions: The device will give you visual or audio instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[1848] Feedback: Based on the data collected during exercise, feedback is sent from the server to the device, and specific advice such as "It would be more effective if you raised the angle of your arms a little more" is displayed.

[1849] Cognitive test examples

[1850] Test instructions: The device will give instructions such as "The next test is a memory test. Please remember the words that appear on the screen," and then carry out the test.

[1851] Presentation of evaluation report: After the user completes the test, the evaluation report analyzed by the server is sent to the terminal and displayed to the user.

[1852] Example prompts for generative AI models

[1853] "When a user asks, 'What would you like to talk about today?' the AI ​​responds naturally with everyday topics."

[1854] "It tells users, 'Today, stretch your upper body. Extend your arms in front of you,' and provides specific exercise advice."

[1855] "During a memory test, you are instructed to 'remember the following words,' and the test results are collected, evaluated, and a report is generated."

[1856] The above is a mode for carrying out the present invention, and shows that natural dialogue, exercise support, and cognitive testing can be effectively provided through specific processing steps.

[1857] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1858] Step 1: Receiving voice input

[1859] The user speaks to the terminal, asking questions such as, "What would you like to talk about today?"

[1860] Input: User's voice data

[1861] Output: Audio data collected by the device's microphone

[1862] Specific operation: The user speaks into the device's microphone, and voice data is input into the device.

[1863] Step 2: Speech to text

[1864] The terminal converts the user's voice data into text data using a voice recognition means (e.g., Google Cloud Speech-to-Text).

[1865] Input: User's voice data

[1866] Output: Text data

[1867] Specific operation: The terminal's voice recognition means analyzes the voice data and converts it into corresponding text.

[1868] Step 3: Send text data

[1869] The terminal transmits the converted text data to the server.

[1870] Input: Text data

[1871] Output: Text data sent to the server

[1872] Specific operation: The device sends text data to a server via the Internet.

[1873] Step 4: Generate response text

[1874] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate response to the received text data.

[1875] Input: Text data

[1876] Output: Response text data

[1877] How it works: The server's AI model analyzes the text data and generates an appropriate response, such as "The weather is nice today. Have you gone out anywhere?"

[1878] Step 5: Convert response text to speech

[1879] The server converts the generated response text into voice data using a speech synthesis means (e.g., Amazon Polly) and sends it to the terminal.

[1880] Input: Response text data

[1881] Output: Audio data

[1882] Specific operation: The server's speech synthesis means converts the response text into speech and transmits the speech data to the terminal.

[1883] Step 6: Play the reply audio

[1884] The terminal reproduces the voice data received from the server and conveys the response to the user.

[1885] Input: Audio data

[1886] Output: Audio played from the device

[1887] Specific operation: The device speaker plays the audio data and conveys the response to the user.

[1888] Step 7: Generate the exercise program

[1889] The server uses an exercise program generating means to generate an individual exercise program based on the expert data.

[1890] Input: Expert data and user feedback data

[1891] Output: personalized exercise program

[1892] Specific operation: The server uses the exercise program generation means to design an individual exercise program based on the collected data.

[1893] Step 8: Providing an exercise program

[1894] The terminal displays the received exercise program to the user visually or audibly, and provides specific instructions.

[1895] Input: Exercise program data

[1896] Output: An exercise program provided to the user

[1897] Specific actions: The device displays an exercise program on the screen or gives instructions to the user by voice. For example, it may say, "Today, let's stretch your upper body. Stretch your arms out in front of you."

[1898] Step 9: Collect exercise data

[1899] While the user exercises, the device collects exercise data from the attached sensor (e.g., Fitbit).

[1900] Input: User's exercise data

[1901] Output: Collected exercise data

[1902] Specific operation: The sensor collects the user's movement data in real time and transmits it to the device.

[1903] Step 10: Analyze the movement data

[1904] The server analyzes the collected movement data and generates appropriate feedback.

[1905] Input: Collected exercise data

[1906] Output: Feedback data

[1907] Specific movements: The server analyzes the movement data and generates specific feedback, such as advice like "It would be more effective if you angled your arms a little higher."

[1908] Step 11: Offer cognitive testing

[1909] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[1910] Input: User context data

[1911] Output: Cognitive test data

[1912] Specific operation: The server selects the appropriate cognitive test from multiple options and sends it to the device.

[1913] Step 12: Conduct cognitive testing

[1914] The terminal presents the received cognitive test to the user and allows the user to perform the test.

[1915] Input: Cognitive test data

[1916] Output: Test result data performed by the user

[1917] Specific operation: The device displays a cognitive test on the screen and asks the user to complete it. For example, it instructs the user, "The next test is a memory test. Please remember the words displayed on the screen."

[1918] Step 13: Collecting test results

[1919] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means.

[1920] Input: User test result data

[1921] Output: Test result data sent to the server

[1922] Specific operation: The device collects the test results and sends them to the server.

[1923] Step 14: Analyze test results and generate evaluation reports

[1924] The server analyzes the results and generates an evaluation report, which includes the user's current cognitive function and areas for improvement.

[1925] Input: Test result data

[1926] Output: Evaluation report

[1927] Specific operation: The server analyzes the test result data and generates an evaluation report.

[1928] Step 15: Provide an evaluation report

[1929] The terminal displays the evaluation report to the user.

[1930] Input: Evaluation report data

[1931] Output: Evaluation report provided to the user

[1932] Specific operation: The device displays the evaluation report on the screen and shows it to the user.

[1933] (Application example 1)

[1934] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1935] Conventional online shopping systems are often cumbersome, requiring complex operations such as product searches, ordering, and return procedures, making them particularly difficult for elderly users and those with low IT literacy. Another issue is the lack of product suggestions based on individual user preferences and body types, making it time-consuming to find the perfect product. Furthermore, there is a lack of systems that can routinely conduct cognitive tests on users with concerns about cognitive decline.

[1936] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1937] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into speech data and playing it back to the user, a means for providing functions that allow the user to search for, order, and process returns of products by voice, and a means for suggesting recommended items based on the user's body type and preferences. This allows the user to easily operate the shop by voice, makes it possible to suggest products that suit the user's individual needs, and also maintains and improves cognitive function.

[1938] A "generative model" is an artificial intelligence model that generates new text or information based on given text or data.

[1939] "User interface means" are technical means that provide an interface through which a user can directly interact with a system.

[1940] "Speech recognition means" refers to technical means for converting speech input into text data.

[1941] "Server means" refers to technical means that provides the functionality of a server that receives text data and generates a response based on a generative model.

[1942] "Speech synthesis means" refers to technical means for converting text data into voice data and playing it back to the user.

[1943] "Means for providing functions for searching for, ordering, and processing returns of products" refers to technical means that allow users to search for, order, and process returns of products using voice commands.

[1944] The "means for proposing recommended items" refers to a technical means for proposing appropriate products according to the user's body type and preferences.

[1945] This invention is a virtual assistant system that integrates natural dialogue using generative models with online shopping functions. The system includes a server means that uses speech recognition, speech synthesis, generative models, and a user interface means.

[1946] Hardware and Software Configuration

[1947] Hardware:

[1948] Smartphone (iOS, Android compatible)

[1949] software:

[1950] Speech Recognition API (Google Cloud Speech-to-Text)

[1951] Generative AI model (GPT-4)

[1952] Speech synthesis API (Google Cloud Text-to-Speech)

[1953] System Operation

[1954] Voice input and recognition:

[1955] When a user speaks into a smartphone, the built-in microphone picks up the voice, and this voice data is converted into text data in real time via a speech recognition API.

[1956] Sending data to the server and generating a model response:

[1957] The converted text data is sent to a server, where a generative AI model (GPT-4) generates natural-sounding responses based on the user's speech. This response text is then converted back into voice data using a speech synthesis API and sent back to the smartphone.

[1958] Response playback in the user interface:

[1959] The smartphone plays back the returned voice data and provides natural and appropriate responses to the user. For example, if the user says, "I'm looking for a black dress," the speech recognition API converts this into text, and the generative AI model responds, "I'll show you the search results for black dresses."

[1960] Product search, suggestions, and ordering:

[1961] The server also allows users to search for products, place orders, and process returns using voice commands, and it also provides recommendations based on the user's body type and preferences, helping users find the perfect product.

[1962] Examples and prompts

[1963] For example, if a user says, "I'm looking for a black dress," a voice-synthesized message will be played in response, such as, "Here are the search results for black dresses. There are several options." This is followed by, "Today's recommended items are also displayed," allowing the user to enjoy shopping more intuitively through voice.

[1964] Example prompt sentence:

[1965] "A virtual shopping assistant, recommending suitable products for users looking for black dresses."

[1966] This system makes online shopping more natural and user-friendly, and is suitable for elderly people and those with low IT literacy. Furthermore, for users who are concerned about declining cognitive function, it also provides a daily cognitive test function, helping them manage their health.

[1967] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1968] Step 1:

[1969] The user speaks into the smartphone. The smartphone's built-in microphone receives the voice and collects this voice data. The input is the user's voice, and the output is voice data.

[1970] Step 2:

[1971] The device sends the collected voice data to a voice recognition API (Google Cloud Speech-to-Text) in real time, which converts the voice data into text data. The input is voice data, and the output is text data. This process recognizes the voice and converts it into text.

[1972] Step 3:

[1973] Text data is sent from the terminal to the server. When the server receives the text data, the input is text data and the output is the text data sent to the server. This process prepares the data to be used in the next step.

[1974] Step 4:

[1975] The server feeds the received text data into a generative AI model (GPT-4) to generate an appropriate response. The input is text data, and the output is the generated response text. This data processing enables natural dialogue.

[1976] Step 5:

[1977] The generated response text is converted into audio data on the server side using a speech synthesis API (Google Cloud Text-to-Speech). The input is the response text and the output is audio data. This process provides the response in audio format.

[1978] Step 6:

[1979] The generated voice data is sent from the server to the terminal. The user can continue the conversation by receiving and playing the voice data. The input is the voice data from the server, and the output is the played voice.

[1980] Step 7:

[1981] When a user wants to search for a product, the server will suggest recommended items based on the user's body type and preferences. The input is the user's detailed information and search criteria, and the output is a list of recommended products. The server generates this and sends it to the device.

[1982] Step 8:

[1983] The terminal displays the received product list to the user, who then uses voice commands to order products or perform other operations. The input is the user's voice commands, and the output is the results of operations such as ordering. Specifically, if the user says, "I'm looking for a black dress," the terminal responds, "I'll show you the search results for black dresses."

[1984] This series of steps allows users to easily shop through voice, with data processing and calculations at each step providing a highly interactive and personalized shopping experience.

[1985] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1986] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[1987] Natural conversation with AI

[1988] User voice input

[1989] The user speaks to the terminal, for example, "What would you like to talk about today?"

[1990] Speech-to-text

[1991] The terminal receives the user's voice through a microphone and acquires it as voice data.

[1992] Speech and Emotion Recognition

[1993] The terminal converts the voice data into text data using a voice recognition means, and simultaneously recognizes the user's emotion using an emotion engine. The converted text and the recognized emotion data are sent to the server.

[1994] Response generation using AI models

[1995] The server generates an appropriate response to the received text data based on the generative model. It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[1996] Playing back replies to the user

[1997] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[1998] Exercise support and advice features

[1999] Providing exercise programs

[2000] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[2001] Exercise instructions

[2002] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[2003] Exercise data collection and analysis

[2004] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[2005] Cognitive Test Function

[2006] Cognitive testing provided

[2007] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[2008] Cognitive testing

[2009] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[2010] Collecting and evaluating test results

[2011] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[2012] Through these functions, this system aims to maintain and improve cognitive function by enabling people with Alzheimer's disease and those at risk of developing dementia, such as mild cognitive impairment (MCI), to communicate naturally at home or in a facility, and by providing exercise and cognitive training. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable conversational experience for the user.

[2013] The above is a description of the preferred embodiment of the present invention, which shows how the invention can be effectively practiced within the scope of the claims.

[2014] The processing flow will be explained below.

[2015] Natural conversation with AI

[2016] Step 1:

[2017] User: Speaks into the device. For example, "What would you like to talk about today?"

[2018] Step 2:

[2019] Terminal: Receives the user's voice via a microphone and captures it as voice data.

[2020] Step 3:

[2021] Terminal: Converts voice data into text data using a voice recognition means.

[2022] Step 4:

[2023] Terminal: Uses an emotion engine to recognize emotions from the user's voice.

[2024] Step 5:

[2025] Terminal: Transmits the converted text data and recognized emotion data to the server.

[2026] Step 6:

[2027] Server: Inputs the received text data into a generative model to generate an appropriate response. Based on the emotional data, the tone and content of the response are adjusted.

[2028] Step 7:

[2029] Server: Uses a voice synthesis means to convert the generated response text into voice data and sends it to the terminal.

[2030] Step 8:

[2031] Terminal: Plays back the audio data and responds to the user. For example, if the user is sad, the response will be something comforting like, "What happened today? Tell me."

[2032] Exercise support and advice features

[2033] Step 1:

[2034] Server: Creates an individual exercise program using an exercise program generation means based on data provided by experts.

[2035] Step 2:

[2036] Server: Sends the generated exercise program to the terminal.

[2037] Step 3:

[2038] Terminal: displays the exercise program to the user visually or audibly and provides exercise instructions.

[2039] Step 4:

[2040] User: Starts exercising according to instructions. For example, the user is instructed to do specific exercises, such as "Today, let's stretch your upper body. Stretch your arms out in front of you."

[2041] Step 5:

[2042] Device: Collects user movement data in real time from sensors.

[2043] Step 6:

[2044] Terminal: Sends collected exercise data to the server.

[2045] Step 7:

[2046] Server: Analyzes the received motion data and generates feedback.

[2047] Step 8:

[2048] Server: Sends the generated feedback to the device.

[2049] Step 9:

[2050] Device: Provides feedback to the user, informing them of the next exercise or improvement. For example, it may advise, "It would be more effective if you angle your arms a little higher."

[2051] Cognitive Test Function

[2052] Step 1:

[2053] Server: Using the cognitive test providing means, selects multiple types of cognitive tests from a database and transmits the test appropriate for the user to the terminal.

[2054] Step 2:

[2055] Device: Presents the cognitive test to the user visually or audibly, for example, "The next test is a memory test. Please remember the words that appear on the screen."

[2056] Step 3:

[2057] User: Follow the instructions to complete the cognitive test.

[2058] Step 4:

[2059] Terminal: Collects the results of the cognitive tests performed by the user through a data collection means and transmits them to the server.

[2060] Step 5:

[2061] Server: Analyzes the received test results and generates an evaluation report.

[2062] Step 6:

[2063] Server: Sends the generated evaluation report to the terminal.

[2064] Step 7:

[2065] Device: The device displays the assessment report visually or audibly to the user, informing them of the current state of cognitive function and areas for improvement. For example, it may say, "Your memory is stable, but you could improve your attention a bit more."

[2066] The above is a specific processing flow of the present invention. By implementing each function of the present invention, Alzheimer's disease patients and people at risk of developing dementia such as MCI can undergo safe and effective rehabilitation at home or in a facility. Furthermore, the introduction of an emotion engine enables more appropriate and human-like responses based on the user's emotions, providing a comfortable dialogue experience for the user.

[2067] Example 2

[2068] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2069] The objective of this invention is to provide a system that integrates natural user interaction, exercise support, cognitive testing, and emotion recognition. Conventional systems often provide these functions separately, resulting in a fragmented user experience. Furthermore, emotion recognition accuracy is low, making it difficult to respond to the user's emotions. As a result, users experience low satisfaction and effectiveness.

[2070] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a voice recognition means for converting a user's voice input into text data, an emotion engine means for recognizing the user's emotions, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, and a voice synthesis means for converting the generated response into voice data and playing it back to the user. This not only enables the user to enjoy a natural dialogue experience, but also enables emotionally sensitive communication. The server also includes an exercise program generation means for collecting exercise data from the user and generating an exercise program provided by an expert, a display means for displaying the exercise program to the user and providing instructions, and a feedback means for transmitting the user's exercise data to the server and providing analysis results as feedback. The server also includes a cognitive test provision means for providing multiple types of cognitive tests, a data collection means for collecting results of the cognitive tests performed by the user, a report generation means for analyzing the test results and generating an evaluation report, and a display means for displaying the evaluation report to the user. This enables comprehensive health management, thereby increasing user satisfaction and effectiveness.

[2071] A "generative model" is an algorithm that uses artificial intelligence to generate natural-looking dialogue and responses based on given input data.

[2072] "User interface means" refers to a device or software that allows a user to interact with a system, including voice input and a display.

[2073] "Speech recognition means" refers to a technology or device that analyzes voice input from a user as digital data and converts it into character data.

[2074] "Emotion engine means" refers to a technology or device that identifies emotions from the user's voice or text and outputs the emotional state as data.

[2075] "Server Means" refers to a centralized computing resource that processes data for the entire system and generates responses based on the Generative Model.

[2076] "Speech synthesis means" refers to a technique or device that converts the generated text data into voice data and plays it back to the user.

[2077] The term "exercise program generation means" refers to a technology or device that creates an individually customized exercise program based on the user's exercise data and expert guidance.

[2078] "Display means" refers to technology or devices for visually presenting exercise instructions, cognitive test results, reports, etc. to the user.

[2079] "Feedback means" refers to technology or devices that analyze the user's exercise data and cognitive test results and provide advice on areas for improvement and next steps.

[2080] A "cognitive test delivery means" refers to a technology or device that delivers tests to a user to assess various cognitive functions.

[2081] "Data collection means" refers to sensors and software that collect user behavior and test results.

[2082] A "report generation means" is a technique or device that analyzes user data and compiles the results into a report.

[2083] MODE FOR CARRYING OUT THE INVENTION

[2084] The present invention provides a system incorporating an emotion engine that recognizes user emotions, in addition to natural dialogue using a generative model, exercise support, and cognitive testing. This system includes a user interface, a speech recognition unit, an emotion engine, a server, a speech synthesis unit, an exercise program generation unit, a display unit, a feedback unit, a cognitive test provision unit, a data collection unit, and a report generation unit.

[2085] Natural conversation with AI

[2086] User voice input

[2087] The user speaks to the terminal, for example, "What would you like to talk about today?"

[2088] Speech-to-text

[2089] The device receives the user's voice through a microphone and acquires it as voice data. The Google Cloud Speech-to-Text API is used for voice recognition.

[2090] Speech and Emotion Recognition

[2091] The device converts voice data into text data using a voice recognition means. At the same time, it recognizes the user's emotions using an emotion engine means. The emotion recognition uses Microsoft Azure's emotion analysis API. The converted text data and the recognized emotion data are then sent to the server.

[2092] Response generation using AI models

[2093] The server generates an appropriate response to the received text data based on a generative AI model (e.g., ChatGPT). It also adjusts the content and tone of the response based on the emotional data. For example, if the user is sad, it generates a more gentle and encouraging response. The generated response text is converted into audio data and sent to the device.

[2094] Playing back replies to the user

[2095] The device then plays back the voice data received from the server and responds to the user, allowing for natural conversation, for example, "The weather is nice today. Have you gone out anywhere?"

[2096] Exercise support and advice features

[2097] Providing exercise programs

[2098] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[2099] Exercise instructions

[2100] The device displays the received exercise program to the user visually or audibly and provides exercise instructions, such as "Today, stretch your upper body. Stretch your arms out in front of you."

[2101] Exercise data collection and analysis

[2102] When a user exercises, the device collects exercise data through sensors attached to the device and sends the data to a server. The server analyzes the collected data and generates appropriate feedback, including specific advice such as "It would be more effective if you raised the angle of your arms a little more."

[2103] Cognitive Test Function

[2104] Cognitive testing provided

[2105] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[2106] Cognitive testing

[2107] The device presents the received cognitive test to the user and has them perform it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[2108] Collecting and evaluating test results

[2109] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[2110] This not only allows users to enjoy a natural conversational experience, but also enables emotionally sensitive communication and comprehensive health management through exercise and cognitive training.

[2111] Prompt Sentence Examples

[2112] "Create prompts that are sensitive to the user's emotions and start the conversation around the weather. For example, 'The weather is nice today. Have you been out anywhere?'"

[2113] The above is an embodiment of the invention.

[2114] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2115] Step 1:

[2116] User voice input

[2117] Input: User speech

[2118] How it works: The user speaks to the device, asking something like, "What would you like to talk about today?"

[2119] Output: User's voice data

[2120] Step 2:

[2121] Speech-to-text

[2122] Input: User's voice data

[2123] How it works: The device captures the user's voice with its built-in microphone and converts the voice data into text using the Google Cloud Speech-to-Text API.

[2124] Output: Text data

[2125] Step 3:

[2126] emotion recognition

[2127] Input: Text data

[2128] How it works: The device uses Microsoft Azure's sentiment analysis API to recognize user emotions from text data.

[2129] Output: Emotion data

[2130] Step 4:

[2131] Sending data to the server

[2132] Input: Text data, emotion data

[2133] Operation: The device sends the converted text data and the recognized emotion data to the server.

[2134] Output: Data sent to the server

[2135] Step 5:

[2136] Response generation using AI models

[2137] Input: Text data and emotion data sent to the server

[2138] How it works: The server uses a generative AI model (e.g., ChatGPT) to generate appropriate responses based on the data it receives, and also adjusts the content and tone of the responses to match the user's emotions.

[2139] Output: Response text data

[2140] Step 6:

[2141] Speech synthesis and reply sending

[2142] Input: Response text data

[2143] Operation: The server converts the response text data into voice data using a voice synthesis means, and transmits the voice data to the terminal.

[2144] Output: Audio data

[2145] Step 7:

[2146] Playing back replies to the user

[2147] Input: Audio data

[2148] Operation: The device receives the voice data from the server, plays it back through the speaker, and responds to the user. For example, it might say, "The weather is nice today. Have you gone out anywhere?"

[2149] Output: The reply speech the user hears

[2150] Step 8:

[2151] Providing exercise programs

[2152] Input: Exercise data provided by experts

[2153] Operation: The server uses the exercise program generation means to generate an individual exercise program based on the data provided by the expert, and transmits it to the terminal.

[2154] Output: Exercise program

[2155] Step 9:

[2156] Exercise instructions

[2157] Input: Exercise program

[2158] Actions: The device displays the received exercise program to the user visually or audibly, and provides specific exercise instructions such as, "Today, stretch your upper body. Extend your arms in front of you."

[2159] Output: Exercise instruction content

[2160] Step 10:

[2161] Exercise data collection and analysis

[2162] Input: User's exercise data

[2163] How it works: When a user exercises, the device collects exercise data through sensors attached to the device and transmits the data to a server, which analyzes the data and generates appropriate feedback.

[2164] Output: Analysis results and feedback

[2165] Step 11:

[2166] Cognitive testing provided

[2167] Input: Cognitive test selection instructions

[2168] Operation: The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal.

[2169] Output: Cognitive test content

[2170] Step 12:

[2171] Cognitive testing

[2172] Input: Cognitive test content

[2173] Operation: The device presents the received cognitive test to the user and has them complete it. For example, it may say, "The next test is a memory test. Please remember the words that appear on the screen."

[2174] Output: User's cognitive test results

[2175] Step 13:

[2176] Collecting and evaluating test results

[2177] Input: User's cognitive test results

[2178] Operation: The device sends the results of the cognitive tests taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the user's current cognitive function and areas for improvement, and is displayed to the user via the device.

[2179] Output: Evaluation report

[2180] (Application example 2)

[2181] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2182] There is a need for an effective method to support the dialogue between customers and store clerks in brick-and-mortar stores. In particular, there is a need for a system that can recognize customers' emotions and respond flexibly and naturally. Another challenge is to improve customer satisfaction by suggesting products that correspond to the customer's interests and emotions.

[2183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2184] In this invention, the server includes a user interface means for conducting natural dialogue using a generative model, a speech recognition means for converting speech input from the user into text data, a server means for transmitting the text data and emotion data to the server and generating a response based on the generative model, a speech synthesis means for converting the generated response into voice data and playing it back to the user, and a suggestion means for making product suggestions in the store. This allows for natural dialogue with customers and enables flexible and appropriate responses based on the customer's emotions. Furthermore, by making product suggestions based on the customer's interests, customer satisfaction can be improved.

[2185] A "generative model" is an artificial intelligence algorithm that generates appropriate responses or results based on input data from a user.

[2186] A "user interface means" is a means that provides an interface for a user to interact with the system.

[2187] The "voice recognition means" is a means having a function of converting voice input from a user into text data.

[2188] "Emotional data" is data collected to describe a user's emotional state.

[2189] The "server means" is a means for receiving text data and emotion data and generating an appropriate response based on the received data.

[2190] The "voice synthesis means" is a means having the function of converting the generated response into voice data and playing it back to the user.

[2191] The "proposal means" is a means having a function for proposing products and services to customers in a store.

[2192] The "exercise program generating means" is a means having a function of generating an appropriate exercise program based on data provided by an expert.

[2193] The "display means" is a means having a function for visually presenting the generated exercise program and other information to the user.

[2194] The "feedback means" is a means having a function for analyzing the user's exercise data and providing appropriate feedback.

[2195] The "cognitive test providing means" is a means having a function for providing multiple types of cognitive tests to the user.

[2196] The "data collection means" is a means having a function for collecting the results of the cognitive tests performed by the user.

[2197] The "report generation means" is a means having a function for analyzing the results of the cognitive test and generating an evaluation report.

[2198] The system that realizes this application example provides natural dialogue, exercise support, and cognitive testing using a generative model, and incorporates an emotion engine. A specific embodiment of this system will be described below.

[2199] Basic system configuration

[2200] The system consists of the following means:

[2201] User Interface Means

[2202] Voice recognition means

[2203] emotion recognition means

[2204] Server Means

[2205] Voice synthesis means

[2206] Proposal means

[2207] Exercise program generation means

[2208] Display means

[2209] Feedback Methods

[2210] Cognitive test delivery methods

[2211] Data collection methods

[2212] Report Generation Method

[2213] Natural conversational features

[2214] User voice input

[2215] The user speaks to a terminal connected to the system. For example, they might say, "Hello, do you have any product recommendations?" The terminal receives the voice through a microphone and captures it as voice data.

[2216] Speech-to-text and emotion recognition

[2217] The device converts the voice data into text data using a speech recognition engine (e.g., Google Speech-to-Text), and simultaneously recognizes the user's emotions using an emotion recognition engine (e.g., Microsoft Azure Emotion API). This data is then sent to the server.

[2218] Response generation using AI models

[2219] The server generates an appropriate response to the received text data based on a generative AI model (e.g., OpenAI GPT-3). It also adjusts the content and tone of the response based on emotional data. For example, if the user shows interest, it generates a friendly response such as, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?"

[2220] Playing back replies to the user

[2221] The generated response text is converted into voice data by a speech synthesis engine, and the device plays this voice data over a speaker to convey the response to the user.

[2222] Exercise support and advice features

[2223] Providing exercise programs

[2224] The server uses an exercise program generating means to generate an individual exercise program based on data provided by the expert and transmits it to the terminal.

[2225] Exercise instructions and feedback

[2226] The device displays the received exercise program to the user visually or audibly and provides exercise instructions. During exercise, exercise data collected through sensors is sent to a server for analysis. Based on the analysis results, the device provides feedback such as, "It would be more effective if you raised the angle of your arms a little more."

[2227] Cognitive Test Function

[2228] Providing and administering cognitive tests

[2229] The server uses the cognitive test providing means to select an appropriate cognitive test according to the user's situation and transmits it to the terminal, for example, by informing the user that "The next test is a memory test. Please remember the words displayed on the screen."

[2230] Analysis and evaluation of test results

[2231] The terminal transmits the results of the cognitive test taken by the user to the server via the data collection means. The server analyzes the results and generates an evaluation report. The evaluation report includes the current state of the user's cognitive function and areas for improvement, and is displayed to the user via the terminal.

[2232] Specific examples

[2233] As examples of use in physical stores, the following specific dialogue scenarios are envisioned.

[2234] Usage

[2235] 1. User (Customer): "Hi, do you have any product recommendations?"

[2236] 2. Device (application running on smartphone):

[2237] Acquisition speech-to-text: "Hi, do you have any product recommendations?"

[2238] Emotion Recognition: "Curious" (Emotion Engine)

[2239] Response generation: "Hello, we have the latest smartwatch that's just for you. Would you like to take a look?" (Generative AI model)

[2240] Speech synthesis: Generated responses are spoken

[2241] Speaker Output: "Hello, we have the latest smartwatch just for you. Would you like to take a look?"

[2242] Prompt Sentence Examples

[2243] Generate relevant product suggestions when customers express interest: "Create an engaging product suggestion based on the following sentence: 'Hi, can you recommend something?' Keep the tone of your suggestion friendly but professional."

[2244] This system enables natural dialogue with customers in physical stores, and flexible and appropriate responses based on customer emotions. In addition, it can improve customer satisfaction by suggesting products based on the customer's interests.

[2245] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2246] Step 1:

[2247] User voice input

[2248] The user speaks to the device, for example, "Hello, do you have any product recommendations?" This speech is picked up by the device's microphone and saved as audio data.

[2249] Input: User's voice

[2250] Output: Audio data

[2251] Step 2:

[2252] Speech-to-text

[2253] The device uses a speech recognition engine to convert the acquired voice data into text data. For this purpose, it uses a speech recognition API (for example, Google Speech-to-Text).

[2254] Input: Audio data

[2255] Output: Text data

[2256] Step 3:

[2257] emotion recognition

[2258] The device sends the text data to an emotion recognition engine to recognize the user's emotions. This uses an emotion recognition API (for example, Microsoft Azure Emotion API). Emotion data reflects the user's psychological state.

[2259] Input: Text data

[2260] Output: Emotion data

[2261] Step 4:

[2262] Response generation based on generative models

[2263] The server receives the text data and emotion data and uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The tone and content of the response are adjusted based on the emotion data. For example, if the user shows interest, the response will be more friendly and engaging.

[2264] Input: Text data and emotion data

[2265] Output: Response text

[2266] Step 5:

[2267] Response text transcription

[2268] The server converts the response text into voice data using a speech synthesis engine. For this process, it uses a speech synthesis API (for example, Amazon Polly).

[2269] Input: Reply text

[2270] Output: Reply audio data

[2271] Step 6:

[2272] Playback of response audio data

[2273] The terminal plays the received response voice data over the speaker, thereby conveying the dialogue response to the user.

[2274] Input: Response audio data

[2275] Output: The audio played to the user

[2276] Step 7:

[2277] Product proposals using proposal methods

[2278] The server then suggests appropriate products based on the user's interests and emotions. For example, it generates a response like, "Hello, we have the latest smartwatch that's just right for you. Would you like to take a look?" This product suggestion is made using a generative AI model that takes into account the user's emotional data.

[2279] Input: User interest and emotion data

[2280] Output: Product suggestion text

[2281] Through these steps, the system can realize natural dialogue with the user, responding flexibly and appropriately according to their emotions. Furthermore, it can improve customer satisfaction by suggesting products based on the user's interests.

[2282] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2283] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2284] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2285] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2286] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2287] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2288] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2289] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2290] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2291] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2292] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2293] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2294] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2295] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2296] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2297] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2298] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2299] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2300] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the in...

Claims

1. a user interface means for performing natural dialogue using a generative model; a speech recognition means for converting speech input from a user into text data; server means for transmitting text data to a server and generating a response based on the generative model; a voice synthesis means for converting the generated response into voice data and playing it back to the user; A system including:

2. an exercise program generating means for collecting exercise data from a user and generating an exercise program provided by an expert; a display means for displaying an exercise program to a user and providing instructions; a feedback means for transmitting the user's exercise data to a server and providing the analysis result as feedback; The system of claim 1 , comprising:

3. a cognitive test providing means for providing a plurality of types of cognitive tests; data collection means for collecting results of cognitive tests administered by users; a report generating means for analyzing the test results and generating an evaluation report; a display means for displaying the evaluation report to a user; The system of claim 1 , comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A