System

The system assists users in achieving their schedules and improving communication by predicting optimal behavioral patterns and identifying inappropriate statements, thus enhancing goal achievement and reducing stress.

JP2026028699APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131315
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Users face challenges in selecting optimal behavioral patterns to achieve their schedules and goals efficiently, and interpersonal communication is often hindered by inappropriate statements, leading to stress and misunderstandings.

Method used

A system that allows users to input plans and goals, collect internal and external information, predict behavioral patterns, record and analyze voice data for redundant or inappropriate statements, and provide pre-alerts and advice to improve communication.

Benefits of technology

Enables users to efficiently achieve their goals and reduce interpersonal troubles by providing optimal behavioral patterns and appropriate speech advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028699000001_ABST
    Figure 2026028699000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting a schedule and a goal of a day by a user; means for collecting internal information and external information of the user; means for predicting a behavior of the user based on the schedule, the goal, the internal information, and the external information; and means for presenting the predicted behavior pattern to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, it is difficult for users to select optimal behavioral patterns to efficiently achieve their schedules and goals. Furthermore, in interpersonal relationships, problems can easily arise depending on the content and timing of statements, which can become a major source of stress in daily life. The present invention aims to provide a system that helps users achieve their schedules and goals and supports smooth communication. [Means for solving the problem]

[0005] The present invention provides a system including a means for a user to input the user's plans and goals for the day, a means for collecting the user's internal and external information, a means for predicting the user's behavior based on the plans, goals, internal and external information, and a means for presenting the predicted behavioral pattern to the user. The system also includes a means for recording the user's daily conversations and collecting voice data, a means for converting the voice data into text, a means for identifying redundant or inappropriate statements from the text data, and a means for generating a pre-alert for a statement and advice on appropriate statements based on the identified statements and notifying the user. Furthermore, the system realizes a system that allows the user to efficiently achieve their goals and reduce interpersonal troubles by having the means for predicting the behavioral pattern calculate an optimal behavioral pattern based on the internal and external information, using a speech recognition engine to convert the voice data into text, and using a natural language processing model to identify redundant or inappropriate statements.

[0006] "User" refers to any individual or legal entity that uses the System.

[0007] "Appointments" refer to actions or events that a user plans to take place at a specific date and time.

[0008] A "goal" is a specific result or objective that a user is trying to achieve.

[0009] "Internal information" refers to information about the user's own physical and psychological state, such as the user's mood or health.

[0010] "External information" refers to information about the user's external environment, such as weather conditions and traffic conditions.

[0011] "Behavioral prediction methods" refer to algorithms or systems that calculate optimal behavioral patterns for a user based on the user's schedule, goals, internal information, and external information.

[0012] A "behavioral pattern" refers to a sequence of specific actions that a user must take to achieve a plan or goal.

[0013] "Means for recording conversations" refers to devices or software that store a user's voice as digital data.

[0014] "Audio Data" means a digital audio file that records a user's speech.

[0015] "Means for converting to text" refers to a speech recognition engine that converts voice data into character data.

[0016] "Redundant statements" refer to statements that are unnecessary for achieving the intended purpose and are unnecessarily long.

[0017] "Inappropriate remarks" refer to remarks that may lead to misunderstandings or cause trouble in interpersonal relationships.

[0018] "Pre-speech alert" refers to a feature that warns users of the risks of inappropriate statements before they make them.

[0019] "Appropriate speech advice" refers to a function that suggests effective speech and expressions to help users communicate better.

[0020] A "natural language processing model" refers to a machine learning algorithm that analyzes text data and understands and classifies its content. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[0043] Embodiment of behavior prediction

[0044] User:

[0045] Users launch the smartphone app and enter their plans and goals for the day. They also enter internal information such as their mood and health status. For example, a user might enter their plan to "go to the gym in the morning and go shopping in the afternoon," select "normal" as their mood, and "good" as their health status.

[0046] Device:

[0047] The schedule, goals, and internal information entered by the user are stored in a local database. Then, external APIs are used to obtain external information such as weather conditions and traffic conditions. For example, a weather API is called to obtain information such as "sunny," and a traffic API is called to obtain information such as "no traffic jams."

[0048] server:

[0049] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. Based on the stored data, an AI model predicts the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, the AI ​​model may generate a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[0050] Device:

[0051] The received prediction results are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[0052] Embodiment of speech prediction

[0053] User:

[0054] Start the smartphone app and record your everyday conversations. For example, record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[0055] Device:

[0056] The recorded voice data is stored in a local database and sent to the server, along with any memo information.

[0057] server:

[0058] The voice data received from the device is converted into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0059] server:

[0060] Based on the analysis results, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent to the device.

[0061] Device:

[0062] The received feedback is notified to the user and displayed within the app, allowing the user to adjust what they say based on the feedback and improve communication. For example, when talking with a friend, a user can refer to the advice to choose a topic and smoothly advance the conversation.

[0063] Specific examples

[0064] Specific examples of behavioral prediction

[0065] The user enters their schedule into the app, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict behavioral patterns such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends this to the device. The device notifies the user of the results, and the user acts accordingly.

[0066] Specific examples of speech prediction

[0067] A user records a conversation with a close friend and leaves a note saying, "I'd like to talk about my next trip." The device then sends the recorded audio data and note to a server. The server converts the audio data into text and uses a natural language processing model to identify "overly long explanations" and "potentially misleading" parts. Based on this, the device generates and sends feedback to the device, such as "mention your next trip, but avoid talking about your previous trip." The device then notifies the user of the feedback, and the user can use the advice to smoothly move the conversation forward.

[0068] This allows users to receive support in achieving their schedules and goals, as well as in communicating smoothly.

[0069] The processing flow will be explained below.

[0070] Program processing of behavior prediction

[0071] Step 1:

[0072] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[0073] Step 2:

[0074] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[0075] Step 3:

[0076] Device: Stores user-entered schedules, goals, and internal information in a local database.

[0077] Step 4:

[0078] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[0079] Step 5:

[0080] Terminal: Send all collected data, including internal and external information, to the server.

[0081] Step 6:

[0082] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[0083] Step 7:

[0084] Server: Based on the stored data, it uses an AI model to predict the user's optimal behavioral patterns.

[0085] Step 8:

[0086] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[0087] Step 9:

[0088] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[0089] Step 10:

[0090] User: Follows the suggested behavioral patterns and achieves the goal.

[0091] Step 11:

[0092] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[0093] Program processing of speech prediction

[0094] Step 1:

[0095] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[0096] Step 2:

[0097] Users: Leave notes during conversations about things that concern them or that need attention.

[0098] Step 3:

[0099] Device: Stores recorded audio data and notes in a local database.

[0100] Step 4:

[0101] Terminal: Sends voice data and memo information to the server.

[0102] Step 5:

[0103] Server: Converts voice data received from the device into text using a speech recognition engine.

[0104] Step 6:

[0105] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[0106] Step 7:

[0107] Server: Based on the analysis results, generate appropriate feedback based on the profile information of the person you are speaking to.

[0108] Step 8:

[0109] Server: Sends the generated feedback (e.g., "mention your next trip, but avoid talking about your previous trip") to the device.

[0110] Step 9:

[0111] On the device: Notify the user of the feedback received and display it visually within the app.

[0112] Step 10:

[0113] Users: Use suggested feedback to adjust conversations and improve communication.

[0114] This allows users to effectively achieve their schedules and goals, as well as communicate smoothly.

[0115] Example 1

[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0117] Conventional systems have difficulty predicting behavior based on internal and external information in addition to users' schedules and goals. Furthermore, there was a lack of systems that notified users of appropriate timing and content for speech in everyday conversations. Furthermore, there was no mechanism in place to incorporate user feedback to improve the system's overall prediction accuracy. This made it difficult for users to achieve their schedules and communicate smoothly.

[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0119] In this invention, the server includes: means for a user to input the day's schedule and goals; means for collecting the user's internal and external information; means for predicting the user's behavior based on the schedule, goals, internal and external information; means for presenting the predicted behavioral pattern to the user; means for collecting voice data and converting it into text; means for identifying redundant or inappropriate comments from the text data; means for generating a pre-alert for a comment and advice on appropriate comments based on the identified comments and notifying the user; means for collecting and storing user feedback data; and means for improving the prediction accuracy of the entire system based on the feedback data, thereby enabling the user to achieve their schedule and communicate smoothly.

[0120] "User" refers to the entity that uses the system, such as the person who inputs schedules and goals, provides feedback, and records voice data.

[0121] "Schedule" refers to the planned actions and tasks that the user intends to carry out on the day.

[0122] "Goals" refer to the specific outcomes or results that a user is trying to achieve on that day.

[0123] "Internal information" refers to information related to the user's internal state, such as their mood or health.

[0124] "External information" refers to information related to the external environment that affects user behavior, such as weather conditions and traffic conditions.

[0125] "Behavioral patterns" refer to the optimal sequence and timing of actions predicted based on the user's schedule, goals, internal information, and external information.

[0126] "Voice data" refers to audio data obtained when a user records everyday conversations.

[0127] "Text data" refers to the result of converting voice data into text information using voice recognition technology.

[0128] "Redundant utterances" refer to utterances in a conversation that are unnecessarily long or overlapping.

[0129] "Inappropriate remarks" refer to remarks that may be misunderstood in a conversation or that may be perceived as offensive by the other person.

[0130] "Advance Alert" refers to a warning or notification to a user about an upcoming action or statement.

[0131] "Feedback data" refers to the actual behavioral results and impressions that users provide in response to predictions and advice from the system.

[0132] "Generative AI models" refer to machine learning models used to predict behavioral patterns and speech content based on user data.

[0133] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[0134] Embodiment of behavior prediction

[0135] User:

[0136] The user launches the smartphone app and inputs their plans and goals for the day. They also input internal information such as their mood and health status. For example, they input specific plans such as "I want to go to the gym in the morning and go shopping in the afternoon," and select "normal" as their mood and "good" as their health status.

[0137] Device:

[0138] The device stores the schedule, goals, and internal information entered by the user in a local database. At the same time, it uses external APIs to obtain external information such as weather conditions and traffic conditions. Specifically, it obtains information such as "sunny" from the weather API and "no traffic jams" from the traffic API.

[0139] server:

[0140] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. Based on the stored data, a generative AI model is used to predict the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, it generates a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[0141] Device:

[0142] The device receives the prediction results from the server and notifies the user, visually displaying them within the app. The user can then act on the results and enter the results as feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[0143] Embodiment of speech prediction

[0144] User:

[0145] Users can launch the smartphone app and record everyday conversations. For example, they can record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[0146] Device:

[0147] The device stores the recorded voice data and memo information in a local database and transmits it to the server.

[0148] server:

[0149] The server converts the voice data received from the device into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, it generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0150] server:

[0151] Based on the analysis results, the server generates appropriate feedback, such as advice like "mention your next trip, but avoid talking about your previous trip," and sends it to the device.

[0152] Device:

[0153] The device notifies the user of the received feedback and displays it in the app. The user can then use the advice to adjust what they say and communicate smoothly. For example, a user can use the advice to choose a topic and smoothly advance the conversation with a friend.

[0154] Prompt Sentence Examples

[0155] Behavioral prediction prompt:

[0156] "The user entered their plan to go to the gym in the morning and do some shopping at the market in the afternoon, and also entered their mood as normal and their health status as good. The device retrieved information from the weather API that it was sunny and from the traffic API that it was clear, and sent this information to the server. Based on this information, the server predicted the optimal behavioral pattern and notified the user to leave for the gym at 9 a.m. and the market at 2 p.m."

[0157] Predictive prompt:

[0158] "A user recorded a conversation with a close friend and left a note saying they wanted to talk about their next trip. The device sent the recorded audio and note to a server. The server converted the audio to text and used a natural language processing model to identify overly long and potentially misleading descriptions. As a result, the server generated feedback and sent it to the device, advising them to mention their next trip but avoid discussing their previous trip."

[0159] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0160] Behavior prediction processing steps

[0161] Step 1: User Data Entry

[0162] input:

[0163] Users launch the smartphone app and enter their plans and goals for the day.

[0164] Users input internal information such as mood and health status.

[0165] Specific behavior:

[0166] The user inputs "I want to go to the gym in the morning and do some shopping in the afternoon," and selects his mood as "normal" and his health status as "good."

[0167] output:

[0168] The input schedule, goals, and internal information are generated.

[0169] Step 2: Saving data on the device and acquiring external information

[0170] input:

[0171] User schedules, goals, and internal information entered into a smartphone app

[0172] Specific behavior:

[0173] The terminal stores the entered information in a local database.

[0174] The device accesses external APIs (weather APIs and traffic APIs) to obtain external information (weather conditions and traffic conditions).

[0175] output:

[0176] User schedules, goals, and internal information stored in a local database

[0177] External information obtained ("Sunny" from the weather API, "No traffic jam" from the traffic API)

[0178] Step 3: Receiving data from the server and making predictions using the AI ​​model

[0179] input:

[0180] Schedules, goals, internal information, and external information sent from the device

[0181] Specific behavior:

[0182] The server receives all data sent by the devices and stores it in a central database.

[0183] The server uses a generative AI model based on the stored data to predict behavioral patterns.

[0184] The server inputs data into an AI model that predicts that the best time to leave for the gym is 9 a.m. and the best time to leave for the market is 2 p.m.

[0185] output:

[0186] Prediction results of optimal behavioral patterns

[0187] Step 4: Notification of results and feedback via device

[0188] input:

[0189] Prediction results sent from the server

[0190] Specific behavior:

[0191] The terminal notifies the user of the prediction result received from the server.

[0192] The device will visually display the prediction results within the app.

[0193] The user acts based on the prediction results and inputs the results into the app as feedback.

[0194] output:

[0195] User notifications and in-app displays

[0196] Feedback Data

[0197] ---

[0198] Speech prediction processing steps

[0199] Step 1: User voice recording and note taking

[0200] input:

[0201] Users launch the smartphone app and record their everyday conversations.

[0202] The user enters notes for the recording.

[0203] Specific behavior:

[0204] For example, a user records a conversation with a friend and enters a note saying, "I want to talk about my next trip."

[0205] output:

[0206] Recorded audio data and memo information

[0207] Step 2: Save and send data from your device

[0208] input:

[0209] Recorded audio data and memo information

[0210] Specific behavior:

[0211] The device stores the recorded audio data and memo information in a local database.

[0212] The terminal transmits the saved data to the server.

[0213] output:

[0214] Audio data and notes stored in a local database

[0215] Data sent to the server

[0216] Step 3: Audio data conversion and analysis on the server

[0217] input:

[0218] Voice data and memo information sent from the device

[0219] Specific behavior:

[0220] The server converts the voice data into text using a voice recognition engine.

[0221] The server analyzes the converted text data using a natural language processing model.

[0222] The server identifies redundant or inappropriate statements and generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0223] output:

[0224] Text data and analysis results

[0225] Step 4: Server feedback generation and sending

[0226] input:

[0227] Analysis results

[0228] Specific behavior:

[0229] The server generates appropriate utterance feedback based on the analysis results.

[0230] For example, advice such as "mention your next trip, but avoid talking about your previous trip" may be generated.

[0231] The server transmits the generated feedback to the terminal.

[0232] output:

[0233] Feedback Data

[0234] Step 5: Feedback notification by device

[0235] input:

[0236] Feedback data sent from the server

[0237] Specific behavior:

[0238] The terminal notifies the user of the feedback data.

[0239] The device displays the feedback within the app, giving users information to adjust what they say.

[0240] output:

[0241] User feedback notification and in-app display

[0242] (Application example 1)

[0243] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0244] Conventional behavior prediction systems and comment prediction systems are limited to supporting actions based on the user's schedule and goals, making it difficult to address diverse user needs. Furthermore, they lack sufficient support for content recommendations and communication improvement, resulting in a lack of improvement in the user experience. Therefore, there is a need for a comprehensive system that provides appropriate advice based not only on the user's daily actions and comments, but also on their content consumption trends and message content.

[0245] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0246] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information and external information, means for presenting the predicted behavioral patterns to the user, means for predicting the user's content consumption tendencies using a generative artificial intelligence model and recommending optimal content, and means for providing advice to improve the user's comments based on the content of the user's message.

[0247] This allows users to receive comprehensive support for their actions, content recommendations, and advice on what to say.

[0248] A "user" is an individual who uses the system to receive assistance with schedule input, information gathering, behavior prediction, and speech advice.

[0249] A "generative artificial intelligence model" is an algorithm that learns various patterns and trends based on data and makes predictions and recommendations.

[0250] "Content consumption trends" refers to patterns such as what types of content a user prefers to consume and what types of content they consume at what times of the day.

[0251] "Message content" refers to all statements and texts made by a User when communicating with others.

[0252] "Advice" means instructions or guidance provided to a user to help them act or speak more appropriately.

[0253] "Schedules" refer to planned actions and events that a user undertakes in their daily life.

[0254] A "goal" is a specific action or outcome that a user aims to achieve.

[0255] "Internal information" refers to information about the user's inner self, such as their emotional state or health status.

[0256] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[0257] "Behavioral prediction" is the calculation of optimal behavioral patterns based on a user's schedule, goals, internal information, and external information.

[0258] A "behavioral pattern" refers to a series of actions that a user takes at what timing.

[0259] "Content" is a general term for information assets consumed by users, such as music, videos, and articles.

[0260] The system for carrying out the present invention integrates various functions for predicting and optimizing user actions and comments. Specific embodiments will be described below.

[0261] Hardware and Software Configuration

[0262] Hardware configuration:

[0263] Devices such as smartphones, tablets, or computers

[0264] Servers (including using cloud computing environments)

[0265] Software configuration:

[0266] Application software (smartphone apps, etc.)

[0267] Databases (local and central)

[0268] External APIs (weather API, traffic API, etc.)

[0269] Speech Recognition Engine

[0270] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)

[0271] Generative artificial intelligence model (AI model)

[0272] Embodiment of behavior prediction

[0273] user:

[0274] Users open the app on their smartphone or tablet and enter their plans and goals for the day, as well as internal information such as their mood and health status.

[0275] Device:

[0276] The device stores user-entered schedules, goals, and internal information in a local database, and then uses external APIs to retrieve external information such as weather and traffic conditions.

[0277] server:

[0278] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. It then uses this data to predict the user's behavior using a generative artificial intelligence model. It then sends the prediction results to the device and notifies the user.

[0279] Embodiment of speech prediction

[0280] user:

[0281] Users start the smartphone app and record their everyday conversations, taking notes on topics they want to talk about.

[0282] Device:

[0283] The recorded voice data is stored in a local database and then sent to the server, along with any memo information.

[0284] server:

[0285] The voice data is converted into text by a speech recognition engine, and the text data is analyzed by a natural language processing model to identify redundant or inappropriate utterances, and based on that, appropriate feedback is generated and sent to the device.

[0286] Device:

[0287] Users receive feedback and adjust what they say based on in-app suggestions.

[0288] Embodiment of content recommendation function

[0289] user:

[0290] Users input their goals, such as "I want to relax in the evening" or "I want to catch up on the latest news on Sunday afternoon."

[0291] Device:

[0292] The device sends data to the server based on internal information and information from external APIs.

[0293] server:

[0294] The server uses a generative artificial intelligence model to predict the user's content consumption habits, recommends the most suitable content (movies, music, news articles, etc.) for the user, and sends the results to the device.

[0295] Device:

[0296] Users receive recommendations and choose the content they see on the app to get the best experience.

[0297] Specific examples

[0298] The following example prompts could be fed into a generative AI model:

[0299] Example prompt sentence:

[0300] The user entered their goal of "I want to relax at night," selected "Relaxed" as their mood, and "Energetic" as their health condition. Weather information obtained from an external API showed that the temperature was below 20 degrees and the traffic situation was clear. Based on this information, please recommend the most suitable relaxation content for the user.

[0301] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0302] Step 1:

[0303] The user launches the app on their smartphone or tablet and inputs their plans and goals for the day, such as "I want to relax in the evening," as well as their mood and health status. The input information is then stored in a local database on the device.

[0304] Input: User's schedule, goals, internal information

[0305] Output: Save data to a local database

[0306] Specific behavior: A user enters information into the app's input form and presses the "Save" button, which saves the data to a local database.

[0307] Step 2:

[0308] The device calls external APIs such as weather APIs and traffic APIs to obtain external information such as weather conditions and traffic conditions. The obtained external information is stored in a local database.

[0309] Input: External API request

[0310] Output: External data such as weather information, traffic information, etc.

[0311] Specific operation: The device periodically sends requests to an external API and stores the data obtained in response in a local database.

[0312] Step 3:

[0313] The device transmits the user's schedule, goals, internal information, and external information from a local database to the server, where the transmitted data is stored in a central database.

[0314] Input: All data in the local database

[0315] Output: Send data to server, store in central database

[0316] Specific operation: The terminal periodically uploads all data to the server, and the server stores the received data in a central database.

[0317] Step 4:

[0318] Based on all the data received by the server, a generative artificial intelligence model is used to predict user behavior. For example, if a user inputs "I want to relax at night," the model predicts appropriate relaxing content.

[0319] Input: All data in the central database

[0320] Output: Behavior prediction results

[0321] Specific operation: The server runs a generative artificial intelligence model to predict optimal actions based on past and new data.

[0322] Step 5:

[0323] The server sends the prediction result to the device, and the device notifies the user. For example, the server generates a prediction result such as "This movie is recommended for relaxing at night" and sends it to the device.

[0324] Input: Behavior prediction result

[0325] Output: User notification

[0326] Specific operation: The device receives the prediction results sent from the server and notifies the user visually within the app.

[0327] Step 6:

[0328] Users can take recommended actions based on the app's notifications, such as watching a recommended movie, and provide feedback to the app, allowing the system to learn from that feedback and improve its prediction accuracy next time.

[0329] Input: User feedback

[0330] Output: Accumulation of feedback data

[0331] What it does: The user enters feedback within the app, such as "I finished watching the movie," and that data is stored in a local database.

[0332] Step 7:

[0333] It records users' everyday conversations and collects audio data. For example, a user records a conversation with a friend and leaves a note saying, "I want to talk about my next trip."

[0334] Input: Audio data, memo information

[0335] Output: Save data to a local database

[0336] What happens: A user uses the app's recording feature to record a conversation and takes notes within the app.

[0337] Step 8:

[0338] The recorded voice data is stored in a local database and sent to a server, along with any memo information. The server converts the voice data into text using a speech recognition engine, and the text data is analyzed using a natural language processing model.

[0339] Input: Audio data, memo information

[0340] Output: Text data, analysis results

[0341] Specific operation: The server uses a speech recognition engine to convert speech to text and analyzes the text data using a natural language processing model.

[0342] Step 9:

[0343] Based on the results of the analysis using the natural language processing model, appropriate feedback (e.g., "Avoid talking about your previous trip") is generated and sent to the device. The device notifies the user of the feedback and displays it appropriately within the app.

[0344] Input: Analysis results

[0345] Output: Feedback results

[0346] Specific operation: The server generates appropriate feedback and sends it to the device, which notifies the user of the feedback and displays it.

[0347] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0348] This invention is a system that combines "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals; "statement prediction," which provides appropriate advice on the content and timing of statements; and "emotion recognition," which recognizes the user's emotions.

[0349] Embodiment of behavior prediction

[0350] User:

[0351] The user launches the smartphone app and inputs the plan and goals for the day. Internal information such as mood and health status is also input. For example, the user can provide information such as "I want to go to the gym in the morning and go shopping in the afternoon" or "I feel normal and my health is good."

[0352] Device:

[0353] The system stores the schedule, goals, and internal information entered by the user in a local database. It then uses an external API to obtain external information such as weather and traffic conditions. For example, it collects information such as "sunny" and "no traffic jam."

[0354] server:

[0355] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. In addition, an emotion engine extracts emotional information from the user's input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[0356] server:

[0357] Generate prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and send them to the device.

[0358] Device:

[0359] The received action prediction results are notified to the user and displayed visually within the app. The user can act based on the results and enter feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[0360] Embodiment of speech prediction

[0361] User:

[0362] Start a smartphone app and record your everyday conversations. For example, record a conversation with your best friend and leave a note saying, "I want to talk about my next trip."

[0363] Device:

[0364] The recorded voice data and notes are stored in a local database and sent to a server, where an emotion engine extracts the user's emotional information from the recorded data and sends it together.

[0365] server:

[0366] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[0367] server:

[0368] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[0369] Device:

[0370] The app notifies users of the feedback it receives and visually displays it within the app, allowing users to adjust what they say based on that feedback and improve communication. For example, when talking with a friend, users can use the advice to choose topics and keep the conversation flowing smoothly.

[0371] Emotion Recognition Embodiment

[0372] User:

[0373] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[0374] Device:

[0375] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[0376] server:

[0377] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[0378] Specific examples

[0379] Specific examples of behavioral prediction

[0380] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[0381] Specific examples of speech prediction

[0382] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[0383] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[0384] The processing flow will be explained below.

[0385] Program processing of behavior prediction

[0386] Step 1:

[0387] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[0388] Step 2:

[0389] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[0390] Step 3:

[0391] Device: Stores user-entered schedules, goals, and internal information in a local database.

[0392] Step 4:

[0393] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[0394] Step 5:

[0395] Terminal: Send all collected data, including internal and external information, to the server.

[0396] Step 6:

[0397] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[0398] Step 7:

[0399] Server: Extracts user emotional information using the emotion engine based on the stored data.

[0400] Step 8:

[0401] Server: Based on the stored data and emotional information, the AI ​​model predicts optimal behavioral patterns.

[0402] Step 9:

[0403] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[0404] Step 10:

[0405] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[0406] Step 11:

[0407] User: Follows the suggested behavioral patterns and achieves the goal.

[0408] Step 12:

[0409] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[0410] ---

[0411] Program processing of speech prediction

[0412] Step 1:

[0413] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[0414] Step 2:

[0415] Users: Leave notes during conversations about things that concern them or that need attention.

[0416] Step 3:

[0417] Device: Stores recorded audio data and notes in a local database.

[0418] Step 4:

[0419] Terminal: Sends voice data and memo information to the server.

[0420] Step 5:

[0421] Server: Converts voice data received from the device into text using a speech recognition engine.

[0422] Step 6:

[0423] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[0424] Step 7:

[0425] Server: Extracts user emotion information from voice and text data using an emotion engine.

[0426] Step 8:

[0427] Server: Generates appropriate feedback based on sentiment information and text analysis results (e.g., "Mention your next trip, but avoid talking about your previous trip").

[0428] Step 9:

[0429] Server: Sends the generated feedback to the device.

[0430] Step 10:

[0431] On the device: Notify the user of the feedback received and display it visually within the app.

[0432] Step 11:

[0433] Users: Use suggested feedback to adjust conversations and improve communication.

[0434] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[0435] Example 2

[0436] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0437] Conventional behavior prediction systems only predict behavior based on a user's schedule and goals, and have the problem of being unable to provide appropriate feedback or advice that takes emotional information into account. Furthermore, speech prediction systems are limited to identifying redundant or inappropriate speech, making it difficult to provide advice that appropriately reflects the user's emotions. This has led to the issue of a lack of effective support for user behavior and speech.

[0438] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0439] In this invention, the server includes: means for a user to input internal information such as the user's schedule and goals for the day, mood, and health status; means for collecting the user's internal information and external information such as weather conditions and traffic conditions; means for predicting the user's behavior based on the schedule, goals, internal information, and external information and generating an optimal behavior pattern based on emotional information; means for notifying the user of the predicted behavior pattern and collecting user feedback; means for recording the user's daily conversations and collecting voice data; means for converting the voice data into text; means for identifying redundant or inappropriate statements from the text data and advising the user on appropriate statements based on emotional information; and means for generating a pre-alert for statements and advice on appropriate statements based on the identified statements and notifying the user. This makes it possible to optimize the user's behavior and statements while taking emotions into consideration, thereby providing effective support.

[0440] "User" refers to an individual or corporation that uses the system, inputs their own schedules, goals, and internal information, and receives feedback from the system.

[0441] "Internal information" is information about a user's internal state, such as their mood, health, or emotions.

[0442] "External information" refers to information about the user's external environment, such as weather conditions and traffic conditions, and is obtained using an external API.

[0443] "Emotion information" is information about the user's emotional state, and is extracted by an emotion engine or the like.

[0444] An "emotion engine" is a software or hardware mechanism for extracting emotional information from user input or voice data.

[0445] A "behavioral pattern" is a schedule or action plan for the user to act optimally, and is the result predicted by the system.

[0446] "Feedback" is information that users input into the system based on the results of their actual actions, and the system uses this information to make its next prediction more accurate.

[0447] "Speech prediction" is a system function that analyzes the user's everyday conversations and advises them on appropriate content to say.

[0448] A "speech recognition engine" is a software or hardware mechanism for converting a user's voice data into text data.

[0449] A "natural language processing model" is a technology for analyzing text data and identifying redundant or inappropriate statements.

[0450] A "pre-alert" is a warning message that the system displays before a user makes an identified inappropriate comment.

[0451] "Advice on appropriate speech" refers to advice on speech content provided by the system to help users communicate more smoothly.

[0452] This invention is a system that combines "behavior prediction" that efficiently predicts a user's daily behavior and helps them achieve their schedules and goals, "statement prediction" that provides appropriate advice on the content and timing of statements, and "emotion recognition" that recognizes the user's emotions. A specific embodiment of this system will be described below.

[0453] Embodiment of behavior prediction

[0454] User:

[0455] Users launch the smartphone app and enter detailed information about the day's plans, goals, mood, health status, etc. For example, they might enter, "I want to go to the gym in the morning and do some shopping in the afternoon," or "I'm feeling normal, and my health is good."

[0456] Device:

[0457] The device stores the user's schedule, goals, mood, and health status in a local database. It then uses external APIs to obtain external information such as weather conditions (e.g., OpenWeatherMap API) and traffic conditions (e.g., Google Maps Traffic API). For example, it collects information such as "sunny" and "no traffic jam" and sends it to the server.

[0458] server:

[0459] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. At the same time, it uses an emotion engine to extract emotional information from the user's input. Based on the stored data and emotional information, the AI ​​model predicts the optimal behavioral pattern. For example, it generates a result such as "It is best to leave for the gym at 9:00 AM and the market at 2:00 PM" and sends it to the device.

[0460] Device:

[0461] The device then notifies the user of the predicted behavior and displays it visually within the app. The user can then act based on the results and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[0462] Embodiment of speech prediction

[0463] User:

[0464] Users can launch the smartphone app and record everyday conversations, such as a conversation with a close friend, and leave a note saying, "I want to talk about my next trip."

[0465] Device:

[0466] The device stores the recorded voice data and notes in a local database, and uses an emotion engine to extract the user's emotional information from the recorded data and send it together with the data to the server.

[0467] server:

[0468] The server converts the voice data into text using a speech recognition engine (e.g., Google Speech-to-Text API), analyzes the converted text data using a natural language processing model, and identifies redundant or inappropriate statements. For example, it generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading." It then generates appropriate feedback based on the emotional information. For example, it generates advice such as "Talk about your next trip, but avoid talking about your previous trip," and sends it to the device.

[0469] Device:

[0470] The device will notify the user of the received feedback and display it visually within the app, allowing the user to adjust what they say based on this feedback and improve communication. For example, when talking with a friend, the user can use the advice to choose topics and keep the conversation flowing smoothly.

[0471] Emotion Recognition Embodiment

[0472] User:

[0473] Users provide emotional information to the app by inputting their mood and emotions and recording changes in facial expressions and voice during conversations.

[0474] Device:

[0475] The device uses an emotion recognition engine to extract emotional information from the recorded voice data and transmits the emotional information to a server.

[0476] server:

[0477] The server uses the emotional information analyzed by the emotion recognition engine and natural language processing model to generate behavioral patterns and speech feedback based on the user's emotions.

[0478] Specific examples

[0479] Specific examples of behavioral prediction

[0480] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "Sunny" from a weather API (e.g., OpenWeatherMap API) and "No traffic jam" from a traffic API (e.g., Google Maps Traffic API), and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "The best time to leave for the gym is 9:00 AM and the market is 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[0481] Example prompt sentence:

[0482] User: I want to go to the gym in the morning and do some shopping in the afternoon. I feel normal and my health is good.

[0483] Terminal: The weather is clear and traffic is smooth.

[0484] Server: The best time to leave for the gym is 9am and the market is 2pm.

[0485] Specific examples of speech prediction

[0486] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text (e.g., Google Speech-to-Text API) and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[0487] Example prompt sentence:

[0488] User: Recording a conversation with a best friend. Want to talk about an upcoming trip.

[0489] Device: Sends voice data, notes, and emotional information to the server.

[0490] Server: Analyzes the text for redundant content and generates appropriate feedback.

[0491] Server: Mention your next trip, but avoid talking about your last trip.

[0492] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0493] Behavior prediction processing steps

[0494] Step 1:

[0495] User: Launches the smartphone app and enters internal information such as the day's plans, goals, mood, and health status.

[0496] Input: Schedule, Goals, Mood, Health

[0497] Output: User input data

[0498] For example: "I go to the gym in the morning and shop at the market in the afternoon." "I feel fine and my health is good."

[0499] Step 2:

[0500] Device: Stores the entered information in a local database, then uses external APIs to retrieve external information such as weather conditions and traffic conditions.

[0501] Input: User-entered data

[0502] Output: Internal and external information

[0503] Examples: "The weather is clear" and "The traffic is smooth."

[0504] Step 3:

[0505] Server: Receives all data sent from the devices and stores it in a central database. Uses an emotion engine to extract emotional information from user input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[0506] Input: User input data, internal information, external information

[0507] Output: Predicted behavior pattern

[0508] For example: "The best time to leave for the gym is 9am and the best time to leave for the market is 2pm."

[0509] Step 4:

[0510] Server: Sends the prediction results to the device.

[0511] Input: Predicted behavior pattern

[0512] Output: Data sent to the terminal

[0513] Step 5:

[0514] On the device: The user is notified of the received predictions and visually displayed within the app, allowing the user to act on them and provide feedback to the app.

[0515] Input: Predicted behavior pattern

[0516] Output: User feedback

[0517] Step 6:

[0518] Device: User feedback is sent to the server to help improve prediction accuracy in future predictions.

[0519] Input: User feedback

[0520] Output: Data sent to the server

[0521] Speech prediction processing steps

[0522] Step 1:

[0523] User: Launches the smartphone app, records everyday conversations, and enters notes.

[0524] Input: Recording data, notes

[0525] Output: Recording data and notes

[0526] Examples: "I want to record a conversation with my best friend" or "I want to talk about my next trip."

[0527] Step 2:

[0528] Device: Recorded voice data and notes are stored in a local database, and emotional information is extracted using an emotion engine, which then sends the data to the server.

[0529] Input: Recording data, notes

[0530] Output: Recording data, notes, emotional information

[0531] Step 3:

[0532] Server: The speech data is converted into text using a speech recognition engine, and a natural language processing model is used to identify redundant or inappropriate speech.

[0533] Input: Audio recording, emotional information

[0534] Output: Text data and detected verbose / inappropriate comments

[0535] For example: "This part is redundant and difficult to understand" or "This part may be misleading"

[0536] Step 4:

[0537] Server: Generates appropriate feedback based on emotion information and sends it to the device.

[0538] Input: Text data, emotion information

[0539] Output: Feedback

[0540] For example: "Talk about your next trip, but avoid talking about your last trip."

[0541] Step 5:

[0542] On-device: The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust their voice accordingly.

[0543] Input: Feedback

[0544] Output: Adjusted speech data

[0545] Processing steps for emotion recognition

[0546] Step 1:

[0547] User: Inputs mood and emotions and records facial expressions and vocal changes during conversations.

[0548] Input: Mood, emotion, facial expression, vocal changes

[0549] Output: Emotion data

[0550] Step 2:

[0551] Terminal: The emotion recognition engine extracts emotional information from the voice data and sends it to the server.

[0552] Input: Recording data

[0553] Output: Emotional information

[0554] Step 3:

[0555] Server: Analyzes emotional information and generates behavioral patterns and speech feedback based on the user's emotions.

[0556] Input: Emotion information

[0557] Output: Behavioral patterns, speech feedback

[0558] (Application example 2)

[0559] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0560] Currently, there is a lack of security systems that understand users' daily behavior and emotions and provide advice on appropriate behavioral patterns and speech based on that understanding. This can make it difficult for users to understand the degree of risk involved in their own behavior and speech, making it difficult to ensure their safety. The present invention aims to provide a system that comprehensively analyzes users' behavior, speech, and emotional information to reduce security risks.

[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0562] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information, and external information, means for presenting the predicted behavioral pattern to the user, means for predicting the user's utterances from the content of the user's conversation and recorded voice data and providing appropriate advice, and means for recognizing the user's emotions and reflecting them in the behavioral pattern and utterance content, thereby enabling the user to understand the degree to which their actions and utterances pose security risks and to take appropriate preventative measures.

[0563] "User" refers to an individual who uses this system.

[0564] "Means for entering the day's plans and goals" refers to the function that allows users to enter their daily plans and goals into the system.

[0565] "Internal information" refers to information about a user's personal state or emotions, such as mood or health.

[0566] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[0567] "Means for predicting behavior" refers to the function of calculating the optimal behavioral pattern for a user based on internal and external information.

[0568] "Means for presenting predicted behavioral patterns" refers to the function of notifying the user of the predicted results and displaying them visually.

[0569] "Means of predicting what will be said based on conversation content and recorded audio data and providing appropriate advice" refers to a function that analyzes a user's everyday conversations and generates dedicated alerts and advice.

[0570] "Means of recognizing emotions and reflecting them in behavioral patterns and speech content" refers to a function that analyzes the user's emotions and provides behavioral patterns and speech advice that take these into consideration.

[0571] A "system that reduces security risks" refers to a system that evaluates the risks associated with users' actions and statements and supports safe behavior and communication.

[0572] This invention combines a system that efficiently predicts a user's daily behavior and supports safe behavior, a system that analyzes conversation content and gives advice on appropriate remarks, and a system that recognizes the user's emotions and assesses risk. The system is intended to be used by users through a smartphone app.

[0573] To implement this system, the following hardware and software are used:

[0574] Hardware: smartphones, servers, cameras and microphones for emotion recognition

[0575] Software: Python, external API (weather API, traffic API), voice recognition AI engine (Hugging Face transformers), emotion recognition AI engine (DeepFace)

[0576] Embodiment of behavior prediction

[0577] Device:

[0578] The user launches the app on their smartphone and enters their plans and goals for the day, as well as internal information such as their mood and health status. The device stores this information in a local database and uses external APIs to obtain external information such as weather and traffic conditions.

[0579] server:

[0580] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. The AI ​​model uses this data to predict optimal behavioral patterns. For example, it generates a result such as "It's sunny and traffic is good, so it's best to leave for the gym at 9:00 AM." The server then sends this prediction to the device.

[0581] Device:

[0582] The predictions received are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app.

[0583] Embodiment of speech prediction

[0584] Device:

[0585] The user launches the app on their smartphone and records their everyday conversations. The recorded audio data is stored in a local database and sent to a server. The emotion engine extracts the user's emotional information from the recorded data and sends it along with the data.

[0586] server:

[0587] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[0588] server:

[0589] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[0590] Device:

[0591] The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust what they say and communicate more effectively.

[0592] Emotion Recognition Embodiment

[0593] user:

[0594] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[0595] Device:

[0596] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[0597] server:

[0598] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[0599] Specific examples

[0600] Specific examples of behavioral prediction

[0601] The user inputs "I plan to go to the gym at 6 PM" and also inputs their mood and health status as additional information. The device obtains "sunny" information from the weather API and "no traffic jams" information from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict that "the best time to leave is 5 PM" and sends the result to the device. The device notifies the user of the result, and the user acts accordingly.

[0602] Specific examples of speech prediction

[0603] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[0604] Prompt Sentence Examples

[0605] Behavioral prediction: "Predict what time a user should go to the gym based on current weather and traffic information."

[0606] Speech Prediction: "Identify sensitive information that users should avoid in conversations with their friends."

[0607] Emotion Recognition: "Perform emotion analysis using the user's current image to conduct risk assessment."

[0608] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0609] Behavior prediction processing steps

[0610] Step 1: User enters schedule and goal

[0611] The user launches the smartphone app and enters the day's plans, goals, and internal information (mood and health status).

[0612] Input: Schedule, goals, internal information

[0613] Output: Data stored in a local database

[0614] Step 2: Obtaining external information

[0615] The device calls external APIs (weather API, traffic API) to obtain external information such as weather conditions and traffic conditions.

[0616] Input: External API endpoint

[0617] Output: Weather conditions, traffic conditions (e.g. sunny, no traffic jams)

[0618] Step 3: Sending data

[0619] The terminal transmits the input internal information and the acquired external information to the server.

[0620] Input: Plan, goal, internal information, external information

[0621] Output: All data sent to the server

[0622] Step 4: Predicting behavioral patterns

[0623] Based on all the data received by the server, an AI model is used to predict optimal behavioral patterns.

[0624] Input: All data and AI model

[0625] Output: Predicted behavioral pattern (e.g., departure at 5pm)

[0626] Step 5: Send prediction results

[0627] The server sends the prediction results to the device.

[0628] Input: predicted behavior pattern

[0629] Output: Prediction results sent to the device

[0630] Step 6: Visualizing the results

[0631] The device notifies the user of the prediction results received and displays them visually within the app.

[0632] Input: predicted behavior pattern

[0633] Output: The result communicated to the user

[0634] Step 7: Provide feedback

[0635] The user provides feedback on the results of the execution to the app.

[0636] Input: Actual action results

[0637] Output: Feedback data

[0638] Speech prediction processing steps

[0639] Step 1: Recording audio data

[0640] Users launch the smartphone app and record their everyday conversations.

[0641] Input: Audio data

[0642] Output: Audio data stored in a local database

[0643] Step 2: Sending audio data

[0644] The device sends the recorded audio data to the server.

[0645] Input: Audio data

[0646] Output: Audio data sent to the server

[0647] Step 3: Voice Recognition

[0648] The server converts the voice data into text using a voice recognition engine.

[0649] Input: Audio data

[0650] Output: Text data

[0651] Step 4: Parsing utterances

[0652] The server analyzes the text data using a natural language processing model to identify inappropriate or redundant statements.

[0653] Input: Text data, NLP model

[0654] Output: Analysis results (identification of inappropriate and redundant parts)

[0655] Step 5: Generate Advice

[0656] The server generates feedback on the speech based on the analysis results and emotional information.

[0657] Input: Analysis results, emotion information

[0658] Output: Feedback (e.g. mention the next trip but avoid talking about the previous trip)

[0659] Step 6: Submit your feedback

[0660] The server generates feedback and sends it to the device.

[0661] Input: Feedback

[0662] Output: Feedback sent to the terminal

[0663] Step 7: Notification of feedback

[0664] Notify the user of the feedback received by the device and display it visually within the app.

[0665] Input: Feedback

[0666] Output: Feedback given to the user

[0667] Processing steps for emotion recognition

[0668] Step 1: Enter emotional information

[0669] Users input their moods and emotions into the app.

[0670] Input: Mood, emotional information

[0671] Output: Emotion information stored in a local database

[0672] Step 2: Analyzing the audio data

[0673] Emotional information is extracted from the audio data recorded by the device and sent to the server.

[0674] Input: Audio data

[0675] Output: Extracted emotion information

[0676] Step 3: Sending emotional information

[0677] The device transmits the emotion information to the server.

[0678] Input: Emotion information

[0679] Output: Emotion information sent to the server

[0680] Step 4: Sentiment-based analysis

[0681] The server uses the emotion information to complement the analysis results of behavior prediction and utterance prediction.

[0682] Input: Emotion information, behavior prediction, speech prediction

[0683] Output: Analysis results (behavior and speech) supplemented based on emotions

[0684] Step 5: Notification of completion results

[0685] The server transmits the completed analysis results to the terminal.

[0686] Input: Completed analysis results

[0687] Output: Completion results sent to the terminal

[0688] Step 6: Visualizing the results

[0689] The device notifies the user of the completion results received and displays them visually within the app.

[0690] Input: Completed analysis results

[0691] Output: Completion result notified to the user

[0692] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0693] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0694] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0695] [Second embodiment]

[0696] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0697] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0698] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0699] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0700] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0701] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0702] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0703] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0704] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0705] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0706] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0707] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0708] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[0709] Embodiment of behavior prediction

[0710] User:

[0711] Users launch the smartphone app and enter their plans and goals for the day. They also enter internal information such as their mood and health status. For example, a user might enter their plan to "go to the gym in the morning and go shopping in the afternoon," select "normal" as their mood, and "good" as their health status.

[0712] Device:

[0713] The schedule, goals, and internal information entered by the user are stored in a local database. Then, external APIs are used to obtain external information such as weather conditions and traffic conditions. For example, a weather API is called to obtain information such as "sunny," and a traffic API is called to obtain information such as "no traffic jams."

[0714] server:

[0715] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. Based on the stored data, an AI model predicts the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, the AI ​​model may generate a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[0716] Device:

[0717] The received prediction results are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[0718] Embodiment of speech prediction

[0719] User:

[0720] Start the smartphone app and record your everyday conversations. For example, record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[0721] Device:

[0722] The recorded voice data is stored in a local database and sent to the server, along with any memo information.

[0723] server:

[0724] The voice data received from the device is converted into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0725] server:

[0726] Based on the analysis results, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent to the device.

[0727] Device:

[0728] The received feedback is notified to the user and displayed within the app, allowing the user to adjust what they say based on the feedback and improve communication. For example, when talking with a friend, a user can refer to the advice to choose a topic and smoothly advance the conversation.

[0729] Specific examples

[0730] Specific examples of behavioral prediction

[0731] The user enters their schedule into the app, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict behavioral patterns such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends this to the device. The device notifies the user of the results, and the user acts accordingly.

[0732] Specific examples of speech prediction

[0733] A user records a conversation with a close friend and leaves a note saying, "I'd like to talk about my next trip." The device then sends the recorded audio data and note to a server. The server converts the audio data into text and uses a natural language processing model to identify "overly long explanations" and "potentially misleading" parts. Based on this, the device generates and sends feedback to the device, such as "mention your next trip, but avoid talking about your previous trip." The device then notifies the user of the feedback, and the user can use the advice to smoothly move the conversation forward.

[0734] This allows users to receive support in achieving their schedules and goals, as well as in communicating smoothly.

[0735] The processing flow will be explained below.

[0736] Program processing of behavior prediction

[0737] Step 1:

[0738] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[0739] Step 2:

[0740] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[0741] Step 3:

[0742] Device: Stores user-entered schedules, goals, and internal information in a local database.

[0743] Step 4:

[0744] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[0745] Step 5:

[0746] Terminal: Send all collected data, including internal and external information, to the server.

[0747] Step 6:

[0748] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[0749] Step 7:

[0750] Server: Based on the stored data, it uses an AI model to predict the user's optimal behavioral patterns.

[0751] Step 8:

[0752] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[0753] Step 9:

[0754] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[0755] Step 10:

[0756] User: Follows the suggested behavioral patterns and achieves the goal.

[0757] Step 11:

[0758] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[0759] Program processing of speech prediction

[0760] Step 1:

[0761] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[0762] Step 2:

[0763] Users: Leave notes during conversations about things that concern them or that need attention.

[0764] Step 3:

[0765] Device: Stores recorded audio data and notes in a local database.

[0766] Step 4:

[0767] Terminal: Sends voice data and memo information to the server.

[0768] Step 5:

[0769] Server: Converts voice data received from the device into text using a speech recognition engine.

[0770] Step 6:

[0771] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[0772] Step 7:

[0773] Server: Based on the analysis results, generate appropriate feedback based on the profile information of the person you are speaking to.

[0774] Step 8:

[0775] Server: Sends the generated feedback (e.g., "mention your next trip, but avoid talking about your previous trip") to the device.

[0776] Step 9:

[0777] On the device: Notify the user of the feedback received and display it visually within the app.

[0778] Step 10:

[0779] Users: Use suggested feedback to adjust conversations and improve communication.

[0780] This allows users to effectively achieve their schedules and goals, as well as communicate smoothly.

[0781] Example 1

[0782] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0783] Conventional systems have difficulty predicting behavior based on internal and external information in addition to users' schedules and goals. Furthermore, there was a lack of systems that notified users of appropriate timing and content for speech in everyday conversations. Furthermore, there was no mechanism in place to incorporate user feedback to improve the system's overall prediction accuracy. This made it difficult for users to achieve their schedules and communicate smoothly.

[0784] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0785] In this invention, the server includes: means for a user to input the day's schedule and goals; means for collecting the user's internal and external information; means for predicting the user's behavior based on the schedule, goals, internal and external information; means for presenting the predicted behavioral pattern to the user; means for collecting voice data and converting it into text; means for identifying redundant or inappropriate comments from the text data; means for generating a pre-alert for a comment and advice on appropriate comments based on the identified comments and notifying the user; means for collecting and storing user feedback data; and means for improving the prediction accuracy of the entire system based on the feedback data, thereby enabling the user to achieve their schedule and communicate smoothly.

[0786] "User" refers to the entity that uses the system, such as the person who inputs schedules and goals, provides feedback, and records voice data.

[0787] "Schedule" refers to the planned actions and tasks that the user intends to carry out on the day.

[0788] "Goals" refer to the specific outcomes or results that a user is trying to achieve on that day.

[0789] "Internal information" refers to information related to the user's internal state, such as their mood or health.

[0790] "External information" refers to information related to the external environment that affects user behavior, such as weather conditions and traffic conditions.

[0791] "Behavioral patterns" refer to the optimal sequence and timing of actions predicted based on the user's schedule, goals, internal information, and external information.

[0792] "Voice data" refers to audio data obtained when a user records everyday conversations.

[0793] "Text data" refers to the result of converting voice data into text information using voice recognition technology.

[0794] "Redundant utterances" refer to utterances in a conversation that are unnecessarily long or overlapping.

[0795] "Inappropriate remarks" refer to remarks that may be misunderstood in a conversation or that may be perceived as offensive by the other person.

[0796] "Advance Alert" refers to a warning or notification to a user about an upcoming action or statement.

[0797] "Feedback data" refers to the actual behavioral results and impressions that users provide in response to predictions and advice from the system.

[0798] "Generative AI models" refer to machine learning models used to predict behavioral patterns and speech content based on user data.

[0799] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[0800] Embodiment of behavior prediction

[0801] User:

[0802] The user launches the smartphone app and inputs their plans and goals for the day. They also input internal information such as their mood and health status. For example, they input specific plans such as "I want to go to the gym in the morning and go shopping in the afternoon," and select "normal" as their mood and "good" as their health status.

[0803] Device:

[0804] The device stores the schedule, goals, and internal information entered by the user in a local database. At the same time, it uses external APIs to obtain external information such as weather conditions and traffic conditions. Specifically, it obtains information such as "sunny" from the weather API and "no traffic jams" from the traffic API.

[0805] server:

[0806] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. Based on the stored data, a generative AI model is used to predict the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, it generates a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[0807] Device:

[0808] The device receives the prediction results from the server and notifies the user, visually displaying them within the app. The user can then act on the results and enter the results as feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[0809] Embodiment of speech prediction

[0810] User:

[0811] Users can launch the smartphone app and record everyday conversations. For example, they can record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[0812] Device:

[0813] The device stores the recorded voice data and memo information in a local database and transmits it to the server.

[0814] server:

[0815] The server converts the voice data received from the device into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, it generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0816] server:

[0817] Based on the analysis results, the server generates appropriate feedback, such as advice like "mention your next trip, but avoid talking about your previous trip," and sends it to the device.

[0818] Device:

[0819] The device notifies the user of the received feedback and displays it in the app. The user can then use the advice to adjust what they say and communicate smoothly. For example, a user can use the advice to choose a topic and smoothly advance the conversation with a friend.

[0820] Prompt Sentence Examples

[0821] Behavioral prediction prompt:

[0822] "The user entered their plan to go to the gym in the morning and do some shopping at the market in the afternoon, and also entered their mood as normal and their health status as good. The device retrieved information from the weather API that it was sunny and from the traffic API that it was clear, and sent this information to the server. Based on this information, the server predicted the optimal behavioral pattern and notified the user to leave for the gym at 9 a.m. and the market at 2 p.m."

[0823] Predictive prompt:

[0824] "A user recorded a conversation with a close friend and left a note saying they wanted to talk about their next trip. The device sent the recorded audio and note to a server. The server converted the audio to text and used a natural language processing model to identify overly long and potentially misleading descriptions. As a result, the server generated feedback and sent it to the device, advising them to mention their next trip but avoid discussing their previous trip."

[0825] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0826] Behavior prediction processing steps

[0827] Step 1: User Data Entry

[0828] input:

[0829] Users launch the smartphone app and enter their plans and goals for the day.

[0830] Users input internal information such as mood and health status.

[0831] Specific behavior:

[0832] The user inputs "I want to go to the gym in the morning and do some shopping in the afternoon," and selects his mood as "normal" and his health status as "good."

[0833] output:

[0834] The input schedule, goals, and internal information are generated.

[0835] Step 2: Saving data on the device and acquiring external information

[0836] input:

[0837] User schedules, goals, and internal information entered into a smartphone app

[0838] Specific behavior:

[0839] The terminal stores the entered information in a local database.

[0840] The device accesses external APIs (weather APIs and traffic APIs) to obtain external information (weather conditions and traffic conditions).

[0841] output:

[0842] User schedules, goals, and internal information stored in a local database

[0843] External information obtained ("Sunny" from the weather API, "No traffic jam" from the traffic API)

[0844] Step 3: Receiving data from the server and making predictions using the AI ​​model

[0845] input:

[0846] Schedules, goals, internal information, and external information sent from the device

[0847] Specific behavior:

[0848] The server receives all data sent by the devices and stores it in a central database.

[0849] The server uses a generative AI model based on the stored data to predict behavioral patterns.

[0850] The server inputs data into an AI model that predicts that the best time to leave for the gym is 9 a.m. and the best time to leave for the market is 2 p.m.

[0851] output:

[0852] Prediction results of optimal behavioral patterns

[0853] Step 4: Notification of results and feedback via device

[0854] input:

[0855] Prediction results sent from the server

[0856] Specific behavior:

[0857] The terminal notifies the user of the prediction result received from the server.

[0858] The device will visually display the prediction results within the app.

[0859] The user acts based on the prediction results and inputs the results into the app as feedback.

[0860] output:

[0861] User notifications and in-app displays

[0862] Feedback Data

[0863] ---

[0864] Speech prediction processing steps

[0865] Step 1: User voice recording and note taking

[0866] input:

[0867] Users launch the smartphone app and record their everyday conversations.

[0868] The user enters notes for the recording.

[0869] Specific behavior:

[0870] For example, a user records a conversation with a friend and enters a note saying, "I want to talk about my next trip."

[0871] output:

[0872] Recorded audio data and memo information

[0873] Step 2: Save and send data from your device

[0874] input:

[0875] Recorded audio data and memo information

[0876] Specific behavior:

[0877] The device stores the recorded audio data and memo information in a local database.

[0878] The terminal transmits the saved data to the server.

[0879] output:

[0880] Audio data and notes stored in a local database

[0881] Data sent to the server

[0882] Step 3: Audio data conversion and analysis on the server

[0883] input:

[0884] Voice data and memo information sent from the device

[0885] Specific behavior:

[0886] The server converts the voice data into text using a voice recognition engine.

[0887] The server analyzes the converted text data using a natural language processing model.

[0888] The server identifies redundant or inappropriate statements and generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[0889] output:

[0890] Text data and analysis results

[0891] Step 4: Server feedback generation and sending

[0892] input:

[0893] Analysis results

[0894] Specific behavior:

[0895] The server generates appropriate utterance feedback based on the analysis results.

[0896] For example, advice such as "mention your next trip, but avoid talking about your previous trip" may be generated.

[0897] The server transmits the generated feedback to the terminal.

[0898] output:

[0899] Feedback Data

[0900] Step 5: Feedback notification by device

[0901] input:

[0902] Feedback data sent from the server

[0903] Specific behavior:

[0904] The terminal notifies the user of the feedback data.

[0905] The device displays the feedback within the app, giving users information to adjust what they say.

[0906] output:

[0907] User feedback notification and in-app display

[0908] (Application example 1)

[0909] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0910] Conventional behavior prediction systems and comment prediction systems are limited to supporting actions based on the user's schedule and goals, making it difficult to address diverse user needs. Furthermore, they lack sufficient support for content recommendations and communication improvement, resulting in a lack of improvement in the user experience. Therefore, there is a need for a comprehensive system that provides appropriate advice based not only on the user's daily actions and comments, but also on their content consumption trends and message content.

[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0912] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information and external information, means for presenting the predicted behavioral patterns to the user, means for predicting the user's content consumption tendencies using a generative artificial intelligence model and recommending optimal content, and means for providing advice to improve the user's comments based on the content of the user's message.

[0913] This allows users to receive comprehensive support for their actions, content recommendations, and advice on what to say.

[0914] A "user" is an individual who uses the system to receive assistance with schedule input, information gathering, behavior prediction, and speech advice.

[0915] A "generative artificial intelligence model" is an algorithm that learns various patterns and trends based on data and makes predictions and recommendations.

[0916] "Content consumption trends" refers to patterns such as what types of content a user prefers to consume and what types of content they consume at what times of the day.

[0917] "Message content" refers to all statements and texts made by a User when communicating with others.

[0918] "Advice" means instructions or guidance provided to a user to help them act or speak more appropriately.

[0919] "Schedules" refer to planned actions and events that a user undertakes in their daily life.

[0920] A "goal" is a specific action or outcome that a user aims to achieve.

[0921] "Internal information" refers to information about the user's inner self, such as their emotional state or health status.

[0922] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[0923] "Behavioral prediction" is the calculation of optimal behavioral patterns based on a user's schedule, goals, internal information, and external information.

[0924] A "behavioral pattern" refers to a series of actions that a user takes at what timing.

[0925] "Content" is a general term for information assets consumed by users, such as music, videos, and articles.

[0926] The system for carrying out the present invention integrates various functions for predicting and optimizing user actions and comments. Specific embodiments will be described below.

[0927] Hardware and Software Configuration

[0928] Hardware configuration:

[0929] Devices such as smartphones, tablets, or computers

[0930] Servers (including using cloud computing environments)

[0931] Software configuration:

[0932] Application software (smartphone apps, etc.)

[0933] Databases (local and central)

[0934] External APIs (weather API, traffic API, etc.)

[0935] Speech Recognition Engine

[0936] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)

[0937] Generative artificial intelligence model (AI model)

[0938] Embodiment of behavior prediction

[0939] user:

[0940] Users open the app on their smartphone or tablet and enter their plans and goals for the day, as well as internal information such as their mood and health status.

[0941] Device:

[0942] The device stores user-entered schedules, goals, and internal information in a local database, and then uses external APIs to retrieve external information such as weather and traffic conditions.

[0943] server:

[0944] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. It then uses this data to predict the user's behavior using a generative artificial intelligence model. It then sends the prediction results to the device and notifies the user.

[0945] Embodiment of speech prediction

[0946] user:

[0947] Users start the smartphone app and record their everyday conversations, taking notes on topics they want to talk about.

[0948] Device:

[0949] The recorded voice data is stored in a local database and then sent to the server, along with any memo information.

[0950] server:

[0951] The voice data is converted into text by a speech recognition engine, and the text data is analyzed by a natural language processing model to identify redundant or inappropriate utterances, and based on that, appropriate feedback is generated and sent to the device.

[0952] Device:

[0953] Users receive feedback and adjust what they say based on in-app suggestions.

[0954] Embodiment of content recommendation function

[0955] user:

[0956] Users input their goals, such as "I want to relax in the evening" or "I want to catch up on the latest news on Sunday afternoon."

[0957] Device:

[0958] The device sends data to the server based on internal information and information from external APIs.

[0959] server:

[0960] The server uses a generative artificial intelligence model to predict the user's content consumption habits, recommends the most suitable content (movies, music, news articles, etc.) for the user, and sends the results to the device.

[0961] Device:

[0962] Users receive recommendations and choose the content they see on the app to get the best experience.

[0963] Specific examples

[0964] The following example prompts could be fed into a generative AI model:

[0965] Example prompt sentence:

[0966] The user entered their goal of "I want to relax at night," selected "Relaxed" as their mood, and "Energetic" as their health condition. Weather information obtained from an external API showed that the temperature was below 20 degrees and the traffic situation was clear. Based on this information, please recommend the most suitable relaxation content for the user.

[0967] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0968] Step 1:

[0969] The user launches the app on their smartphone or tablet and inputs their plans and goals for the day, such as "I want to relax in the evening," as well as their mood and health status. The input information is then stored in a local database on the device.

[0970] Input: User's schedule, goals, internal information

[0971] Output: Save data to a local database

[0972] Specific behavior: A user enters information into the app's input form and presses the "Save" button, which saves the data to a local database.

[0973] Step 2:

[0974] The device calls external APIs such as weather APIs and traffic APIs to obtain external information such as weather conditions and traffic conditions. The obtained external information is stored in a local database.

[0975] Input: External API request

[0976] Output: External data such as weather information, traffic information, etc.

[0977] Specific operation: The device periodically sends requests to an external API and stores the data obtained in response in a local database.

[0978] Step 3:

[0979] The device transmits the user's schedule, goals, internal information, and external information from a local database to the server, where the transmitted data is stored in a central database.

[0980] Input: All data in the local database

[0981] Output: Send data to server, store in central database

[0982] Specific operation: The terminal periodically uploads all data to the server, and the server stores the received data in a central database.

[0983] Step 4:

[0984] Based on all the data received by the server, a generative artificial intelligence model is used to predict user behavior. For example, if a user inputs "I want to relax at night," the model predicts appropriate relaxing content.

[0985] Input: All data in the central database

[0986] Output: Behavior prediction results

[0987] Specific operation: The server runs a generative artificial intelligence model to predict optimal actions based on past and new data.

[0988] Step 5:

[0989] The server sends the prediction result to the device, and the device notifies the user. For example, the server generates a prediction result such as "This movie is recommended for relaxing at night" and sends it to the device.

[0990] Input: Behavior prediction result

[0991] Output: User notification

[0992] Specific operation: The device receives the prediction results sent from the server and notifies the user visually within the app.

[0993] Step 6:

[0994] Users can take recommended actions based on the app's notifications, such as watching a recommended movie, and provide feedback to the app, allowing the system to learn from that feedback and improve its prediction accuracy next time.

[0995] Input: User feedback

[0996] Output: Accumulation of feedback data

[0997] What it does: The user enters feedback within the app, such as "I finished watching the movie," and that data is stored in a local database.

[0998] Step 7:

[0999] It records users' everyday conversations and collects audio data. For example, a user records a conversation with a friend and leaves a note saying, "I want to talk about my next trip."

[1000] Input: Audio data, memo information

[1001] Output: Save data to a local database

[1002] What happens: A user uses the app's recording feature to record a conversation and takes notes within the app.

[1003] Step 8:

[1004] The recorded voice data is stored in a local database and sent to a server, along with any memo information. The server converts the voice data into text using a speech recognition engine, and the text data is analyzed using a natural language processing model.

[1005] Input: Audio data, memo information

[1006] Output: Text data, analysis results

[1007] Specific operation: The server uses a speech recognition engine to convert speech to text and analyzes the text data using a natural language processing model.

[1008] Step 9:

[1009] Based on the results of the analysis using the natural language processing model, appropriate feedback (e.g., "Avoid talking about your previous trip") is generated and sent to the device. The device notifies the user of the feedback and displays it appropriately within the app.

[1010] Input: Analysis results

[1011] Output: Feedback results

[1012] Specific operation: The server generates appropriate feedback and sends it to the device, which notifies the user of the feedback and displays it.

[1013] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1014] This invention is a system that combines "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals; "statement prediction," which provides appropriate advice on the content and timing of statements; and "emotion recognition," which recognizes the user's emotions.

[1015] Embodiment of behavior prediction

[1016] User:

[1017] The user launches the smartphone app and inputs the plan and goals for the day. Internal information such as mood and health status is also input. For example, the user can provide information such as "I want to go to the gym in the morning and go shopping in the afternoon" or "I feel normal and my health is good."

[1018] Device:

[1019] The system stores the schedule, goals, and internal information entered by the user in a local database. It then uses an external API to obtain external information such as weather and traffic conditions. For example, it collects information such as "sunny" and "no traffic jam."

[1020] server:

[1021] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. In addition, an emotion engine extracts emotional information from the user's input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[1022] server:

[1023] Generate prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and send them to the device.

[1024] Device:

[1025] The received action prediction results are notified to the user and displayed visually within the app. The user can act based on the results and enter feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[1026] Embodiment of speech prediction

[1027] User:

[1028] Start a smartphone app and record your everyday conversations. For example, record a conversation with your best friend and leave a note saying, "I want to talk about my next trip."

[1029] Device:

[1030] The recorded voice data and notes are stored in a local database and sent to a server, where an emotion engine extracts the user's emotional information from the recorded data and sends it together.

[1031] server:

[1032] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[1033] server:

[1034] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[1035] Device:

[1036] The app notifies users of the feedback it receives and visually displays it within the app, allowing users to adjust what they say based on that feedback and improve communication. For example, when talking with a friend, users can use the advice to choose topics and keep the conversation flowing smoothly.

[1037] Emotion Recognition Embodiment

[1038] User:

[1039] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[1040] Device:

[1041] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[1042] server:

[1043] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[1044] Specific examples

[1045] Specific examples of behavioral prediction

[1046] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[1047] Specific examples of speech prediction

[1048] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1049] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[1050] The processing flow will be explained below.

[1051] Program processing of behavior prediction

[1052] Step 1:

[1053] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[1054] Step 2:

[1055] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[1056] Step 3:

[1057] Device: Stores user-entered schedules, goals, and internal information in a local database.

[1058] Step 4:

[1059] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[1060] Step 5:

[1061] Terminal: Send all collected data, including internal and external information, to the server.

[1062] Step 6:

[1063] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[1064] Step 7:

[1065] Server: Extracts user emotional information using the emotion engine based on the stored data.

[1066] Step 8:

[1067] Server: Based on the stored data and emotional information, the AI ​​model predicts optimal behavioral patterns.

[1068] Step 9:

[1069] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[1070] Step 10:

[1071] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[1072] Step 11:

[1073] User: Follows the suggested behavioral patterns and achieves the goal.

[1074] Step 12:

[1075] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[1076] ---

[1077] Program processing of speech prediction

[1078] Step 1:

[1079] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[1080] Step 2:

[1081] Users: Leave notes during conversations about things that concern them or that need attention.

[1082] Step 3:

[1083] Device: Stores recorded audio data and notes in a local database.

[1084] Step 4:

[1085] Terminal: Sends voice data and memo information to the server.

[1086] Step 5:

[1087] Server: Converts voice data received from the device into text using a speech recognition engine.

[1088] Step 6:

[1089] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[1090] Step 7:

[1091] Server: Extracts user emotion information from voice and text data using an emotion engine.

[1092] Step 8:

[1093] Server: Generates appropriate feedback based on sentiment information and text analysis results (e.g., "Mention your next trip, but avoid talking about your previous trip").

[1094] Step 9:

[1095] Server: Sends the generated feedback to the device.

[1096] Step 10:

[1097] On the device: Notify the user of the feedback received and display it visually within the app.

[1098] Step 11:

[1099] Users: Use suggested feedback to adjust conversations and improve communication.

[1100] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[1101] Example 2

[1102] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1103] Conventional behavior prediction systems only predict behavior based on a user's schedule and goals, and have the problem of being unable to provide appropriate feedback or advice that takes emotional information into account. Furthermore, speech prediction systems are limited to identifying redundant or inappropriate speech, making it difficult to provide advice that appropriately reflects the user's emotions. This has led to the issue of a lack of effective support for user behavior and speech.

[1104] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1105] In this invention, the server includes: means for a user to input internal information such as the user's schedule and goals for the day, mood, and health status; means for collecting the user's internal information and external information such as weather conditions and traffic conditions; means for predicting the user's behavior based on the schedule, goals, internal information, and external information and generating an optimal behavior pattern based on emotional information; means for notifying the user of the predicted behavior pattern and collecting user feedback; means for recording the user's daily conversations and collecting voice data; means for converting the voice data into text; means for identifying redundant or inappropriate statements from the text data and advising the user on appropriate statements based on emotional information; and means for generating a pre-alert for statements and advice on appropriate statements based on the identified statements and notifying the user. This makes it possible to optimize the user's behavior and statements while taking emotions into consideration, thereby providing effective support.

[1106] "User" refers to an individual or corporation that uses the system, inputs their own schedules, goals, and internal information, and receives feedback from the system.

[1107] "Internal information" is information about a user's internal state, such as their mood, health, or emotions.

[1108] "External information" refers to information about the user's external environment, such as weather conditions and traffic conditions, and is obtained using an external API.

[1109] "Emotion information" is information about the user's emotional state, and is extracted by an emotion engine or the like.

[1110] An "emotion engine" is a software or hardware mechanism for extracting emotional information from user input or voice data.

[1111] A "behavioral pattern" is a schedule or action plan for the user to act optimally, and is the result predicted by the system.

[1112] "Feedback" is information that users input into the system based on the results of their actual actions, and the system uses this information to make its next prediction more accurate.

[1113] "Speech prediction" is a system function that analyzes the user's everyday conversations and advises them on appropriate content to say.

[1114] A "speech recognition engine" is a software or hardware mechanism for converting a user's voice data into text data.

[1115] A "natural language processing model" is a technology for analyzing text data and identifying redundant or inappropriate statements.

[1116] A "pre-alert" is a warning message that the system displays before a user makes an identified inappropriate comment.

[1117] "Advice on appropriate speech" refers to advice on speech content provided by the system to help users communicate more smoothly.

[1118] This invention is a system that combines "behavior prediction" that efficiently predicts a user's daily behavior and helps them achieve their schedules and goals, "statement prediction" that provides appropriate advice on the content and timing of statements, and "emotion recognition" that recognizes the user's emotions. A specific embodiment of this system will be described below.

[1119] Embodiment of behavior prediction

[1120] User:

[1121] Users launch the smartphone app and enter detailed information about the day's plans, goals, mood, health status, etc. For example, they might enter, "I want to go to the gym in the morning and do some shopping in the afternoon," or "I'm feeling normal, and my health is good."

[1122] Device:

[1123] The device stores the user's schedule, goals, mood, and health status in a local database. It then uses external APIs to obtain external information such as weather conditions (e.g., OpenWeatherMap API) and traffic conditions (e.g., Google Maps Traffic API). For example, it collects information such as "sunny" and "no traffic jam" and sends it to the server.

[1124] server:

[1125] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. At the same time, it uses an emotion engine to extract emotional information from the user's input. Based on the stored data and emotional information, the AI ​​model predicts the optimal behavioral pattern. For example, it generates a result such as "It is best to leave for the gym at 9:00 AM and the market at 2:00 PM" and sends it to the device.

[1126] Device:

[1127] The device then notifies the user of the predicted behavior and displays it visually within the app. The user can then act based on the results and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[1128] Embodiment of speech prediction

[1129] User:

[1130] Users can launch the smartphone app and record everyday conversations, such as a conversation with a close friend, and leave a note saying, "I want to talk about my next trip."

[1131] Device:

[1132] The device stores the recorded voice data and notes in a local database, and uses an emotion engine to extract the user's emotional information from the recorded data and send it together with the data to the server.

[1133] server:

[1134] The server converts the voice data into text using a speech recognition engine (e.g., Google Speech-to-Text API), analyzes the converted text data using a natural language processing model, and identifies redundant or inappropriate statements. For example, it generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading." It then generates appropriate feedback based on the emotional information. For example, it generates advice such as "Talk about your next trip, but avoid talking about your previous trip," and sends it to the device.

[1135] Device:

[1136] The device will notify the user of the received feedback and display it visually within the app, allowing the user to adjust what they say based on this feedback and improve communication. For example, when talking with a friend, the user can use the advice to choose topics and keep the conversation flowing smoothly.

[1137] Emotion Recognition Embodiment

[1138] User:

[1139] Users provide emotional information to the app by inputting their mood and emotions and recording changes in facial expressions and voice during conversations.

[1140] Device:

[1141] The device uses an emotion recognition engine to extract emotional information from the recorded voice data and transmits the emotional information to a server.

[1142] server:

[1143] The server uses the emotional information analyzed by the emotion recognition engine and natural language processing model to generate behavioral patterns and speech feedback based on the user's emotions.

[1144] Specific examples

[1145] Specific examples of behavioral prediction

[1146] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "Sunny" from a weather API (e.g., OpenWeatherMap API) and "No traffic jam" from a traffic API (e.g., Google Maps Traffic API), and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "The best time to leave for the gym is 9:00 AM and the market is 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[1147] Example prompt sentence:

[1148] User: I want to go to the gym in the morning and do some shopping in the afternoon. I feel normal and my health is good.

[1149] Terminal: The weather is clear and traffic is smooth.

[1150] Server: The best time to leave for the gym is 9am and the market is 2pm.

[1151] Specific examples of speech prediction

[1152] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text (e.g., Google Speech-to-Text API) and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1153] Example prompt sentence:

[1154] User: Recording a conversation with a best friend. Want to talk about an upcoming trip.

[1155] Device: Sends voice data, notes, and emotional information to the server.

[1156] Server: Analyzes the text for redundant content and generates appropriate feedback.

[1157] Server: Mention your next trip, but avoid talking about your last trip.

[1158] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1159] Behavior prediction processing steps

[1160] Step 1:

[1161] User: Launches the smartphone app and enters internal information such as the day's plans, goals, mood, and health status.

[1162] Input: Schedule, Goals, Mood, Health

[1163] Output: User input data

[1164] For example: "I go to the gym in the morning and shop at the market in the afternoon." "I feel fine and my health is good."

[1165] Step 2:

[1166] Device: Stores the entered information in a local database, then uses external APIs to retrieve external information such as weather conditions and traffic conditions.

[1167] Input: User-entered data

[1168] Output: Internal and external information

[1169] Examples: "The weather is clear" and "The traffic is smooth."

[1170] Step 3:

[1171] Server: Receives all data sent from the devices and stores it in a central database. Uses an emotion engine to extract emotional information from user input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[1172] Input: User input data, internal information, external information

[1173] Output: Predicted behavior pattern

[1174] For example: "The best time to leave for the gym is 9am and the best time to leave for the market is 2pm."

[1175] Step 4:

[1176] Server: Sends the prediction results to the device.

[1177] Input: Predicted behavior pattern

[1178] Output: Data sent to the terminal

[1179] Step 5:

[1180] On the device: The user is notified of the received predictions and visually displayed within the app, allowing the user to act on them and provide feedback to the app.

[1181] Input: Predicted behavior pattern

[1182] Output: User feedback

[1183] Step 6:

[1184] Device: User feedback is sent to the server to help improve prediction accuracy in future predictions.

[1185] Input: User feedback

[1186] Output: Data sent to the server

[1187] Speech prediction processing steps

[1188] Step 1:

[1189] User: Launches the smartphone app, records everyday conversations, and enters notes.

[1190] Input: Recording data, notes

[1191] Output: Recording data and notes

[1192] Examples: "I want to record a conversation with my best friend" or "I want to talk about my next trip."

[1193] Step 2:

[1194] Device: Recorded voice data and notes are stored in a local database, and emotional information is extracted using an emotion engine, which then sends the data to the server.

[1195] Input: Recording data, notes

[1196] Output: Recording data, notes, emotional information

[1197] Step 3:

[1198] Server: The speech data is converted into text using a speech recognition engine, and a natural language processing model is used to identify redundant or inappropriate speech.

[1199] Input: Audio recording, emotional information

[1200] Output: Text data and detected verbose / inappropriate comments

[1201] For example: "This part is redundant and difficult to understand" or "This part may be misleading"

[1202] Step 4:

[1203] Server: Generates appropriate feedback based on emotion information and sends it to the device.

[1204] Input: Text data, emotion information

[1205] Output: Feedback

[1206] For example: "Talk about your next trip, but avoid talking about your last trip."

[1207] Step 5:

[1208] On-device: The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust their voice accordingly.

[1209] Input: Feedback

[1210] Output: Adjusted speech data

[1211] Processing steps for emotion recognition

[1212] Step 1:

[1213] User: Inputs mood and emotions and records facial expressions and vocal changes during conversations.

[1214] Input: Mood, emotion, facial expression, vocal changes

[1215] Output: Emotion data

[1216] Step 2:

[1217] Terminal: The emotion recognition engine extracts emotional information from the voice data and sends it to the server.

[1218] Input: Recording data

[1219] Output: Emotional information

[1220] Step 3:

[1221] Server: Analyzes emotional information and generates behavioral patterns and speech feedback based on the user's emotions.

[1222] Input: Emotion information

[1223] Output: Behavioral patterns, speech feedback

[1224] (Application example 2)

[1225] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1226] Currently, there is a lack of security systems that understand users' daily behavior and emotions and provide advice on appropriate behavioral patterns and speech based on that understanding. This can make it difficult for users to understand the degree of risk involved in their own behavior and speech, making it difficult to ensure their safety. The present invention aims to provide a system that comprehensively analyzes users' behavior, speech, and emotional information to reduce security risks.

[1227] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1228] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information, and external information, means for presenting the predicted behavioral pattern to the user, means for predicting the user's utterances from the content of the user's conversation and recorded voice data and providing appropriate advice, and means for recognizing the user's emotions and reflecting them in the behavioral pattern and utterance content, thereby enabling the user to understand the degree to which their actions and utterances pose security risks and to take appropriate preventative measures.

[1229] "User" refers to an individual who uses this system.

[1230] "Means for entering the day's plans and goals" refers to the function that allows users to enter their daily plans and goals into the system.

[1231] "Internal information" refers to information about a user's personal state or emotions, such as mood or health.

[1232] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[1233] "Means for predicting behavior" refers to the function of calculating the optimal behavioral pattern for a user based on internal and external information.

[1234] "Means for presenting predicted behavioral patterns" refers to the function of notifying the user of the predicted results and displaying them visually.

[1235] "Means of predicting what will be said based on conversation content and recorded audio data and providing appropriate advice" refers to a function that analyzes a user's everyday conversations and generates dedicated alerts and advice.

[1236] "Means of recognizing emotions and reflecting them in behavioral patterns and speech content" refers to a function that analyzes the user's emotions and provides behavioral patterns and speech advice that take these into consideration.

[1237] A "system that reduces security risks" refers to a system that evaluates the risks associated with users' actions and statements and supports safe behavior and communication.

[1238] This invention combines a system that efficiently predicts a user's daily behavior and supports safe behavior, a system that analyzes conversation content and gives advice on appropriate remarks, and a system that recognizes the user's emotions and assesses risk. The system is intended to be used by users through a smartphone app.

[1239] To implement this system, the following hardware and software are used:

[1240] Hardware: smartphones, servers, cameras and microphones for emotion recognition

[1241] Software: Python, external API (weather API, traffic API), voice recognition AI engine (Hugging Face transformers), emotion recognition AI engine (DeepFace)

[1242] Embodiment of behavior prediction

[1243] Device:

[1244] The user launches the app on their smartphone and enters their plans and goals for the day, as well as internal information such as their mood and health status. The device stores this information in a local database and uses external APIs to obtain external information such as weather and traffic conditions.

[1245] server:

[1246] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. The AI ​​model uses this data to predict optimal behavioral patterns. For example, it generates a result such as "It's sunny and traffic is good, so it's best to leave for the gym at 9:00 AM." The server then sends this prediction to the device.

[1247] Device:

[1248] The predictions received are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app.

[1249] Embodiment of speech prediction

[1250] Device:

[1251] The user launches the app on their smartphone and records their everyday conversations. The recorded audio data is stored in a local database and sent to a server. The emotion engine extracts the user's emotional information from the recorded data and sends it along with the data.

[1252] server:

[1253] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[1254] server:

[1255] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[1256] Device:

[1257] The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust what they say and communicate more effectively.

[1258] Emotion Recognition Embodiment

[1259] user:

[1260] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[1261] Device:

[1262] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[1263] server:

[1264] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[1265] Specific examples

[1266] Specific examples of behavioral prediction

[1267] The user inputs "I plan to go to the gym at 6 PM" and also inputs their mood and health status as additional information. The device obtains "sunny" information from the weather API and "no traffic jams" information from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict that "the best time to leave is 5 PM" and sends the result to the device. The device notifies the user of the result, and the user acts accordingly.

[1268] Specific examples of speech prediction

[1269] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1270] Prompt Sentence Examples

[1271] Behavioral prediction: "Predict what time a user should go to the gym based on current weather and traffic information."

[1272] Speech Prediction: "Identify sensitive information that users should avoid in conversations with their friends."

[1273] Emotion Recognition: "Perform emotion analysis using the user's current image to conduct risk assessment."

[1274] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1275] Behavior prediction processing steps

[1276] Step 1: User enters schedule and goal

[1277] The user launches the smartphone app and enters the day's plans, goals, and internal information (mood and health status).

[1278] Input: Schedule, goals, internal information

[1279] Output: Data stored in a local database

[1280] Step 2: Obtaining external information

[1281] The device calls external APIs (weather API, traffic API) to obtain external information such as weather conditions and traffic conditions.

[1282] Input: External API endpoint

[1283] Output: Weather conditions, traffic conditions (e.g. sunny, no traffic jams)

[1284] Step 3: Sending data

[1285] The terminal transmits the input internal information and the acquired external information to the server.

[1286] Input: Plan, goal, internal information, external information

[1287] Output: All data sent to the server

[1288] Step 4: Predicting behavioral patterns

[1289] Based on all the data received by the server, an AI model is used to predict optimal behavioral patterns.

[1290] Input: All data and AI model

[1291] Output: Predicted behavioral pattern (e.g., departure at 5pm)

[1292] Step 5: Send prediction results

[1293] The server sends the prediction results to the device.

[1294] Input: predicted behavior pattern

[1295] Output: Prediction results sent to the device

[1296] Step 6: Visualizing the results

[1297] The device notifies the user of the prediction results received and displays them visually within the app.

[1298] Input: predicted behavior pattern

[1299] Output: The result communicated to the user

[1300] Step 7: Provide feedback

[1301] The user provides feedback on the results of the execution to the app.

[1302] Input: Actual action results

[1303] Output: Feedback data

[1304] Speech prediction processing steps

[1305] Step 1: Recording audio data

[1306] Users launch the smartphone app and record their everyday conversations.

[1307] Input: Audio data

[1308] Output: Audio data stored in a local database

[1309] Step 2: Sending audio data

[1310] The device sends the recorded audio data to the server.

[1311] Input: Audio data

[1312] Output: Audio data sent to the server

[1313] Step 3: Voice Recognition

[1314] The server converts the voice data into text using a voice recognition engine.

[1315] Input: Audio data

[1316] Output: Text data

[1317] Step 4: Parsing utterances

[1318] The server analyzes the text data using a natural language processing model to identify inappropriate or redundant statements.

[1319] Input: Text data, NLP model

[1320] Output: Analysis results (identification of inappropriate and redundant parts)

[1321] Step 5: Generate Advice

[1322] The server generates feedback on the speech based on the analysis results and emotional information.

[1323] Input: Analysis results, emotion information

[1324] Output: Feedback (e.g. mention the next trip but avoid talking about the previous trip)

[1325] Step 6: Submit your feedback

[1326] The server generates feedback and sends it to the device.

[1327] Input: Feedback

[1328] Output: Feedback sent to the terminal

[1329] Step 7: Notification of feedback

[1330] Notify the user of the feedback received by the device and display it visually within the app.

[1331] Input: Feedback

[1332] Output: Feedback given to the user

[1333] Processing steps for emotion recognition

[1334] Step 1: Enter emotional information

[1335] Users input their moods and emotions into the app.

[1336] Input: Mood, emotional information

[1337] Output: Emotion information stored in a local database

[1338] Step 2: Analyzing the audio data

[1339] Emotional information is extracted from the audio data recorded by the device and sent to the server.

[1340] Input: Audio data

[1341] Output: Extracted emotion information

[1342] Step 3: Sending emotional information

[1343] The device transmits the emotion information to the server.

[1344] Input: Emotion information

[1345] Output: Emotion information sent to the server

[1346] Step 4: Sentiment-based analysis

[1347] The server uses the emotion information to complement the analysis results of behavior prediction and utterance prediction.

[1348] Input: Emotion information, behavior prediction, speech prediction

[1349] Output: Analysis results (behavior and speech) supplemented based on emotions

[1350] Step 5: Notification of completion results

[1351] The server transmits the completed analysis results to the terminal.

[1352] Input: Completed analysis results

[1353] Output: Completion results sent to the terminal

[1354] Step 6: Visualizing the results

[1355] The device notifies the user of the completion results received and displays them visually within the app.

[1356] Input: Completed analysis results

[1357] Output: Completion result notified to the user

[1358] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1359] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1360] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1361] [Third embodiment]

[1362] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1363] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1364] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1365] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1366] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1367] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1368] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1369] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1370] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1371] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1372] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1373] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1374] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[1375] Embodiment of behavior prediction

[1376] User:

[1377] Users launch the smartphone app and enter their plans and goals for the day. They also enter internal information such as their mood and health status. For example, a user might enter their plan to "go to the gym in the morning and go shopping in the afternoon," select "normal" as their mood, and "good" as their health status.

[1378] Device:

[1379] The schedule, goals, and internal information entered by the user are stored in a local database. Then, external APIs are used to obtain external information such as weather conditions and traffic conditions. For example, a weather API is called to obtain information such as "sunny," and a traffic API is called to obtain information such as "no traffic jams."

[1380] server:

[1381] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. Based on the stored data, an AI model predicts the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, the AI ​​model may generate a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[1382] Device:

[1383] The received prediction results are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[1384] Embodiment of speech prediction

[1385] User:

[1386] Start the smartphone app and record your everyday conversations. For example, record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[1387] Device:

[1388] The recorded voice data is stored in a local database and sent to the server, along with any memo information.

[1389] server:

[1390] The voice data received from the device is converted into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[1391] server:

[1392] Based on the analysis results, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent to the device.

[1393] Device:

[1394] The received feedback is notified to the user and displayed within the app, allowing the user to adjust what they say based on the feedback and improve communication. For example, when talking with a friend, a user can refer to the advice to choose a topic and smoothly advance the conversation.

[1395] Specific examples

[1396] Specific examples of behavioral prediction

[1397] The user enters their schedule into the app, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict behavioral patterns such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends this to the device. The device notifies the user of the results, and the user acts accordingly.

[1398] Specific examples of speech prediction

[1399] A user records a conversation with a close friend and leaves a note saying, "I'd like to talk about my next trip." The device then sends the recorded audio data and note to a server. The server converts the audio data into text and uses a natural language processing model to identify "overly long explanations" and "potentially misleading" parts. Based on this, the device generates and sends feedback to the device, such as "mention your next trip, but avoid talking about your previous trip." The device then notifies the user of the feedback, and the user can use the advice to smoothly move the conversation forward.

[1400] This allows users to receive support in achieving their schedules and goals, as well as in communicating smoothly.

[1401] The processing flow will be explained below.

[1402] Program processing of behavior prediction

[1403] Step 1:

[1404] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[1405] Step 2:

[1406] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[1407] Step 3:

[1408] Device: Stores user-entered schedules, goals, and internal information in a local database.

[1409] Step 4:

[1410] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[1411] Step 5:

[1412] Terminal: Send all collected data, including internal and external information, to the server.

[1413] Step 6:

[1414] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[1415] Step 7:

[1416] Server: Based on the stored data, it uses an AI model to predict the user's optimal behavioral patterns.

[1417] Step 8:

[1418] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[1419] Step 9:

[1420] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[1421] Step 10:

[1422] User: Follows the suggested behavioral patterns and achieves the goal.

[1423] Step 11:

[1424] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[1425] Program processing of speech prediction

[1426] Step 1:

[1427] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[1428] Step 2:

[1429] Users: Leave notes during conversations about things that concern them or that need attention.

[1430] Step 3:

[1431] Device: Stores recorded audio data and notes in a local database.

[1432] Step 4:

[1433] Terminal: Sends voice data and memo information to the server.

[1434] Step 5:

[1435] Server: Converts voice data received from the device into text using a speech recognition engine.

[1436] Step 6:

[1437] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[1438] Step 7:

[1439] Server: Based on the analysis results, generate appropriate feedback based on the profile information of the person you are speaking to.

[1440] Step 8:

[1441] Server: Sends the generated feedback (e.g., "mention your next trip, but avoid talking about your previous trip") to the device.

[1442] Step 9:

[1443] On the device: Notify the user of the feedback received and display it visually within the app.

[1444] Step 10:

[1445] Users: Use suggested feedback to adjust conversations and improve communication.

[1446] This allows users to effectively achieve their schedules and goals, as well as communicate smoothly.

[1447] Example 1

[1448] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1449] Conventional systems have difficulty predicting behavior based on internal and external information in addition to users' schedules and goals. Furthermore, there was a lack of systems that notified users of appropriate timing and content for speech in everyday conversations. Furthermore, there was no mechanism in place to incorporate user feedback to improve the system's overall prediction accuracy. This made it difficult for users to achieve their schedules and communicate smoothly.

[1450] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1451] In this invention, the server includes: means for a user to input the day's schedule and goals; means for collecting the user's internal and external information; means for predicting the user's behavior based on the schedule, goals, internal and external information; means for presenting the predicted behavioral pattern to the user; means for collecting voice data and converting it into text; means for identifying redundant or inappropriate comments from the text data; means for generating a pre-alert for a comment and advice on appropriate comments based on the identified comments and notifying the user; means for collecting and storing user feedback data; and means for improving the prediction accuracy of the entire system based on the feedback data, thereby enabling the user to achieve their schedule and communicate smoothly.

[1452] "User" refers to the entity that uses the system, such as the person who inputs schedules and goals, provides feedback, and records voice data.

[1453] "Schedule" refers to the planned actions and tasks that the user intends to carry out on the day.

[1454] "Goals" refer to the specific outcomes or results that a user is trying to achieve on that day.

[1455] "Internal information" refers to information related to the user's internal state, such as their mood or health.

[1456] "External information" refers to information related to the external environment that affects user behavior, such as weather conditions and traffic conditions.

[1457] "Behavioral patterns" refer to the optimal sequence and timing of actions predicted based on the user's schedule, goals, internal information, and external information.

[1458] "Voice data" refers to audio data obtained when a user records everyday conversations.

[1459] "Text data" refers to the result of converting voice data into text information using voice recognition technology.

[1460] "Redundant utterances" refer to utterances in a conversation that are unnecessarily long or overlapping.

[1461] "Inappropriate remarks" refer to remarks that may be misunderstood in a conversation or that may be perceived as offensive by the other person.

[1462] "Advance Alert" refers to a warning or notification to a user about an upcoming action or statement.

[1463] "Feedback data" refers to the actual behavioral results and impressions that users provide in response to predictions and advice from the system.

[1464] "Generative AI models" refer to machine learning models used to predict behavioral patterns and speech content based on user data.

[1465] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[1466] Embodiment of behavior prediction

[1467] User:

[1468] The user launches the smartphone app and inputs their plans and goals for the day. They also input internal information such as their mood and health status. For example, they input specific plans such as "I want to go to the gym in the morning and go shopping in the afternoon," and select "normal" as their mood and "good" as their health status.

[1469] Device:

[1470] The device stores the schedule, goals, and internal information entered by the user in a local database. At the same time, it uses external APIs to obtain external information such as weather conditions and traffic conditions. Specifically, it obtains information such as "sunny" from the weather API and "no traffic jams" from the traffic API.

[1471] server:

[1472] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. Based on the stored data, a generative AI model is used to predict the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, it generates a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[1473] Device:

[1474] The device receives the prediction results from the server and notifies the user, visually displaying them within the app. The user can then act on the results and enter the results as feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[1475] Embodiment of speech prediction

[1476] User:

[1477] Users can launch the smartphone app and record everyday conversations. For example, they can record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[1478] Device:

[1479] The device stores the recorded voice data and memo information in a local database and transmits it to the server.

[1480] server:

[1481] The server converts the voice data received from the device into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, it generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[1482] server:

[1483] Based on the analysis results, the server generates appropriate feedback, such as advice like "mention your next trip, but avoid talking about your previous trip," and sends it to the device.

[1484] Device:

[1485] The device notifies the user of the received feedback and displays it in the app. The user can then use the advice to adjust what they say and communicate smoothly. For example, a user can use the advice to choose a topic and smoothly advance the conversation with a friend.

[1486] Prompt Sentence Examples

[1487] Behavioral prediction prompt:

[1488] "The user entered their plan to go to the gym in the morning and do some shopping at the market in the afternoon, and also entered their mood as normal and their health status as good. The device retrieved information from the weather API that it was sunny and from the traffic API that it was clear, and sent this information to the server. Based on this information, the server predicted the optimal behavioral pattern and notified the user to leave for the gym at 9 a.m. and the market at 2 p.m."

[1489] Predictive prompt:

[1490] "A user recorded a conversation with a close friend and left a note saying they wanted to talk about their next trip. The device sent the recorded audio and note to a server. The server converted the audio to text and used a natural language processing model to identify overly long and potentially misleading descriptions. As a result, the server generated feedback and sent it to the device, advising them to mention their next trip but avoid discussing their previous trip."

[1491] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1492] Behavior prediction processing steps

[1493] Step 1: User Data Entry

[1494] input:

[1495] Users launch the smartphone app and enter their plans and goals for the day.

[1496] Users input internal information such as mood and health status.

[1497] Specific behavior:

[1498] The user inputs "I want to go to the gym in the morning and do some shopping in the afternoon," and selects his mood as "normal" and his health status as "good."

[1499] output:

[1500] The input schedule, goals, and internal information are generated.

[1501] Step 2: Saving data on the device and acquiring external information

[1502] input:

[1503] User schedules, goals, and internal information entered into a smartphone app

[1504] Specific behavior:

[1505] The terminal stores the entered information in a local database.

[1506] The device accesses external APIs (weather APIs and traffic APIs) to obtain external information (weather conditions and traffic conditions).

[1507] output:

[1508] User schedules, goals, and internal information stored in a local database

[1509] External information obtained ("Sunny" from the weather API, "No traffic jam" from the traffic API)

[1510] Step 3: Receiving data from the server and making predictions using the AI ​​model

[1511] input:

[1512] Schedules, goals, internal information, and external information sent from the device

[1513] Specific behavior:

[1514] The server receives all data sent by the devices and stores it in a central database.

[1515] The server uses a generative AI model based on the stored data to predict behavioral patterns.

[1516] The server inputs data into an AI model that predicts that the best time to leave for the gym is 9 a.m. and the best time to leave for the market is 2 p.m.

[1517] output:

[1518] Prediction results of optimal behavioral patterns

[1519] Step 4: Notification of results and feedback via device

[1520] input:

[1521] Prediction results sent from the server

[1522] Specific behavior:

[1523] The terminal notifies the user of the prediction result received from the server.

[1524] The device will visually display the prediction results within the app.

[1525] The user acts based on the prediction results and inputs the results into the app as feedback.

[1526] output:

[1527] User notifications and in-app displays

[1528] Feedback Data

[1529] ---

[1530] Speech prediction processing steps

[1531] Step 1: User voice recording and note taking

[1532] input:

[1533] Users launch the smartphone app and record their everyday conversations.

[1534] The user enters notes for the recording.

[1535] Specific behavior:

[1536] For example, a user records a conversation with a friend and enters a note saying, "I want to talk about my next trip."

[1537] output:

[1538] Recorded audio data and memo information

[1539] Step 2: Save and send data from your device

[1540] input:

[1541] Recorded audio data and memo information

[1542] Specific behavior:

[1543] The device stores the recorded audio data and memo information in a local database.

[1544] The terminal transmits the saved data to the server.

[1545] output:

[1546] Audio data and notes stored in a local database

[1547] Data sent to the server

[1548] Step 3: Audio data conversion and analysis on the server

[1549] input:

[1550] Voice data and memo information sent from the device

[1551] Specific behavior:

[1552] The server converts the voice data into text using a voice recognition engine.

[1553] The server analyzes the converted text data using a natural language processing model.

[1554] The server identifies redundant or inappropriate statements and generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[1555] output:

[1556] Text data and analysis results

[1557] Step 4: Server feedback generation and sending

[1558] input:

[1559] Analysis results

[1560] Specific behavior:

[1561] The server generates appropriate utterance feedback based on the analysis results.

[1562] For example, advice such as "mention your next trip, but avoid talking about your previous trip" may be generated.

[1563] The server transmits the generated feedback to the terminal.

[1564] output:

[1565] Feedback Data

[1566] Step 5: Feedback notification by device

[1567] input:

[1568] Feedback data sent from the server

[1569] Specific behavior:

[1570] The terminal notifies the user of the feedback data.

[1571] The device displays the feedback within the app, giving users information to adjust what they say.

[1572] output:

[1573] User feedback notification and in-app display

[1574] (Application example 1)

[1575] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1576] Conventional behavior prediction systems and comment prediction systems are limited to supporting actions based on the user's schedule and goals, making it difficult to address diverse user needs. Furthermore, they lack sufficient support for content recommendations and communication improvement, resulting in a lack of improvement in the user experience. Therefore, there is a need for a comprehensive system that provides appropriate advice based not only on the user's daily actions and comments, but also on their content consumption trends and message content.

[1577] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1578] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information and external information, means for presenting the predicted behavioral patterns to the user, means for predicting the user's content consumption tendencies using a generative artificial intelligence model and recommending optimal content, and means for providing advice to improve the user's comments based on the content of the user's message.

[1579] This allows users to receive comprehensive support for their actions, content recommendations, and advice on what to say.

[1580] A "user" is an individual who uses the system to receive assistance with schedule input, information gathering, behavior prediction, and speech advice.

[1581] A "generative artificial intelligence model" is an algorithm that learns various patterns and trends based on data and makes predictions and recommendations.

[1582] "Content consumption trends" refers to patterns such as what types of content a user prefers to consume and what types of content they consume at what times of the day.

[1583] "Message content" refers to all statements and texts made by a User when communicating with others.

[1584] "Advice" means instructions or guidance provided to a user to help them act or speak more appropriately.

[1585] "Schedules" refer to planned actions and events that a user undertakes in their daily life.

[1586] A "goal" is a specific action or outcome that a user aims to achieve.

[1587] "Internal information" refers to information about the user's inner self, such as their emotional state or health status.

[1588] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[1589] "Behavioral prediction" is the calculation of optimal behavioral patterns based on a user's schedule, goals, internal information, and external information.

[1590] A "behavioral pattern" refers to a series of actions that a user takes at what timing.

[1591] "Content" is a general term for information assets consumed by users, such as music, videos, and articles.

[1592] The system for carrying out the present invention integrates various functions for predicting and optimizing user actions and comments. Specific embodiments will be described below.

[1593] Hardware and Software Configuration

[1594] Hardware configuration:

[1595] Devices such as smartphones, tablets, or computers

[1596] Servers (including using cloud computing environments)

[1597] Software configuration:

[1598] Application software (smartphone apps, etc.)

[1599] Databases (local and central)

[1600] External APIs (weather API, traffic API, etc.)

[1601] Speech Recognition Engine

[1602] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)

[1603] Generative artificial intelligence model (AI model)

[1604] Embodiment of behavior prediction

[1605] user:

[1606] Users open the app on their smartphone or tablet and enter their plans and goals for the day, as well as internal information such as their mood and health status.

[1607] Device:

[1608] The device stores user-entered schedules, goals, and internal information in a local database, and then uses external APIs to retrieve external information such as weather and traffic conditions.

[1609] server:

[1610] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. It then uses this data to predict the user's behavior using a generative artificial intelligence model. It then sends the prediction results to the device and notifies the user.

[1611] Embodiment of speech prediction

[1612] user:

[1613] Users start the smartphone app and record their everyday conversations, taking notes on topics they want to talk about.

[1614] Device:

[1615] The recorded voice data is stored in a local database and then sent to the server, along with any memo information.

[1616] server:

[1617] The voice data is converted into text by a speech recognition engine, and the text data is analyzed by a natural language processing model to identify redundant or inappropriate utterances, and based on that, appropriate feedback is generated and sent to the device.

[1618] Device:

[1619] Users receive feedback and adjust what they say based on in-app suggestions.

[1620] Embodiment of content recommendation function

[1621] user:

[1622] Users input their goals, such as "I want to relax in the evening" or "I want to catch up on the latest news on Sunday afternoon."

[1623] Device:

[1624] The device sends data to the server based on internal information and information from external APIs.

[1625] server:

[1626] The server uses a generative artificial intelligence model to predict the user's content consumption habits, recommends the most suitable content (movies, music, news articles, etc.) for the user, and sends the results to the device.

[1627] Device:

[1628] Users receive recommendations and choose the content they see on the app to get the best experience.

[1629] Specific examples

[1630] The following example prompts could be fed into a generative AI model:

[1631] Example prompt sentence:

[1632] The user entered their goal of "I want to relax at night," selected "Relaxed" as their mood, and "Energetic" as their health condition. Weather information obtained from an external API showed that the temperature was below 20 degrees and the traffic situation was clear. Based on this information, please recommend the most suitable relaxation content for the user.

[1633] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1634] Step 1:

[1635] The user launches the app on their smartphone or tablet and inputs their plans and goals for the day, such as "I want to relax in the evening," as well as their mood and health status. The input information is then stored in a local database on the device.

[1636] Input: User's schedule, goals, internal information

[1637] Output: Save data to a local database

[1638] Specific behavior: A user enters information into the app's input form and presses the "Save" button, which saves the data to a local database.

[1639] Step 2:

[1640] The device calls external APIs such as weather APIs and traffic APIs to obtain external information such as weather conditions and traffic conditions. The obtained external information is stored in a local database.

[1641] Input: External API request

[1642] Output: External data such as weather information, traffic information, etc.

[1643] Specific operation: The device periodically sends requests to an external API and stores the data obtained in response in a local database.

[1644] Step 3:

[1645] The device transmits the user's schedule, goals, internal information, and external information from a local database to the server, where the transmitted data is stored in a central database.

[1646] Input: All data in the local database

[1647] Output: Send data to server, store in central database

[1648] Specific operation: The terminal periodically uploads all data to the server, and the server stores the received data in a central database.

[1649] Step 4:

[1650] Based on all the data received by the server, a generative artificial intelligence model is used to predict user behavior. For example, if a user inputs "I want to relax at night," the model predicts appropriate relaxing content.

[1651] Input: All data in the central database

[1652] Output: Behavior prediction results

[1653] Specific operation: The server runs a generative artificial intelligence model to predict optimal actions based on past and new data.

[1654] Step 5:

[1655] The server sends the prediction result to the device, and the device notifies the user. For example, the server generates a prediction result such as "This movie is recommended for relaxing at night" and sends it to the device.

[1656] Input: Behavior prediction result

[1657] Output: User notification

[1658] Specific operation: The device receives the prediction results sent from the server and notifies the user visually within the app.

[1659] Step 6:

[1660] Users can take recommended actions based on the app's notifications, such as watching a recommended movie, and provide feedback to the app, allowing the system to learn from that feedback and improve its prediction accuracy next time.

[1661] Input: User feedback

[1662] Output: Accumulation of feedback data

[1663] What it does: The user enters feedback within the app, such as "I finished watching the movie," and that data is stored in a local database.

[1664] Step 7:

[1665] It records users' everyday conversations and collects audio data. For example, a user records a conversation with a friend and leaves a note saying, "I want to talk about my next trip."

[1666] Input: Audio data, memo information

[1667] Output: Save data to a local database

[1668] What happens: A user uses the app's recording feature to record a conversation and takes notes within the app.

[1669] Step 8:

[1670] The recorded voice data is stored in a local database and sent to a server, along with any memo information. The server converts the voice data into text using a speech recognition engine, and the text data is analyzed using a natural language processing model.

[1671] Input: Audio data, memo information

[1672] Output: Text data, analysis results

[1673] Specific operation: The server uses a speech recognition engine to convert speech to text and analyzes the text data using a natural language processing model.

[1674] Step 9:

[1675] Based on the results of the analysis using the natural language processing model, appropriate feedback (e.g., "Avoid talking about your previous trip") is generated and sent to the device. The device notifies the user of the feedback and displays it appropriately within the app.

[1676] Input: Analysis results

[1677] Output: Feedback results

[1678] Specific operation: The server generates appropriate feedback and sends it to the device, which notifies the user of the feedback and displays it.

[1679] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1680] This invention is a system that combines "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals; "statement prediction," which provides appropriate advice on the content and timing of statements; and "emotion recognition," which recognizes the user's emotions.

[1681] Embodiment of behavior prediction

[1682] User:

[1683] The user launches the smartphone app and inputs the plan and goals for the day. Internal information such as mood and health status is also input. For example, the user can provide information such as "I want to go to the gym in the morning and go shopping in the afternoon" or "I feel normal and my health is good."

[1684] Device:

[1685] The system stores the schedule, goals, and internal information entered by the user in a local database. It then uses an external API to obtain external information such as weather and traffic conditions. For example, it collects information such as "sunny" and "no traffic jam."

[1686] server:

[1687] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. In addition, an emotion engine extracts emotional information from the user's input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[1688] server:

[1689] Generate prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and send them to the device.

[1690] Device:

[1691] The received action prediction results are notified to the user and displayed visually within the app. The user can act based on the results and enter feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[1692] Embodiment of speech prediction

[1693] User:

[1694] Start a smartphone app and record your everyday conversations. For example, record a conversation with your best friend and leave a note saying, "I want to talk about my next trip."

[1695] Device:

[1696] The recorded voice data and notes are stored in a local database and sent to a server, where an emotion engine extracts the user's emotional information from the recorded data and sends it together.

[1697] server:

[1698] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[1699] server:

[1700] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[1701] Device:

[1702] The app notifies users of the feedback it receives and visually displays it within the app, allowing users to adjust what they say based on that feedback and improve communication. For example, when talking with a friend, users can use the advice to choose topics and keep the conversation flowing smoothly.

[1703] Emotion Recognition Embodiment

[1704] User:

[1705] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[1706] Device:

[1707] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[1708] server:

[1709] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[1710] Specific examples

[1711] Specific examples of behavioral prediction

[1712] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[1713] Specific examples of speech prediction

[1714] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1715] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[1716] The processing flow will be explained below.

[1717] Program processing of behavior prediction

[1718] Step 1:

[1719] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[1720] Step 2:

[1721] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[1722] Step 3:

[1723] Device: Stores user-entered schedules, goals, and internal information in a local database.

[1724] Step 4:

[1725] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[1726] Step 5:

[1727] Terminal: Send all collected data, including internal and external information, to the server.

[1728] Step 6:

[1729] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[1730] Step 7:

[1731] Server: Extracts user emotional information using the emotion engine based on the stored data.

[1732] Step 8:

[1733] Server: Based on the stored data and emotional information, the AI ​​model predicts optimal behavioral patterns.

[1734] Step 9:

[1735] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[1736] Step 10:

[1737] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[1738] Step 11:

[1739] User: Follows the suggested behavioral patterns and achieves the goal.

[1740] Step 12:

[1741] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[1742] ---

[1743] Program processing of speech prediction

[1744] Step 1:

[1745] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[1746] Step 2:

[1747] Users: Leave notes during conversations about things that concern them or that need attention.

[1748] Step 3:

[1749] Device: Stores recorded audio data and notes in a local database.

[1750] Step 4:

[1751] Terminal: Sends voice data and memo information to the server.

[1752] Step 5:

[1753] Server: Converts voice data received from the device into text using a speech recognition engine.

[1754] Step 6:

[1755] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[1756] Step 7:

[1757] Server: Extracts user emotion information from voice and text data using an emotion engine.

[1758] Step 8:

[1759] Server: Generates appropriate feedback based on sentiment information and text analysis results (e.g., "Mention your next trip, but avoid talking about your previous trip").

[1760] Step 9:

[1761] Server: Sends the generated feedback to the device.

[1762] Step 10:

[1763] On the device: Notify the user of the feedback received and display it visually within the app.

[1764] Step 11:

[1765] Users: Use suggested feedback to adjust conversations and improve communication.

[1766] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[1767] Example 2

[1768] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1769] Conventional behavior prediction systems only predict behavior based on a user's schedule and goals, and have the problem of being unable to provide appropriate feedback or advice that takes emotional information into account. Furthermore, speech prediction systems are limited to identifying redundant or inappropriate speech, making it difficult to provide advice that appropriately reflects the user's emotions. This has led to the issue of a lack of effective support for user behavior and speech.

[1770] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1771] In this invention, the server includes: means for a user to input internal information such as the user's schedule and goals for the day, mood, and health status; means for collecting the user's internal information and external information such as weather conditions and traffic conditions; means for predicting the user's behavior based on the schedule, goals, internal information, and external information and generating an optimal behavior pattern based on emotional information; means for notifying the user of the predicted behavior pattern and collecting user feedback; means for recording the user's daily conversations and collecting voice data; means for converting the voice data into text; means for identifying redundant or inappropriate statements from the text data and advising the user on appropriate statements based on emotional information; and means for generating a pre-alert for statements and advice on appropriate statements based on the identified statements and notifying the user. This makes it possible to optimize the user's behavior and statements while taking emotions into consideration, thereby providing effective support.

[1772] "User" refers to an individual or corporation that uses the system, inputs their own schedules, goals, and internal information, and receives feedback from the system.

[1773] "Internal information" is information about a user's internal state, such as their mood, health, or emotions.

[1774] "External information" refers to information about the user's external environment, such as weather conditions and traffic conditions, and is obtained using an external API.

[1775] "Emotion information" is information about the user's emotional state, and is extracted by an emotion engine or the like.

[1776] An "emotion engine" is a software or hardware mechanism for extracting emotional information from user input or voice data.

[1777] A "behavioral pattern" is a schedule or action plan for the user to act optimally, and is the result predicted by the system.

[1778] "Feedback" is information that users input into the system based on the results of their actual actions, and the system uses this information to make its next prediction more accurate.

[1779] "Speech prediction" is a system function that analyzes the user's everyday conversations and advises them on appropriate content to say.

[1780] A "speech recognition engine" is a software or hardware mechanism for converting a user's voice data into text data.

[1781] A "natural language processing model" is a technology for analyzing text data and identifying redundant or inappropriate statements.

[1782] A "pre-alert" is a warning message that the system displays before a user makes an identified inappropriate comment.

[1783] "Advice on appropriate speech" refers to advice on speech content provided by the system to help users communicate more smoothly.

[1784] This invention is a system that combines "behavior prediction" that efficiently predicts a user's daily behavior and helps them achieve their schedules and goals, "statement prediction" that provides appropriate advice on the content and timing of statements, and "emotion recognition" that recognizes the user's emotions. A specific embodiment of this system will be described below.

[1785] Embodiment of behavior prediction

[1786] User:

[1787] Users launch the smartphone app and enter detailed information about the day's plans, goals, mood, health status, etc. For example, they might enter, "I want to go to the gym in the morning and do some shopping in the afternoon," or "I'm feeling normal, and my health is good."

[1788] Device:

[1789] The device stores the user's schedule, goals, mood, and health status in a local database. It then uses external APIs to obtain external information such as weather conditions (e.g., OpenWeatherMap API) and traffic conditions (e.g., Google Maps Traffic API). For example, it collects information such as "sunny" and "no traffic jam" and sends it to the server.

[1790] server:

[1791] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. At the same time, it uses an emotion engine to extract emotional information from the user's input. Based on the stored data and emotional information, the AI ​​model predicts the optimal behavioral pattern. For example, it generates a result such as "It is best to leave for the gym at 9:00 AM and the market at 2:00 PM" and sends it to the device.

[1792] Device:

[1793] The device then notifies the user of the predicted behavior and displays it visually within the app. The user can then act based on the results and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[1794] Embodiment of speech prediction

[1795] User:

[1796] Users can launch the smartphone app and record everyday conversations, such as a conversation with a close friend, and leave a note saying, "I want to talk about my next trip."

[1797] Device:

[1798] The device stores the recorded voice data and notes in a local database, and uses an emotion engine to extract the user's emotional information from the recorded data and send it together with the data to the server.

[1799] server:

[1800] The server converts the voice data into text using a speech recognition engine (e.g., Google Speech-to-Text API), analyzes the converted text data using a natural language processing model, and identifies redundant or inappropriate statements. For example, it generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading." It then generates appropriate feedback based on the emotional information. For example, it generates advice such as "Talk about your next trip, but avoid talking about your previous trip," and sends it to the device.

[1801] Device:

[1802] The device will notify the user of the received feedback and display it visually within the app, allowing the user to adjust what they say based on this feedback and improve communication. For example, when talking with a friend, the user can use the advice to choose topics and keep the conversation flowing smoothly.

[1803] Emotion Recognition Embodiment

[1804] User:

[1805] Users provide emotional information to the app by inputting their mood and emotions and recording changes in facial expressions and voice during conversations.

[1806] Device:

[1807] The device uses an emotion recognition engine to extract emotional information from the recorded voice data and transmits the emotional information to a server.

[1808] server:

[1809] The server uses the emotional information analyzed by the emotion recognition engine and natural language processing model to generate behavioral patterns and speech feedback based on the user's emotions.

[1810] Specific examples

[1811] Specific examples of behavioral prediction

[1812] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "Sunny" from a weather API (e.g., OpenWeatherMap API) and "No traffic jam" from a traffic API (e.g., Google Maps Traffic API), and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "The best time to leave for the gym is 9:00 AM and the market is 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[1813] Example prompt sentence:

[1814] User: I want to go to the gym in the morning and do some shopping in the afternoon. I feel normal and my health is good.

[1815] Terminal: The weather is clear and traffic is smooth.

[1816] Server: The best time to leave for the gym is 9am and the market is 2pm.

[1817] Specific examples of speech prediction

[1818] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text (e.g., Google Speech-to-Text API) and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1819] Example prompt sentence:

[1820] User: Recording a conversation with a best friend. Want to talk about an upcoming trip.

[1821] Device: Sends voice data, notes, and emotional information to the server.

[1822] Server: Analyzes the text for redundant content and generates appropriate feedback.

[1823] Server: Mention your next trip, but avoid talking about your last trip.

[1824] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1825] Behavior prediction processing steps

[1826] Step 1:

[1827] User: Launches the smartphone app and enters internal information such as the day's plans, goals, mood, and health status.

[1828] Input: Schedule, Goals, Mood, Health

[1829] Output: User input data

[1830] For example: "I go to the gym in the morning and shop at the market in the afternoon." "I feel fine and my health is good."

[1831] Step 2:

[1832] Device: Stores the entered information in a local database, then uses external APIs to retrieve external information such as weather conditions and traffic conditions.

[1833] Input: User-entered data

[1834] Output: Internal and external information

[1835] Examples: "The weather is clear" and "The traffic is smooth."

[1836] Step 3:

[1837] Server: Receives all data sent from the devices and stores it in a central database. Uses an emotion engine to extract emotional information from user input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[1838] Input: User input data, internal information, external information

[1839] Output: Predicted behavior pattern

[1840] For example: "The best time to leave for the gym is 9am and the best time to leave for the market is 2pm."

[1841] Step 4:

[1842] Server: Sends the prediction results to the device.

[1843] Input: Predicted behavior pattern

[1844] Output: Data sent to the terminal

[1845] Step 5:

[1846] On the device: The user is notified of the received predictions and visually displayed within the app, allowing the user to act on them and provide feedback to the app.

[1847] Input: Predicted behavior pattern

[1848] Output: User feedback

[1849] Step 6:

[1850] Device: User feedback is sent to the server to help improve prediction accuracy in future predictions.

[1851] Input: User feedback

[1852] Output: Data sent to the server

[1853] Speech prediction processing steps

[1854] Step 1:

[1855] User: Launches the smartphone app, records everyday conversations, and enters notes.

[1856] Input: Recording data, notes

[1857] Output: Recording data and notes

[1858] Examples: "I want to record a conversation with my best friend" or "I want to talk about my next trip."

[1859] Step 2:

[1860] Device: Recorded voice data and notes are stored in a local database, and emotional information is extracted using an emotion engine, which then sends the data to the server.

[1861] Input: Recording data, notes

[1862] Output: Recording data, notes, emotional information

[1863] Step 3:

[1864] Server: The speech data is converted into text using a speech recognition engine, and a natural language processing model is used to identify redundant or inappropriate speech.

[1865] Input: Audio recording, emotional information

[1866] Output: Text data and detected verbose / inappropriate comments

[1867] For example: "This part is redundant and difficult to understand" or "This part may be misleading"

[1868] Step 4:

[1869] Server: Generates appropriate feedback based on emotion information and sends it to the device.

[1870] Input: Text data, emotion information

[1871] Output: Feedback

[1872] For example: "Talk about your next trip, but avoid talking about your last trip."

[1873] Step 5:

[1874] On-device: The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust their voice accordingly.

[1875] Input: Feedback

[1876] Output: Adjusted speech data

[1877] Processing steps for emotion recognition

[1878] Step 1:

[1879] User: Inputs mood and emotions and records facial expressions and vocal changes during conversations.

[1880] Input: Mood, emotion, facial expression, vocal changes

[1881] Output: Emotion data

[1882] Step 2:

[1883] Terminal: The emotion recognition engine extracts emotional information from the voice data and sends it to the server.

[1884] Input: Recording data

[1885] Output: Emotional information

[1886] Step 3:

[1887] Server: Analyzes emotional information and generates behavioral patterns and speech feedback based on the user's emotions.

[1888] Input: Emotion information

[1889] Output: Behavioral patterns, speech feedback

[1890] (Application example 2)

[1891] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1892] Currently, there is a lack of security systems that understand users' daily behavior and emotions and provide advice on appropriate behavioral patterns and speech based on that understanding. This can make it difficult for users to understand the degree of risk involved in their own behavior and speech, making it difficult to ensure their safety. The present invention aims to provide a system that comprehensively analyzes users' behavior, speech, and emotional information to reduce security risks.

[1893] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1894] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information, and external information, means for presenting the predicted behavioral pattern to the user, means for predicting the user's utterances from the content of the user's conversation and recorded voice data and providing appropriate advice, and means for recognizing the user's emotions and reflecting them in the behavioral pattern and utterance content, thereby enabling the user to understand the degree to which their actions and utterances pose security risks and to take appropriate preventative measures.

[1895] "User" refers to an individual who uses this system.

[1896] "Means for entering the day's plans and goals" refers to the function that allows users to enter their daily plans and goals into the system.

[1897] "Internal information" refers to information about a user's personal state or emotions, such as mood or health.

[1898] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[1899] "Means for predicting behavior" refers to the function of calculating the optimal behavioral pattern for a user based on internal and external information.

[1900] "Means for presenting predicted behavioral patterns" refers to the function of notifying the user of the predicted results and displaying them visually.

[1901] "Means of predicting what will be said based on conversation content and recorded audio data and providing appropriate advice" refers to a function that analyzes a user's everyday conversations and generates dedicated alerts and advice.

[1902] "Means of recognizing emotions and reflecting them in behavioral patterns and speech content" refers to a function that analyzes the user's emotions and provides behavioral patterns and speech advice that take these into consideration.

[1903] A "system that reduces security risks" refers to a system that evaluates the risks associated with users' actions and statements and supports safe behavior and communication.

[1904] This invention combines a system that efficiently predicts a user's daily behavior and supports safe behavior, a system that analyzes conversation content and gives advice on appropriate remarks, and a system that recognizes the user's emotions and assesses risk. The system is intended to be used by users through a smartphone app.

[1905] To implement this system, the following hardware and software are used:

[1906] Hardware: smartphones, servers, cameras and microphones for emotion recognition

[1907] Software: Python, external API (weather API, traffic API), voice recognition AI engine (Hugging Face transformers), emotion recognition AI engine (DeepFace)

[1908] Embodiment of behavior prediction

[1909] Device:

[1910] The user launches the app on their smartphone and enters their plans and goals for the day, as well as internal information such as their mood and health status. The device stores this information in a local database and uses external APIs to obtain external information such as weather and traffic conditions.

[1911] server:

[1912] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. The AI ​​model uses this data to predict optimal behavioral patterns. For example, it generates a result such as "It's sunny and traffic is good, so it's best to leave for the gym at 9:00 AM." The server then sends this prediction to the device.

[1913] Device:

[1914] The predictions received are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app.

[1915] Embodiment of speech prediction

[1916] Device:

[1917] The user launches the app on their smartphone and records their everyday conversations. The recorded audio data is stored in a local database and sent to a server. The emotion engine extracts the user's emotional information from the recorded data and sends it along with the data.

[1918] server:

[1919] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[1920] server:

[1921] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[1922] Device:

[1923] The feedback received is notified to the user and displayed visually within the app, allowing the user to adjust what they say and communicate more effectively.

[1924] Emotion Recognition Embodiment

[1925] user:

[1926] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[1927] Device:

[1928] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[1929] server:

[1930] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[1931] Specific examples

[1932] Specific examples of behavioral prediction

[1933] The user inputs "I plan to go to the gym at 6 PM" and also inputs their mood and health status as additional information. The device obtains "sunny" information from the weather API and "no traffic jams" information from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict that "the best time to leave is 5 PM" and sends the result to the device. The device notifies the user of the result, and the user acts accordingly.

[1934] Specific examples of speech prediction

[1935] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[1936] Prompt Sentence Examples

[1937] Behavioral prediction: "Predict what time a user should go to the gym based on current weather and traffic information."

[1938] Speech Prediction: "Identify sensitive information that users should avoid in conversations with their friends."

[1939] Emotion Recognition: "Perform emotion analysis using the user's current image to conduct risk assessment."

[1940] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1941] Behavior prediction processing steps

[1942] Step 1: User enters schedule and goal

[1943] The user launches the smartphone app and enters the day's plans, goals, and internal information (mood and health status).

[1944] Input: Schedule, goals, internal information

[1945] Output: Data stored in a local database

[1946] Step 2: Obtaining external information

[1947] The device calls external APIs (weather API, traffic API) to obtain external information such as weather conditions and traffic conditions.

[1948] Input: External API endpoint

[1949] Output: Weather conditions, traffic conditions (e.g. sunny, no traffic jams)

[1950] Step 3: Sending data

[1951] The terminal transmits the input internal information and the acquired external information to the server.

[1952] Input: Plan, goal, internal information, external information

[1953] Output: All data sent to the server

[1954] Step 4: Predicting behavioral patterns

[1955] Based on all the data received by the server, an AI model is used to predict optimal behavioral patterns.

[1956] Input: All data and AI model

[1957] Output: Predicted behavioral pattern (e.g., departure at 5pm)

[1958] Step 5: Send prediction results

[1959] The server sends the prediction results to the device.

[1960] Input: predicted behavior pattern

[1961] Output: Prediction results sent to the device

[1962] Step 6: Visualizing the results

[1963] The device notifies the user of the prediction results received and displays them visually within the app.

[1964] Input: predicted behavior pattern

[1965] Output: The result communicated to the user

[1966] Step 7: Provide feedback

[1967] The user provides feedback on the results of the execution to the app.

[1968] Input: Actual action results

[1969] Output: Feedback data

[1970] Speech prediction processing steps

[1971] Step 1: Recording audio data

[1972] Users launch the smartphone app and record their everyday conversations.

[1973] Input: Audio data

[1974] Output: Audio data stored in a local database

[1975] Step 2: Sending audio data

[1976] The device sends the recorded audio data to the server.

[1977] Input: Audio data

[1978] Output: Audio data sent to the server

[1979] Step 3: Voice Recognition

[1980] The server converts the voice data into text using a voice recognition engine.

[1981] Input: Audio data

[1982] Output: Text data

[1983] Step 4: Parsing utterances

[1984] The server analyzes the text data using a natural language processing model to identify inappropriate or redundant statements.

[1985] Input: Text data, NLP model

[1986] Output: Analysis results (identification of inappropriate and redundant parts)

[1987] Step 5: Generate Advice

[1988] The server generates feedback on the speech based on the analysis results and emotional information.

[1989] Input: Analysis results, emotion information

[1990] Output: Feedback (e.g. mention the next trip but avoid talking about the previous trip)

[1991] Step 6: Submit your feedback

[1992] The server generates feedback and sends it to the device.

[1993] Input: Feedback

[1994] Output: Feedback sent to the terminal

[1995] Step 7: Notification of feedback

[1996] Notify the user of the feedback received by the device and display it visually within the app.

[1997] Input: Feedback

[1998] Output: Feedback given to the user

[1999] Processing steps for emotion recognition

[2000] Step 1: Enter emotional information

[2001] Users input their moods and emotions into the app.

[2002] Input: Mood, emotional information

[2003] Output: Emotion information stored in a local database

[2004] Step 2: Analyzing the audio data

[2005] Emotional information is extracted from the audio data recorded by the device and sent to the server.

[2006] Input: Audio data

[2007] Output: Extracted emotion information

[2008] Step 3: Sending emotional information

[2009] The device transmits the emotion information to the server.

[2010] Input: Emotion information

[2011] Output: Emotion information sent to the server

[2012] Step 4: Sentiment-based analysis

[2013] The server uses the emotion information to complement the analysis results of behavior prediction and utterance prediction.

[2014] Input: Emotion information, behavior prediction, speech prediction

[2015] Output: Analysis results (behavior and speech) supplemented based on emotions

[2016] Step 5: Notification of completion results

[2017] The server transmits the completed analysis results to the terminal.

[2018] Input: Completed analysis results

[2019] Output: Completion results sent to the terminal

[2020] Step 6: Visualizing the results

[2021] The device notifies the user of the completion results received and displays them visually within the app.

[2022] Input: Completed analysis results

[2023] Output: Completion result notified to the user

[2024] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[2025] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2026] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[2027] [Fourth embodiment]

[2028] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[2029] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2031] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[2032] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[2033] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[2034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[2035] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[2036] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[2037] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2039] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[2040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2041] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[2042] Embodiment of behavior prediction

[2043] User:

[2044] Users launch the smartphone app and enter their plans and goals for the day. They also enter internal information such as their mood and health status. For example, a user might enter their plan to "go to the gym in the morning and go shopping in the afternoon," select "normal" as their mood, and "good" as their health status.

[2045] Device:

[2046] The schedule, goals, and internal information entered by the user are stored in a local database. Then, external APIs are used to obtain external information such as weather conditions and traffic conditions. For example, a weather API is called to obtain information such as "sunny," and a traffic API is called to obtain information such as "no traffic jams."

[2047] server:

[2048] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. Based on the stored data, an AI model predicts the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, the AI ​​model may generate a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[2049] Device:

[2050] The received prediction results are notified to the user and displayed visually within the app, allowing the user to act on them and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[2051] Embodiment of speech prediction

[2052] User:

[2053] Start the smartphone app and record your everyday conversations. For example, record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[2054] Device:

[2055] The recorded voice data is stored in a local database and sent to the server, along with any memo information.

[2056] server:

[2057] The voice data received from the device is converted into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[2058] server:

[2059] Based on the analysis results, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent to the device.

[2060] Device:

[2061] The received feedback is notified to the user and displayed within the app, allowing the user to adjust what they say based on the feedback and improve communication. For example, when talking with a friend, a user can refer to the advice to choose a topic and smoothly advance the conversation.

[2062] Specific examples

[2063] Specific examples of behavioral prediction

[2064] The user enters their schedule into the app, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model to predict behavioral patterns such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends this to the device. The device notifies the user of the results, and the user acts accordingly.

[2065] Specific examples of speech prediction

[2066] A user records a conversation with a close friend and leaves a note saying, "I'd like to talk about my next trip." The device then sends the recorded audio data and note to a server. The server converts the audio data into text and uses a natural language processing model to identify "overly long explanations" and "potentially misleading" parts. Based on this, the device generates and sends feedback to the device, such as "mention your next trip, but avoid talking about your previous trip." The device then notifies the user of the feedback, and the user can use the advice to smoothly move the conversation forward.

[2067] This allows users to receive support in achieving their schedules and goals, as well as in communicating smoothly.

[2068] The processing flow will be explained below.

[2069] Program processing of behavior prediction

[2070] Step 1:

[2071] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[2072] Step 2:

[2073] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[2074] Step 3:

[2075] Device: Stores user-entered schedules, goals, and internal information in a local database.

[2076] Step 4:

[2077] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[2078] Step 5:

[2079] Terminal: Send all collected data, including internal and external information, to the server.

[2080] Step 6:

[2081] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[2082] Step 7:

[2083] Server: Based on the stored data, it uses an AI model to predict the user's optimal behavioral patterns.

[2084] Step 8:

[2085] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[2086] Step 9:

[2087] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[2088] Step 10:

[2089] User: Follows the suggested behavioral patterns and achieves the goal.

[2090] Step 11:

[2091] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[2092] Program processing of speech prediction

[2093] Step 1:

[2094] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[2095] Step 2:

[2096] Users: Leave notes during conversations about things that concern them or that need attention.

[2097] Step 3:

[2098] Device: Stores recorded audio data and notes in a local database.

[2099] Step 4:

[2100] Terminal: Sends voice data and memo information to the server.

[2101] Step 5:

[2102] Server: Converts voice data received from the device into text using a speech recognition engine.

[2103] Step 6:

[2104] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[2105] Step 7:

[2106] Server: Based on the analysis results, generate appropriate feedback based on the profile information of the person you are speaking to.

[2107] Step 8:

[2108] Server: Sends the generated feedback (e.g., "mention your next trip, but avoid talking about your previous trip") to the device.

[2109] Step 9:

[2110] On the device: Notify the user of the feedback received and display it visually within the app.

[2111] Step 10:

[2112] Users: Use suggested feedback to adjust conversations and improve communication.

[2113] This allows users to effectively achieve their schedules and goals, as well as communicate smoothly.

[2114] Example 1

[2115] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2116] Conventional systems have difficulty predicting behavior based on internal and external information in addition to users' schedules and goals. Furthermore, there was a lack of systems that notified users of appropriate timing and content for speech in everyday conversations. Furthermore, there was no mechanism in place to incorporate user feedback to improve the system's overall prediction accuracy. This made it difficult for users to achieve their schedules and communicate smoothly.

[2117] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2118] In this invention, the server includes: means for a user to input the day's schedule and goals; means for collecting the user's internal and external information; means for predicting the user's behavior based on the schedule, goals, internal and external information; means for presenting the predicted behavioral pattern to the user; means for collecting voice data and converting it into text; means for identifying redundant or inappropriate comments from the text data; means for generating a pre-alert for a comment and advice on appropriate comments based on the identified comments and notifying the user; means for collecting and storing user feedback data; and means for improving the prediction accuracy of the entire system based on the feedback data, thereby enabling the user to achieve their schedule and communicate smoothly.

[2119] "User" refers to the entity that uses the system, such as the person who inputs schedules and goals, provides feedback, and records voice data.

[2120] "Schedule" refers to the planned actions and tasks that the user intends to carry out on the day.

[2121] "Goals" refer to the specific outcomes or results that a user is trying to achieve on that day.

[2122] "Internal information" refers to information related to the user's internal state, such as their mood or health.

[2123] "External information" refers to information related to the external environment that affects user behavior, such as weather conditions and traffic conditions.

[2124] "Behavioral patterns" refer to the optimal sequence and timing of actions predicted based on the user's schedule, goals, internal information, and external information.

[2125] "Voice data" refers to audio data obtained when a user records everyday conversations.

[2126] "Text data" refers to the result of converting voice data into text information using voice recognition technology.

[2127] "Redundant utterances" refer to utterances in a conversation that are unnecessarily long or overlapping.

[2128] "Inappropriate remarks" refer to remarks that may be misunderstood in a conversation or that may be perceived as offensive by the other person.

[2129] "Advance Alert" refers to a warning or notification to a user about an upcoming action or statement.

[2130] "Feedback data" refers to the actual behavioral results and impressions that users provide in response to predictions and advice from the system.

[2131] "Generative AI models" refer to machine learning models used to predict behavioral patterns and speech content based on user data.

[2132] This invention is a system that provides two main functions: "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals, and "statement prediction," which provides appropriate advice on the content and timing of statements.

[2133] Embodiment of behavior prediction

[2134] User:

[2135] The user launches the smartphone app and inputs their plans and goals for the day. They also input internal information such as their mood and health status. For example, they input specific plans such as "I want to go to the gym in the morning and go shopping in the afternoon," and select "normal" as their mood and "good" as their health status.

[2136] Device:

[2137] The device stores the schedule, goals, and internal information entered by the user in a local database. At the same time, it uses external APIs to obtain external information such as weather conditions and traffic conditions. Specifically, it obtains information such as "sunny" from the weather API and "no traffic jams" from the traffic API.

[2138] server:

[2139] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. Based on the stored data, a generative AI model is used to predict the user's behavior. Based on this prediction, the optimal behavioral pattern is calculated and the result is sent to the device. For example, it generates a prediction that "it is best to leave for the gym at 9:00 AM and the market at 2:00 PM."

[2140] Device:

[2141] The device receives the prediction results from the server and notifies the user, visually displaying them within the app. The user can then act on the results and enter the results as feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[2142] Embodiment of speech prediction

[2143] User:

[2144] Users can launch the smartphone app and record everyday conversations. For example, they can record a conversation with a friend and leave a note saying, "I want to talk about my next trip."

[2145] Device:

[2146] The device stores the recorded voice data and memo information in a local database and transmits it to the server.

[2147] server:

[2148] The server converts the voice data received from the device into text using a speech recognition engine. The converted text data is then analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, it generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[2149] server:

[2150] Based on the analysis results, the server generates appropriate feedback, such as advice like "mention your next trip, but avoid talking about your previous trip," and sends it to the device.

[2151] Device:

[2152] The device notifies the user of the received feedback and displays it in the app. The user can then use the advice to adjust what they say and communicate smoothly. For example, a user can use the advice to choose a topic and smoothly advance the conversation with a friend.

[2153] Prompt Sentence Examples

[2154] Behavioral prediction prompt:

[2155] "The user entered their plan to go to the gym in the morning and do some shopping at the market in the afternoon, and also entered their mood as normal and their health status as good. The device retrieved information from the weather API that it was sunny and from the traffic API that it was clear, and sent this information to the server. Based on this information, the server predicted the optimal behavioral pattern and notified the user to leave for the gym at 9 a.m. and the market at 2 p.m."

[2156] Predictive prompt:

[2157] "A user recorded a conversation with a close friend and left a note saying they wanted to talk about their next trip. The device sent the recorded audio and note to a server. The server converted the audio to text and used a natural language processing model to identify overly long and potentially misleading descriptions. As a result, the server generated feedback and sent it to the device, advising them to mention their next trip but avoid discussing their previous trip."

[2158] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2159] Behavior prediction processing steps

[2160] Step 1: User Data Entry

[2161] input:

[2162] Users launch the smartphone app and enter their plans and goals for the day.

[2163] Users input internal information such as mood and health status.

[2164] Specific behavior:

[2165] The user inputs "I want to go to the gym in the morning and do some shopping in the afternoon," and selects his mood as "normal" and his health status as "good."

[2166] output:

[2167] The input schedule, goals, and internal information are generated.

[2168] Step 2: Saving data on the device and acquiring external information

[2169] input:

[2170] User schedules, goals, and internal information entered into a smartphone app

[2171] Specific behavior:

[2172] The terminal stores the entered information in a local database.

[2173] The device accesses external APIs (weather APIs and traffic APIs) to obtain external information (weather conditions and traffic conditions).

[2174] output:

[2175] User schedules, goals, and internal information stored in a local database

[2176] External information obtained ("Sunny" from the weather API, "No traffic jam" from the traffic API)

[2177] Step 3: Receiving data from the server and making predictions using the AI ​​model

[2178] input:

[2179] Schedules, goals, internal information, and external information sent from the device

[2180] Specific behavior:

[2181] The server receives all data sent by the devices and stores it in a central database.

[2182] The server uses a generative AI model based on the stored data to predict behavioral patterns.

[2183] The server inputs data into an AI model that predicts that the best time to leave for the gym is 9 a.m. and the best time to leave for the market is 2 p.m.

[2184] output:

[2185] Prediction results of optimal behavioral patterns

[2186] Step 4: Notification of results and feedback via device

[2187] input:

[2188] Prediction results sent from the server

[2189] Specific behavior:

[2190] The terminal notifies the user of the prediction result received from the server.

[2191] The device will visually display the prediction results within the app.

[2192] The user acts based on the prediction results and inputs the results into the app as feedback.

[2193] output:

[2194] User notifications and in-app displays

[2195] Feedback Data

[2196] ---

[2197] Speech prediction processing steps

[2198] Step 1: User voice recording and note taking

[2199] input:

[2200] Users launch the smartphone app and record their everyday conversations.

[2201] The user enters notes for the recording.

[2202] Specific behavior:

[2203] For example, a user records a conversation with a friend and enters a note saying, "I want to talk about my next trip."

[2204] output:

[2205] Recorded audio data and memo information

[2206] Step 2: Save and send data from your device

[2207] input:

[2208] Recorded audio data and memo information

[2209] Specific behavior:

[2210] The device stores the recorded audio data and memo information in a local database.

[2211] The terminal transmits the saved data to the server.

[2212] output:

[2213] Audio data and notes stored in a local database

[2214] Data sent to the server

[2215] Step 3: Audio data conversion and analysis on the server

[2216] input:

[2217] Voice data and memo information sent from the device

[2218] Specific behavior:

[2219] The server converts the voice data into text using a voice recognition engine.

[2220] The server analyzes the converted text data using a natural language processing model.

[2221] The server identifies redundant or inappropriate statements and generates analysis results such as "This part is too long and difficult to understand" or "This part may be misleading."

[2222] output:

[2223] Text data and analysis results

[2224] Step 4: Server feedback generation and sending

[2225] input:

[2226] Analysis results

[2227] Specific behavior:

[2228] The server generates appropriate utterance feedback based on the analysis results.

[2229] For example, advice such as "mention your next trip, but avoid talking about your previous trip" may be generated.

[2230] The server transmits the generated feedback to the terminal.

[2231] output:

[2232] Feedback Data

[2233] Step 5: Feedback notification by device

[2234] input:

[2235] Feedback data sent from the server

[2236] Specific behavior:

[2237] The terminal notifies the user of the feedback data.

[2238] The device displays the feedback within the app, giving users information to adjust what they say.

[2239] output:

[2240] User feedback notification and in-app display

[2241] (Application example 1)

[2242] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2243] Conventional behavior prediction systems and comment prediction systems are limited to supporting actions based on the user's schedule and goals, making it difficult to address diverse user needs. Furthermore, they lack sufficient support for content recommendations and communication improvement, resulting in a lack of improvement in the user experience. Therefore, there is a need for a comprehensive system that provides appropriate advice based not only on the user's daily actions and comments, but also on their content consumption trends and message content.

[2244] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2245] In this invention, the server includes means for the user to input the day's schedule and goals, means for collecting the user's internal information and external information, means for predicting the user's behavior based on the schedule, goals, internal information and external information, means for presenting the predicted behavioral patterns to the user, means for predicting the user's content consumption tendencies using a generative artificial intelligence model and recommending optimal content, and means for providing advice to improve the user's comments based on the content of the user's message.

[2246] This allows users to receive comprehensive support for their actions, content recommendations, and advice on what to say.

[2247] A "user" is an individual who uses the system to receive assistance with schedule input, information gathering, behavior prediction, and speech advice.

[2248] A "generative artificial intelligence model" is an algorithm that learns various patterns and trends based on data and makes predictions and recommendations.

[2249] "Content consumption trends" refers to patterns such as what types of content a user prefers to consume and what types of content they consume at what times of the day.

[2250] "Message content" refers to all statements and texts made by a User when communicating with others.

[2251] "Advice" means instructions or guidance provided to a user to help them act or speak more appropriately.

[2252] "Schedules" refer to planned actions and events that a user undertakes in their daily life.

[2253] A "goal" is a specific action or outcome that a user aims to achieve.

[2254] "Internal information" refers to information about the user's inner self, such as their emotional state or health status.

[2255] "External information" refers to information about the user's external environment, such as weather and traffic conditions.

[2256] "Behavioral prediction" is the calculation of optimal behavioral patterns based on a user's schedule, goals, internal information, and external information.

[2257] A "behavioral pattern" refers to a series of actions that a user takes at what timing.

[2258] "Content" is a general term for information assets consumed by users, such as music, videos, and articles.

[2259] The system for carrying out the present invention integrates various functions for predicting and optimizing user actions and comments. Specific embodiments will be described below.

[2260] Hardware and Software Configuration

[2261] Hardware configuration:

[2262] Devices such as smartphones, tablets, or computers

[2263] Servers (including using cloud computing environments)

[2264] Software configuration:

[2265] Application software (smartphone apps, etc.)

[2266] Databases (local and central)

[2267] External APIs (weather API, traffic API, etc.)

[2268] Speech Recognition Engine

[2269] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)

[2270] Generative artificial intelligence model (AI model)

[2271] Embodiment of behavior prediction

[2272] user:

[2273] Users open the app on their smartphone or tablet and enter their plans and goals for the day, as well as internal information such as their mood and health status.

[2274] Device:

[2275] The device stores user-entered schedules, goals, and internal information in a local database, and then uses external APIs to retrieve external information such as weather and traffic conditions.

[2276] server:

[2277] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. It then uses this data to predict the user's behavior using a generative artificial intelligence model. It then sends the prediction results to the device and notifies the user.

[2278] Embodiment of speech prediction

[2279] user:

[2280] Users start the smartphone app and record their everyday conversations, taking notes on topics they want to talk about.

[2281] Device:

[2282] The recorded voice data is stored in a local database and then sent to the server, along with any memo information.

[2283] server:

[2284] The voice data is converted into text by a speech recognition engine, and the text data is analyzed by a natural language processing model to identify redundant or inappropriate utterances, and based on that, appropriate feedback is generated and sent to the device.

[2285] Device:

[2286] Users receive feedback and adjust what they say based on in-app suggestions.

[2287] Embodiment of content recommendation function

[2288] user:

[2289] Users input their goals, such as "I want to relax in the evening" or "I want to catch up on the latest news on Sunday afternoon."

[2290] Device:

[2291] The device sends data to the server based on internal information and information from external APIs.

[2292] server:

[2293] The server uses a generative artificial intelligence model to predict the user's content consumption habits, recommends the most suitable content (movies, music, news articles, etc.) for the user, and sends the results to the device.

[2294] Device:

[2295] Users receive recommendations and choose the content they see on the app to get the best experience.

[2296] Specific examples

[2297] The following example prompts could be fed into a generative AI model:

[2298] Example prompt sentence:

[2299] The user entered their goal of "I want to relax at night," selected "Relaxed" as their mood, and "Energetic" as their health condition. Weather information obtained from an external API showed that the temperature was below 20 degrees and the traffic situation was clear. Based on this information, please recommend the most suitable relaxation content for the user.

[2300] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2301] Step 1:

[2302] The user launches the app on their smartphone or tablet and inputs their plans and goals for the day, such as "I want to relax in the evening," as well as their mood and health status. The input information is then stored in a local database on the device.

[2303] Input: User's schedule, goals, internal information

[2304] Output: Save data to a local database

[2305] Specific behavior: A user enters information into the app's input form and presses the "Save" button, which saves the data to a local database.

[2306] Step 2:

[2307] The device calls external APIs such as weather APIs and traffic APIs to obtain external information such as weather conditions and traffic conditions. The obtained external information is stored in a local database.

[2308] Input: External API request

[2309] Output: External data such as weather information, traffic information, etc.

[2310] Specific operation: The device periodically sends requests to an external API and stores the data obtained in response in a local database.

[2311] Step 3:

[2312] The device transmits the user's schedule, goals, internal information, and external information from a local database to the server, where the transmitted data is stored in a central database.

[2313] Input: All data in the local database

[2314] Output: Send data to server, store in central database

[2315] Specific operation: The terminal periodically uploads all data to the server, and the server stores the received data in a central database.

[2316] Step 4:

[2317] Based on all the data received by the server, a generative artificial intelligence model is used to predict user behavior. For example, if a user inputs "I want to relax at night," the model predicts appropriate relaxing content.

[2318] Input: All data in the central database

[2319] Output: Behavior prediction results

[2320] Specific operation: The server runs a generative artificial intelligence model to predict optimal actions based on past and new data.

[2321] Step 5:

[2322] The server sends the prediction result to the device, and the device notifies the user. For example, the server generates a prediction result such as "This movie is recommended for relaxing at night" and sends it to the device.

[2323] Input: Behavior prediction result

[2324] Output: User notification

[2325] Specific operation: The device receives the prediction results sent from the server and notifies the user visually within the app.

[2326] Step 6:

[2327] Users can take recommended actions based on the app's notifications, such as watching a recommended movie, and provide feedback to the app, allowing the system to learn from that feedback and improve its prediction accuracy next time.

[2328] Input: User feedback

[2329] Output: Accumulation of feedback data

[2330] What it does: The user enters feedback within the app, such as "I finished watching the movie," and that data is stored in a local database.

[2331] Step 7:

[2332] It records users' everyday conversations and collects audio data. For example, a user records a conversation with a friend and leaves a note saying, "I want to talk about my next trip."

[2333] Input: Audio data, memo information

[2334] Output: Save data to a local database

[2335] What happens: A user uses the app's recording feature to record a conversation and takes notes within the app.

[2336] Step 8:

[2337] The recorded voice data is stored in a local database and sent to a server, along with any memo information. The server converts the voice data into text using a speech recognition engine, and the text data is analyzed using a natural language processing model.

[2338] Input: Audio data, memo information

[2339] Output: Text data, analysis results

[2340] Specific operation: The server uses a speech recognition engine to convert speech to text and analyzes the text data using a natural language processing model.

[2341] Step 9:

[2342] Based on the results of the analysis using the natural language processing model, appropriate feedback (e.g., "Avoid talking about your previous trip") is generated and sent to the device. The device notifies the user of the feedback and displays it appropriately within the app.

[2343] Input: Analysis results

[2344] Output: Feedback results

[2345] Specific operation: The server generates appropriate feedback and sends it to the device, which notifies the user of the feedback and displays it.

[2346] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2347] This invention is a system that combines "behavior prediction," which efficiently predicts a user's daily behavior and helps them achieve their plans and goals; "statement prediction," which provides appropriate advice on the content and timing of statements; and "emotion recognition," which recognizes the user's emotions.

[2348] Embodiment of behavior prediction

[2349] User:

[2350] The user launches the smartphone app and inputs the plan and goals for the day. Internal information such as mood and health status is also input. For example, the user can provide information such as "I want to go to the gym in the morning and go shopping in the afternoon" or "I feel normal and my health is good."

[2351] Device:

[2352] The system stores the schedule, goals, and internal information entered by the user in a local database. It then uses an external API to obtain external information such as weather and traffic conditions. For example, it collects information such as "sunny" and "no traffic jam."

[2353] server:

[2354] All data sent from the device (schedules, goals, internal information, external information) is received and stored in a central database. In addition, an emotion engine extracts emotional information from the user's input. Based on the stored data and emotional information, an AI model predicts optimal behavioral patterns.

[2355] server:

[2356] Generate prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and send them to the device.

[2357] Device:

[2358] The received action prediction results are notified to the user and displayed visually within the app. The user can act based on the results and enter feedback into the app. For example, if the user acts as suggested and achieves their goal, they can enter the results as feedback.

[2359] Embodiment of speech prediction

[2360] User:

[2361] Start a smartphone app and record your everyday conversations. For example, record a conversation with your best friend and leave a note saying, "I want to talk about my next trip."

[2362] Device:

[2363] The recorded voice data and notes are stored in a local database and sent to a server, where an emotion engine extracts the user's emotional information from the recorded data and sends it together.

[2364] server:

[2365] The voice data is converted into text using a speech recognition engine. The converted text data is analyzed using a natural language processing model to identify redundant or inappropriate statements. For example, the system generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading."

[2366] server:

[2367] Based on the emotional information, appropriate feedback is generated, such as advice such as "mention your next trip, but avoid talking about your previous trip," and sent along with the message.

[2368] Device:

[2369] The app notifies users of the feedback it receives and visually displays it within the app, allowing users to adjust what they say based on that feedback and improve communication. For example, when talking with a friend, users can use the advice to choose topics and keep the conversation flowing smoothly.

[2370] Emotion Recognition Embodiment

[2371] User:

[2372] In addition to inputting moods and emotions, emotional information is provided to the app by recording facial expressions and changes in voice during conversations.

[2373] Device:

[2374] An emotion recognition engine is used to extract emotional information from the recorded voice data, and the emotional information is sent to the server.

[2375] server:

[2376] Emotional information analyzed by an emotion recognition engine and natural language processing model is used to generate behavioral patterns and speech feedback based on the user's emotions.

[2377] Specific examples

[2378] Specific examples of behavioral prediction

[2379] The user inputs their plan, such as "Go to the gym in the morning and shop at the market in the afternoon," along with their mood and health status as additional information. The device obtains information such as "sunny" from the weather API and "no traffic" from the traffic API, and sends this information to the server. Based on this information, the server uses an AI model and emotion engine to predict behavioral patterns, such as "optimal time to leave for the gym at 9:00 AM and the market at 2:00 PM," and sends the results to the device. The device then notifies the user of the results, and the user acts accordingly.

[2380] Specific examples of speech prediction

[2381] A user records a conversation with a close friend and leaves a note saying, "I want to talk about my next trip." The emotion engine extracts the user's emotions, such as excitement and anticipation, from the voice. The device sends the voice data, note, and emotion information to the server. The server converts the voice data into text and uses a natural language processing model to identify parts that are "too long" or "potentially misleading." Based on this, feedback is generated and sent to the device, such as "mention your next trip, but avoid talking about your previous trip." The device notifies the user of the feedback and displays it visually within the app.

[2382] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[2383] The processing flow will be explained below.

[2384] Program processing of behavior prediction

[2385] Step 1:

[2386] User: Launches the smartphone app and enters the plan for the day (e.g., "Gym in the morning, shopping in the afternoon") and goal (e.g., "Walk 10,000 steps").

[2387] Step 2:

[2388] User: Enters internal information into the app, such as mood (e.g., "normal") or health status (e.g., "good")

[2389] Step 3:

[2390] Device: Stores user-entered schedules, goals, and internal information in a local database.

[2391] Step 4:

[2392] Terminal: Obtain external information such as weather conditions (e.g., "sunny" using a weather API) and traffic conditions (e.g., "no traffic jams" using a traffic API).

[2393] Step 5:

[2394] Terminal: Send all collected data, including internal and external information, to the server.

[2395] Step 6:

[2396] Server: Receives all data sent from the devices (schedules, goals, internal information, external information) and stores them in a central database.

[2397] Step 7:

[2398] Server: Extracts user emotional information using the emotion engine based on the stored data.

[2399] Step 8:

[2400] Server: Based on the stored data and emotional information, the AI ​​model predicts optimal behavioral patterns.

[2401] Step 9:

[2402] Server: Generates prediction results (e.g., "The best time to leave for the gym is 9:00 AM and the best time to leave for the market is 2:00 PM") and sends them to the device.

[2403] Step 10:

[2404] Device: The received behavior prediction results are notified to the user and displayed visually within the app.

[2405] Step 11:

[2406] User: Follows the suggested behavioral patterns and achieves the goal.

[2407] Step 12:

[2408] User: Enters behavioral feedback into the app (e.g., "I went to the gym as planned and finished my shopping").

[2409] ---

[2410] Program processing of speech prediction

[2411] Step 1:

[2412] User: Launches the smartphone app and records everyday conversations (e.g., "Talking about traveling with my best friend").

[2413] Step 2:

[2414] Users: Leave notes during conversations about things that concern them or that need attention.

[2415] Step 3:

[2416] Device: Stores recorded audio data and notes in a local database.

[2417] Step 4:

[2418] Terminal: Sends voice data and memo information to the server.

[2419] Step 5:

[2420] Server: Converts voice data received from the device into text using a speech recognition engine.

[2421] Step 6:

[2422] Server: Analyzes text data using natural language processing (NLP) models to identify redundant or inappropriate statements.

[2423] Step 7:

[2424] Server: Extracts user emotion information from voice and text data using an emotion engine.

[2425] Step 8:

[2426] Server: Generates appropriate feedback based on sentiment information and text analysis results (e.g., "Mention your next trip, but avoid talking about your previous trip").

[2427] Step 9:

[2428] Server: Sends the generated feedback to the device.

[2429] Step 10:

[2430] On the device: Notify the user of the feedback received and display it visually within the app.

[2431] Step 11:

[2432] Users: Use suggested feedback to adjust conversations and improve communication.

[2433] This allows users to effectively achieve their plans and goals, as well as communicate smoothly, while also taking emotions into consideration.

[2434] Example 2

[2435] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2436] Conventional behavior prediction systems only predict behavior based on a user's schedule and goals, and have the problem of being unable to provide appropriate feedback or advice that takes emotional information into account. Furthermore, speech prediction systems are limited to identifying redundant or inappropriate speech, making it difficult to provide advice that appropriately reflects the user's emotions. This has led to the issue of a lack of effective support for user behavior and speech.

[2437] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2438] In this invention, the server includes: means for a user to input internal information such as the user's schedule and goals for the day, mood, and health status; means for collecting the user's internal information and external information such as weather conditions and traffic conditions; means for predicting the user's behavior based on the schedule, goals, internal information, and external information and generating an optimal behavior pattern based on emotional information; means for notifying the user of the predicted behavior pattern and collecting user feedback; means for recording the user's daily conversations and collecting voice data; means for converting the voice data into text; means for identifying redundant or inappropriate statements from the text data and advising the user on appropriate statements based on emotional information; and means for generating a pre-alert for statements and advice on appropriate statements based on the identified statements and notifying the user. This makes it possible to optimize the user's behavior and statements while taking emotions into consideration, thereby providing effective support.

[2439] "User" refers to an individual or corporation that uses the system, inputs their own schedules, goals, and internal information, and receives feedback from the system.

[2440] "Internal information" is information about a user's internal state, such as their mood, health, or emotions.

[2441] "External information" refers to information about the user's external environment, such as weather conditions and traffic conditions, and is obtained using an external API.

[2442] "Emotion information" is information about the user's emotional state, and is extracted by an emotion engine or the like.

[2443] An "emotion engine" is a software or hardware mechanism for extracting emotional information from user input or voice data.

[2444] A "behavioral pattern" is a schedule or action plan for the user to act optimally, and is the result predicted by the system.

[2445] "Feedback" is information that users input into the system based on the results of their actual actions, and the system uses this information to make its next prediction more accurate.

[2446] "Speech prediction" is a system function that analyzes the user's everyday conversations and advises them on appropriate content to say.

[2447] A "speech recognition engine" is a software or hardware mechanism for converting a user's voice data into text data.

[2448] A "natural language processing model" is a technology for analyzing text data and identifying redundant or inappropriate statements.

[2449] A "pre-alert" is a warning message that the system displays before a user makes an identified inappropriate comment.

[2450] "Advice on appropriate speech" refers to advice on speech content provided by the system to help users communicate more smoothly.

[2451] This invention is a system that combines "behavior prediction" that efficiently predicts a user's daily behavior and helps them achieve their schedules and goals, "statement prediction" that provides appropriate advice on the content and timing of statements, and "emotion recognition" that recognizes the user's emotions. A specific embodiment of this system will be described below.

[2452] Embodiment of behavior prediction

[2453] User:

[2454] Users launch the smartphone app and enter detailed information about the day's plans, goals, mood, health status, etc. For example, they might enter, "I want to go to the gym in the morning and do some shopping in the afternoon," or "I'm feeling normal, and my health is good."

[2455] Device:

[2456] The device stores the user's schedule, goals, mood, and health status in a local database. It then uses external APIs to obtain external information such as weather conditions (e.g., OpenWeatherMap API) and traffic conditions (e.g., Google Maps Traffic API). For example, it collects information such as "sunny" and "no traffic jam" and sends it to the server.

[2457] server:

[2458] The server receives all data (schedules, goals, internal information, external information) sent from the device and stores it in a central database. At the same time, it uses an emotion engine to extract emotional information from the user's input. Based on the stored data and emotional information, the AI ​​model predicts the optimal behavioral pattern. For example, it generates a result such as "It is best to leave for the gym at 9:00 AM and the market at 2:00 PM" and sends it to the device.

[2459] Device:

[2460] The device then notifies the user of the predicted behavior and displays it visually within the app. The user can then act based on the results and provide feedback to the app. For example, if the user acts as suggested and achieves their goal, they can provide feedback.

[2461] Embodiment of speech prediction

[2462] User:

[2463] Users can launch the smartphone app and record everyday conversations, such as a conversation with a close friend, and leave a note saying, "I want to talk about my next trip."

[2464] Device:

[2465] The device stores the recorded voice data and notes in a local database, and uses an emotion engine to extract the user's emotional information from the recorded data and send it together with the data to the server.

[2466] server:

[2467] The server converts the voice data into text using a speech recognition engine (e.g., Google Speech-to-Text API), analyzes the converted text data using a natural language processing model, and identifies redundant or inappropriate statements. For example, it generates analysis results such as "This part is redundant and difficult to understand" or "This part may be misleading." It then generates appropriate feedback based on the emotional information. For example, it generates advice such as "Talk about your next trip, but avoid talking about your previous trip," and sends it to the device.

[2468] Device:

[2469] The device will notify the user of the received feedback and display it visually within the app, allowing the user to adjust what they say based on this feedback and improve communication. For example, when talking with a friend, the user can use the advice to choose topics and keep the conversation flowing smoothly.

[2470] Emotion Recognition Embodiment

[2471] User:

[2472] Users provide emotional information to the app by inputting their mood and emotions and recording changes in facial expressions and voice during conversations.

[2473] Device:

[2474] The device uses an emotion recognition engine to extract emotional information from the recorded voice data and transmits the emotional information to a server.

[2475] server:

[2476] The server uses the emotional information analyzed by the emotion recognition engine and natural language processing model to generate behavioral patterns and speech feedback based on the user's emotions.

[2477] Specific examples

[2478] Specific examples of behavioral prediction

[2479] The user inputs...

Claims

1. A way for users to input their plans and goals for the day, means of collecting internal and external information about users; means for predicting user behavior based on said schedule, goals, internal information, and external information; means for presenting the predicted behavioral pattern to a user.

2. A means for recording users' daily conversations and collecting voice data; means for converting the voice data into text; means for identifying redundant or inappropriate comments from the text data; The system includes a means for generating a pre-alert of a speech and an appropriate speech advice based on the identified speech and notifying the user of the pre-alert.

3. The system according to claim 1, wherein the means for predicting a user's behavioral pattern calculates an optimal behavioral pattern based on internal information such as the user's mood and health condition, as well as external information such as weather conditions and traffic conditions.

4. 3. The system of claim 2, wherein the means for converting the recorded data of the user's daily conversation into text uses a voice recognition engine.

5. 3. The system according to claim 2, wherein the means for identifying redundant or inappropriate statements uses a natural language processing model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A