System

A system that collects and analyzes parent-child conversation and facial expression data to provide real-time learning advice addresses emotional conflicts, enhancing learning outcomes and relationships by offering immediate guidance based on emotional and psychological states.

JP2026030573APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133557
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Emotional conflicts between parents and children during study sessions lead to a decline in learning outcomes and negatively impact parent-child relationships, necessitating a solution to provide effective learning support and improve communication.

Method used

A system that collects parent-child conversational voices and facial expression data in real-time, analyzes these data using a multimodal language model, determines psychological states and emotions, and generates appropriate learning advice to be displayed on a user interface, thereby reducing emotional conflicts and enhancing learning effectiveness.

Benefits of technology

The system effectively reduces emotional conflicts and improves parent-child relationships by providing real-time learning advice based on psychological and emotional analysis, ensuring effective learning sessions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030573000001_ABST
    Figure 2026030573000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting conversation voices of a parent and a child; means for collecting facial expression data of the parent and the child; means for analyzing the voice data and the facial expression data in real time; means for determining psychological states and emotions of the parent and the child based on an analysis result; means for generating learning advice appropriate for the parent and the child based on a determination result; and means for providing the generated advice to the parent and the child.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When parents and children study together, emotional conflicts tend to arise between them, which can lead to a decline in learning outcomes. This problem is common in many families. Furthermore, emotional conflicts between parents and children can have a negative impact on the parent-child relationship, so there is a need to find a solution to this problem. [Means for solving the problem]

[0005] The present invention is a system that includes means for collecting parent-child conversational voices, means for collecting parent-child facial expression data, means for analyzing the voice data and facial expression data in real time, means for judging the psychological state and emotions of the parent and child based on the analysis results, means for generating appropriate learning advice for the parent and child based on the judgment results, and means for providing the generated advice to the parent and child, thereby making it possible to reduce emotional conflict between parent and child, provide effective learning support, and build a good parent-child relationship.

[0006] "Means for collecting parent-child conversation audio" refers to a function or device that records the conversation audio between parent and child using a device such as a microphone and saves it as digital data.

[0007] The "means for collecting facial expression data of parents and children" refers to a function or device that uses a device such as a camera or sensor to capture changes in the facial expressions of parents and children in real time and record that data.

[0008] "Means for analyzing voice data and facial expression data in real time" refers to technologies and algorithms that simultaneously process collected voice data and facial expression data and instantly analyze their content and characteristics.

[0009] "Means for determining the psychological state and emotions of parents and children" refers to technologies and algorithms that evaluate the psychological state and emotions of parents and children based on analyzed voice data and facial expression data, and display the results in numerical or text form.

[0010] The "means for generating appropriate learning advice" refers to functions or algorithms for proposing optimal learning support and guidance to parents and children based on the determined psychological state and emotions.

[0011] "Means for providing generated advice to parents and children" refers to output means or display devices for communicating the generated learning advice to parents and children, and in many cases, this involves presenting information via a user interface. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] Overall system configuration

[0034] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[0035] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[0036] 3. The server uses a large-scale multimodal language model to analyze the received data.

[0037] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[0038] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[0039] Specific Embodiments of the System

[0040] The user launches the app

[0041] The user launches the app on their device and the app's initial screen is displayed.

[0042] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[0043] Data collection

[0044] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[0045] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0046] Sending data

[0047] The device transmits the collected voice data and facial expression data to the server in real time.

[0048] For example, this is done by transferring a data packet to the server every 30 seconds.

[0049] Data analysis

[0050] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0051] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0052] Generating Advice

[0053] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0054] This advice will help users create a viable and effective learning environment on the fly.

[0055] Sending Advice

[0056] The server transmits the generated advice to the terminal, and the user receives it.

[0057] For example, the server sends an advice message to the terminal, which receives and prepares it.

[0058] Displaying Advice

[0059] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0060] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[0061] User executes advice

[0062] The user checks the displayed advice and acts accordingly.

[0063] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0064] Specific examples

[0065] Scenario: You're doing math homework together.

[0066] 1. A user launches the app and begins a math homework session for their child.

[0067] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0068] 3. The device sends the recorded data to the server every 30 seconds.

[0069] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0070] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0071] 6. The server sends this advice to the device, which displays it on the screen.

[0072] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0073] This system makes parent-child study sessions more effective, avoids emotional conflicts, and is expected to improve parent-child relationships.

[0074] The processing flow will be explained below.

[0075] Step 1:

[0076] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[0077] Step 2:

[0078] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0079] Step 3:

[0080] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[0081] Step 4:

[0082] The server inputs the received voice and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0083] Step 5:

[0084] Based on the analysis results, the server generates optimal advice, such as a message like, "Parents, please speak softer and explain with concrete examples." This advice helps users create an effective learning environment that they can implement on the spot.

[0085] Step 6:

[0086] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0087] Step 7:

[0088] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[0089] Step 8:

[0090] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[0091] Example 1

[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0093] In conventional parent-child study sessions, parental guidance often leads to emotional conflict, making it difficult to provide effective learning support. In particular, continuing guidance without properly understanding the psychological and emotional states of parents and children can lead to a decline in the child's motivation to learn. The present invention aims to provide a system that analyzes the psychological and emotional states of parents and children in real time and provides appropriate learning advice based on the results, so that parents and children can study more effectively and while avoiding emotional conflict.

[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0095] In this invention, the server includes a device for collecting parent-child conversation voices, a device for collecting parent-child facial expression data, a device for analyzing the voice data and facial expression data in real time, a device for generating appropriate learning advice for the parent and child based on the judgment results, a device for providing the generated advice to the parent and child, a device for using a multimodal large-scale language model to analyze the psychological states and emotions of the parent and child, a device for displaying the generated advice on a user interface of a terminal, a device for collecting data in real time, and a device for transmitting the collected data to the server in real time, thereby enabling the parent-child learning session to be conducted effectively and avoiding emotional conflicts between the parent and child.

[0096] The "device for collecting parent-child conversation voices" is a device for collecting voice data of the speech and conversation between parents and children.

[0097] The "device for collecting facial expression data of parents and children" is a device for collecting facial expressions of parents and children in the form of images or videos and storing them as digital data.

[0098] The "device for analyzing voice data and facial expression data in real time" is a device that instantly analyzes collected voice data and facial expression data to analyze the psychological state and emotions of parents and children.

[0099] The "device that generates appropriate learning advice for parents and children based on the results of its judgment" is a device that generates specific instructions and advice for parents and children to effectively advance their studies based on the results of analyzing voice data and facial expression data.

[0100] The "device for providing generated advice to parent and child" is a device for presenting generated study advice so that the parent and child can confirm it.

[0101] The "device that uses a multimodal large-scale language model to analyze the psychological state and emotions of parents and children" is a device that receives voice data and facial expression data as input and analyzes them using advanced algorithms and machine learning models.

[0102] The "device that displays the generated advice on the user interface of the terminal" is a device that displays the generated advice on a terminal such as a smartphone or tablet.

[0103] The "real-time data collection device" is a device that instantly collects parent-child conversation voice and facial expression data.

[0104] "Devices that transmit collected data to a server in real time" are devices that instantly transfer collected voice data and facial expression data to a server.

[0105] The present invention is a system for conducting effective parent-child study sessions. This system collects and analyzes parent-child conversation voice and facial expression data in real time, and provides appropriate study advice.

[0106] Hardware and software used

[0107] This system uses the following hardware and software:

[0108] Device: A portable electronic device such as a smartphone or tablet. This device is equipped with a camera and microphone and collects voice and facial expression data from parent-child conversations.

[0109] Server: A high-performance computing system that hosts large-scale language models and performs analysis of speech and facial expression data.

[0110] Multimodal large-scale language model: For example, large-scale language models from Google Cloud AI, Microsoft Azure Cognitive Services, etc. are used. This model performs advanced analysis of voice and facial expression data to determine the psychological state and emotions of parents and children.

[0111] Network: An internet connection for data communication between your device and the server.

[0112] Data processing and calculation

[0113] 1. Data collection: When users (parents and children) launch the app on their devices and allow the use of the camera and microphone, the device collects conversational voice and facial expression data in real time. For example, it records the situation where a parent is explaining something to a child while they are solving a math problem.

[0114] 2. Data transmission: The collected data is sent from the device to the server. The data is batched and transmitted every 30 seconds. For example, data packets are sent to the server using an HTTP POST request.

[0115] 3. Data analysis: The server receives the voice and facial expression data and inputs it into a large-scale language model for analysis, determining the parent's tone of voice and the child's emotional state, such as whether they are confused or not.

[0116] 4. Advice generation: Based on the analysis results, the server generates effective learning advice, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0117] 5. Sending and displaying advice: The generated advice is sent from the server to the terminal and displayed on the terminal's user interface. The displayed advice can be immediately checked by the user.

[0118] Specific examples

[0119] Scenario: You're doing math homework together.

[0120] A user launches the app to begin a math homework session for their child.

[0121] The device activates the camera and microphone to record conversations and facial expressions between parent and child.

[0122] The device sends the recorded data to the server every 30 seconds.

[0123] The server analyzes the data and detects when the parent's tone is harsh and the child is in trouble.

[0124] The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0125] The server sends this advice to the terminal, which displays it on the screen.

[0126] The user (parent) follows the advice and gently resumes the explanation, using concrete examples.

[0127] Prompt Sentence Examples

[0128] "What advice would you give to a child who is confused when their parent explains things to them in a harsh tone?"

[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0130] System program processing flow

[0131] Step 1: Launching the app and initial settings

[0132] 1. The user launches the app on their device.

[0133] Enter: Tap the app icon on your device.

[0134] Output: The initial screen of the app is displayed.

[0135] Specific behavior: You tap the app icon and the application loads into memory.

[0136] 2. The device will ask for permission to use the camera and microphone.

[0137] Input: The permission dialog appears on the initial screen of the app.

[0138] Output: User's choice (allow / deny) on the permission dialog.

[0139] What happens: Selecting "Allow" in the dialog grants the app camera and microphone permissions.

[0140] Step 2: Data collection

[0141] 3. The user allows use of the camera and microphone.

[0142] Input: Select "Allow" in the permission dialog.

[0143] Output: Camera and microphone are activated.

[0144] What happens: Permission is granted and the camera and microphone are automatically turned on.

[0145] 4. The device collects the parent-child conversation voice and facial expression data.

[0146] Input: A parent-child study session begins with the camera and microphone activated.

[0147] Output: Real-time collected voice and facial expression data.

[0148] Specific operation: Audio data is collected from the microphone, and facial expression data is captured from the camera.

[0149] Step 3: Send data

[0150] 5. The device sends the collected data to the server every 30 seconds.

[0151] Input: A batch of speech and facial expression data collected in real time.

[0152] Output: The data sent to the server in the HTTP POST request.

[0153] What it does: Data is batched at regular intervals and sent over the internet to a server.

[0154] 6. The server receives the voice data and facial expression data.

[0155] Input: Voice and facial expression data sent from the device.

[0156] Output: Acknowledgement message.

[0157] Specific operation: The server saves the received data and sends a confirmation message to the device confirming successful reception.

[0158] Step 4: Data analysis

[0159] 7. The server formats the received voice and facial expression data for analysis.

[0160] Input: The raw data received.

[0161] Output: Data formatted in a way that can be fed into a large-scale language model.

[0162] Specific operations: Converts voice data into text and facial expression data into a format for image analysis.

[0163] 8. The server inputs the data into a large-scale language model.

[0164] Input: Formatted speech and facial expression data.

[0165] Output: Analysis results (mental state and emotions of parent and child).

[0166] What it does: Query the model with data to get a predicted state of mind or emotion.

[0167] Step 5: Advice Generation

[0168] 9. The server generates appropriate advice based on the analysis results.

[0169] Input: Analysis results (mental state and emotions of parent and child).

[0170] Output: A study advice message.

[0171] Specific behavior: Creates a newly generated advice message.

[0172] 10. The server sends the generated advice to the device.

[0173] Input: Study advice message.

[0174] Output: Advice sent to the terminal.

[0175] Specific operation: Sends an advice message to the device via an HTTP POST request.

[0176] Step 6: Display Advice

[0177] 11. The device acknowledges receipt of the advice and displays it on the user interface.

[0178] Input: The advice message sent by the server.

[0179] Output: Advice displayed in the user interface.

[0180] Specific operation: Display the received advice message on the screen.

[0181] Step 7: User Actions

[0182] 12. The user checks the advice displayed on the device and corrects their behavior.

[0183] Input: The advice message displayed on the terminal.

[0184] Output: Modified parent behavior.

[0185] Specific actions: Based on the advice, take action such as explaining in a gentle tone with concrete examples.

[0186] (Application example 1)

[0187] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0188] While existing parent-child learning support systems have the means to collect and analyze parent-child conversation voice and facial expression data, they face the problem of difficulty in providing appropriate learning advice in real time. Furthermore, there is a need for systems that can accurately analyze the psychological state and emotions of parents and children and provide appropriate feedback immediately to facilitate communication between parents and children and improve learning effectiveness.

[0189] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0190] In this invention, the server includes means for periodically transmitting data collected by the terminal, means for generating advice in real time based on the analysis results of the server and transmitting the advice to the terminal, and means for displaying the advice on a user interface, thereby making it possible to provide appropriate learning advice in real time based on parent-child conversation voice and facial expression data.

[0191] A "means for collecting parent-child conversation audio" is a device or software used to record a conversation between a parent and a child.

[0192] "Means for collecting parent and child facial expression data" refers to cameras and sensors for capturing the facial expressions of parents and children.

[0193] The "means for analyzing voice data and facial expression data in real time" refers to software or hardware that can instantly analyze collected voice data and facial expression data.

[0194] The "means for determining the psychological state and emotions of the parent and child based on the analysis results" is an algorithm for evaluating the psychological state and emotions of the parent and child using the analysis results of the voice data and facial expression data.

[0195] The "means for generating appropriate learning advice for parents and children based on the judgment results" is software that creates optimal learning advice based on the psychological state and emotions of parents and children.

[0196] The "means for providing generated advice to parent and child" is a device or interface for communicating the generated advice to the parent and child.

[0197] "Means for periodically transmitting data collected by the terminal to a server" refers to a function or protocol for periodically transferring voice and facial expression data to a central server.

[0198] The "means for generating advice in real time based on the analysis results at the server and transmitting the advice to the terminal" is a function in which the server analyzes data, generates advice in real time, and transmits the advice to the terminal.

[0199] The "means for displaying advice on a user interface" refers to a screen or application that displays the generated advice so that the user can check it.

[0200] This invention relates to a parent-child learning support system that collects and analyzes parent-child conversation voice and facial expression data to provide appropriate learning advice in real time. This system is mainly composed of a terminal and a server.

[0201] Overall system configuration

[0202] 1. Users (parents and children) access the application using a device (smartphone or tablet).

[0203] 2. The device collects conversational audio and facial expression data between parent and child through a camera and microphone, and transmits the data to a server in real time.

[0204] 3. The server uses a multimodal large-scale language model (e.g., pipeline from the Transformers library) to analyze the collected data, generates appropriate advice in real time based on the analysis results, and sends it to the device.

[0205] 4. The device displays the advice obtained from the server on the user interface, enabling the user to receive appropriate learning support.

[0206] Specific Embodiments of the System

[0207] The user launches the app

[0208] When a user launches the app on their device, the app's initial screen is displayed.

[0209] The device will ask for permission to use the camera and microphone to ensure the parent-child learning session is ready.

[0210] Data collection

[0211] The device activates the camera and microphone to collect voice and facial expression data of parent-child conversations in real time.

[0212] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0213] Sending data

[0214] The terminal periodically (for example, every 30 seconds) transmits the collected voice data and facial expression data to the server.

[0215] Data analysis

[0216] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0217] For example, it detects that the parent's tone of voice has become harsher and the child is confused.

[0218] Generating Advice

[0219] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0220] Sending Advice

[0221] The server transmits the generated advice to the terminal, and the user receives it.

[0222] Displaying Advice

[0223] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0224] User follows advice

[0225] The user can then review the advice provided and act accordingly.

[0226] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0227] Specific examples

[0228] Scenario: You're doing math homework together.

[0229] 1. A user launches the app and begins a math homework session for their child.

[0230] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0231] 3. The device sends the recorded data to the server every 30 seconds.

[0232] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0233] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0234] 6. The server sends this advice to the device, which displays it on the screen.

[0235] 7. The user (parent) follows the advice and gently resumes the explanation with examples.

[0236] Prompt Sentence Examples

[0237] "Currently, a parent and child are doing math homework together, but the parent's tone is harsh, which is confusing the child. Please provide appropriate advice to parents."

[0238] This system will make parent-child learning sessions more effective, avoid emotional conflicts, and is expected to improve parent-child relationships.

[0239] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0240] Step 1:

[0241] The user launches the app

[0242] When a user launches the app on their device to begin a learning session, the device requests permission to use the camera and microphone. The input is the user's operation, and the output is the completion of camera and microphone settings.

[0243] Step 2:

[0244] Data collection

[0245] The device activates a camera and microphone to collect parent-child conversation voice and facial expression data in real time. Specifically, the camera captures facial expression data and the microphone records the conversation voice. The input is the parent-child conversation and facial expression, and the output is the collected voice and facial expression data.

[0246] Step 3:

[0247] Sending data

[0248] The device periodically (for example, every 30 seconds) transmits the collected voice and facial expression data to the server. The input is the collected voice and facial expression data, which are then transmitted to the server as output.

[0249] Step 4:

[0250] Data analysis

[0251] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. Specifically, it uses an AI model to analyze changes in voice tone and facial expression. The input is the received voice data and facial expression data, and the output is the analysis results.

[0252] Step 5:

[0253] Generating Advice

[0254] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please speak softer and explain with concrete examples." The input is the analysis results, and the output is the generated advice.

[0255] Step 6:

[0256] Sending Advice

[0257] The server sends the generated advice to the terminal, and the user receives it. The input is the generated advice, and the output is the transmission to the terminal.

[0258] Step 7:

[0259] Displaying Advice

[0260] The terminal displays the received advice on the user interface so that the user can immediately check it. The input is the received advice, and the output is the screen on which the advice is displayed.

[0261] Step 8:

[0262] User follows advice

[0263] The user reviews the displayed advice and acts accordingly. For example, a parent might explain the advice to their child again in a slow, gentle tone, using specific examples. The input is the displayed advice, and the output is the user's action.

[0264] These processing steps enable the system to provide appropriate learning advice to parents and children in real time and support effective learning.

[0265] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0266] Overall system configuration

[0267] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[0268] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[0269] 3. The server analyzes the received data using a multimodal large-scale language model and emotion engine.

[0270] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[0271] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[0272] Specific Embodiments of the System

[0273] The user launches the app

[0274] The user launches the app on their device and the app's initial screen is displayed.

[0275] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[0276] Data collection

[0277] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[0278] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0279] Sending data

[0280] The device transmits the collected voice data and facial expression data to the server in real time.

[0281] For example, this is done by transferring a data packet to the server every 30 seconds.

[0282] Data analysis

[0283] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0284] The server also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results supplementally.

[0285] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0286] Generating Advice

[0287] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0288] The analysis results of the emotion engine are also taken into consideration, providing more accurate, situation-specific advice.

[0289] This advice will help users create a viable and effective learning environment on the fly.

[0290] Sending Advice

[0291] The server transmits the generated advice to the terminal, and the user receives it.

[0292] For example, the server sends an advice message to the terminal, which receives and prepares it.

[0293] Displaying Advice

[0294] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0295] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[0296] User executes advice

[0297] The user checks the displayed advice and acts accordingly.

[0298] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0299] Specific examples

[0300] Scenario: You're doing math homework together.

[0301] 1. A user launches the app and begins a math homework session for their child.

[0302] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0303] 3. The device sends the recorded data to the server every 30 seconds.

[0304] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0305] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, try softening your tone of voice and explaining with examples."

[0306] 6. The server sends this advice to the device, which displays it on the screen.

[0307] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0308] This system will make parent-child study sessions more effective, avoid emotional conflicts, and improve parent-child relationships. The combination of an emotion engine will further improve the accuracy of analysis and the quality of advice.

[0309] The processing flow will be explained below.

[0310] Step 1:

[0311] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[0312] Step 2:

[0313] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0314] Step 3:

[0315] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[0316] Step 4:

[0317] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. It also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results as a supplement. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0318] Step 5:

[0319] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please soften your tone of voice and explain with specific examples." Since the analysis results of the emotion engine are also taken into consideration, more accurate advice tailored to the situation is provided.

[0320] Step 6:

[0321] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0322] Step 7:

[0323] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[0324] Step 8:

[0325] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[0326] Example 2

[0327] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0328] Until now, there has been no system that can analyze psychological and emotional issues that arise when parents and children study together in real time and provide effective advice. As a result, not only can parent-child study sessions become inefficient, but communication between parents and children can also deteriorate. To solve these problems, a system is needed that can analyze parent-child conversation voice and facial expression data in real time and provide appropriate advice.

[0329] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0330] In this invention, the server includes means for collecting parent-child conversation voices, means for collecting parent-child facial expression data, means for analyzing the voice data and facial expression data in real time, means for multimodally analyzing the voice data and facial expression data using a generative AI model, means for determining the psychological state and emotions of the parent and child based on the analysis results, means for generating appropriate study advice for the parent and child based on the determination results, means for providing the generated advice to the parent and child, means for obtaining permission to use the device's camera and microphone, means for preprocessing the voice data and facial expression data, and means for transmitting data to the server via a network. This makes parent-child study sessions more effective, avoiding emotional conflicts and improving the parent-child relationship.

[0331] "Means for collecting parent-child conversation audio" is a function for recording parent-child conversations and acquiring that data.

[0332] "Means for collecting facial expression data of parents and children" refers to a function that uses cameras and sensors to record the facial expressions of parents and children and acquire that data.

[0333] "Means for analyzing voice data and facial expression data in real time" refers to a function for instantly processing collected voice and facial expression data and analyzing its contents.

[0334] "Means for multimodal analysis of voice data and facial expression data using a generative AI model" is a function that uses a generative artificial intelligence model to combine and analyze voice data and facial expression data in an integrated manner.

[0335] "Means for determining the psychological state and emotions of parents and children based on the analysis results" is a function that infers and determines the psychological state and emotions of parents and children based on the analyzed data.

[0336] The "means for generating appropriate learning advice for parents and children based on the judgment results" is a function for creating optimal advice for parents and children based on the judged psychological state and emotions.

[0337] The "means for providing generated advice to parent and child" is a function for distributing the generated advice so that parent and child can check it.

[0338] The "means for obtaining permission to use the camera and microphone of the terminal" is a function for requesting permission to access the camera and microphone of the terminal and permitting their use.

[0339] "Means for preprocessing voice data and facial expression data" refers to a function that converts and filters voice data and facial expression data into an appropriate format before analysis.

[0340] "Means for transmitting data to a server using a network" refers to a function for transferring collected data to a server via a network.

[0341] This invention is a system that effectively supports parent-child study sessions. The system collects parent-child conversational voice and facial expression data, analyzes them using a generative AI model, and provides appropriate study advice in real time. Specific embodiments are described below.

[0342] Hardware and software configuration

[0343] Hardware:

[0344] Device (smartphone, tablet, etc.): A device used by parents and children that is equipped with a camera and microphone.

[0345] Server: A central system for analyzing data. Equipped with high-performance CPUs, GPUs, and sufficient memory.

[0346] software:

[0347] Generative AI model: Includes a large-scale language model and emotion engine, which analyzes voice and facial expression data.

[0348] Front-end application: An app for smartphones and tablets that provides the user interface.

[0349] Back-end system: Software that runs on a server for data analysis and advice generation.

[0350] Data processing and calculation

[0351] The device:

[0352] The device obtains permission to use the camera and microphone and begins collecting parent-child conversation audio and facial expression data.

[0353] The collected voice data is compressed in real time, and the facial expression data is pre-processed.

[0354] The device generates a data packet every 30 seconds and sends it to the server over the network.

[0355] The server:

[0356] The server first preprocesses the received voice and facial expression data, which includes noise removal and conversion to the required data format.

[0357] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[0358] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[0359] Based on the analysis results, advice is generated, such as a message like, "Parents, please speak softer and explain using concrete examples."

[0360] Convert the advice into an appropriate format and send it to the device.

[0361] Specific examples

[0362] Scenario: You're doing math homework together.

[0363] 1. The user launches the app and begins a math homework session for their child.

[0364] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0365] 3. The device sends the recorded data to the server every 30 seconds.

[0366] 4. The server analyzes the data and detects that the parent's tone of voice is becoming harsher and the child is in distress.

[0367] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, please soften your tone of voice and try explaining with examples."

[0368] 6. The server sends this advice to the terminal, where it is displayed on the screen.

[0369] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0370] Prompt Sentence Examples

[0371] An example of a prompt sentence input to the generative AI model is shown below.

[0372] "When a parent's tone becomes harsh during a parent-child study session, consider the child's confusion and generate advice to soften the tone of your voice."

[0373] This format makes parent-child study sessions more effective, and is expected to improve parent-child relationships while avoiding emotional conflicts.

[0374] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0375] Step 1:

[0376] The user launches the app.

[0377] Specific behavior:

[0378] A user taps to launch an app on their smartphone or tablet.

[0379] Input: User action (launching the app)

[0380] Output: The initial screen of the app is displayed.

[0381] The app's initial screen will appear, asking for permission to use the camera and microphone.

[0382] Step 2:

[0383] The device will turn on the camera and microphone and begin collecting data.

[0384] Specific behavior:

[0385] The device activates a camera and microphone to collect parent-child conversation audio and facial expression data.

[0386] Input: User permission (permission to use camera and microphone)

[0387] Output: Camera and microphone are activated and begin collecting data.

[0388] The audio data is compressed in real time, and the facial expression data is pre-processed using image processing algorithms.

[0389] Step 3:

[0390] The terminal transmits the collected data to the server.

[0391] Specific behavior:

[0392] The voice data and facial expression data collected by the terminal are sent to the server in batch processing.

[0393] Input: Collected voice and facial expression data

[0394] Output: Data packets sent to the server (e.g., data packets every 30 seconds)

[0395] The device checks the network connection and verifies that the data was sent successfully.

[0396] Step 4:

[0397] The server analyzes the data.

[0398] Specific behavior:

[0399] The server first preprocesses the received voice data and facial expression data.

[0400] Input: Voice data and facial expression data sent to the server

[0401] Output: Preprocessed data

[0402] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[0403] Input: Preprocessed speech and facial expression data

[0404] Output: Analysis results (mental state and emotional information)

[0405] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[0406] Step 5:

[0407] The server generates the advice.

[0408] Specific behavior:

[0409] The server generates optimal advice based on the analysis results.

[0410] Input: Analysis results

[0411] Output: The generated advice

[0412] A prompt sentence is input into the generative AI model to generate advice text.

[0413] Example: Create a message that reads, "Parents, please speak softly and provide examples."

[0414] Step 6:

[0415] The server sends the advice to the terminal.

[0416] Specific behavior:

[0417] Based on the advice generated by the server, a data packet is created in an appropriate format and sent to the terminal.

[0418] Input: Generated advice

[0419] Output: Data packets sent to the device (e.g., JSON formatted messages)

[0420] The server logs the status of the transmission and confirms that the transmission completed successfully.

[0421] Step 7:

[0422] The device displays the advice.

[0423] Specific behavior:

[0424] The advice received by the terminal is displayed on the user interface.

[0425] Input: Advice data sent from the server

[0426] Output: Advisory message displayed on the user interface

[0427] Example: The message on the screen reads, "Parents, please speak softly and use examples."

[0428] Step 8:

[0429] The user executes the advice.

[0430] Specific behavior:

[0431] The user reviews the displayed advice and acts accordingly.

[0432] Input: Advice displayed on screen

[0433] Output: User action (repeated in a gentle tone with examples)

[0434] The parent continues to explain using specific examples based on the advice.

[0435] (Application example 2)

[0436] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0437] In modern factories, collaboration between workers and robots has become an important element, but there is still a lack of systems that can provide appropriate feedback from robots based on the psychological state and emotions of workers. This has led to issues such as reduced work efficiency and increased worker stress.

[0438] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting the voice of the worker, means for collecting facial expression data of the worker, and means for analyzing the voice data and facial expression data in real time. This makes it possible to determine the psychological state and emotions of the worker and generate and provide appropriate work advice.

[0439] The "means for collecting worker voice" refers to a device or system that acquires voice information such as instructions and comments from workers at the factory work site in real time.

[0440] "Means for collecting facial expression data of workers" refers to devices or systems that capture the facial expressions of workers in real time using cameras or sensors and collect that data.

[0441] The "means for analyzing voice data and facial expression data in real time" refers to software and hardware that instantly processes acquired voice data and facial expression data and derives the analysis results.

[0442] The "means for determining the worker's psychological state and emotions based on the analysis results" refers to a device or algorithm that uses the results of voice and facial expression analysis to evaluate and make a judgment on the worker's psychological state and emotions.

[0443] The "means for generating appropriate work advice for the worker based on the judgment results" refers to a system or program that creates appropriate advice to improve work efficiency and reduce the worker's stress, depending on the judged psychological state and emotions.

[0444] "Means for providing generated advice to a worker" refers to a display, audio output, or other user interface for communicating generated advice to a worker.

[0445] Overall system configuration

[0446] This system is designed to support efficient communication between workers and robots at factory work sites. A specific embodiment is shown below.

[0447] The user launches the app

[0448] The user (worker) starts the system and the initial screen is displayed. The terminal requests permission to use the camera and microphone to confirm that the system is operational.

[0449] Data collection

[0450] The device collects voice and facial expression data from the worker in real time through a camera and microphone. For example, when a worker is giving instructions to a robot, the device records the voice instructions and the worker's facial expression.

[0451] Sending data

[0452] The device transmits the collected voice and facial expression data to the server in real time, for example, by transferring a data packet to the server every 30 seconds.

[0453] Data analysis

[0454] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time and supplements the analysis results. For example, the server may detect that the worker's voice tone has become harsher and that the worker's facial expression is tense.

[0455] Generating Advice

[0456] The server generates optimal advice based on the analysis results. For example, it creates a message such as "Slow down your work pace a little." Since the analysis results of the emotion engine are also taken into account, more accurate advice tailored to the situation is provided. This advice helps users create an effective work environment that can be implemented on the spot.

[0457] Sending Advice

[0458] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0459] Displaying Advice

[0460] The device displays the received advice on the user interface so that the user can immediately check it. For example, the device may display a message on the screen saying, "Try to slow down your work pace a little."

[0461] User executes advice

[0462] The user checks the displayed advice and acts accordingly. For example, a worker may slow down the pace of work based on the advice.

[0463] Specific examples and prompts

[0464] Imagine a worker in a factory is giving instructions to a robot. If this application is installed on the robot, the robot will capture the worker's voice and facial expressions in real time and send them to a server. Emotion analysis is performed on the server side, and it is determined that the worker is feeling stressed. In this case, the robot will display advice such as "Slow down your work pace a little," thereby reducing the worker's stress and improving work efficiency.

[0465] Example of input prompt for generative AI model

[0466] If the user's voice tone is judged to be "harsh": Parents, please soften your voice and provide specific examples.

[0467] "If the user is confused: Slowly review the steps. Provide specific instructions."

[0468] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0469] Step 1:

[0470] The user launches the app. The device requests permission to use the camera and microphone. This completes the initial setup and starts the system's data collection. The user has given permission as input, and the camera and microphone are available as output.

[0471] Step 2:

[0472] The terminal collects the worker's voice and facial expression data in real time through a camera and microphone. At this stage, the user (worker) gives instructions to the robot, and the voice instructions and facial expressions are captured as data. The input is the user's voice and facial expression, and the output is voice data and facial expression data.

[0473] Step 3:

[0474] The device transmits the collected voice and facial expression data to the server in real time. This transmission is performed periodically, for example, every 30 seconds, and the data is transferred to the server as a data packet. The input is the data collected in step 2, and the output is the data packet sent to the server.

[0475] Step 4:

[0476] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time. For example, it detects that the worker's voice tone has become harsher and their facial expression is tense. The input is the data sent to the server, and the output is the analysis result.

[0477] Step 5:

[0478] The server generates optimal advice based on the analysis results. For example, it creates advice such as "Slow down your work pace a little" depending on the psychological state and emotions obtained from the analysis. The input is the analysis result from step 4, and the output is the generated advice.

[0479] Step 6:

[0480] The server sends the generated advice to the terminal. The terminal receives this advice and prepares it. The input is the advice sent from the server, and the output is the advice message prepared by the terminal.

[0481] Step 7:

[0482] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Try to slow down a bit." The input is the advice message, and the output is the displayed advice.

[0483] Step 8:

[0484] The user checks the displayed advice and acts accordingly. For example, a worker slows down the pace of work based on the advice. The input is the displayed advice, and the output is the user's actual behavior.

[0485] Through this series of processing steps, the system analyzes the worker's psychological state and emotions in real time and provides appropriate feedback, thereby improving work efficiency.

[0486] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0487] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0488] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0489] [Second embodiment]

[0490] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0491] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0492] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0493] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0494] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0495] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0496] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0497] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0498] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0499] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0500] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0501] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0502] Overall system configuration

[0503] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[0504] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[0505] 3. The server uses a large-scale multimodal language model to analyze the received data.

[0506] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[0507] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[0508] Specific Embodiments of the System

[0509] The user launches the app

[0510] The user launches the app on their device and the app's initial screen is displayed.

[0511] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[0512] Data collection

[0513] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[0514] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0515] Sending data

[0516] The device transmits the collected voice data and facial expression data to the server in real time.

[0517] For example, this is done by transferring a data packet to the server every 30 seconds.

[0518] Data analysis

[0519] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0520] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0521] Generating Advice

[0522] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0523] This advice will help users create a viable and effective learning environment on the fly.

[0524] Sending Advice

[0525] The server transmits the generated advice to the terminal, and the user receives it.

[0526] For example, the server sends an advice message to the terminal, which receives and prepares it.

[0527] Displaying Advice

[0528] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0529] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[0530] User executes advice

[0531] The user checks the displayed advice and acts accordingly.

[0532] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0533] Specific examples

[0534] Scenario: You're doing math homework together.

[0535] 1. A user launches the app and begins a math homework session for their child.

[0536] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0537] 3. The device sends the recorded data to the server every 30 seconds.

[0538] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0539] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0540] 6. The server sends this advice to the device, which displays it on the screen.

[0541] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0542] This system makes parent-child study sessions more effective, avoids emotional conflicts, and is expected to improve parent-child relationships.

[0543] The processing flow will be explained below.

[0544] Step 1:

[0545] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[0546] Step 2:

[0547] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0548] Step 3:

[0549] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[0550] Step 4:

[0551] The server inputs the received voice and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0552] Step 5:

[0553] Based on the analysis results, the server generates optimal advice, such as a message like, "Parents, please speak softer and explain with concrete examples." This advice helps users create an effective learning environment that they can implement on the spot.

[0554] Step 6:

[0555] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0556] Step 7:

[0557] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[0558] Step 8:

[0559] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[0560] Example 1

[0561] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0562] In conventional parent-child study sessions, parental guidance often leads to emotional conflict, making it difficult to provide effective learning support. In particular, continuing guidance without properly understanding the psychological and emotional states of parents and children can lead to a decline in the child's motivation to learn. The present invention aims to provide a system that analyzes the psychological and emotional states of parents and children in real time and provides appropriate learning advice based on the results, so that parents and children can study more effectively and while avoiding emotional conflict.

[0563] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0564] In this invention, the server includes a device for collecting parent-child conversation voices, a device for collecting parent-child facial expression data, a device for analyzing the voice data and facial expression data in real time, a device for generating appropriate learning advice for the parent and child based on the judgment results, a device for providing the generated advice to the parent and child, a device for using a multimodal large-scale language model to analyze the psychological states and emotions of the parent and child, a device for displaying the generated advice on a user interface of a terminal, a device for collecting data in real time, and a device for transmitting the collected data to the server in real time, thereby enabling the parent-child learning session to be conducted effectively and avoiding emotional conflicts between the parent and child.

[0565] The "device for collecting parent-child conversation voices" is a device for collecting voice data of the speech and conversation between parents and children.

[0566] The "device for collecting facial expression data of parents and children" is a device for collecting facial expressions of parents and children in the form of images or videos and storing them as digital data.

[0567] The "device for analyzing voice data and facial expression data in real time" is a device that instantly analyzes collected voice data and facial expression data to analyze the psychological state and emotions of parents and children.

[0568] The "device that generates appropriate learning advice for parents and children based on the results of its judgment" is a device that generates specific instructions and advice for parents and children to effectively advance their studies based on the results of analyzing voice data and facial expression data.

[0569] The "device for providing generated advice to parent and child" is a device for presenting generated study advice so that the parent and child can confirm it.

[0570] The "device that uses a multimodal large-scale language model to analyze the psychological state and emotions of parents and children" is a device that receives voice data and facial expression data as input and analyzes them using advanced algorithms and machine learning models.

[0571] The "device that displays the generated advice on the user interface of the terminal" is a device that displays the generated advice on a terminal such as a smartphone or tablet.

[0572] The "real-time data collection device" is a device that instantly collects parent-child conversation voice and facial expression data.

[0573] "Devices that transmit collected data to a server in real time" are devices that instantly transfer collected voice data and facial expression data to a server.

[0574] The present invention is a system for conducting effective parent-child study sessions. This system collects and analyzes parent-child conversation voice and facial expression data in real time, and provides appropriate study advice.

[0575] Hardware and software used

[0576] This system uses the following hardware and software:

[0577] Device: A portable electronic device such as a smartphone or tablet. This device is equipped with a camera and microphone and collects voice and facial expression data from parent-child conversations.

[0578] Server: A high-performance computing system that hosts large-scale language models and performs analysis of speech and facial expression data.

[0579] Multimodal large-scale language model: For example, large-scale language models from Google Cloud AI, Microsoft Azure Cognitive Services, etc. are used. This model performs advanced analysis of voice and facial expression data to determine the psychological state and emotions of parents and children.

[0580] Network: An internet connection for data communication between your device and the server.

[0581] Data processing and calculation

[0582] 1. Data collection: When users (parents and children) launch the app on their devices and allow the use of the camera and microphone, the device collects conversational voice and facial expression data in real time. For example, it records the situation where a parent is explaining something to a child while they are solving a math problem.

[0583] 2. Data transmission: The collected data is sent from the device to the server. The data is batched and transmitted every 30 seconds. For example, data packets are sent to the server using an HTTP POST request.

[0584] 3. Data analysis: The server receives the voice and facial expression data and inputs it into a large-scale language model for analysis, determining the parent's tone of voice and the child's emotional state, such as whether they are confused or not.

[0585] 4. Advice generation: Based on the analysis results, the server generates effective learning advice, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0586] 5. Sending and displaying advice: The generated advice is sent from the server to the terminal and displayed on the terminal's user interface. The displayed advice can be immediately checked by the user.

[0587] Specific examples

[0588] Scenario: You're doing math homework together.

[0589] A user launches the app to begin a math homework session for their child.

[0590] The device activates the camera and microphone to record conversations and facial expressions between parent and child.

[0591] The device sends the recorded data to the server every 30 seconds.

[0592] The server analyzes the data and detects when the parent's tone is harsh and the child is in trouble.

[0593] The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0594] The server sends this advice to the terminal, which displays it on the screen.

[0595] The user (parent) follows the advice and gently resumes the explanation, using concrete examples.

[0596] Prompt Sentence Examples

[0597] "What advice would you give to a child who is confused when their parent explains things to them in a harsh tone?"

[0598] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0599] System program processing flow

[0600] Step 1: Launching the app and initial settings

[0601] 1. The user launches the app on their device.

[0602] Enter: Tap the app icon on your device.

[0603] Output: The initial screen of the app is displayed.

[0604] Specific behavior: You tap the app icon and the application loads into memory.

[0605] 2. The device will ask for permission to use the camera and microphone.

[0606] Input: The permission dialog appears on the initial screen of the app.

[0607] Output: User's choice (allow / deny) on the permission dialog.

[0608] What happens: Selecting "Allow" in the dialog grants the app camera and microphone permissions.

[0609] Step 2: Data collection

[0610] 3. The user allows use of the camera and microphone.

[0611] Input: Select "Allow" in the permission dialog.

[0612] Output: Camera and microphone are activated.

[0613] What happens: Permission is granted and the camera and microphone are automatically turned on.

[0614] 4. The device collects the parent-child conversation voice and facial expression data.

[0615] Input: A parent-child study session begins with the camera and microphone activated.

[0616] Output: Real-time collected voice and facial expression data.

[0617] Specific operation: Audio data is collected from the microphone, and facial expression data is captured from the camera.

[0618] Step 3: Send data

[0619] 5. The device sends the collected data to the server every 30 seconds.

[0620] Input: A batch of speech and facial expression data collected in real time.

[0621] Output: The data sent to the server in the HTTP POST request.

[0622] What it does: Data is batched at regular intervals and sent over the internet to a server.

[0623] 6. The server receives the voice data and facial expression data.

[0624] Input: Voice and facial expression data sent from the device.

[0625] Output: Acknowledgement message.

[0626] Specific operation: The server saves the received data and sends a confirmation message to the device confirming successful reception.

[0627] Step 4: Data analysis

[0628] 7. The server formats the received voice and facial expression data for analysis.

[0629] Input: The raw data received.

[0630] Output: Data formatted in a way that can be fed into a large-scale language model.

[0631] Specific operations: Converts voice data into text and facial expression data into a format for image analysis.

[0632] 8. The server inputs the data into a large-scale language model.

[0633] Input: Formatted speech and facial expression data.

[0634] Output: Analysis results (mental state and emotions of parent and child).

[0635] What it does: Query the model with data to get a predicted state of mind or emotion.

[0636] Step 5: Advice Generation

[0637] 9. The server generates appropriate advice based on the analysis results.

[0638] Input: Analysis results (mental state and emotions of parent and child).

[0639] Output: A study advice message.

[0640] Specific behavior: Creates a newly generated advice message.

[0641] 10. The server sends the generated advice to the device.

[0642] Input: Study advice message.

[0643] Output: Advice sent to the terminal.

[0644] Specific operation: Sends an advice message to the device via an HTTP POST request.

[0645] Step 6: Display Advice

[0646] 11. The device acknowledges receipt of the advice and displays it on the user interface.

[0647] Input: The advice message sent by the server.

[0648] Output: Advice displayed in the user interface.

[0649] Specific operation: Display the received advice message on the screen.

[0650] Step 7: User Actions

[0651] 12. The user checks the advice displayed on the device and corrects their behavior.

[0652] Input: The advice message displayed on the terminal.

[0653] Output: Modified parent behavior.

[0654] Specific actions: Based on the advice, take action such as explaining in a gentle tone with concrete examples.

[0655] (Application example 1)

[0656] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0657] While existing parent-child learning support systems have the means to collect and analyze parent-child conversation voice and facial expression data, they face the problem of difficulty in providing appropriate learning advice in real time. Furthermore, there is a need for systems that can accurately analyze the psychological state and emotions of parents and children and provide appropriate feedback immediately to facilitate communication between parents and children and improve learning effectiveness.

[0658] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0659] In this invention, the server includes means for periodically transmitting data collected by the terminal, means for generating advice in real time based on the analysis results of the server and transmitting the advice to the terminal, and means for displaying the advice on a user interface, thereby making it possible to provide appropriate learning advice in real time based on parent-child conversation voice and facial expression data.

[0660] A "means for collecting parent-child conversation audio" is a device or software used to record a conversation between a parent and a child.

[0661] "Means for collecting parent and child facial expression data" refers to cameras and sensors for capturing the facial expressions of parents and children.

[0662] The "means for analyzing voice data and facial expression data in real time" refers to software or hardware that can instantly analyze collected voice data and facial expression data.

[0663] The "means for determining the psychological state and emotions of the parent and child based on the analysis results" is an algorithm for evaluating the psychological state and emotions of the parent and child using the analysis results of the voice data and facial expression data.

[0664] The "means for generating appropriate learning advice for parents and children based on the judgment results" is software that creates optimal learning advice based on the psychological state and emotions of parents and children.

[0665] The "means for providing generated advice to parent and child" is a device or interface for communicating the generated advice to the parent and child.

[0666] "Means for periodically transmitting data collected by the terminal to a server" refers to a function or protocol for periodically transferring voice and facial expression data to a central server.

[0667] The "means for generating advice in real time based on the analysis results at the server and transmitting the advice to the terminal" is a function in which the server analyzes data, generates advice in real time, and transmits the advice to the terminal.

[0668] The "means for displaying advice on a user interface" refers to a screen or application that displays the generated advice so that the user can check it.

[0669] This invention relates to a parent-child learning support system that collects and analyzes parent-child conversation voice and facial expression data to provide appropriate learning advice in real time. This system is mainly composed of a terminal and a server.

[0670] Overall system configuration

[0671] 1. Users (parents and children) access the application using a device (smartphone or tablet).

[0672] 2. The device collects conversational audio and facial expression data between parent and child through a camera and microphone, and transmits the data to a server in real time.

[0673] 3. The server uses a multimodal large-scale language model (e.g., pipeline from the Transformers library) to analyze the collected data, generates appropriate advice in real time based on the analysis results, and sends it to the device.

[0674] 4. The device displays the advice obtained from the server on the user interface, enabling the user to receive appropriate learning support.

[0675] Specific Embodiments of the System

[0676] The user launches the app

[0677] When a user launches the app on their device, the app's initial screen is displayed.

[0678] The device will ask for permission to use the camera and microphone to ensure the parent-child learning session is ready.

[0679] Data collection

[0680] The device activates the camera and microphone to collect voice and facial expression data of parent-child conversations in real time.

[0681] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0682] Sending data

[0683] The terminal periodically (for example, every 30 seconds) transmits the collected voice data and facial expression data to the server.

[0684] Data analysis

[0685] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0686] For example, it detects that the parent's tone of voice has become harsher and the child is confused.

[0687] Generating Advice

[0688] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0689] Sending Advice

[0690] The server transmits the generated advice to the terminal, and the user receives it.

[0691] Displaying Advice

[0692] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0693] User follows advice

[0694] The user can then review the advice provided and act accordingly.

[0695] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0696] Specific examples

[0697] Scenario: You're doing math homework together.

[0698] 1. A user launches the app and begins a math homework session for their child.

[0699] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0700] 3. The device sends the recorded data to the server every 30 seconds.

[0701] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0702] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[0703] 6. The server sends this advice to the device, which displays it on the screen.

[0704] 7. The user (parent) follows the advice and gently resumes the explanation with examples.

[0705] Prompt Sentence Examples

[0706] "Currently, a parent and child are doing math homework together, but the parent's tone is harsh, which is confusing the child. Please provide appropriate advice to parents."

[0707] This system will make parent-child learning sessions more effective, avoid emotional conflicts, and is expected to improve parent-child relationships.

[0708] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0709] Step 1:

[0710] The user launches the app

[0711] When a user launches the app on their device to begin a learning session, the device requests permission to use the camera and microphone. The input is the user's operation, and the output is the completion of camera and microphone settings.

[0712] Step 2:

[0713] Data collection

[0714] The device activates a camera and microphone to collect parent-child conversation voice and facial expression data in real time. Specifically, the camera captures facial expression data and the microphone records the conversation voice. The input is the parent-child conversation and facial expression, and the output is the collected voice and facial expression data.

[0715] Step 3:

[0716] Sending data

[0717] The device periodically (for example, every 30 seconds) transmits the collected voice and facial expression data to the server. The input is the collected voice and facial expression data, which are then transmitted to the server as output.

[0718] Step 4:

[0719] Data analysis

[0720] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. Specifically, it uses an AI model to analyze changes in voice tone and facial expression. The input is the received voice data and facial expression data, and the output is the analysis results.

[0721] Step 5:

[0722] Generating Advice

[0723] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please speak softer and explain with concrete examples." The input is the analysis results, and the output is the generated advice.

[0724] Step 6:

[0725] Sending Advice

[0726] The server sends the generated advice to the terminal, and the user receives it. The input is the generated advice, and the output is the transmission to the terminal.

[0727] Step 7:

[0728] Displaying Advice

[0729] The terminal displays the received advice on the user interface so that the user can immediately check it. The input is the received advice, and the output is the screen on which the advice is displayed.

[0730] Step 8:

[0731] User follows advice

[0732] The user reviews the displayed advice and acts accordingly. For example, a parent might explain the advice to their child again in a slow, gentle tone, using specific examples. The input is the displayed advice, and the output is the user's action.

[0733] These processing steps enable the system to provide appropriate learning advice to parents and children in real time and support effective learning.

[0734] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0735] Overall system configuration

[0736] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[0737] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[0738] 3. The server analyzes the received data using a multimodal large-scale language model and emotion engine.

[0739] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[0740] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[0741] Specific Embodiments of the System

[0742] The user launches the app

[0743] The user launches the app on their device and the app's initial screen is displayed.

[0744] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[0745] Data collection

[0746] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[0747] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0748] Sending data

[0749] The device transmits the collected voice data and facial expression data to the server in real time.

[0750] For example, this is done by transferring a data packet to the server every 30 seconds.

[0751] Data analysis

[0752] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0753] The server also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results supplementally.

[0754] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0755] Generating Advice

[0756] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0757] The analysis results of the emotion engine are also taken into consideration, providing more accurate, situation-specific advice.

[0758] This advice will help users create a viable and effective learning environment on the fly.

[0759] Sending Advice

[0760] The server transmits the generated advice to the terminal, and the user receives it.

[0761] For example, the server sends an advice message to the terminal, which receives and prepares it.

[0762] Displaying Advice

[0763] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0764] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[0765] User executes advice

[0766] The user checks the displayed advice and acts accordingly.

[0767] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[0768] Specific examples

[0769] Scenario: You're doing math homework together.

[0770] 1. A user launches the app and begins a math homework session for their child.

[0771] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0772] 3. The device sends the recorded data to the server every 30 seconds.

[0773] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[0774] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, try softening your tone of voice and explaining with examples."

[0775] 6. The server sends this advice to the device, which displays it on the screen.

[0776] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0777] This system will make parent-child study sessions more effective, avoid emotional conflicts, and improve parent-child relationships. The combination of an emotion engine will further improve the accuracy of analysis and the quality of advice.

[0778] The processing flow will be explained below.

[0779] Step 1:

[0780] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[0781] Step 2:

[0782] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0783] Step 3:

[0784] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[0785] Step 4:

[0786] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. It also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results as a supplement. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0787] Step 5:

[0788] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please soften your tone of voice and explain with specific examples." Since the analysis results of the emotion engine are also taken into consideration, more accurate advice tailored to the situation is provided.

[0789] Step 6:

[0790] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0791] Step 7:

[0792] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[0793] Step 8:

[0794] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[0795] Example 2

[0796] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0797] Until now, there has been no system that can analyze psychological and emotional issues that arise when parents and children study together in real time and provide effective advice. As a result, not only can parent-child study sessions become inefficient, but communication between parents and children can also deteriorate. To solve these problems, a system is needed that can analyze parent-child conversation voice and facial expression data in real time and provide appropriate advice.

[0798] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0799] In this invention, the server includes means for collecting parent-child conversation voices, means for collecting parent-child facial expression data, means for analyzing the voice data and facial expression data in real time, means for multimodally analyzing the voice data and facial expression data using a generative AI model, means for determining the psychological state and emotions of the parent and child based on the analysis results, means for generating appropriate study advice for the parent and child based on the determination results, means for providing the generated advice to the parent and child, means for obtaining permission to use the device's camera and microphone, means for preprocessing the voice data and facial expression data, and means for transmitting data to the server via a network. This makes parent-child study sessions more effective, avoiding emotional conflicts and improving the parent-child relationship.

[0800] "Means for collecting parent-child conversation audio" is a function for recording parent-child conversations and acquiring that data.

[0801] "Means for collecting facial expression data of parents and children" refers to a function that uses cameras and sensors to record the facial expressions of parents and children and acquire that data.

[0802] "Means for analyzing voice data and facial expression data in real time" refers to a function for instantly processing collected voice and facial expression data and analyzing its contents.

[0803] "Means for multimodal analysis of voice data and facial expression data using a generative AI model" is a function that uses a generative artificial intelligence model to combine and analyze voice data and facial expression data in an integrated manner.

[0804] "Means for determining the psychological state and emotions of parents and children based on the analysis results" is a function that infers and determines the psychological state and emotions of parents and children based on the analyzed data.

[0805] The "means for generating appropriate learning advice for parents and children based on the judgment results" is a function for creating optimal advice for parents and children based on the judged psychological state and emotions.

[0806] The "means for providing generated advice to parent and child" is a function for distributing the generated advice so that parent and child can check it.

[0807] The "means for obtaining permission to use the camera and microphone of the terminal" is a function for requesting permission to access the camera and microphone of the terminal and permitting their use.

[0808] "Means for preprocessing voice data and facial expression data" refers to a function that converts and filters voice data and facial expression data into an appropriate format before analysis.

[0809] "Means for transmitting data to a server using a network" refers to a function for transferring collected data to a server via a network.

[0810] This invention is a system that effectively supports parent-child study sessions. The system collects parent-child conversational voice and facial expression data, analyzes them using a generative AI model, and provides appropriate study advice in real time. Specific embodiments are described below.

[0811] Hardware and software configuration

[0812] Hardware:

[0813] Device (smartphone, tablet, etc.): A device used by parents and children that is equipped with a camera and microphone.

[0814] Server: A central system for analyzing data. Equipped with high-performance CPUs, GPUs, and sufficient memory.

[0815] software:

[0816] Generative AI model: Includes a large-scale language model and emotion engine, which analyzes voice and facial expression data.

[0817] Front-end application: An app for smartphones and tablets that provides the user interface.

[0818] Back-end system: Software that runs on a server for data analysis and advice generation.

[0819] Data processing and calculation

[0820] The device:

[0821] The device obtains permission to use the camera and microphone and begins collecting parent-child conversation audio and facial expression data.

[0822] The collected voice data is compressed in real time, and the facial expression data is pre-processed.

[0823] The device generates a data packet every 30 seconds and sends it to the server over the network.

[0824] The server:

[0825] The server first preprocesses the received voice and facial expression data, which includes noise removal and conversion to the required data format.

[0826] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[0827] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[0828] Based on the analysis results, advice is generated, such as a message like, "Parents, please speak softer and explain using concrete examples."

[0829] Convert the advice into an appropriate format and send it to the device.

[0830] Specific examples

[0831] Scenario: You're doing math homework together.

[0832] 1. The user launches the app and begins a math homework session for their child.

[0833] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[0834] 3. The device sends the recorded data to the server every 30 seconds.

[0835] 4. The server analyzes the data and detects that the parent's tone of voice is becoming harsher and the child is in distress.

[0836] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, please soften your tone of voice and try explaining with examples."

[0837] 6. The server sends this advice to the terminal, where it is displayed on the screen.

[0838] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[0839] Prompt Sentence Examples

[0840] An example of a prompt sentence input to the generative AI model is shown below.

[0841] "When a parent's tone becomes harsh during a parent-child study session, consider the child's confusion and generate advice to soften the tone of your voice."

[0842] This format makes parent-child study sessions more effective, and is expected to improve parent-child relationships while avoiding emotional conflicts.

[0843] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0844] Step 1:

[0845] The user launches the app.

[0846] Specific behavior:

[0847] A user taps to launch an app on their smartphone or tablet.

[0848] Input: User action (launching the app)

[0849] Output: The initial screen of the app is displayed.

[0850] The app's initial screen will appear, asking for permission to use the camera and microphone.

[0851] Step 2:

[0852] The device will turn on the camera and microphone and begin collecting data.

[0853] Specific behavior:

[0854] The device activates a camera and microphone to collect parent-child conversation audio and facial expression data.

[0855] Input: User permission (permission to use camera and microphone)

[0856] Output: Camera and microphone are activated and begin collecting data.

[0857] The audio data is compressed in real time, and the facial expression data is pre-processed using image processing algorithms.

[0858] Step 3:

[0859] The terminal transmits the collected data to the server.

[0860] Specific behavior:

[0861] The voice data and facial expression data collected by the terminal are sent to the server in batch processing.

[0862] Input: Collected voice and facial expression data

[0863] Output: Data packets sent to the server (e.g., data packets every 30 seconds)

[0864] The device checks the network connection and verifies that the data was sent successfully.

[0865] Step 4:

[0866] The server analyzes the data.

[0867] Specific behavior:

[0868] The server first preprocesses the received voice data and facial expression data.

[0869] Input: Voice data and facial expression data sent to the server

[0870] Output: Preprocessed data

[0871] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[0872] Input: Preprocessed speech and facial expression data

[0873] Output: Analysis results (mental state and emotional information)

[0874] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[0875] Step 5:

[0876] The server generates the advice.

[0877] Specific behavior:

[0878] The server generates optimal advice based on the analysis results.

[0879] Input: Analysis results

[0880] Output: The generated advice

[0881] A prompt sentence is input into the generative AI model to generate advice text.

[0882] Example: Create a message that reads, "Parents, please speak softly and provide examples."

[0883] Step 6:

[0884] The server sends the advice to the terminal.

[0885] Specific behavior:

[0886] Based on the advice generated by the server, a data packet is created in an appropriate format and sent to the terminal.

[0887] Input: Generated advice

[0888] Output: Data packets sent to the device (e.g., JSON formatted messages)

[0889] The server logs the status of the transmission and confirms that the transmission completed successfully.

[0890] Step 7:

[0891] The device displays the advice.

[0892] Specific behavior:

[0893] The advice received by the terminal is displayed on the user interface.

[0894] Input: Advice data sent from the server

[0895] Output: Advisory message displayed on the user interface

[0896] Example: The message on the screen reads, "Parents, please speak softly and use examples."

[0897] Step 8:

[0898] The user executes the advice.

[0899] Specific behavior:

[0900] The user reviews the displayed advice and acts accordingly.

[0901] Input: Advice displayed on screen

[0902] Output: User action (repeated in a gentle tone with examples)

[0903] The parent continues to explain using specific examples based on the advice.

[0904] (Application example 2)

[0905] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0906] In modern factories, collaboration between workers and robots has become an important element, but there is still a lack of systems that can provide appropriate feedback from robots based on the psychological state and emotions of workers. This has led to issues such as reduced work efficiency and increased worker stress.

[0907] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting the voice of the worker, means for collecting facial expression data of the worker, and means for analyzing the voice data and facial expression data in real time. This makes it possible to determine the psychological state and emotions of the worker and generate and provide appropriate work advice.

[0908] The "means for collecting worker voice" refers to a device or system that acquires voice information such as instructions and comments from workers at the factory work site in real time.

[0909] "Means for collecting facial expression data of workers" refers to devices or systems that capture the facial expressions of workers in real time using cameras or sensors and collect that data.

[0910] The "means for analyzing voice data and facial expression data in real time" refers to software and hardware that instantly processes acquired voice data and facial expression data and derives the analysis results.

[0911] The "means for determining the worker's psychological state and emotions based on the analysis results" refers to a device or algorithm that uses the results of voice and facial expression analysis to evaluate and make a judgment on the worker's psychological state and emotions.

[0912] The "means for generating appropriate work advice for the worker based on the judgment results" refers to a system or program that creates appropriate advice to improve work efficiency and reduce the worker's stress, depending on the judged psychological state and emotions.

[0913] "Means for providing generated advice to a worker" refers to a display, audio output, or other user interface for communicating generated advice to a worker.

[0914] Overall system configuration

[0915] This system is designed to support efficient communication between workers and robots at factory work sites. A specific embodiment is shown below.

[0916] The user launches the app

[0917] The user (worker) starts the system and the initial screen is displayed. The terminal requests permission to use the camera and microphone to confirm that the system is operational.

[0918] Data collection

[0919] The device collects voice and facial expression data from the worker in real time through a camera and microphone. For example, when a worker is giving instructions to a robot, the device records the voice instructions and the worker's facial expression.

[0920] Sending data

[0921] The device transmits the collected voice and facial expression data to the server in real time, for example, by transferring a data packet to the server every 30 seconds.

[0922] Data analysis

[0923] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time and supplements the analysis results. For example, the server may detect that the worker's voice tone has become harsher and that the worker's facial expression is tense.

[0924] Generating Advice

[0925] The server generates optimal advice based on the analysis results. For example, it creates a message such as "Slow down your work pace a little." Since the analysis results of the emotion engine are also taken into account, more accurate advice tailored to the situation is provided. This advice helps users create an effective work environment that can be implemented on the spot.

[0926] Sending Advice

[0927] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[0928] Displaying Advice

[0929] The device displays the received advice on the user interface so that the user can immediately check it. For example, the device may display a message on the screen saying, "Try to slow down your work pace a little."

[0930] User executes advice

[0931] The user checks the displayed advice and acts accordingly. For example, a worker may slow down the pace of work based on the advice.

[0932] Specific examples and prompts

[0933] Imagine a worker in a factory is giving instructions to a robot. If this application is installed on the robot, the robot will capture the worker's voice and facial expressions in real time and send them to a server. Emotion analysis is performed on the server side, and it is determined that the worker is feeling stressed. In this case, the robot will display advice such as "Slow down your work pace a little," thereby reducing the worker's stress and improving work efficiency.

[0934] Example of input prompt for generative AI model

[0935] If the user's voice tone is judged to be "harsh": Parents, please soften your voice and provide specific examples.

[0936] "If the user is confused: Slowly review the steps. Provide specific instructions."

[0937] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0938] Step 1:

[0939] The user launches the app. The device requests permission to use the camera and microphone. This completes the initial setup and starts the system's data collection. The user has given permission as input, and the camera and microphone are available as output.

[0940] Step 2:

[0941] The terminal collects the worker's voice and facial expression data in real time through a camera and microphone. At this stage, the user (worker) gives instructions to the robot, and the voice instructions and facial expressions are captured as data. The input is the user's voice and facial expression, and the output is voice data and facial expression data.

[0942] Step 3:

[0943] The device transmits the collected voice and facial expression data to the server in real time. This transmission is performed periodically, for example, every 30 seconds, and the data is transferred to the server as a data packet. The input is the data collected in step 2, and the output is the data packet sent to the server.

[0944] Step 4:

[0945] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time. For example, it detects that the worker's voice tone has become harsher and their facial expression is tense. The input is the data sent to the server, and the output is the analysis result.

[0946] Step 5:

[0947] The server generates optimal advice based on the analysis results. For example, it creates advice such as "Slow down your work pace a little" depending on the psychological state and emotions obtained from the analysis. The input is the analysis result from step 4, and the output is the generated advice.

[0948] Step 6:

[0949] The server sends the generated advice to the terminal. The terminal receives this advice and prepares it. The input is the advice sent from the server, and the output is the advice message prepared by the terminal.

[0950] Step 7:

[0951] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Try to slow down a bit." The input is the advice message, and the output is the displayed advice.

[0952] Step 8:

[0953] The user checks the displayed advice and acts accordingly. For example, a worker slows down the pace of work based on the advice. The input is the displayed advice, and the output is the user's actual behavior.

[0954] Through this series of processing steps, the system analyzes the worker's psychological state and emotions in real time and provides appropriate feedback, thereby improving work efficiency.

[0955] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0956] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0957] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0958] [Third embodiment]

[0959] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0960] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0961] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0962] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0963] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0964] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0965] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0966] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0967] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0968] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0969] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0970] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0971] Overall system configuration

[0972] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[0973] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[0974] 3. The server uses a large-scale multimodal language model to analyze the received data.

[0975] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[0976] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[0977] Specific Embodiments of the System

[0978] The user launches the app

[0979] The user launches the app on their device and the app's initial screen is displayed.

[0980] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[0981] Data collection

[0982] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[0983] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[0984] Sending data

[0985] The device transmits the collected voice data and facial expression data to the server in real time.

[0986] For example, this is done by transferring a data packet to the server every 30 seconds.

[0987] Data analysis

[0988] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[0989] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[0990] Generating Advice

[0991] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[0992] This advice will help users create a viable and effective learning environment on the fly.

[0993] Sending Advice

[0994] The server transmits the generated advice to the terminal, and the user receives it.

[0995] For example, the server sends an advice message to the terminal, which receives and prepares it.

[0996] Displaying Advice

[0997] The terminal displays the received advice on the user interface so that the user can immediately check it.

[0998] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[0999] User executes advice

[1000] The user checks the displayed advice and acts accordingly.

[1001] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1002] Specific examples

[1003] Scenario: You're doing math homework together.

[1004] 1. A user launches the app and begins a math homework session for their child.

[1005] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1006] 3. The device sends the recorded data to the server every 30 seconds.

[1007] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1008] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1009] 6. The server sends this advice to the device, which displays it on the screen.

[1010] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1011] This system makes parent-child study sessions more effective, avoids emotional conflicts, and is expected to improve parent-child relationships.

[1012] The processing flow will be explained below.

[1013] Step 1:

[1014] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[1015] Step 2:

[1016] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1017] Step 3:

[1018] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[1019] Step 4:

[1020] The server inputs the received voice and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1021] Step 5:

[1022] Based on the analysis results, the server generates optimal advice, such as a message like, "Parents, please speak softer and explain with concrete examples." This advice helps users create an effective learning environment that they can implement on the spot.

[1023] Step 6:

[1024] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1025] Step 7:

[1026] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[1027] Step 8:

[1028] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[1029] Example 1

[1030] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1031] In conventional parent-child study sessions, parental guidance often leads to emotional conflict, making it difficult to provide effective learning support. In particular, continuing guidance without properly understanding the psychological and emotional states of parents and children can lead to a decline in the child's motivation to learn. The present invention aims to provide a system that analyzes the psychological and emotional states of parents and children in real time and provides appropriate learning advice based on the results, so that parents and children can study more effectively and while avoiding emotional conflict.

[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1033] In this invention, the server includes a device for collecting parent-child conversation voices, a device for collecting parent-child facial expression data, a device for analyzing the voice data and facial expression data in real time, a device for generating appropriate learning advice for the parent and child based on the judgment results, a device for providing the generated advice to the parent and child, a device for using a multimodal large-scale language model to analyze the psychological states and emotions of the parent and child, a device for displaying the generated advice on a user interface of a terminal, a device for collecting data in real time, and a device for transmitting the collected data to the server in real time, thereby enabling the parent-child learning session to be conducted effectively and avoiding emotional conflicts between the parent and child.

[1034] The "device for collecting parent-child conversation voices" is a device for collecting voice data of the speech and conversation between parents and children.

[1035] The "device for collecting facial expression data of parents and children" is a device for collecting facial expressions of parents and children in the form of images or videos and storing them as digital data.

[1036] The "device for analyzing voice data and facial expression data in real time" is a device that instantly analyzes collected voice data and facial expression data to analyze the psychological state and emotions of parents and children.

[1037] The "device that generates appropriate learning advice for parents and children based on the results of its judgment" is a device that generates specific instructions and advice for parents and children to effectively advance their studies based on the results of analyzing voice data and facial expression data.

[1038] The "device for providing generated advice to parent and child" is a device for presenting generated study advice so that the parent and child can confirm it.

[1039] The "device that uses a multimodal large-scale language model to analyze the psychological state and emotions of parents and children" is a device that receives voice data and facial expression data as input and analyzes them using advanced algorithms and machine learning models.

[1040] The "device that displays the generated advice on the user interface of the terminal" is a device that displays the generated advice on a terminal such as a smartphone or tablet.

[1041] The "real-time data collection device" is a device that instantly collects parent-child conversation voice and facial expression data.

[1042] "Devices that transmit collected data to a server in real time" are devices that instantly transfer collected voice data and facial expression data to a server.

[1043] The present invention is a system for conducting effective parent-child study sessions. This system collects and analyzes parent-child conversation voice and facial expression data in real time, and provides appropriate study advice.

[1044] Hardware and software used

[1045] This system uses the following hardware and software:

[1046] Device: A portable electronic device such as a smartphone or tablet. This device is equipped with a camera and microphone and collects voice and facial expression data from parent-child conversations.

[1047] Server: A high-performance computing system that hosts large-scale language models and performs analysis of speech and facial expression data.

[1048] Multimodal large-scale language model: For example, large-scale language models from Google Cloud AI, Microsoft Azure Cognitive Services, etc. are used. This model performs advanced analysis of voice and facial expression data to determine the psychological state and emotions of parents and children.

[1049] Network: An internet connection for data communication between your device and the server.

[1050] Data processing and calculation

[1051] 1. Data collection: When users (parents and children) launch the app on their devices and allow the use of the camera and microphone, the device collects conversational voice and facial expression data in real time. For example, it records the situation where a parent is explaining something to a child while they are solving a math problem.

[1052] 2. Data transmission: The collected data is sent from the device to the server. The data is batched and transmitted every 30 seconds. For example, data packets are sent to the server using an HTTP POST request.

[1053] 3. Data analysis: The server receives the voice and facial expression data and inputs it into a large-scale language model for analysis, determining the parent's tone of voice and the child's emotional state, such as whether they are confused or not.

[1054] 4. Advice generation: Based on the analysis results, the server generates effective learning advice, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1055] 5. Sending and displaying advice: The generated advice is sent from the server to the terminal and displayed on the terminal's user interface. The displayed advice can be immediately checked by the user.

[1056] Specific examples

[1057] Scenario: You're doing math homework together.

[1058] A user launches the app to begin a math homework session for their child.

[1059] The device activates the camera and microphone to record conversations and facial expressions between parent and child.

[1060] The device sends the recorded data to the server every 30 seconds.

[1061] The server analyzes the data and detects when the parent's tone is harsh and the child is in trouble.

[1062] The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1063] The server sends this advice to the terminal, which displays it on the screen.

[1064] The user (parent) follows the advice and gently resumes the explanation, using concrete examples.

[1065] Prompt Sentence Examples

[1066] "What advice would you give to a child who is confused when their parent explains things to them in a harsh tone?"

[1067] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1068] System program processing flow

[1069] Step 1: Launching the app and initial settings

[1070] 1. The user launches the app on their device.

[1071] Enter: Tap the app icon on your device.

[1072] Output: The initial screen of the app is displayed.

[1073] Specific behavior: You tap the app icon and the application loads into memory.

[1074] 2. The device will ask for permission to use the camera and microphone.

[1075] Input: The permission dialog appears on the initial screen of the app.

[1076] Output: User's choice (allow / deny) on the permission dialog.

[1077] What happens: Selecting "Allow" in the dialog grants the app camera and microphone permissions.

[1078] Step 2: Data collection

[1079] 3. The user allows use of the camera and microphone.

[1080] Input: Select "Allow" in the permission dialog.

[1081] Output: Camera and microphone are activated.

[1082] What happens: Permission is granted and the camera and microphone are automatically turned on.

[1083] 4. The device collects the parent-child conversation voice and facial expression data.

[1084] Input: A parent-child study session begins with the camera and microphone activated.

[1085] Output: Real-time collected voice and facial expression data.

[1086] Specific operation: Audio data is collected from the microphone, and facial expression data is captured from the camera.

[1087] Step 3: Send data

[1088] 5. The device sends the collected data to the server every 30 seconds.

[1089] Input: A batch of speech and facial expression data collected in real time.

[1090] Output: The data sent to the server in the HTTP POST request.

[1091] What it does: Data is batched at regular intervals and sent over the internet to a server.

[1092] 6. The server receives the voice data and facial expression data.

[1093] Input: Voice and facial expression data sent from the device.

[1094] Output: Acknowledgement message.

[1095] Specific operation: The server saves the received data and sends a confirmation message to the device confirming successful reception.

[1096] Step 4: Data analysis

[1097] 7. The server formats the received voice and facial expression data for analysis.

[1098] Input: The raw data received.

[1099] Output: Data formatted in a way that can be fed into a large-scale language model.

[1100] Specific operations: Converts voice data into text and facial expression data into a format for image analysis.

[1101] 8. The server inputs the data into a large-scale language model.

[1102] Input: Formatted speech and facial expression data.

[1103] Output: Analysis results (mental state and emotions of parent and child).

[1104] What it does: Query the model with data to get a predicted state of mind or emotion.

[1105] Step 5: Advice Generation

[1106] 9. The server generates appropriate advice based on the analysis results.

[1107] Input: Analysis results (mental state and emotions of parent and child).

[1108] Output: A study advice message.

[1109] Specific behavior: Creates a newly generated advice message.

[1110] 10. The server sends the generated advice to the device.

[1111] Input: Study advice message.

[1112] Output: Advice sent to the terminal.

[1113] Specific operation: Sends an advice message to the device via an HTTP POST request.

[1114] Step 6: Display Advice

[1115] 11. The device acknowledges receipt of the advice and displays it on the user interface.

[1116] Input: The advice message sent by the server.

[1117] Output: Advice displayed in the user interface.

[1118] Specific operation: Display the received advice message on the screen.

[1119] Step 7: User Actions

[1120] 12. The user checks the advice displayed on the device and corrects their behavior.

[1121] Input: The advice message displayed on the terminal.

[1122] Output: Modified parent behavior.

[1123] Specific actions: Based on the advice, take action such as explaining in a gentle tone with concrete examples.

[1124] (Application example 1)

[1125] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1126] While existing parent-child learning support systems have the means to collect and analyze parent-child conversation voice and facial expression data, they face the problem of difficulty in providing appropriate learning advice in real time. Furthermore, there is a need for systems that can accurately analyze the psychological state and emotions of parents and children and provide appropriate feedback immediately to facilitate communication between parents and children and improve learning effectiveness.

[1127] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1128] In this invention, the server includes means for periodically transmitting data collected by the terminal, means for generating advice in real time based on the analysis results of the server and transmitting the advice to the terminal, and means for displaying the advice on a user interface, thereby making it possible to provide appropriate learning advice in real time based on parent-child conversation voice and facial expression data.

[1129] A "means for collecting parent-child conversation audio" is a device or software used to record a conversation between a parent and a child.

[1130] "Means for collecting parent and child facial expression data" refers to cameras and sensors for capturing the facial expressions of parents and children.

[1131] The "means for analyzing voice data and facial expression data in real time" refers to software or hardware that can instantly analyze collected voice data and facial expression data.

[1132] The "means for determining the psychological state and emotions of the parent and child based on the analysis results" is an algorithm for evaluating the psychological state and emotions of the parent and child using the analysis results of the voice data and facial expression data.

[1133] The "means for generating appropriate learning advice for parents and children based on the judgment results" is software that creates optimal learning advice based on the psychological state and emotions of parents and children.

[1134] The "means for providing generated advice to parent and child" is a device or interface for communicating the generated advice to the parent and child.

[1135] "Means for periodically transmitting data collected by the terminal to a server" refers to a function or protocol for periodically transferring voice and facial expression data to a central server.

[1136] The "means for generating advice in real time based on the analysis results at the server and transmitting the advice to the terminal" is a function in which the server analyzes data, generates advice in real time, and transmits the advice to the terminal.

[1137] The "means for displaying advice on a user interface" refers to a screen or application that displays the generated advice so that the user can check it.

[1138] This invention relates to a parent-child learning support system that collects and analyzes parent-child conversation voice and facial expression data to provide appropriate learning advice in real time. This system is mainly composed of a terminal and a server.

[1139] Overall system configuration

[1140] 1. Users (parents and children) access the application using a device (smartphone or tablet).

[1141] 2. The device collects conversational audio and facial expression data between parent and child through a camera and microphone, and transmits the data to a server in real time.

[1142] 3. The server uses a multimodal large-scale language model (e.g., pipeline from the Transformers library) to analyze the collected data, generates appropriate advice in real time based on the analysis results, and sends it to the device.

[1143] 4. The device displays the advice obtained from the server on the user interface, enabling the user to receive appropriate learning support.

[1144] Specific Embodiments of the System

[1145] The user launches the app

[1146] When a user launches the app on their device, the app's initial screen is displayed.

[1147] The device will ask for permission to use the camera and microphone to ensure the parent-child learning session is ready.

[1148] Data collection

[1149] The device activates the camera and microphone to collect voice and facial expression data of parent-child conversations in real time.

[1150] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1151] Sending data

[1152] The terminal periodically (for example, every 30 seconds) transmits the collected voice data and facial expression data to the server.

[1153] Data analysis

[1154] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[1155] For example, it detects that the parent's tone of voice has become harsher and the child is confused.

[1156] Generating Advice

[1157] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1158] Sending Advice

[1159] The server transmits the generated advice to the terminal, and the user receives it.

[1160] Displaying Advice

[1161] The terminal displays the received advice on the user interface so that the user can immediately check it.

[1162] User follows advice

[1163] The user can then review the advice provided and act accordingly.

[1164] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1165] Specific examples

[1166] Scenario: You're doing math homework together.

[1167] 1. A user launches the app and begins a math homework session for their child.

[1168] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1169] 3. The device sends the recorded data to the server every 30 seconds.

[1170] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1171] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1172] 6. The server sends this advice to the device, which displays it on the screen.

[1173] 7. The user (parent) follows the advice and gently resumes the explanation with examples.

[1174] Prompt Sentence Examples

[1175] "Currently, a parent and child are doing math homework together, but the parent's tone is harsh, which is confusing the child. Please provide appropriate advice to parents."

[1176] This system will make parent-child learning sessions more effective, avoid emotional conflicts, and is expected to improve parent-child relationships.

[1177] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1178] Step 1:

[1179] The user launches the app

[1180] When a user launches the app on their device to begin a learning session, the device requests permission to use the camera and microphone. The input is the user's operation, and the output is the completion of camera and microphone settings.

[1181] Step 2:

[1182] Data collection

[1183] The device activates a camera and microphone to collect parent-child conversation voice and facial expression data in real time. Specifically, the camera captures facial expression data and the microphone records the conversation voice. The input is the parent-child conversation and facial expression, and the output is the collected voice and facial expression data.

[1184] Step 3:

[1185] Sending data

[1186] The device periodically (for example, every 30 seconds) transmits the collected voice and facial expression data to the server. The input is the collected voice and facial expression data, which are then transmitted to the server as output.

[1187] Step 4:

[1188] Data analysis

[1189] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. Specifically, it uses an AI model to analyze changes in voice tone and facial expression. The input is the received voice data and facial expression data, and the output is the analysis results.

[1190] Step 5:

[1191] Generating Advice

[1192] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please speak softer and explain with concrete examples." The input is the analysis results, and the output is the generated advice.

[1193] Step 6:

[1194] Sending Advice

[1195] The server sends the generated advice to the terminal, and the user receives it. The input is the generated advice, and the output is the transmission to the terminal.

[1196] Step 7:

[1197] Displaying Advice

[1198] The terminal displays the received advice on the user interface so that the user can immediately check it. The input is the received advice, and the output is the screen on which the advice is displayed.

[1199] Step 8:

[1200] User follows advice

[1201] The user reviews the displayed advice and acts accordingly. For example, a parent might explain the advice to their child again in a slow, gentle tone, using specific examples. The input is the displayed advice, and the output is the user's action.

[1202] These processing steps enable the system to provide appropriate learning advice to parents and children in real time and support effective learning.

[1203] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1204] Overall system configuration

[1205] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[1206] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[1207] 3. The server analyzes the received data using a multimodal large-scale language model and emotion engine.

[1208] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[1209] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[1210] Specific Embodiments of the System

[1211] The user launches the app

[1212] The user launches the app on their device and the app's initial screen is displayed.

[1213] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[1214] Data collection

[1215] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[1216] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1217] Sending data

[1218] The device transmits the collected voice data and facial expression data to the server in real time.

[1219] For example, this is done by transferring a data packet to the server every 30 seconds.

[1220] Data analysis

[1221] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[1222] The server also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results supplementally.

[1223] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1224] Generating Advice

[1225] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1226] The analysis results of the emotion engine are also taken into consideration, providing more accurate, situation-specific advice.

[1227] This advice will help users create a viable and effective learning environment on the fly.

[1228] Sending Advice

[1229] The server transmits the generated advice to the terminal, and the user receives it.

[1230] For example, the server sends an advice message to the terminal, which receives and prepares it.

[1231] Displaying Advice

[1232] The terminal displays the received advice on the user interface so that the user can immediately check it.

[1233] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[1234] User executes advice

[1235] The user checks the displayed advice and acts accordingly.

[1236] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1237] Specific examples

[1238] Scenario: You're doing math homework together.

[1239] 1. A user launches the app and begins a math homework session for their child.

[1240] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1241] 3. The device sends the recorded data to the server every 30 seconds.

[1242] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1243] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, try softening your tone of voice and explaining with examples."

[1244] 6. The server sends this advice to the device, which displays it on the screen.

[1245] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1246] This system will make parent-child study sessions more effective, avoid emotional conflicts, and improve parent-child relationships. The combination of an emotion engine will further improve the accuracy of analysis and the quality of advice.

[1247] The processing flow will be explained below.

[1248] Step 1:

[1249] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[1250] Step 2:

[1251] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1252] Step 3:

[1253] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[1254] Step 4:

[1255] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. It also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results as a supplement. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1256] Step 5:

[1257] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please soften your tone of voice and explain with specific examples." Since the analysis results of the emotion engine are also taken into consideration, more accurate advice tailored to the situation is provided.

[1258] Step 6:

[1259] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1260] Step 7:

[1261] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[1262] Step 8:

[1263] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[1264] Example 2

[1265] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1266] Until now, there has been no system that can analyze psychological and emotional issues that arise when parents and children study together in real time and provide effective advice. As a result, not only can parent-child study sessions become inefficient, but communication between parents and children can also deteriorate. To solve these problems, a system is needed that can analyze parent-child conversation voice and facial expression data in real time and provide appropriate advice.

[1267] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1268] In this invention, the server includes means for collecting parent-child conversation voices, means for collecting parent-child facial expression data, means for analyzing the voice data and facial expression data in real time, means for multimodally analyzing the voice data and facial expression data using a generative AI model, means for determining the psychological state and emotions of the parent and child based on the analysis results, means for generating appropriate study advice for the parent and child based on the determination results, means for providing the generated advice to the parent and child, means for obtaining permission to use the device's camera and microphone, means for preprocessing the voice data and facial expression data, and means for transmitting data to the server via a network. This makes parent-child study sessions more effective, avoiding emotional conflicts and improving the parent-child relationship.

[1269] "Means for collecting parent-child conversation audio" is a function for recording parent-child conversations and acquiring that data.

[1270] "Means for collecting facial expression data of parents and children" refers to a function that uses cameras and sensors to record the facial expressions of parents and children and acquire that data.

[1271] "Means for analyzing voice data and facial expression data in real time" refers to a function for instantly processing collected voice and facial expression data and analyzing its contents.

[1272] "Means for multimodal analysis of voice data and facial expression data using a generative AI model" is a function that uses a generative artificial intelligence model to combine and analyze voice data and facial expression data in an integrated manner.

[1273] "Means for determining the psychological state and emotions of parents and children based on the analysis results" is a function that infers and determines the psychological state and emotions of parents and children based on the analyzed data.

[1274] The "means for generating appropriate learning advice for parents and children based on the judgment results" is a function for creating optimal advice for parents and children based on the judged psychological state and emotions.

[1275] The "means for providing generated advice to parent and child" is a function for distributing the generated advice so that parent and child can check it.

[1276] The "means for obtaining permission to use the camera and microphone of the terminal" is a function for requesting permission to access the camera and microphone of the terminal and permitting their use.

[1277] "Means for preprocessing voice data and facial expression data" refers to a function that converts and filters voice data and facial expression data into an appropriate format before analysis.

[1278] "Means for transmitting data to a server using a network" refers to a function for transferring collected data to a server via a network.

[1279] This invention is a system that effectively supports parent-child study sessions. The system collects parent-child conversational voice and facial expression data, analyzes them using a generative AI model, and provides appropriate study advice in real time. Specific embodiments are described below.

[1280] Hardware and software configuration

[1281] Hardware:

[1282] Device (smartphone, tablet, etc.): A device used by parents and children that is equipped with a camera and microphone.

[1283] Server: A central system for analyzing data. Equipped with high-performance CPUs, GPUs, and sufficient memory.

[1284] software:

[1285] Generative AI model: Includes a large-scale language model and emotion engine, which analyzes voice and facial expression data.

[1286] Front-end application: An app for smartphones and tablets that provides the user interface.

[1287] Back-end system: Software that runs on a server for data analysis and advice generation.

[1288] Data processing and calculation

[1289] The device:

[1290] The device obtains permission to use the camera and microphone and begins collecting parent-child conversation audio and facial expression data.

[1291] The collected voice data is compressed in real time, and the facial expression data is pre-processed.

[1292] The device generates a data packet every 30 seconds and sends it to the server over the network.

[1293] The server:

[1294] The server first preprocesses the received voice and facial expression data, which includes noise removal and conversion to the required data format.

[1295] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[1296] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[1297] Based on the analysis results, advice is generated, such as a message like, "Parents, please speak softer and explain using concrete examples."

[1298] Convert the advice into an appropriate format and send it to the device.

[1299] Specific examples

[1300] Scenario: You're doing math homework together.

[1301] 1. The user launches the app and begins a math homework session for their child.

[1302] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1303] 3. The device sends the recorded data to the server every 30 seconds.

[1304] 4. The server analyzes the data and detects that the parent's tone of voice is becoming harsher and the child is in distress.

[1305] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, please soften your tone of voice and try explaining with examples."

[1306] 6. The server sends this advice to the terminal, where it is displayed on the screen.

[1307] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1308] Prompt Sentence Examples

[1309] An example of a prompt sentence input to the generative AI model is shown below.

[1310] "When a parent's tone becomes harsh during a parent-child study session, consider the child's confusion and generate advice to soften the tone of your voice."

[1311] This format makes parent-child study sessions more effective, and is expected to improve parent-child relationships while avoiding emotional conflicts.

[1312] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1313] Step 1:

[1314] The user launches the app.

[1315] Specific behavior:

[1316] A user taps to launch an app on their smartphone or tablet.

[1317] Input: User action (launching the app)

[1318] Output: The initial screen of the app is displayed.

[1319] The app's initial screen will appear, asking for permission to use the camera and microphone.

[1320] Step 2:

[1321] The device will turn on the camera and microphone and begin collecting data.

[1322] Specific behavior:

[1323] The device activates a camera and microphone to collect parent-child conversation audio and facial expression data.

[1324] Input: User permission (permission to use camera and microphone)

[1325] Output: Camera and microphone are activated and begin collecting data.

[1326] The audio data is compressed in real time, and the facial expression data is pre-processed using image processing algorithms.

[1327] Step 3:

[1328] The terminal transmits the collected data to the server.

[1329] Specific behavior:

[1330] The voice data and facial expression data collected by the terminal are sent to the server in batch processing.

[1331] Input: Collected voice and facial expression data

[1332] Output: Data packets sent to the server (e.g., data packets every 30 seconds)

[1333] The device checks the network connection and verifies that the data was sent successfully.

[1334] Step 4:

[1335] The server analyzes the data.

[1336] Specific behavior:

[1337] The server first preprocesses the received voice data and facial expression data.

[1338] Input: Voice data and facial expression data sent to the server

[1339] Output: Preprocessed data

[1340] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[1341] Input: Preprocessed speech and facial expression data

[1342] Output: Analysis results (mental state and emotional information)

[1343] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[1344] Step 5:

[1345] The server generates the advice.

[1346] Specific behavior:

[1347] The server generates optimal advice based on the analysis results.

[1348] Input: Analysis results

[1349] Output: The generated advice

[1350] A prompt sentence is input into the generative AI model to generate advice text.

[1351] Example: Create a message that reads, "Parents, please speak softly and provide examples."

[1352] Step 6:

[1353] The server sends the advice to the terminal.

[1354] Specific behavior:

[1355] Based on the advice generated by the server, a data packet is created in an appropriate format and sent to the terminal.

[1356] Input: Generated advice

[1357] Output: Data packets sent to the device (e.g., JSON formatted messages)

[1358] The server logs the status of the transmission and confirms that the transmission completed successfully.

[1359] Step 7:

[1360] The device displays the advice.

[1361] Specific behavior:

[1362] The advice received by the terminal is displayed on the user interface.

[1363] Input: Advice data sent from the server

[1364] Output: Advisory message displayed on the user interface

[1365] Example: The message on the screen reads, "Parents, please speak softly and use examples."

[1366] Step 8:

[1367] The user executes the advice.

[1368] Specific behavior:

[1369] The user reviews the displayed advice and acts accordingly.

[1370] Input: Advice displayed on screen

[1371] Output: User action (repeated in a gentle tone with examples)

[1372] The parent continues to explain using specific examples based on the advice.

[1373] (Application example 2)

[1374] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1375] In modern factories, collaboration between workers and robots has become an important element, but there is still a lack of systems that can provide appropriate feedback from robots based on the psychological state and emotions of workers. This has led to issues such as reduced work efficiency and increased worker stress.

[1376] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting the voice of the worker, means for collecting facial expression data of the worker, and means for analyzing the voice data and facial expression data in real time. This makes it possible to determine the psychological state and emotions of the worker and generate and provide appropriate work advice.

[1377] The "means for collecting worker voice" refers to a device or system that acquires voice information such as instructions and comments from workers at the factory work site in real time.

[1378] "Means for collecting facial expression data of workers" refers to devices or systems that capture the facial expressions of workers in real time using cameras or sensors and collect that data.

[1379] The "means for analyzing voice data and facial expression data in real time" refers to software and hardware that instantly processes acquired voice data and facial expression data and derives the analysis results.

[1380] The "means for determining the worker's psychological state and emotions based on the analysis results" refers to a device or algorithm that uses the results of voice and facial expression analysis to evaluate and make a judgment on the worker's psychological state and emotions.

[1381] The "means for generating appropriate work advice for the worker based on the judgment results" refers to a system or program that creates appropriate advice to improve work efficiency and reduce the worker's stress, depending on the judged psychological state and emotions.

[1382] "Means for providing generated advice to a worker" refers to a display, audio output, or other user interface for communicating generated advice to a worker.

[1383] Overall system configuration

[1384] This system is designed to support efficient communication between workers and robots at factory work sites. A specific embodiment is shown below.

[1385] The user launches the app

[1386] The user (worker) starts the system and the initial screen is displayed. The terminal requests permission to use the camera and microphone to confirm that the system is operational.

[1387] Data collection

[1388] The device collects voice and facial expression data from the worker in real time through a camera and microphone. For example, when a worker is giving instructions to a robot, the device records the voice instructions and the worker's facial expression.

[1389] Sending data

[1390] The device transmits the collected voice and facial expression data to the server in real time, for example, by transferring a data packet to the server every 30 seconds.

[1391] Data analysis

[1392] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time and supplements the analysis results. For example, the server may detect that the worker's voice tone has become harsher and that the worker's facial expression is tense.

[1393] Generating Advice

[1394] The server generates optimal advice based on the analysis results. For example, it creates a message such as "Slow down your work pace a little." Since the analysis results of the emotion engine are also taken into account, more accurate advice tailored to the situation is provided. This advice helps users create an effective work environment that can be implemented on the spot.

[1395] Sending Advice

[1396] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1397] Displaying Advice

[1398] The device displays the received advice on the user interface so that the user can immediately check it. For example, the device may display a message on the screen saying, "Try to slow down your work pace a little."

[1399] User executes advice

[1400] The user checks the displayed advice and acts accordingly. For example, a worker may slow down the pace of work based on the advice.

[1401] Specific examples and prompts

[1402] Imagine a worker in a factory is giving instructions to a robot. If this application is installed on the robot, the robot will capture the worker's voice and facial expressions in real time and send them to a server. Emotion analysis is performed on the server side, and it is determined that the worker is feeling stressed. In this case, the robot will display advice such as "Slow down your work pace a little," thereby reducing the worker's stress and improving work efficiency.

[1403] Example of input prompt for generative AI model

[1404] If the user's voice tone is judged to be "harsh": Parents, please soften your voice and provide specific examples.

[1405] "If the user is confused: Slowly review the steps. Provide specific instructions."

[1406] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1407] Step 1:

[1408] The user launches the app. The device requests permission to use the camera and microphone. This completes the initial setup and starts the system's data collection. The user has given permission as input, and the camera and microphone are available as output.

[1409] Step 2:

[1410] The terminal collects the worker's voice and facial expression data in real time through a camera and microphone. At this stage, the user (worker) gives instructions to the robot, and the voice instructions and facial expressions are captured as data. The input is the user's voice and facial expression, and the output is voice data and facial expression data.

[1411] Step 3:

[1412] The device transmits the collected voice and facial expression data to the server in real time. This transmission is performed periodically, for example, every 30 seconds, and the data is transferred to the server as a data packet. The input is the data collected in step 2, and the output is the data packet sent to the server.

[1413] Step 4:

[1414] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time. For example, it detects that the worker's voice tone has become harsher and their facial expression is tense. The input is the data sent to the server, and the output is the analysis result.

[1415] Step 5:

[1416] The server generates optimal advice based on the analysis results. For example, it creates advice such as "Slow down your work pace a little" depending on the psychological state and emotions obtained from the analysis. The input is the analysis result from step 4, and the output is the generated advice.

[1417] Step 6:

[1418] The server sends the generated advice to the terminal. The terminal receives this advice and prepares it. The input is the advice sent from the server, and the output is the advice message prepared by the terminal.

[1419] Step 7:

[1420] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Try to slow down a bit." The input is the advice message, and the output is the displayed advice.

[1421] Step 8:

[1422] The user checks the displayed advice and acts accordingly. For example, a worker slows down the pace of work based on the advice. The input is the displayed advice, and the output is the user's actual behavior.

[1423] Through this series of processing steps, the system analyzes the worker's psychological state and emotions in real time and provides appropriate feedback, thereby improving work efficiency.

[1424] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1425] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1426] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1427] [Fourth embodiment]

[1428] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1429] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1430] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1431] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1432] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1433] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1434] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1435] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1436] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1437] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1438] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1439] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1440] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1441] Overall system configuration

[1442] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[1443] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[1444] 3. The server uses a large-scale multimodal language model to analyze the received data.

[1445] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[1446] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[1447] Specific Embodiments of the System

[1448] The user launches the app

[1449] The user launches the app on their device and the app's initial screen is displayed.

[1450] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[1451] Data collection

[1452] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[1453] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1454] Sending data

[1455] The device transmits the collected voice data and facial expression data to the server in real time.

[1456] For example, this is done by transferring a data packet to the server every 30 seconds.

[1457] Data analysis

[1458] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[1459] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1460] Generating Advice

[1461] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1462] This advice will help users create a viable and effective learning environment on the fly.

[1463] Sending Advice

[1464] The server transmits the generated advice to the terminal, and the user receives it.

[1465] For example, the server sends an advice message to the terminal, which receives and prepares it.

[1466] Displaying Advice

[1467] The terminal displays the received advice on the user interface so that the user can immediately check it.

[1468] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[1469] User executes advice

[1470] The user checks the displayed advice and acts accordingly.

[1471] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1472] Specific examples

[1473] Scenario: You're doing math homework together.

[1474] 1. A user launches the app and begins a math homework session for their child.

[1475] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1476] 3. The device sends the recorded data to the server every 30 seconds.

[1477] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1478] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1479] 6. The server sends this advice to the device, which displays it on the screen.

[1480] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1481] This system makes parent-child study sessions more effective, avoids emotional conflicts, and is expected to improve parent-child relationships.

[1482] The processing flow will be explained below.

[1483] Step 1:

[1484] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[1485] Step 2:

[1486] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1487] Step 3:

[1488] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[1489] Step 4:

[1490] The server inputs the received voice and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1491] Step 5:

[1492] Based on the analysis results, the server generates optimal advice, such as a message like, "Parents, please speak softer and explain with concrete examples." This advice helps users create an effective learning environment that they can implement on the spot.

[1493] Step 6:

[1494] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1495] Step 7:

[1496] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[1497] Step 8:

[1498] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[1499] Example 1

[1500] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1501] In conventional parent-child study sessions, parental guidance often leads to emotional conflict, making it difficult to provide effective learning support. In particular, continuing guidance without properly understanding the psychological and emotional states of parents and children can lead to a decline in the child's motivation to learn. The present invention aims to provide a system that analyzes the psychological and emotional states of parents and children in real time and provides appropriate learning advice based on the results, so that parents and children can study more effectively and while avoiding emotional conflict.

[1502] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1503] In this invention, the server includes a device for collecting parent-child conversation voices, a device for collecting parent-child facial expression data, a device for analyzing the voice data and facial expression data in real time, a device for generating appropriate learning advice for the parent and child based on the judgment results, a device for providing the generated advice to the parent and child, a device for using a multimodal large-scale language model to analyze the psychological states and emotions of the parent and child, a device for displaying the generated advice on a user interface of a terminal, a device for collecting data in real time, and a device for transmitting the collected data to the server in real time, thereby enabling the parent-child learning session to be conducted effectively and avoiding emotional conflicts between the parent and child.

[1504] The "device for collecting parent-child conversation voices" is a device for collecting voice data of the speech and conversation between parents and children.

[1505] The "device for collecting facial expression data of parents and children" is a device for collecting facial expressions of parents and children in the form of images or videos and storing them as digital data.

[1506] The "device for analyzing voice data and facial expression data in real time" is a device that instantly analyzes collected voice data and facial expression data to analyze the psychological state and emotions of parents and children.

[1507] The "device that generates appropriate learning advice for parents and children based on the results of its judgment" is a device that generates specific instructions and advice for parents and children to effectively advance their studies based on the results of analyzing voice data and facial expression data.

[1508] The "device for providing generated advice to parent and child" is a device for presenting generated study advice so that the parent and child can confirm it.

[1509] The "device that uses a multimodal large-scale language model to analyze the psychological state and emotions of parents and children" is a device that receives voice data and facial expression data as input and analyzes them using advanced algorithms and machine learning models.

[1510] The "device that displays the generated advice on the user interface of the terminal" is a device that displays the generated advice on a terminal such as a smartphone or tablet.

[1511] The "real-time data collection device" is a device that instantly collects parent-child conversation voice and facial expression data.

[1512] "Devices that transmit collected data to a server in real time" are devices that instantly transfer collected voice data and facial expression data to a server.

[1513] The present invention is a system for conducting effective parent-child study sessions. This system collects and analyzes parent-child conversation voice and facial expression data in real time, and provides appropriate study advice.

[1514] Hardware and software used

[1515] This system uses the following hardware and software:

[1516] Device: A portable electronic device such as a smartphone or tablet. This device is equipped with a camera and microphone and collects voice and facial expression data from parent-child conversations.

[1517] Server: A high-performance computing system that hosts large-scale language models and performs analysis of speech and facial expression data.

[1518] Multimodal large-scale language model: For example, large-scale language models from Google Cloud AI, Microsoft Azure Cognitive Services, etc. are used. This model performs advanced analysis of voice and facial expression data to determine the psychological state and emotions of parents and children.

[1519] Network: An internet connection for data communication between your device and the server.

[1520] Data processing and calculation

[1521] 1. Data collection: When users (parents and children) launch the app on their devices and allow the use of the camera and microphone, the device collects conversational voice and facial expression data in real time. For example, it records the situation where a parent is explaining something to a child while they are solving a math problem.

[1522] 2. Data transmission: The collected data is sent from the device to the server. The data is batched and transmitted every 30 seconds. For example, data packets are sent to the server using an HTTP POST request.

[1523] 3. Data analysis: The server receives the voice and facial expression data and inputs it into a large-scale language model for analysis, determining the parent's tone of voice and the child's emotional state, such as whether they are confused or not.

[1524] 4. Advice generation: Based on the analysis results, the server generates effective learning advice, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1525] 5. Sending and displaying advice: The generated advice is sent from the server to the terminal and displayed on the terminal's user interface. The displayed advice can be immediately checked by the user.

[1526] Specific examples

[1527] Scenario: You're doing math homework together.

[1528] A user launches the app to begin a math homework session for their child.

[1529] The device activates the camera and microphone to record conversations and facial expressions between parent and child.

[1530] The device sends the recorded data to the server every 30 seconds.

[1531] The server analyzes the data and detects when the parent's tone is harsh and the child is in trouble.

[1532] The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1533] The server sends this advice to the terminal, which displays it on the screen.

[1534] The user (parent) follows the advice and gently resumes the explanation, using concrete examples.

[1535] Prompt Sentence Examples

[1536] "What advice would you give to a child who is confused when their parent explains things to them in a harsh tone?"

[1537] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1538] System program processing flow

[1539] Step 1: Launching the app and initial settings

[1540] 1. The user launches the app on their device.

[1541] Enter: Tap the app icon on your device.

[1542] Output: The initial screen of the app is displayed.

[1543] Specific behavior: You tap the app icon and the application loads into memory.

[1544] 2. The device will ask for permission to use the camera and microphone.

[1545] Input: The permission dialog appears on the initial screen of the app.

[1546] Output: User's choice (allow / deny) on the permission dialog.

[1547] What happens: Selecting "Allow" in the dialog grants the app camera and microphone permissions.

[1548] Step 2: Data collection

[1549] 3. The user allows use of the camera and microphone.

[1550] Input: Select "Allow" in the permission dialog.

[1551] Output: Camera and microphone are activated.

[1552] What happens: Permission is granted and the camera and microphone are automatically turned on.

[1553] 4. The device collects the parent-child conversation voice and facial expression data.

[1554] Input: A parent-child study session begins with the camera and microphone activated.

[1555] Output: Real-time collected voice and facial expression data.

[1556] Specific operation: Audio data is collected from the microphone, and facial expression data is captured from the camera.

[1557] Step 3: Send data

[1558] 5. The device sends the collected data to the server every 30 seconds.

[1559] Input: A batch of speech and facial expression data collected in real time.

[1560] Output: The data sent to the server in the HTTP POST request.

[1561] What it does: Data is batched at regular intervals and sent over the internet to a server.

[1562] 6. The server receives the voice data and facial expression data.

[1563] Input: Voice and facial expression data sent from the device.

[1564] Output: Acknowledgement message.

[1565] Specific operation: The server saves the received data and sends a confirmation message to the device confirming successful reception.

[1566] Step 4: Data analysis

[1567] 7. The server formats the received voice and facial expression data for analysis.

[1568] Input: The raw data received.

[1569] Output: Data formatted in a way that can be fed into a large-scale language model.

[1570] Specific operations: Converts voice data into text and facial expression data into a format for image analysis.

[1571] 8. The server inputs the data into a large-scale language model.

[1572] Input: Formatted speech and facial expression data.

[1573] Output: Analysis results (mental state and emotions of parent and child).

[1574] What it does: Query the model with data to get a predicted state of mind or emotion.

[1575] Step 5: Advice Generation

[1576] 9. The server generates appropriate advice based on the analysis results.

[1577] Input: Analysis results (mental state and emotions of parent and child).

[1578] Output: A study advice message.

[1579] Specific behavior: Creates a newly generated advice message.

[1580] 10. The server sends the generated advice to the device.

[1581] Input: Study advice message.

[1582] Output: Advice sent to the terminal.

[1583] Specific operation: Sends an advice message to the device via an HTTP POST request.

[1584] Step 6: Display Advice

[1585] 11. The device acknowledges receipt of the advice and displays it on the user interface.

[1586] Input: The advice message sent by the server.

[1587] Output: Advice displayed in the user interface.

[1588] Specific operation: Display the received advice message on the screen.

[1589] Step 7: User Actions

[1590] 12. The user checks the advice displayed on the device and corrects their behavior.

[1591] Input: The advice message displayed on the terminal.

[1592] Output: Modified parent behavior.

[1593] Specific actions: Based on the advice, take action such as explaining in a gentle tone with concrete examples.

[1594] (Application example 1)

[1595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1596] While existing parent-child learning support systems have the means to collect and analyze parent-child conversation voice and facial expression data, they face the problem of difficulty in providing appropriate learning advice in real time. Furthermore, there is a need for systems that can accurately analyze the psychological state and emotions of parents and children and provide appropriate feedback immediately to facilitate communication between parents and children and improve learning effectiveness.

[1597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1598] In this invention, the server includes means for periodically transmitting data collected by the terminal, means for generating advice in real time based on the analysis results of the server and transmitting the advice to the terminal, and means for displaying the advice on a user interface, thereby making it possible to provide appropriate learning advice in real time based on parent-child conversation voice and facial expression data.

[1599] A "means for collecting parent-child conversation audio" is a device or software used to record a conversation between a parent and a child.

[1600] "Means for collecting parent and child facial expression data" refers to cameras and sensors for capturing the facial expressions of parents and children.

[1601] The "means for analyzing voice data and facial expression data in real time" refers to software or hardware that can instantly analyze collected voice data and facial expression data.

[1602] The "means for determining the psychological state and emotions of the parent and child based on the analysis results" is an algorithm for evaluating the psychological state and emotions of the parent and child using the analysis results of the voice data and facial expression data.

[1603] The "means for generating appropriate learning advice for parents and children based on the judgment results" is software that creates optimal learning advice based on the psychological state and emotions of parents and children.

[1604] The "means for providing generated advice to parent and child" is a device or interface for communicating the generated advice to the parent and child.

[1605] "Means for periodically transmitting data collected by the terminal to a server" refers to a function or protocol for periodically transferring voice and facial expression data to a central server.

[1606] The "means for generating advice in real time based on the analysis results at the server and transmitting the advice to the terminal" is a function in which the server analyzes data, generates advice in real time, and transmits the advice to the terminal.

[1607] The "means for displaying advice on a user interface" refers to a screen or application that displays the generated advice so that the user can check it.

[1608] This invention relates to a parent-child learning support system that collects and analyzes parent-child conversation voice and facial expression data to provide appropriate learning advice in real time. This system is mainly composed of a terminal and a server.

[1609] Overall system configuration

[1610] 1. Users (parents and children) access the application using a device (smartphone or tablet).

[1611] 2. The device collects conversational audio and facial expression data between parent and child through a camera and microphone, and transmits the data to a server in real time.

[1612] 3. The server uses a multimodal large-scale language model (e.g., pipeline from the Transformers library) to analyze the collected data, generates appropriate advice in real time based on the analysis results, and sends it to the device.

[1613] 4. The device displays the advice obtained from the server on the user interface, enabling the user to receive appropriate learning support.

[1614] Specific Embodiments of the System

[1615] The user launches the app

[1616] When a user launches the app on their device, the app's initial screen is displayed.

[1617] The device will ask for permission to use the camera and microphone to ensure the parent-child learning session is ready.

[1618] Data collection

[1619] The device activates the camera and microphone to collect voice and facial expression data of parent-child conversations in real time.

[1620] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1621] Sending data

[1622] The terminal periodically (for example, every 30 seconds) transmits the collected voice data and facial expression data to the server.

[1623] Data analysis

[1624] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[1625] For example, it detects that the parent's tone of voice has become harsher and the child is confused.

[1626] Generating Advice

[1627] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1628] Sending Advice

[1629] The server transmits the generated advice to the terminal, and the user receives it.

[1630] Displaying Advice

[1631] The terminal displays the received advice on the user interface so that the user can immediately check it.

[1632] User follows advice

[1633] The user can then review the advice provided and act accordingly.

[1634] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1635] Specific examples

[1636] Scenario: You're doing math homework together.

[1637] 1. A user launches the app and begins a math homework session for their child.

[1638] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1639] 3. The device sends the recorded data to the server every 30 seconds.

[1640] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1641] 5. The server generates advice such as, "Parents, please speak softer and try explaining with concrete examples."

[1642] 6. The server sends this advice to the device, which displays it on the screen.

[1643] 7. The user (parent) follows the advice and gently resumes the explanation with examples.

[1644] Prompt Sentence Examples

[1645] "Currently, a parent and child are doing math homework together, but the parent's tone is harsh, which is confusing the child. Please provide appropriate advice to parents."

[1646] This system will make parent-child learning sessions more effective, avoid emotional conflicts, and is expected to improve parent-child relationships.

[1647] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1648] Step 1:

[1649] The user launches the app

[1650] When a user launches the app on their device to begin a learning session, the device requests permission to use the camera and microphone. The input is the user's operation, and the output is the completion of camera and microphone settings.

[1651] Step 2:

[1652] Data collection

[1653] The device activates a camera and microphone to collect parent-child conversation voice and facial expression data in real time. Specifically, the camera captures facial expression data and the microphone records the conversation voice. The input is the parent-child conversation and facial expression, and the output is the collected voice and facial expression data.

[1654] Step 3:

[1655] Sending data

[1656] The device periodically (for example, every 30 seconds) transmits the collected voice and facial expression data to the server. The input is the collected voice and facial expression data, which are then transmitted to the server as output.

[1657] Step 4:

[1658] Data analysis

[1659] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. Specifically, it uses an AI model to analyze changes in voice tone and facial expression. The input is the received voice data and facial expression data, and the output is the analysis results.

[1660] Step 5:

[1661] Generating Advice

[1662] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please speak softer and explain with concrete examples." The input is the analysis results, and the output is the generated advice.

[1663] Step 6:

[1664] Sending Advice

[1665] The server sends the generated advice to the terminal, and the user receives it. The input is the generated advice, and the output is the transmission to the terminal.

[1666] Step 7:

[1667] Displaying Advice

[1668] The terminal displays the received advice on the user interface so that the user can immediately check it. The input is the received advice, and the output is the screen on which the advice is displayed.

[1669] Step 8:

[1670] User follows advice

[1671] The user reviews the displayed advice and acts accordingly. For example, a parent might explain the advice to their child again in a slow, gentle tone, using specific examples. The input is the displayed advice, and the output is the user's action.

[1672] These processing steps enable the system to provide appropriate learning advice to parents and children in real time and support effective learning.

[1673] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1674] Overall system configuration

[1675] 1. Users (parents and children) access the app using a device (smartphone or tablet).

[1676] 2. The device collects the parent-child conversation audio and facial expressions through a camera and microphone, and transmits the data to a server in real time.

[1677] 3. The server analyzes the received data using a multimodal large-scale language model and emotion engine.

[1678] 4. Based on the analysis results, the server generates appropriate advice in real time and sends it to the device.

[1679] 5. The device displays the advice obtained from the server to the user, enabling the user to receive appropriate learning support.

[1680] Specific Embodiments of the System

[1681] The user launches the app

[1682] The user launches the app on their device and the app's initial screen is displayed.

[1683] The device will ask for permission to use the camera and microphone to ensure parents and children are ready for a study session.

[1684] Data collection

[1685] The device activates the camera and microphone to collect parent-child conversation audio and facial expression data in real time.

[1686] For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1687] Sending data

[1688] The device transmits the collected voice data and facial expression data to the server in real time.

[1689] For example, this is done by transferring a data packet to the server every 30 seconds.

[1690] Data analysis

[1691] The server inputs the received voice data and facial expression data into a large-scale language model and analyzes the psychological state and emotions of the parent and child.

[1692] The server also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results supplementally.

[1693] For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1694] Generating Advice

[1695] The server generates optimal advice based on the analysis results, such as a message like, "Parents, please speak softer and explain with concrete examples."

[1696] The analysis results of the emotion engine are also taken into consideration, providing more accurate, situation-specific advice.

[1697] This advice will help users create a viable and effective learning environment on the fly.

[1698] Sending Advice

[1699] The server transmits the generated advice to the terminal, and the user receives it.

[1700] For example, the server sends an advice message to the terminal, which receives and prepares it.

[1701] Displaying Advice

[1702] The terminal displays the received advice on the user interface so that the user can immediately check it.

[1703] For example, the screen will display a message saying, "Parents, please speak in a softer tone and explain using specific examples."

[1704] User executes advice

[1705] The user checks the displayed advice and acts accordingly.

[1706] For example, the parent may start explaining to the child again, based on the advice, using specific examples and in a slow, gentle tone.

[1707] Specific examples

[1708] Scenario: You're doing math homework together.

[1709] 1. A user launches the app and begins a math homework session for their child.

[1710] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1711] 3. The device sends the recorded data to the server every 30 seconds.

[1712] 4. The server analyzes the data and detects that the parent's tone is harsh and the child is in trouble.

[1713] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, try softening your tone of voice and explaining with examples."

[1714] 6. The server sends this advice to the device, which displays it on the screen.

[1715] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1716] This system will make parent-child study sessions more effective, avoid emotional conflicts, and improve parent-child relationships. The combination of an emotion engine will further improve the accuracy of analysis and the quality of advice.

[1717] The processing flow will be explained below.

[1718] Step 1:

[1719] The user launches the app on their device, the initial app screen appears, and the device asks for permission to use the camera and microphone to ensure the parent and child are ready for a study session.

[1720] Step 2:

[1721] The device activates a camera and microphone to collect voice and facial expression data in real time between parent and child. For example, if a parent tries to explain something to a child while they are working on a math problem, the device will record the conversation and facial expressions of both parties.

[1722] Step 3:

[1723] The device transmits the collected voice and facial expression data to the server in real time, for example, by sending a data packet to the server every 30 seconds.

[1724] Step 4:

[1725] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the psychological state and emotions of the parent and child. It also uses an emotion engine to analyze the user's emotions in real time and uses the analysis results as a supplement. For example, the analysis may detect that the parent's tone of voice has become harsher and the child is confused.

[1726] Step 5:

[1727] The server generates optimal advice based on the analysis results. For example, it creates a message such as, "Parents, please soften your tone of voice and explain with specific examples." Since the analysis results of the emotion engine are also taken into consideration, more accurate advice tailored to the situation is provided.

[1728] Step 6:

[1729] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1730] Step 7:

[1731] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Parents, please speak in a softer voice and try explaining with examples."

[1732] Step 8:

[1733] The user checks the displayed advice and acts accordingly. For example, a parent may resume explaining to their child based on the advice in a slow, gentle tone using specific examples.

[1734] Example 2

[1735] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1736] Until now, there has been no system that can analyze psychological and emotional issues that arise when parents and children study together in real time and provide effective advice. As a result, not only can parent-child study sessions become inefficient, but communication between parents and children can also deteriorate. To solve these problems, a system is needed that can analyze parent-child conversation voice and facial expression data in real time and provide appropriate advice.

[1737] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1738] In this invention, the server includes means for collecting parent-child conversation voices, means for collecting parent-child facial expression data, means for analyzing the voice data and facial expression data in real time, means for multimodally analyzing the voice data and facial expression data using a generative AI model, means for determining the psychological state and emotions of the parent and child based on the analysis results, means for generating appropriate study advice for the parent and child based on the determination results, means for providing the generated advice to the parent and child, means for obtaining permission to use the device's camera and microphone, means for preprocessing the voice data and facial expression data, and means for transmitting data to the server via a network. This makes parent-child study sessions more effective, avoiding emotional conflicts and improving the parent-child relationship.

[1739] "Means for collecting parent-child conversation audio" is a function for recording parent-child conversations and acquiring that data.

[1740] "Means for collecting facial expression data of parents and children" refers to a function that uses cameras and sensors to record the facial expressions of parents and children and acquire that data.

[1741] "Means for analyzing voice data and facial expression data in real time" refers to a function for instantly processing collected voice and facial expression data and analyzing its contents.

[1742] "Means for multimodal analysis of voice data and facial expression data using a generative AI model" is a function that uses a generative artificial intelligence model to combine and analyze voice data and facial expression data in an integrated manner.

[1743] "Means for determining the psychological state and emotions of parents and children based on the analysis results" is a function that infers and determines the psychological state and emotions of parents and children based on the analyzed data.

[1744] The "means for generating appropriate learning advice for parents and children based on the judgment results" is a function for creating optimal advice for parents and children based on the judged psychological state and emotions.

[1745] The "means for providing generated advice to parent and child" is a function for distributing the generated advice so that parent and child can check it.

[1746] The "means for obtaining permission to use the camera and microphone of the terminal" is a function for requesting permission to access the camera and microphone of the terminal and permitting their use.

[1747] "Means for preprocessing voice data and facial expression data" refers to a function that converts and filters voice data and facial expression data into an appropriate format before analysis.

[1748] "Means for transmitting data to a server using a network" refers to a function for transferring collected data to a server via a network.

[1749] This invention is a system that effectively supports parent-child study sessions. The system collects parent-child conversational voice and facial expression data, analyzes them using a generative AI model, and provides appropriate study advice in real time. Specific embodiments are described below.

[1750] Hardware and software configuration

[1751] Hardware:

[1752] Device (smartphone, tablet, etc.): A device used by parents and children that is equipped with a camera and microphone.

[1753] Server: A central system for analyzing data. Equipped with high-performance CPUs, GPUs, and sufficient memory.

[1754] software:

[1755] Generative AI model: Includes a large-scale language model and emotion engine, which analyzes voice and facial expression data.

[1756] Front-end application: An app for smartphones and tablets that provides the user interface.

[1757] Back-end system: Software that runs on a server for data analysis and advice generation.

[1758] Data processing and calculation

[1759] The device:

[1760] The device obtains permission to use the camera and microphone and begins collecting parent-child conversation audio and facial expression data.

[1761] The collected voice data is compressed in real time, and the facial expression data is pre-processed.

[1762] The device generates a data packet every 30 seconds and sends it to the server over the network.

[1763] The server:

[1764] The server first preprocesses the received voice and facial expression data, which includes noise removal and conversion to the required data format.

[1765] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[1766] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[1767] Based on the analysis results, advice is generated, such as a message like, "Parents, please speak softer and explain using concrete examples."

[1768] Convert the advice into an appropriate format and send it to the device.

[1769] Specific examples

[1770] Scenario: You're doing math homework together.

[1771] 1. The user launches the app and begins a math homework session for their child.

[1772] 2. The device will activate its camera and microphone to record the conversation and facial expressions between parent and child.

[1773] 3. The device sends the recorded data to the server every 30 seconds.

[1774] 4. The server analyzes the data and detects that the parent's tone of voice is becoming harsher and the child is in distress.

[1775] 5. The server uses its emotion engine to further assess the situation and generate advice such as, "Parents, please soften your tone of voice and try explaining with examples."

[1776] 6. The server sends this advice to the terminal, where it is displayed on the screen.

[1777] 7. The user (parent) follows the advice and gently resumes the explanation with concrete examples.

[1778] Prompt Sentence Examples

[1779] An example of a prompt sentence input to the generative AI model is shown below.

[1780] "When a parent's tone becomes harsh during a parent-child study session, consider the child's confusion and generate advice to soften the tone of your voice."

[1781] This format makes parent-child study sessions more effective, and is expected to improve parent-child relationships while avoiding emotional conflicts.

[1782] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1783] Step 1:

[1784] The user launches the app.

[1785] Specific behavior:

[1786] A user taps to launch an app on their smartphone or tablet.

[1787] Input: User action (launching the app)

[1788] Output: The initial screen of the app is displayed.

[1789] The app's initial screen will appear, asking for permission to use the camera and microphone.

[1790] Step 2:

[1791] The device will turn on the camera and microphone and begin collecting data.

[1792] Specific behavior:

[1793] The device activates a camera and microphone to collect parent-child conversation audio and facial expression data.

[1794] Input: User permission (permission to use camera and microphone)

[1795] Output: Camera and microphone are activated and begin collecting data.

[1796] The audio data is compressed in real time, and the facial expression data is pre-processed using image processing algorithms.

[1797] Step 3:

[1798] The terminal transmits the collected data to the server.

[1799] Specific behavior:

[1800] The voice data and facial expression data collected by the terminal are sent to the server in batch processing.

[1801] Input: Collected voice and facial expression data

[1802] Output: Data packets sent to the server (e.g., data packets every 30 seconds)

[1803] The device checks the network connection and verifies that the data was sent successfully.

[1804] Step 4:

[1805] The server analyzes the data.

[1806] Specific behavior:

[1807] The server first preprocesses the received voice data and facial expression data.

[1808] Input: Voice data and facial expression data sent to the server

[1809] Output: Preprocessed data

[1810] The server inputs data into the generated AI model, which analyzes the psychological state and emotions from the content of the voice and facial expressions.

[1811] Input: Preprocessed speech and facial expression data

[1812] Output: Analysis results (mental state and emotional information)

[1813] An emotion engine works in tandem to extract more detailed emotional information from the parent's tone of voice and the child's facial expressions.

[1814] Step 5:

[1815] The server generates the advice.

[1816] Specific behavior:

[1817] The server generates optimal advice based on the analysis results.

[1818] Input: Analysis results

[1819] Output: The generated advice

[1820] A prompt sentence is input into the generative AI model to generate advice text.

[1821] Example: Create a message that reads, "Parents, please speak softly and provide examples."

[1822] Step 6:

[1823] The server sends the advice to the terminal.

[1824] Specific behavior:

[1825] Based on the advice generated by the server, a data packet is created in an appropriate format and sent to the terminal.

[1826] Input: Generated advice

[1827] Output: Data packets sent to the device (e.g., JSON formatted messages)

[1828] The server logs the status of the transmission and confirms that the transmission completed successfully.

[1829] Step 7:

[1830] The device displays the advice.

[1831] Specific behavior:

[1832] The advice received by the terminal is displayed on the user interface.

[1833] Input: Advice data sent from the server

[1834] Output: Advisory message displayed on the user interface

[1835] Example: The message on the screen reads, "Parents, please speak softly and use examples."

[1836] Step 8:

[1837] The user executes the advice.

[1838] Specific behavior:

[1839] The user reviews the displayed advice and acts accordingly.

[1840] Input: Advice displayed on screen

[1841] Output: User action (repeated in a gentle tone with examples)

[1842] The parent continues to explain using specific examples based on the advice.

[1843] (Application example 2)

[1844] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1845] In modern factories, collaboration between workers and robots has become an important element, but there is still a lack of systems that can provide appropriate feedback from robots based on the psychological state and emotions of workers. This has led to issues such as reduced work efficiency and increased worker stress.

[1846] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting the voice of the worker, means for collecting facial expression data of the worker, and means for analyzing the voice data and facial expression data in real time. This makes it possible to determine the psychological state and emotions of the worker and generate and provide appropriate work advice.

[1847] The "means for collecting worker voice" refers to a device or system that acquires voice information such as instructions and comments from workers at the factory work site in real time.

[1848] "Means for collecting facial expression data of workers" refers to devices or systems that capture the facial expressions of workers in real time using cameras or sensors and collect that data.

[1849] The "means for analyzing voice data and facial expression data in real time" refers to software and hardware that instantly processes acquired voice data and facial expression data and derives the analysis results.

[1850] The "means for determining the worker's psychological state and emotions based on the analysis results" refers to a device or algorithm that uses the results of voice and facial expression analysis to evaluate and make a judgment on the worker's psychological state and emotions.

[1851] The "means for generating appropriate work advice for the worker based on the judgment results" refers to a system or program that creates appropriate advice to improve work efficiency and reduce the worker's stress, depending on the judged psychological state and emotions.

[1852] "Means for providing generated advice to a worker" refers to a display, audio output, or other user interface for communicating generated advice to a worker.

[1853] Overall system configuration

[1854] This system is designed to support efficient communication between workers and robots at factory work sites. A specific embodiment is shown below.

[1855] The user launches the app

[1856] The user (worker) starts the system and the initial screen is displayed. The terminal requests permission to use the camera and microphone to confirm that the system is operational.

[1857] Data collection

[1858] The device collects voice and facial expression data from the worker in real time through a camera and microphone. For example, when a worker is giving instructions to a robot, the device records the voice instructions and the worker's facial expression.

[1859] Sending data

[1860] The device transmits the collected voice and facial expression data to the server in real time, for example, by transferring a data packet to the server every 30 seconds.

[1861] Data analysis

[1862] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time and supplements the analysis results. For example, the server may detect that the worker's voice tone has become harsher and that the worker's facial expression is tense.

[1863] Generating Advice

[1864] The server generates optimal advice based on the analysis results. For example, it creates a message such as "Slow down your work pace a little." Since the analysis results of the emotion engine are also taken into account, more accurate advice tailored to the situation is provided. This advice helps users create an effective work environment that can be implemented on the spot.

[1865] Sending Advice

[1866] The server sends the generated advice to the terminal, which receives it. For example, the server sends an advice message to the terminal, which receives and prepares it.

[1867] Displaying Advice

[1868] The device displays the received advice on the user interface so that the user can immediately check it. For example, the device may display a message on the screen saying, "Try to slow down your work pace a little."

[1869] User executes advice

[1870] The user checks the displayed advice and acts accordingly. For example, a worker may slow down the pace of work based on the advice.

[1871] Specific examples and prompts

[1872] Imagine a worker in a factory is giving instructions to a robot. If this application is installed on the robot, the robot will capture the worker's voice and facial expressions in real time and send them to a server. Emotion analysis is performed on the server side, and it is determined that the worker is feeling stressed. In this case, the robot will display advice such as "Slow down your work pace a little," thereby reducing the worker's stress and improving work efficiency.

[1873] Example of input prompt for generative AI model

[1874] If the user's voice tone is judged to be "harsh": Parents, please soften your voice and provide specific examples.

[1875] "If the user is confused: Slowly review the steps. Provide specific instructions."

[1876] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1877] Step 1:

[1878] The user launches the app. The device requests permission to use the camera and microphone. This completes the initial setup and starts the system's data collection. The user has given permission as input, and the camera and microphone are available as output.

[1879] Step 2:

[1880] The terminal collects the worker's voice and facial expression data in real time through a camera and microphone. At this stage, the user (worker) gives instructions to the robot, and the voice instructions and facial expressions are captured as data. The input is the user's voice and facial expression, and the output is voice data and facial expression data.

[1881] Step 3:

[1882] The device transmits the collected voice and facial expression data to the server in real time. This transmission is performed periodically, for example, every 30 seconds, and the data is transferred to the server as a data packet. The input is the data collected in step 2, and the output is the data packet sent to the server.

[1883] Step 4:

[1884] The server inputs the received voice data and facial expression data into a large-scale language model to analyze the worker's psychological state and emotions. The server also uses an emotion engine to analyze the user's emotions in real time. For example, it detects that the worker's voice tone has become harsher and their facial expression is tense. The input is the data sent to the server, and the output is the analysis result.

[1885] Step 5:

[1886] The server generates optimal advice based on the analysis results. For example, it creates advice such as "Slow down your work pace a little" depending on the psychological state and emotions obtained from the analysis. The input is the analysis result from step 4, and the output is the generated advice.

[1887] Step 6:

[1888] The server sends the generated advice to the terminal. The terminal receives this advice and prepares it. The input is the advice sent from the server, and the output is the advice message prepared by the terminal.

[1889] Step 7:

[1890] The device displays the received advice on the user interface so that the user can immediately check it. For example, the screen might say, "Try to slow down a bit." The input is the advice message, and the output is the displayed advice.

[1891] Step 8:

[1892] The user checks the displayed advice and acts accordingly. For example, a worker slows down the pace of work based on the advice. The input is the displayed advice, and the output is the user's actual behavior.

[1893] Through this series of processing steps, the system analyzes the worker's psychological state and emotions in real time and provides appropriate feedback, thereby improving work efficiency.

[1894] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1895] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1896] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1897] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1898] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1899] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1900] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1901] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1902] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1903] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1904] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1905] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1906] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1907] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1908] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1909] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1910] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1911] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1912] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1913] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1914] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1915] The following is further disclosed regarding the above embodiment.

[1916] (Claim 1)

[1917] A means for collecting parent-child conversation audio;

[1918] A means for collecting facial expression data from parents and children;

[1919] means for analyzing voice data and facial expression data in real time;

[1920] a means for determining the psychological state and emotions of the parent and child based on the analysis results;

[1921] A means for generating appropriate learning advice for parents and children based on the judgment results;

[1922] a means for providing the generated advice to the parent and child;

[1923] A system including:

[1924] (Claim 2)

[1925] 10. The system of claim 1, wherein a multimodal large-scale language model is used to analyze the psychological states and emotions of parents and children.

[1926] (Claim 3)

[1927] The system according to claim 1, characterized in that the generated advice is displayed on a user interface of the terminal.

[1928] "Example 1"

[1929] (Claim 1)

[1930] A device for collecting parent-child conversation audio;

[1931] A device for collecting facial expression data of parents and children;

[1932] A device for analyzing voice data and facial expression data in real time;

[1933] a device for determining the psychological state and emotions of the parent and child based on the analysis results;

[1934] a device for generating appropriate learning advice for parents and children based on the judgment results;

[1935] a device for providing the generated advice to the parent and child;

[1936] A device that uses a multimodal large-scale language model to analyze the psychological states and emotions of parents and children;

[1937] a device for displaying the generated advice on a user interface of a terminal;

[1938] A system including:

[1939] (Claim 2)

[1940] 10. The system of claim 1, further comprising a device for real-time data collection.

[1941] (Claim 3)

[1942] 2. The system according to claim 1, further comprising a device for transmitting collected data to a server in real time.

[1943] "Application Example 1"

[1944] (Claim 1)

[1945] A means for collecting parent-child conversation audio;

[1946] A means for collecting facial expression data from parents and children;

[1947] means for analyzing voice data and facial expression data in real time;

[1948] a means for determining the psychological state and emotions of the parent and child based on the analysis results;

[1949] A means for generating appropriate learning advice for parents and children based on the judgment results;

[1950] a means for providing the generated advice to the parent and child;

[1951] means for periodically transmitting data collected by the terminal to a server;

[1952] A means for generating advice in real time based on the analysis results on the server and transmitting the advice to the terminal;

[1953] means for displaying advice in a user interface;

[1954] A system including:

[1955] (Claim 2)

[1956] 10. The system of claim 1, wherein a multimodal large-scale language model is used to analyze the psychological states and emotions of parents and children.

[1957] (Claim 3)

[1958] 2. The system according to claim 1, wherein the generated advice is displayed on a user interface of the terminal, and the history of previous advice can be checked.

[1959] "Example 2: Combining Emotion Engines"

[1960] (Claim 1)

[1961] A means for collecting parent-child conversation audio;

[1962] A means for collecting facial expression data from parents and children;

[1963] means for analyzing voice data and facial expression data in real time;

[1964] A means for multimodally analyzing voice data and facial expression data using a generative AI model;

[1965] a means for determining the psychological state and emotions of the parent and child based on the analysis results;

[1966] A means for generating appropriate learning advice for parents and children based on the judgment results;

[1967] a means for providing the generated advice to the parent and child;

[1968] A means for obtaining permission to use the device's camera and microphone;

[1969] means for preprocessing speech data and facial expression data;

[1970] means for transmitting data to a server using a network;

[1971] A system including:

[1972] (Claim 2)

[1973] 10. The system of claim 1, wherein a multimodal large-scale language model is used to analyze the psychological states and emotions of parents and children.

[1974] (Claim 3)

[1975] The system according to claim 1, characterized in that the generated advice is displayed on a user interface of the terminal.

[1976] "Application example 2 when combining emotion engines"

[1977] (Claim 1)

[1978] A means for collecting voices of workers;

[1979] A means for collecting facial expression data of a worker;

[1980] means for analyzing voice data and facial expression data in real time;

[1981] a means for determining the worker's mental state and emotions based on the analysis results;

[1982] a means for generating appropriate work advice for the worker based on the judgment result;

[1983] a means for providing the generated advice to a worker;

[1984] A system including:

[1985] (Claim 2)

[1986] 10. The system of claim 1, wherein a multimodal large-scale language model is used to analyze the worker's mental state and emotions.

[1987] (Claim 3)

[1988] The system according to claim 1, characterized in that the generated advice is displayed on a user interface of the terminal. [Explanation of symbols]

[1989] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for collecting parent-child conversation audio; A means for collecting facial expression data from parents and children; means for analyzing voice data and facial expression data in real time; a means for determining the psychological state and emotions of the parent and child based on the analysis results; A means for generating appropriate learning advice for parents and children based on the judgment results; a means for providing the generated advice to the parent and child; A system including:

2. 10. The system of claim 1, wherein a multimodal large-scale language model is used to analyze the psychological states and emotions of parents and children.

3. 2. The system according to claim 1, wherein the generated advice is displayed on a user interface of the terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A