system

The system addresses the lack of comprehensive support for children's development by generating educational content, monitoring mental states, managing indoor environments, and responding to emergencies, thereby ensuring effective growth and safety.

JP2026060632APending Publication Date: 2026-04-08SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Existing systems fail to provide comprehensive support for children's development, including effective learning assistance, mental state monitoring, indoor environment management, and rapid emergency response, particularly for parents with limited time.

Method used

A system that captures children's voices to generate educational content, monitors their mental states through conversation and facial expressions, adjusts indoor environments, and detects emergencies using sensors, all while providing personalized learning and immediate notifications to parents.

Benefits of technology

The system effectively supports children's growth and safety by offering seamless learning assistance, mental state sharing, environment management, and rapid emergency response, ensuring a secure and supportive environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026060632000001_ABST
    Figure 2026060632000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means to capture the voice of a child when they speak and generate read-aloud or educational content using a generative model, A means of playing the generated content, A means of collecting conversation data with users and inferring their mental state, Means of notifying parents of the suspected mental state, A means of using sensors to monitor the indoor environment, A means of automatically adjusting the environment based on monitoring data, A system that includes a means to send an emergency notification to parents when a dangerous situation is detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , ,

[0005] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] "Audio capture" refers to the process of collecting a user's voice as digital data using a device such as a microphone.

[0008] A "generative model" refers to an algorithm or program that automatically generates new content based on collected data.

[0009] "Read-aloud and educational content" refers to information including audio data for reading stories aloud to children, and educational materials to support learning.

[0010] "Conversation data" refers to audio information exchanged between the user and the system, as well as data including facial expressions and actions during those conversations.

[0011] "Inferring mental state" refers to analyzing and determining the user's psychological state (e.g., stress, anxiety, joy) based on collected data.

[0012] "Notifying parents" refers to sending information about the suspected mental state or emergency to the parents' smartphones or other devices in real time.

[0013] "Using sensors" refers to using sensing devices to measure physical environmental data such as temperature, humidity, and CO2 concentration.

[0014] "Monitoring data" refers to information about the physical environment collected by sensors (such as temperature, humidity, and CO2 concentration).

[0015] "Automatic environmental adjustment" refers to optimizing the indoor environment by automatically operating devices such as air conditioners and humidifiers based on monitoring data.

[0016] Detecting a "dangerous situation" refers to sensing abnormal events that could potentially threaten a child's safety, such as the activation of a fire alarm or the detection of carbon monoxide.

[0017] "Emergency notification" refers to the immediate sending of an alert to parents or guardians when a dangerous situation is detected. [Brief explanation of the drawing]

[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0022] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications. The embodiments of each function in this system are described below.

[0040] 1. Audio capture and content generation

[0041] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and uses a generative model to generate appropriate content (e.g., a story or study question). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[0042] Specific example:

[0043] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[0044] 2. Understanding and sharing the child's mental state

[0045] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative model to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety and notifies the parent.

[0046] Specific example:

[0047] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[0048] 3. Monitoring and automatic adjustment of the indoor environment

[0049] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[0050] Specific example:

[0051] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0052] 4. Detection of dangerous situations and emergency notification

[0053] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0054] Specific example:

[0055] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0056] 5. Personalized learning tailored to the user's age.

[0057] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0058] Specific example:

[0059] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0060] The above describes the embodiments of the present invention. By using this system, support for children's growth and ensuring their safety can be effectively achieved.

[0061] The following describes the processing flow.

[0062] 1. Audio capture and content generation

[0063] Step 1:

[0064] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[0065] Step 2:

[0066] The device's microphone captures the user's voice.

[0067] Step 3:

[0068] The device converts the captured audio into text data and sends that data to the server.

[0069] Step 4:

[0070] The server analyzes the text data it receives to understand the content of the request.

[0071] Step 5:

[0072] The server's generative model generates appropriate content (stories and learning questions).

[0073] Step 6:

[0074] The server generates audio data and sends it to the terminal.

[0075] Step 7:

[0076] The device plays back received audio data to provide the user with read-alouds and study support.

[0077] 2. Understanding and sharing the child's mental state

[0078] Step 1:

[0079] The device's microphone and camera collect data on the user's conversation and facial expressions.

[0080] Step 2:

[0081] The device sends the collected data to the server.

[0082] Step 3:

[0083] The server analyzes the received data and infers the user's mental state based on their tone of voice, the words they use, and their facial expressions.

[0084] Step 4:

[0085] The server notifies the parent's smartphone app of the estimated mental state.

[0086] 3. Monitoring and automatic adjustment of the indoor environment

[0087] Step 1:

[0088] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[0089] Step 2:

[0090] The device sends the collected data to the server.

[0091] Step 3:

[0092] The server analyzes the environmental data it receives and compares it to the set baseline values.

[0093] Step 4:

[0094] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[0095] Step 5:

[0096] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[0097] 4. Detection of dangerous situations and emergency notification

[0098] Step 1:

[0099] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[0100] Step 2:

[0101] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[0102] Step 3:

[0103] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[0104] 5. Personalized learning tailored to the user's age.

[0105] Step 1:

[0106] The server retrieves the user's lunar age data and past learning progress data.

[0107] Step 2:

[0108] The server's generation model generates a new learning program based on the lunar phase and past data.

[0109] Step 3:

[0110] The server sends the generated learning program to the terminal.

[0111] Step 4:

[0112] The device supports the user's learning through voice guidance based on the learning program.

[0113] Step 5:

[0114] The device periodically sends learning progress data to the server.

[0115] Step 6:

[0116] The server analyzes the received progress data and adjusts the learning program as needed.

[0117] The above outlines the specific processing steps for each function.

[0118] (Example 1)

[0119] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0120] In modern families, there is a need for systems that effectively support children's development and allow them to grow up in a safe and secure environment. In particular, support for children's learning, monitoring of their mental state, monitoring of the indoor environment, and rapid response in emergencies are crucial. However, existing systems struggle to provide these functions in a unified and effective manner.

[0121] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0122] In this invention, the server includes means for capturing the user's voice and performing speech recognition, means for converting the recognized voice into text data, means for inputting the text data into a generative AI model and generating appropriate content, means for synthesizing the generated voice content, means for transmitting the synthesized content to a terminal, means for collecting conversation data and facial expression data with the user and inferring their mental state, means for notifying the parent of the inferred mental state, means for using sensors to monitor the indoor environment, means for automatically adjusting the environment based on the monitoring data, means for sending an emergency notification to the parent when a dangerous situation is detected, and means for playing the content regenerated by the generative model on the terminal. This enables seamless support for the child's development, safety, and learning.

[0123] "Children" refers to the youngest members of a family and users of this system.

[0124] "Methods for capturing audio" refer to methods of collecting user-generated audio as digital data using microphones, speech recognition devices, etc.

[0125] A "generative model" refers to an algorithm or system that automatically generates new content using machine learning or artificial intelligence technologies.

[0126] "Reading aloud" refers to the act of providing stories or educational content to users in audio format.

[0127] "Study content" refers to educational materials such as problems and explanations provided to support user learning.

[0128] "Means for playing generated content" refers to methods of providing users with viewable digital content generated by a generative model.

[0129] "User conversation data" refers to linguistic information that users communicate with the system, and is collected by a speech recognition system.

[0130] "Facial expression data" refers to data about a user's facial expressions collected using cameras and sensors.

[0131] "Means of inferring mental state" refers to algorithms and systems that analyze collected voice data and facial expression data to estimate the user's emotions and mental state.

[0132] "Methods for notifying parents" refers to methods of sending notifications to parents' smartphones or tablets to inform them of the user's mental state or emergency situation.

[0133] A "sensor for monitoring the indoor environment" refers to a sensor device used to measure temperature, humidity, CO2 concentration, etc.

[0134] "Means of automatically adjusting the environment" refers to methods of controlling environmental devices such as air conditioners and humidifiers based on monitoring data.

[0135] A "dangerous situation" refers to any condition that could potentially threaten the safety of the user or those around them, such as a fire or a carbon monoxide leak.

[0136] "Means of emergency notification" refers to a method of promptly sending notifications to parents or appropriate personnel when a dangerous situation is detected.

[0137] "Speech recognition" refers to a technology that analyzes a user's voice as digital data and converts it into text data.

[0138] "Text data" refers to data that stores speech information analyzed by a speech recognition system as text.

[0139] "Speech synthesis" refers to the technology that generates natural-sounding speech based on text data.

[0140] A "generative AI model" refers to a model that uses artificial intelligence algorithms to automatically generate new content such as speech and text.

[0141] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications.

[0142] Audio capture and content generation

[0143] When a user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies," the device's microphone captures the user's voice and sends that voice as text data to the server. The server analyzes the received text data and uses a generative AI model to generate appropriate content (for example, a story or study problem). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[0144] Specific example:

[0145] When the user says, "Tell me a story," the device captures the voice and sends it to the server. The server uses a generative AI model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, "Once upon a time, in a certain place..."

[0146] Example of a prompt:

[0147] "Please create a short children's story that children will enjoy listening to. The theme is animal friendship."

[0148] Understanding and sharing a child's mental state

[0149] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative AI model (e.g., an emotion analysis model) to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety.

[0150] Specific example:

[0151] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[0152] Indoor environment monitoring and automatic adjustment

[0153] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[0154] Specific example:

[0155] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0156] Detection of dangerous situations and emergency notifications

[0157] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0158] Specific example:

[0159] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0160] Personalized learning tailored to the child's age.

[0161] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0162] Specific example:

[0163] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0164] The above describes the embodiments of the present invention. By using this system, support for child development and safety can be achieved effectively and in a unified manner.

[0165] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0166] Audio capture and content generation

[0167] Step 1:

[0168] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[0169] Input: User's voice

[0170] Output: Captured audio data

[0171] Step 2:

[0172] The device's microphone captures the user's voice and saves that audio as digital data.

[0173] Input: Captured audio data

[0174] Output: Digital audio data

[0175] Step 3:

[0176] The device sends the captured digital audio data to the server.

[0177] Input: Digital audio data

[0178] Output: Data to send to the server

[0179] Step 4:

[0180] The server uses speech recognition software (e.g., Google® Speech-to-Text) to convert the audio data into text data.

[0181] Input: Audio data

[0182] Output: Text data

[0183] Step 5:

[0184] The server inputs text data into an AI model (e.g., GPT-4®) and generates appropriate content as a prompt.

[0185] Input: Text data

[0186] Output: Generated text content

[0187] Step 6:

[0188] The server converts the generated text content into speech data using speech synthesis software (e.g., Amazon Polly).

[0189] Input: Generated text content

[0190] Output: Generated audio data

[0191] Step 7:

[0192] The server sends the generated audio data to the terminal.

[0193] Input: Generated audio data

[0194] Output: Data to send to the terminal

[0195] Step 8:

[0196] The device plays back the received audio data, providing the user with read-alouds and study support.

[0197] Input: Generated audio data

[0198] Output: Audio to be played back to the user

[0199] Understanding and sharing a child's mental state

[0200] Step 1:

[0201] The device's microphone and camera capture the user's voice and facial expression data.

[0202] Input: User's voice and facial expressions

[0203] Output: Audio data and facial expression data

[0204] Step 2:

[0205] The device sends the captured audio and facial expression data to the server.

[0206] Input: Voice data and facial expression data

[0207] Output: Data to send to the server

[0208] Step 3:

[0209] The server uses voice analysis software and emotion analysis models (e.g., IBM Watson® emotion analysis) to analyze voice data and facial expression data and estimate the user's mental state.

[0210] Input: Voice data and facial expression data

[0211] Output: Inferred mental state

[0212] Step 4:

[0213] The server notifies the parent's smartphone app of the prediction result.

[0214] Input: Estimated mental state

[0215] Output: Notification data sent to the parent's smartphone app

[0216] Indoor environment monitoring and automatic adjustment

[0217] Step 1:

[0218] The device acquires indoor environmental data using temperature, humidity, and CO2 sensors installed on the terminal.

[0219] Input: Indoor environment

[0220] Output: Environmental sensor data

[0221] Step 2:

[0222] The device sends the environmental sensor data it acquires to the server.

[0223] Input: Environmental sensor data

[0224] Output: Data to send to the server

[0225] Step 3:

[0226] The server analyzes environmental sensor data and generates necessary adjustment instructions.

[0227] Input: Environmental sensor data

[0228] Output: Environmental adjustment instructions

[0229] Step 4:

[0230] The server sends the generated adjustment instructions to the terminal.

[0231] Input: Environmental adjustment instructions

[0232] Output: Data to send to the terminal

[0233] Step 5:

[0234] The terminal follows the adjustment instructions and operates devices such as air conditioners and humidifiers to adjust the indoor environment.

[0235] Input: Environmental adjustment instructions

[0236] Output: Adjusted indoor environment

[0237] Detection of dangerous situations and emergency notifications

[0238] Step 1:

[0239] The device's built-in fire alarm and carbon monoxide sensor detect dangerous situations.

[0240] Input: Indoor environment

[0241] Output: Hazard detection data

[0242] Step 2:

[0243] The device sends the detected data to the server.

[0244] Input: Hazard detection data

[0245] Output: Data to send to the server

[0246] Step 3:

[0247] The server quickly analyzes the data and generates an emergency notification.

[0248] Input: Hazard detection data

[0249] Output: Emergency notification data

[0250] Step 4:

[0251] The server sends an emergency notification to the parent's smartphone app, prompting them to take immediate action.

[0252] Input: Emergency notification data

[0253] Output: Notification data sent to the parent's smartphone app

[0254] Personalized learning tailored to the child's age.

[0255] Step 1:

[0256] The server retrieves the user's lunar age data and past learning progress data.

[0257] Input: Lunar age data and learning progress data

[0258] Output: Analysis data

[0259] Step 2:

[0260] The server uses the generated AI model to create a new learning program.

[0261] Input: Analysis data

[0262] Output: New learning program

[0263] Step 3:

[0264] The server sends the generated learning program to the terminal.

[0265] Input: New learning program

[0266] Output: Data to send to the terminal

[0267] Step 4:

[0268] The device provides users with learning support content through voice guidance.

[0269] Input: New learning program

[0270] Output: Learning content for the user

[0271] Step 5:

[0272] The device feeds back user learning progress data to the server, which is then used to inform the next learning session.

[0273] Input: Learning progress data

[0274] Output: Feedback data to the server

[0275] (Application Example 1)

[0276] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0277] Supporting the development and ensuring the safety of children in modern society are crucial challenges. Existing systems do not adequately address the detailed understanding of children's mental states, provide appropriate learning support, or manage their environment. Furthermore, monitoring the health of factory workers is insufficient, making it difficult to provide appropriate breaks and preventative measures in a timely manner. To solve these problems, this invention aims to provide a system that effectively supports the development and safety of children, as well as monitors the health of factory workers.

[0278] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0279] In this invention, the server includes means for capturing the voice when a child speaks, generating reading and learning content using a generation model; means for playing back the generated content; means for collecting conversation data with the user and inferring the user's mental state; means for notifying the parent of the inferred mental state; means for using sensors for monitoring the indoor environment; means for automatically adjusting the environment based on the monitoring data; means for sending an emergency notification to the parent when a dangerous situation is detected; and means for capturing the voice and facial expression data when an operator speaks, inferring the health status of the operator using a generation model, and notifying the administrator. This enables support for the growth of children and ensuring safety, as well as monitoring and appropriate response to the health status of operators in the factory.

[0280] "Voice capture" is the process of collecting the voice when a user speaks using an input device such as a microphone.

[0281] A "generation model" is a system that generates new content and information based on the input data using machine learning and artificial intelligence algorithms.

[0282] "Content generation" is a means of creating new reading and learning content based on the collected data.

[0283] "Means for playback" refers to a voice output device or program for playing the generated content to the user.

[0284] "Inference of mental state" is the process of analyzing voice and facial expression data to judge the user's emotions and psychological state.

[0285] "Means for notifying the parent" is a communication system device for informing the parent or guardian of the inferred mental state or emergency situation.

[0286] "Monitoring of indoor environment" is the process of measuring and collecting environmental data such as temperature, humidity, and CO2 using sensors.

[0287] The "means for automatically adjusting the environment" is a system that automatically operates an air conditioner or a ventilation device based on data of the indoor environment.

[0288] The "means for making an emergency notification" is a means for detecting dangerous situations such as a fire or an abnormal carbon monoxide concentration and promptly notifying the relevant parties.

[0289] The "estimation of health status" is a process of analyzing voice and facial expression data of an operator and determining the health status such as fatigue and stress.

[0290] The "means for notifying the manager" is a communication system device for notifying the manager of the estimated health status of the operator.

[0291] The present invention is a system that effectively supports the growth and ensures the safety of children, and monitors the health status of workers in a factory. This system is composed of the following main components.

[0292] Voice capture and content generation

[0293] The terminal of the system captures the voice spoken by the user (child or operator) with a microphone. This voice data is text-converted and transmitted to the server. The server uses a generation AI model to generate appropriate content (e.g., a story to be read aloud, learning questions, an assessment of the health status of the operator) based on the input text data. For example, when a child says "Tell me a story," the generated story is transmitted to the terminal and played as voice.

[0294] Estimation and notification of mental state

[0295] ​The device is equipped with voice capture capabilities and a camera to collect user conversation and facial expression data. This data is sent to a server, where a generative AI model is used to infer the user's mental state. For example, if a child says, "I had a bad day," the server analyzes the voice and facial expression to infer stress and anxiety, and notifies the parent's smartphone app. Similarly, in a factory, the voice and facial expressions of workers are captured to infer their health status and notify the manager.

[0296] Indoor environment monitoring and automatic adjustment

[0297] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and issues environmental adjustment instructions to the terminal as needed. For example, if the temperature is too high, it sends an instruction to the terminal to turn on the air conditioner, and the air conditioner starts operating.

[0298] Detection of dangerous situations and emergency notifications

[0299] The device is equipped with a fire alarm and carbon monoxide sensor, and if a dangerous situation is detected, it immediately notifies the server. The server quickly analyzes the situation and sends an emergency notification to the parent's or administrator's smartphone app. For example, if smoke is detected, it has a function to notify the parent that there is a possibility of fire.

[0300] Children's learning support

[0301] The server generates a new learning program based on the user's (child's) age data and past learning progress. This program is sent to the device, which then assists with learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0302] Specific example

[0303] For example, when an operator in a factory talks to a robot saying "I'm tired today", the robot captures this voice and facial expression data and sends it to the server. If the server analyzes the data using a generative AI model and determines that the operator's fatigue level is high, it notifies this information to the administrator's tablet. Based on this information, the administrator can instruct the corresponding operator to take appropriate breaks.

[0304] Example of prompt sentence

[0305] As an example of a prompt sentence, it is designed as follows.

[0306] Create a program for an app of a robot that monitors the health status of operators using a voice capture system, analyzes the fatigue level of operators using voice recognition, and notifies the administrator. For example, when an operator talks saying "I'm tired today", the robot captures this, estimates the fatigue level using a generative AI model, and notifies the administrator's tablet.

[0307] Hardware and software

[0308] The hardware to be used includes a microphone for voice capture, a camera for facial expression recognition, temperature, humidity, CO2 sensors for environmental monitoring, a fire alarm, and a carbon monoxide sensor. For software, a voice recognition library (e.g., speech_recognition), a generative AI model (e.g., "GPT" or "BERT"), and a web API for data analysis are used.

[0309] By combining these, a system can be realized that comprehensively provides child growth support, ensures safety, and monitors the health status of workers in the factory.

[0310] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0311] Step 1:

[0312] [[ID=Audio data capture

[0313] The device's microphone captures the user's (child or worker's) voice input. The captured audio data is temporarily stored on the device and then sent to the speech recognition engine.

[0314] Input: Audio data

[0315] Output: Audio data sent to the speech recognition engine

[0316] Step 2:

[0317] Converting audio data to text

[0318] The device converts the captured audio data into text using a speech recognition library. The converted text data is then sent to the server.

[0319] Input: Voice data sent to the speech recognition engine

[0320] Output: Text data sent to the server

[0321] Step 3:

[0322] Text data analysis and content generation

[0323] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., stories, learning questions, health assessments). GPT and BERT are used as generative AI models for this analysis.

[0324] Input: Text data sent to the server

[0325] Output: Content data generated using a generative AI model

[0326] Step 4:

[0327] Sending and playing content data

[0328] The generated content data is sent from the server to the terminal, and the terminal provides it to the user via audio using its audio playback function.

[0329] Input: Content data generated using a generation AI model.

[0330] Output: Audio content provided using the audio playback function.

[0331] Step 5:

[0332] Inference of mental state

[0333] The device uses cameras for voice and facial expression capture to collect user conversation and facial expression data. The collected data is sent to a server. The server uses a generative AI model to analyze voice tone, words used, and facial expression data to infer the user's mental state.

[0334] Input: Audio data and facial expression data collected by the camera and microphone.

[0335] Output: Estimated mental state data

[0336] Step 6:

[0337] Notification of mental state

[0338] The estimated mental state data is sent from the server to a smartphone app for parents and administrators. This allows parents and administrators to quickly understand the situation.

[0339] Input: Estimated mental state data

[0340] Output: Notification to the parent's or administrator's smartphone app

[0341] Step 7:

[0342] Environmental data monitoring

[0343] The device's temperature, humidity, and CO2 sensors periodically collect indoor environmental data and transmit it to the server.

[0344] Input: Environmental data from temperature sensor, humidity sensor, and CO2 sensor.

[0345] Output: Monitoring data sent to the server

[0346] Step 8:

[0347] Automatic environment adjustment

[0348] The server analyzes the monitoring data and sends instructions to the terminal to adjust the air conditioning and ventilation systems as needed. The terminal then automatically adjusts the environment according to these instructions.

[0349] Input: Monitoring data sent to the server

[0350] Output: Instructions for adjusting air conditioning and ventilation systems

[0351] Step 9:

[0352] Detection of dangerous situations and emergency notifications

[0353] The device is equipped with a fire alarm and a carbon monoxide sensor, and immediately sends data to a server if a dangerous situation is detected. The server quickly analyzes this data and sends an emergency notification to the parent's or administrator's smartphone app.

[0354] Input: Data from fire alarms and carbon monoxide sensors

[0355] Output: Sending emergency notification data to the parent's or administrator's smartphone app.

[0356] Step 10:

[0357] Health status monitoring and notification

[0358] When a worker in the factory speaks into a terminal, their voice and facial expression data are captured and sent to a server. The server uses a generative AI model to estimate the worker's health status and notifies the administrator of the results.

[0359] Input: Voice data and facial expression data of factory workers

[0360] Output: Estimated health status data and notification to administrator

[0361] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0362] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[0363] 1. Audio capture and content generation

[0364] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and understands the request. Next, a generative model generates appropriate content (stories or learning questions) and sends the audio data to the device. The device then plays the audio data to provide the user with storytelling or study assistance.

[0365] Specific example:

[0366] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[0367] 2. Understanding and sharing the child's mental state

[0368] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[0369] Specific example:

[0370] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[0371] 3. Monitoring and automatic adjustment of the indoor environment

[0372] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[0373] Specific example:

[0374] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0375] 4. Detection of dangerous situations and emergency notification

[0376] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0377] Specific example:

[0378] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0379] 5. Personalized learning tailored to the user's age.

[0380] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0381] Specific example:

[0382] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0383] 6. Use of the Emotion Engine

[0384] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[0385] Specific example:

[0386] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[0387] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[0388] The following describes the processing flow.

[0389] 1. Audio capture and content generation

[0390] Step 1:

[0391] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[0392] Step 2:

[0393] The device's microphone captures the user's voice.

[0394] Step 3:

[0395] The device converts the captured audio into text data and sends that data to the server.

[0396] Step 4:

[0397] The server analyzes the text data it receives to understand the content of the request.

[0398] Step 5:

[0399] The server's generative model generates appropriate content (stories and learning questions).

[0400] Step 6:

[0401] The server generates audio data and sends it to the terminal.

[0402] Step 7:

[0403] The device plays back received audio data to provide the user with read-alouds and study support.

[0404] 2. Understanding and sharing the child's mental state

[0405] Step 1:

[0406] The device's microphone and camera collect data on the user's conversation and facial expressions.

[0407] Step 2:

[0408] The device sends the collected data to the server.

[0409] Step 3:

[0410] The server's generative model analyzes the received data and infers the user's mental state based on voice tone and the words used.

[0411] Step 4:

[0412] The server's emotion engine analyzes voice and facial expression data to recognize the user's emotions.

[0413] Step 5:

[0414] The server notifies the parent's smartphone app of their estimated mental state and emotions.

[0415] 3. Monitoring and automatic adjustment of the indoor environment

[0416] Step 1:

[0417] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[0418] Step 2:

[0419] The device sends the collected data to the server.

[0420] Step 3:

[0421] The server analyzes the environmental data it receives and compares it to the set baseline values.

[0422] Step 4:

[0423] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[0424] Step 5:

[0425] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[0426] 4. Detection of dangerous situations and emergency notification

[0427] Step 1:

[0428] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[0429] Step 2:

[0430] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[0431] Step 3:

[0432] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[0433] 5. Personalized learning tailored to the user's age.

[0434] Step 1:

[0435] The server retrieves the user's lunar age data and past learning progress data.

[0436] Step 2:

[0437] The server's generation model generates a new learning program based on the lunar phase and past data.

[0438] Step 3:

[0439] The server sends the generated learning program to the terminal.

[0440] Step 4:

[0441] The device supports the user's learning through voice guidance based on the learning program.

[0442] Step 5:

[0443] The device periodically sends learning progress data to the server.

[0444] Step 6:

[0445] The server analyzes the received progress data and adjusts the learning program as needed.

[0446] 6. Use of the Emotion Engine

[0447] Step 1:

[0448] The device sends the collected audio and facial expression data to the server.

[0449] Step 2:

[0450] The server's emotion engine analyzes the data to recognize the user's emotions.

[0451] Step 3:

[0452] Based on the sentiment data recognized by the server, the generative model individually adjusts the content.

[0453] Step 4:

[0454] Based on the data in which the server recognizes emotions, it fine-tunes the content of the notifications sent to the parents.

[0455] The above outlines the specific processing steps for each function.

[0456] (Example 2)

[0457] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0458] To effectively promote and monitor children's development and provide a safe and comfortable environment, a system is needed that collects and analyzes various information in real time and optimizes operations based on the results. Conventional systems often perform tasks such as voice capture, content generation, mental state assessment, emergency response, and environmental monitoring and adjustment individually, and few systems integrate these functions, making it difficult for parents to watch over their children with peace of mind. The present invention aims to provide a comprehensive system to solve these problems.

[0459] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data and facial expression data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for monitoring the indoor environment using temperature, humidity, and CO2 sensors; means for automatically adjusting the environment based on the monitoring data; and means for detecting dangerous situations using fire alarms and carbon monoxide sensors and notifying the parent of an emergency. This enables support for the child's growth, emotionally responsive care, environment optimization, and rapid response in emergencies.

[0460] "Voice capture" refers to recording the audio when a user speaks.

[0461] A "generative model" is an algorithm that generates text or content based on received data.

[0462] "Content generation" refers to creating information using a predefined algorithm.

[0463] "Content playback" refers to the process by which a device delivers generated information to the user.

[0464] "Conversation data" refers to audio data generated when a user speaks to the system.

[0465] "Facial expression data" refers to data recorded by a camera or other device capturing a user's facial expressions.

[0466] "Inferring mental state" refers to judging a user's emotions and psychological state based on collected data.

[0467] A "notification" is a message used to convey specific information to an administrator, such as a parent.

[0468] A "temperature sensor" is a device used to measure the temperature of a room.

[0469] A "humidity sensor" is a device used to measure the humidity in a room.

[0470] A "CO2 sensor" is a device used to measure the concentration of carbon dioxide in a room.

[0471] "Monitoring" refers to the continuous observation of specific environmental conditions and the collection of data.

[0472] "Automatic adjustment" refers to the automatic modification of the device's operation based on the received data.

[0473] "Detecting dangerous situations" means sensing emergencies such as fires or carbon monoxide poisoning.

[0474] An "emergency notification" is a system that sends out an immediate warning when danger is detected.

[0475] A "personalized learning program" refers to learning content tailored to each individual user.

[0476] "Progress data" refers to data that records the progress of learning or activities.

[0477] An "emotion engine" is a system that analyzes voice and facial expression data to recognize emotions.

[0478] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[0479] Audio capture and content generation

[0480] The user requests:

[0481] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice, converts the data into text, and sends it to the server. The server uses a generative AI model to analyze the voice data, understand the request, and generate appropriate content. This generated content is saved as text data, converted back into audio data, and sent back to the device. The device then plays the audio data, providing storytelling or study assistance.

[0482] Specific example:

[0483] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[0484] Examples of prompts to input into a generative AI model:

[0485] The user requested, "Tell me a story." Please generate a Japanese children's fairy tale.

[0486] Understanding and sharing a child's mental state

[0487] Device features:

[0488] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[0489] Specific example:

[0490] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[0491] Examples of prompts to input into a generative AI model:

[0492] The user said, "Something unpleasant happened at school today." Based on this statement, identify their emotions and determine whether they are feeling anxious.

[0493] Indoor environment monitoring and automatic adjustment

[0494] Device environment monitoring:

[0495] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[0496] Specific example:

[0497] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0498] Examples of prompts to input into a generative AI model:

[0499] The indoor temperature is over 30°C. Please generate a command to turn on the air conditioner.

[0500] Detection of dangerous situations and emergency notifications

[0501] Device risk detection function:

[0502] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0503] Specific example:

[0504] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0505] Examples of prompts to input into a generative AI model:

[0506] The device has detected smoke. Please generate an emergency notification message for your parents.

[0507] Personalized learning tailored to the child's age.

[0508] Server's learning program generation function:

[0509] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0510] Specific example:

[0511] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0512] Examples of prompts to input into a generative AI model:

[0513] Please generate a program that teaches the hiragana character "い" based on the user's age in months and past learning data.

[0514] Use of the emotion engine

[0515] Emotion recognition feature of the device:

[0516] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[0517] Specific example:

[0518] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[0519] Examples of prompts to input into a generative AI model:

[0520] The user said, "I had fun today!" Use the emotion engine to generate content that the child will find enjoyable, once the emotion of joy is recognized.

[0521] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[0522] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0523] Step 1:

[0524] The user requests:

[0525] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice and collects that audio data. Specifically, when the user says "Tell me a story," the microphone acquires audio data, and that data is processed in the next step.

[0526] Input: User's voice

[0527] Output: Collected audio data

[0528] Step 2:

[0529] Converting audio data to text:

[0530] The device uses speech recognition software to convert the captured audio data into text data. Specifically, the speech recognition engine analyzes the audio and generates text data such as "Tell me your story." This text data is then sent to the server in the next step.

[0531] Input: Audio data

[0532] Output: Text data

[0533] Step 3:

[0534] Sending text data:

[0535] The terminal sends the converted text data to the server. The server receives the text data and prepares it for analysis. Specifically, the text data is sent to the server via the internet.

[0536] Input: Text data

[0537] Output: Text data sent to the server

[0538] Step 4:

[0539] Content generation:

[0540] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., fairy tales or learning questions). Specifically, the generative AI model processes a request like "Tell me a story" and generates a Japanese fairy tale. The generated text data is then converted into audio data.

[0541] Input: Text data

[0542] Output: Generated content (text data)

[0543] Step 5:

[0544] Generating audio data:

[0545] The generated text data is converted into speech data by the server's text-to-speech engine. Specifically, the text-to-speech engine generates speech data such as "Once upon a time, in a certain place..." This speech data is then sent to the terminal in the next step.

[0546] Input: Generated content (text data)

[0547] Output: Audio data

[0548] Step 6:

[0549] Sending audio data:

[0550] The server sends audio data to the terminal. The terminal prepares to play the received audio data. Specifically, the audio data is sent to the terminal via the internet.

[0551] Input: Audio data

[0552] Output: Audio data sent to the terminal

[0553] Step 7:

[0554] Play content:

[0555] The device plays the audio data it receives and provides it to the user. Specifically, the device's speaker plays the audio "Once upon a time, in a certain place..." This allows the user to listen to a fairy tale.

[0556] Input: Audio data sent to the device

[0557] Output: Audio data that the user will listen to.

[0558] ---

[0559] The above outlines the specific processing flow for voice capture and content generation in this system. This allows users to seamlessly receive content tailored to their requests.

[0560] (Application Example 2)

[0561] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0562] In modern society, ensuring family safety and a comfortable shopping experience in physical stores is a crucial challenge. However, existing systems struggle to comprehensively address issues such as providing engaging content for children, understanding their emotional state, ensuring a comfortable environment, and responding quickly to emergencies. Furthermore, there is a lack of means for parents to know their child's current location and status in real time. This leads to problems where parental peace of mind and child safety are not adequately ensured.

[0563] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for emergency notification to the parent when a dangerous situation is detected; means for understanding the location information of the parent and child within the physical store and ensuring the child's safety; and means for providing various information within the store. This makes it possible for families to enjoy shopping safely and comfortably in physical stores.

[0564] "Voice capture" is a technology that captures the voice of a child when they speak and converts it into data that can be analyzed.

[0565] A "generative model" is an algorithm used to generate content based on input data.

[0566] "Read-aloud content" refers to stories and narratives intended to be read aloud to children.

[0567] "Study content" refers to problems and learning materials created to support children's learning.

[0568] "User conversation data" refers to audio and text data recorded during interactions with children.

[0569] "Inferring mental state" refers to judging a user's emotions and mood from their voice and facial expressions.

[0570] "Notification to parents" refers to a means of communication to inform parents about their child's presumed mental state or emergency situation.

[0571] "Indoor environment monitoring" refers to the act of measuring indoor environmental information such as temperature, humidity, and carbon dioxide concentration using sensors.

[0572] "Automatic environmental adjustment" refers to the automatic operation of devices such as air conditioners and humidifiers based on monitoring data to maintain an appropriate indoor environment.

[0573] "Detection of dangerous situations" refers to the act of using sensors to detect emergencies such as fires or carbon monoxide emissions.

[0574] "Emergency notification" refers to a notification function that quickly informs parents when a dangerous situation occurs.

[0575] "Location tracking" refers to technology that tracks and records a child's current location within a physical store.

[0576] "Means for ensuring safety" refers to systems that use location information to prevent children from getting lost and to respond to emergencies.

[0577] "In-store information guidance" refers to a function that informs users of information provided within the store, such as the location of restrooms and the location of sales areas.

[0578] In order to implement this invention, the following system configuration and processing are necessary.

[0579] System configuration:

[0580] 1. Voice capture device: This includes a microphone for capturing the child's voice and converting it into digital data.

[0581] 2. Generative Model Server: Receives audio data as text data, analyzes it, and generates appropriate content (stories and learning questions). For this purpose, it uses a database and AI algorithms (e.g., generative AI models such as GPT).

[0582] 3. Content playback device: Equipped with speakers for playing the generated content as audio.

[0583] 4. Emotion Recognition Engine: A device that analyzes conversation data and facial expression data with a user to infer their mental state, and includes a camera and emotion recognition software (e.g., emotion recognition API).

[0584] 5. Notification system: Equipped with a communication module to notify the parent's smartphone of the suspected mental state.

[0585] 6. Environmental monitoring sensors: Equipped with sensors (e.g., DHT22 or MQ-135 sensors) for measuring indoor temperature, humidity, CO2 concentration, etc.

[0586] 7. Automatic environmental control device: A device for controlling environmental control devices such as air conditioners and humidifiers.

[0587] 8. Emergency notification system: A device that detects fires, carbon monoxide emissions, and other emergencies and sends emergency notifications to the parent's smartphone.

[0588] 9. Location tracking system: GPS device and tracking software for determining a child's current location in real time.

[0589] Processing procedure:

[0590] This system supports children's safety and development through a multi-layered process.

[0591] 1. Audio capture and content generation:

[0592] When the user (child) says something like "Tell me a story" or "Help me with my studies," the device's microphone captures the audio and sends it to the generative model server.

[0593] The server analyzes the audio data as text data and generates appropriate content using a generative model.

[0594] For example, if a user says, "Tell me a story," the server generates a fairy tale, and the device plays the audio data.

[0595] As an example of a prompt message using a generative AI model, you can input the instruction, "Generate a fairy tale suitable for children."

[0596] 2. Emotion recognition and notification to parents:

[0597] The system collects conversation data and facial expression data from users and sends them to the server.

[0598] The server uses an emotion recognition engine to analyze the data and infer the user's mental state.

[0599] If the estimated mental state is "anxiety" or "joy," a notification will be sent to the parent's smartphone stating, "Your child is feeling anxious" or "Your child is very happy."

[0600] For example, if a user says, "Something unpleasant happened at school today...", the device sends the audio and facial expression data to the server, infers the user's anxiety, and sends a notification.

[0601] 3. Monitoring and automatic adjustment of the indoor environment:

[0602] The device measures temperature, humidity, and CO2 concentration at regular intervals and sends the data to the server.

[0603] The server analyzes the data and sends adjustment instructions to the terminal to maintain a comfortable environment.

[0604] For example, when the server detects a room temperature of 30°C, it sends a command to the terminal to turn on the air conditioner, and the terminal follows that command to lower the room temperature.

[0605] 4. Detection of dangerous situations and emergency notification:

[0606] If the terminal detects a fire or an increase in carbon monoxide concentration, it will immediately send data to the server.

[0607] The server analyzes the information and sends a notification to the parent's smartphone saying, "Emergency: Potential fire."

[0608] For example, one possible implementation is to immediately send an emergency notification to the parent when smoke is detected.

[0609] 5. Location information tracking and guidance within physical stores:

[0610] The system tracks the child's location in real time and displays it on the parent's smartphone.

[0611] Parents can know their child's current location, allowing them to enjoy shopping with peace of mind.

[0612] For example, if a child says, "I need to go to the toilet," the system can guide them to the location of the toilet within the store.

[0613] In summary, this invention comprehensively supports the safety and development of children and provides parents with an environment in which they can feel at ease.

[0614] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0615] Step 1:

[0616] Voice capture and text conversion:

[0617] The user (child) speaks aloud, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures this audio, and voice capture software converts the audio data into text data. The input is audio data, and the output is text data.

[0618] Step 2:

[0619] Sending text data to the server:

[0620] The terminal sends the captured text data to the server. The input is text data, and the output is the data sent to the server. Specifically, the data is sent to the server using an HTTP request.

[0621] Step 3:

[0622] Content generation:

[0623] The server analyzes the received text data. Using a generative AI model (e.g., GPT-3®), it generates appropriate content (fairy tales or learning questions) in response to the request. During this process, prompts are input to the generative AI model, and the generated stories or questions are output as prompt results. For example, the prompt "Generate a fairy tale suitable for children" might be used.

[0624] Step 4:

[0625] Converting generated content into audio data:

[0626] The generated content is converted into audio data. The server uses a text-to-speech (TTS) engine to convert the generated text content into audio data. The input is the generated text data, and the output is audio data.

[0627] Step 5:

[0628] Sending audio data to the terminal:

[0629] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the data sent to the terminal. Specifically, the data is sent to the terminal using an HTTP response.

[0630] Step 6:

[0631] Play content:

[0632] The device plays the received audio data. It uses a speaker to let the user (child) listen to the content aloud. The input is audio data, and the output is audio playback. For example, the device starts reading a fairy tale, such as "Once upon a time, in a certain place..."

[0633] Step 7:

[0634] Collection and inference of emotional data:

[0635] The system collects conversation data and facial expression data from the user using the device's microphone and camera. This data is sent to a server, which analyzes it using an emotion recognition engine. The input is voice and facial expression data, and the output is estimated emotion data. For example, from a conversation like, "Something unpleasant happened at school today...", the system might infer "anxiety".

[0636] Step 8:

[0637] Notification to parents regarding their child's mental state:

[0638] The server sends a notification to the parent based on the inferred emotion data. The notification is sent via a smartphone application. The input is emotion data, and the output is a notification message. For example, a notification saying "Your child is feeling anxious" is sent.

[0639] Step 9:

[0640] Environmental data monitoring:

[0641] The terminal is equipped with sensors to measure temperature, humidity, and CO2 concentration. This environmental data is transmitted to the server at regular intervals. The input is sensor data, and the output is data transmitted to the server.

[0642] Step 10:

[0643] Automatic environment adjustment:

[0644] The server analyzes the received environmental data and issues instructions for necessary adjustments. It remotely controls devices such as air conditioners and humidifiers to automatically adjust the environment. The input is environmental data, and the output is instructions for operating the devices. For example, if it detects a room temperature of 30°C, it will issue an instruction to turn on the air conditioner.

[0645] Step 11:

[0646] Detection of dangerous situations and emergency notifications:

[0647] The device has sensors to detect fire and elevated carbon monoxide levels. When these dangerous situations are detected, it immediately sends data to a server. The server then sends an emergency notification to the parent's smartphone. The input is sensor data, and the output is an emergency notification message. For example, if smoke is detected, a notification such as "Emergency: Potential fire" is sent.

[0648] Step 12:

[0649] Location tracking and guidance:

[0650] The device tracks the child's current location in real time and displays the location information on the parent's smartphone. The input is GPS data, and the output is a display of location information. For example, if the child says, "I need to go to the toilet," the system will guide them to the location of the toilet.

[0651] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0652] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0653] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0654] [Second Embodiment]

[0655] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0656] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0657] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0658] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0659] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0660] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0661] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0662] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0663] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0664] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0665] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0666] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0667] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications. The embodiments of each function in this system are described below.

[0668] 1. Audio capture and content generation

[0669] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and uses a generative model to generate appropriate content (e.g., a story or study question). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[0670] Specific example:

[0671] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[0672] 2. Understanding and sharing the child's mental state

[0673] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative model to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety and notifies the parent.

[0674] Specific example:

[0675] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[0676] 3. Monitoring and automatic adjustment of the indoor environment

[0677] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[0678] Specific example:

[0679] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0680] 4. Detection of dangerous situations and emergency notification

[0681] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0682] Specific example:

[0683] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0684] 5. Personalized learning tailored to the user's age.

[0685] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0686] Specific example:

[0687] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0688] The above describes the embodiments of the present invention. By using this system, support for children's growth and ensuring their safety can be effectively achieved.

[0689] The following describes the processing flow.

[0690] 1. Audio capture and content generation

[0691] Step 1:

[0692] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[0693] Step 2:

[0694] The device's microphone captures the user's voice.

[0695] Step 3:

[0696] The device converts the captured audio into text data and sends that data to the server.

[0697] Step 4:

[0698] The server analyzes the text data it receives to understand the content of the request.

[0699] Step 5:

[0700] The server's generative model generates appropriate content (stories and learning questions).

[0701] Step 6:

[0702] The server generates audio data and sends it to the terminal.

[0703] Step 7:

[0704] The device plays back received audio data to provide the user with read-alouds and study support.

[0705] 2. Understanding and sharing the child's mental state

[0706] Step 1:

[0707] The device's microphone and camera collect data on the user's conversation and facial expressions.

[0708] Step 2:

[0709] The device sends the collected data to the server.

[0710] Step 3:

[0711] The server analyzes the received data and infers the user's mental state based on their tone of voice, the words they use, and their facial expressions.

[0712] Step 4:

[0713] The server notifies the parent's smartphone app of the estimated mental state.

[0714] 3. Monitoring and automatic adjustment of the indoor environment

[0715] Step 1:

[0716] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[0717] Step 2:

[0718] The device sends the collected data to the server.

[0719] Step 3:

[0720] The server analyzes the environmental data it receives and compares it to the set baseline values.

[0721] Step 4:

[0722] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[0723] Step 5:

[0724] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[0725] 4. Detection of dangerous situations and emergency notification

[0726] Step 1:

[0727] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[0728] Step 2:

[0729] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[0730] Step 3:

[0731] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[0732] 5. Personalized learning tailored to the user's age.

[0733] Step 1:

[0734] The server retrieves the user's lunar age data and past learning progress data.

[0735] Step 2:

[0736] The server's generation model generates a new learning program based on the lunar phase and past data.

[0737] Step 3:

[0738] The server sends the generated learning program to the terminal.

[0739] Step 4:

[0740] The device supports the user's learning through voice guidance based on the learning program.

[0741] Step 5:

[0742] The device periodically sends learning progress data to the server.

[0743] Step 6:

[0744] The server analyzes the received progress data and adjusts the learning program as needed.

[0745] The above outlines the specific processing steps for each function.

[0746] (Example 1)

[0747] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0748] In modern families, there is a need for systems that effectively support children's development and allow them to grow up in a safe and secure environment. In particular, support for children's learning, monitoring of their mental state, monitoring of the indoor environment, and rapid response in emergencies are crucial. However, existing systems struggle to provide these functions in a unified and effective manner.

[0749] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0750] In this invention, the server includes means for capturing the user's voice and performing speech recognition, means for converting the recognized voice into text data, means for inputting the text data into a generative AI model and generating appropriate content, means for synthesizing the generated voice content, means for transmitting the synthesized content to a terminal, means for collecting conversation data and facial expression data with the user and inferring their mental state, means for notifying the parent of the inferred mental state, means for using sensors to monitor the indoor environment, means for automatically adjusting the environment based on the monitoring data, means for sending an emergency notification to the parent when a dangerous situation is detected, and means for playing the content regenerated by the generative model on the terminal. This enables seamless support for the child's development, safety, and learning.

[0751] "Children" refers to the youngest members of a family and users of this system.

[0752] "Methods for capturing audio" refer to methods of collecting user-generated audio as digital data using microphones, speech recognition devices, etc.

[0753] A "generative model" refers to an algorithm or system that automatically generates new content using machine learning or artificial intelligence technologies.

[0754] "Reading aloud" refers to the act of providing stories or educational content to users in audio format.

[0755] "Study content" refers to educational materials such as problems and explanations provided to support user learning.

[0756] "Means for playing generated content" refers to methods of providing users with viewable digital content generated by a generative model.

[0757] "User conversation data" refers to linguistic information that users communicate with the system, and is collected by a speech recognition system.

[0758] "Facial expression data" refers to data about a user's facial expressions collected using cameras and sensors.

[0759] "Means of inferring mental state" refers to algorithms and systems that analyze collected voice data and facial expression data to estimate the user's emotions and mental state.

[0760] "Methods for notifying parents" refers to methods of sending notifications to parents' smartphones or tablets to inform them of the user's mental state or emergency situation.

[0761] A "sensor for monitoring the indoor environment" refers to a sensor device used to measure temperature, humidity, CO2 concentration, etc.

[0762] "Means of automatically adjusting the environment" refers to methods of controlling environmental devices such as air conditioners and humidifiers based on monitoring data.

[0763] A "dangerous situation" refers to any condition that could potentially threaten the safety of the user or those around them, such as a fire or a carbon monoxide leak.

[0764] "Means of emergency notification" refers to a method of promptly sending notifications to parents or appropriate personnel when a dangerous situation is detected.

[0765] "Speech recognition" refers to a technology that analyzes a user's voice as digital data and converts it into text data.

[0766] "Text data" refers to data that stores speech information analyzed by a speech recognition system as text.

[0767] "Speech synthesis" refers to the technology that generates natural-sounding speech based on text data.

[0768] A "generative AI model" refers to a model that uses artificial intelligence algorithms to automatically generate new content such as speech and text.

[0769] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications.

[0770] Audio capture and content generation

[0771] When a user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies," the device's microphone captures the user's voice and sends that voice as text data to the server. The server analyzes the received text data and uses a generative AI model to generate appropriate content (for example, a story or study problem). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[0772] Specific example:

[0773] When the user says, "Tell me a story," the device captures the voice and sends it to the server. The server uses a generative AI model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, "Once upon a time, in a certain place..."

[0774] Example of a prompt:

[0775] "Please create a short children's story that children will enjoy listening to. The theme is animal friendship."

[0776] Understanding and sharing a child's mental state

[0777] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative AI model (e.g., an emotion analysis model) to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety.

[0778] Specific example:

[0779] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[0780] Indoor environment monitoring and automatic adjustment

[0781] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[0782] Specific example:

[0783] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[0784] Detection of dangerous situations and emergency notifications

[0785] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[0786] Specific example:

[0787] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[0788] Personalized learning tailored to the child's age.

[0789] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0790] Specific example:

[0791] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[0792] The above describes the embodiments of the present invention. By using this system, support for child development and safety can be achieved effectively and in a unified manner.

[0793] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0794] Audio capture and content generation

[0795] Step 1:

[0796] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[0797] Input: User's voice

[0798] Output: Captured audio data

[0799] Step 2:

[0800] The device's microphone captures the user's voice and saves that audio as digital data.

[0801] Input: Captured audio data

[0802] Output: Digital audio data

[0803] Step 3:

[0804] The device sends the captured digital audio data to the server.

[0805] Input: Digital audio data

[0806] Output: Data to send to the server

[0807] Step 4:

[0808] The server uses speech recognition software (e.g., Google Speech-to-Text) to convert the audio data into text data.

[0809] Input: Audio data

[0810] Output: Text data

[0811] Step 5:

[0812] The server inputs text data into an AI model (e.g., GPT-4) and generates appropriate content as a prompt.

[0813] Input: Text data

[0814] Output: Generated text content

[0815] Step 6:

[0816] The server converts the generated text content into speech data using speech synthesis software (e.g., Amazon Polly).

[0817] Input: Generated text content

[0818] Output: Generated audio data

[0819] Step 7:

[0820] The server sends the generated audio data to the terminal.

[0821] Input: Generated audio data

[0822] Output: Data to send to the terminal

[0823] Step 8:

[0824] The device plays back the received audio data, providing the user with read-alouds and study support.

[0825] Input: Generated audio data

[0826] Output: Audio to be played back to the user

[0827] Understanding and sharing a child's mental state

[0828] Step 1:

[0829] The device's microphone and camera capture the user's voice and facial expression data.

[0830] Input: User's voice and facial expressions

[0831] Output: Audio data and facial expression data

[0832] Step 2:

[0833] The device sends the captured audio and facial expression data to the server.

[0834] Input: Voice data and facial expression data

[0835] Output: Data to send to the server

[0836] Step 3:

[0837] The server uses voice analysis software and emotion analysis models (e.g., IBM Watson's emotion analysis) to analyze voice data and facial expression data and infer the user's mental state.

[0838] Input: Voice data and facial expression data

[0839] Output: Inferred mental state

[0840] Step 4:

[0841] The server notifies the parent's smartphone app of the prediction result.

[0842] Input: Estimated mental state

[0843] Output: Notification data sent to the parent's smartphone app

[0844] Indoor environment monitoring and automatic adjustment

[0845] Step 1:

[0846] The device acquires indoor environmental data using temperature, humidity, and CO2 sensors installed on the terminal.

[0847] Input: Indoor environment

[0848] Output: Environmental sensor data

[0849] Step 2:

[0850] The device sends the environmental sensor data it acquires to the server.

[0851] Input: Environmental sensor data

[0852] Output: Data to send to the server

[0853] Step 3:

[0854] The server analyzes environmental sensor data and generates necessary adjustment instructions.

[0855] Input: Environmental sensor data

[0856] Output: Environmental adjustment instructions

[0857] Step 4:

[0858] The server sends the generated adjustment instructions to the terminal.

[0859] Input: Environmental adjustment instructions

[0860] Output: Data to send to the terminal

[0861] Step 5:

[0862] The terminal follows the adjustment instructions and operates devices such as air conditioners and humidifiers to adjust the indoor environment.

[0863] Input: Environmental adjustment instructions

[0864] Output: Adjusted indoor environment

[0865] Detection of dangerous situations and emergency notifications

[0866] Step 1:

[0867] The device's built-in fire alarm and carbon monoxide sensor detect dangerous situations.

[0868] Input: Indoor environment

[0869] Output: Hazard detection data

[0870] Step 2:

[0871] The device sends the detected data to the server.

[0872] Input: Hazard detection data

[0873] Output: Data to send to the server

[0874] Step 3:

[0875] The server quickly analyzes the data and generates an emergency notification.

[0876] Input: Hazard detection data

[0877] Output: Emergency notification data

[0878] Step 4:

[0879] The server sends an emergency notification to the parent's smartphone app, prompting them to take immediate action.

[0880] Input: Emergency notification data

[0881] Output: Notification data sent to the parent's smartphone app

[0882] Personalized learning tailored to the child's age.

[0883] Step 1:

[0884] The server retrieves the user's lunar age data and past learning progress data.

[0885] Input: Lunar age data and learning progress data

[0886] Output: Analysis data

[0887] Step 2:

[0888] The server uses the generated AI model to create a new learning program.

[0889] Input: Analysis data

[0890] Output: New learning program

[0891] Step 3:

[0892] The server sends the generated learning program to the terminal.

[0893] Input: New learning program

[0894] Output: Data to send to the terminal

[0895] Step 4:

[0896] The device provides users with learning support content through voice guidance.

[0897] Input: New learning program

[0898] Output: Learning content for the user

[0899] Step 5:

[0900] The device feeds back user learning progress data to the server, which is then used to inform the next learning session.

[0901] Input: Learning progress data

[0902] Output: Feedback data to the server

[0903] (Application Example 1)

[0904] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0905] Supporting the development and ensuring the safety of children in modern society are crucial challenges. Existing systems do not adequately address the detailed understanding of children's mental states, provide appropriate learning support, or manage their environment. Furthermore, monitoring the health of factory workers is insufficient, making it difficult to provide appropriate breaks and preventative measures in a timely manner. To solve these problems, this invention aims to provide a system that effectively supports the development and safety of children, as well as monitors the health of factory workers.

[0906] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0907] In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or educational content using a generative model; means for playing the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for making an emergency notification to the parent when a dangerous situation is detected; and means for capturing the voice and facial expression data of a worker when they speak, inferring the worker's health status using a generative model, and notifying the manager. This enables support for the development and safety of children, as well as monitoring the health status of workers in the factory and taking appropriate action.

[0908] "Voice capture" is the process of collecting the voice of a user when they speak using an input device such as a microphone.

[0909] A "generative model" is a system that uses machine learning or artificial intelligence algorithms to generate new content or information based on input data.

[0910] "Content generation" refers to the process of creating new read-aloud or educational content based on collected data.

[0911] "Means of playback" refers to audio output devices or programs that allow users to listen to the generated content.

[0912] "Mental state estimation" is a process that analyzes voice and facial expression data to determine the user's emotions and psychological state.

[0913] "Means of notifying parents" refers to communication systems and devices used to inform parents or guardians of suspected mental states or emergencies.

[0914] "Indoor environment monitoring" is the process of measuring and collecting environmental data such as temperature, humidity, and CO2 using sensors.

[0915] "Means for automatically adjusting the environment" refers to a system that automatically operates air conditioners and ventilation devices based on data about the indoor environment.

[0916] "Means of emergency notification" refers to means of detecting dangerous situations such as fires or abnormal carbon monoxide levels and quickly informing relevant parties.

[0917] "Health status estimation" is a process that analyzes the worker's voice and facial expression data to determine their health status, such as fatigue and stress.

[0918] "Means of notifying the administrator" refers to a communication system device for informing the administrator of the estimated health status of the worker.

[0919] This invention provides a system that effectively supports the growth and safety of children, and monitors the health status of workers in a factory. The system consists of the following main components:

[0920] Audio capture and content generation

[0921] The system's terminal captures the voice spoken by the user (child or worker) using a microphone. This voice data is converted to text and sent to the server. The server uses a generative AI model to generate appropriate content (e.g., read-aloud stories, learning questions, or health assessments for workers) based on the input text data. For example, if a child says, "Tell me a story," the generated story is sent to the terminal and played back as audio.

[0922] Mental state assessment and notification

[0923] The device is equipped with voice capture capabilities and a camera to collect user conversation and facial expression data. This data is sent to a server, where a generative AI model is used to infer the user's mental state. For example, if a child says, "I had a bad day," the server analyzes the voice and facial expression to infer stress and anxiety, and notifies the parent's smartphone app. Similarly, in a factory, the voice and facial expressions of workers are captured to infer their health status and notify the manager.

[0924] Indoor environment monitoring and automatic adjustment

[0925] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and issues environmental adjustment instructions to the terminal as needed. For example, if the temperature is too high, it sends an instruction to the terminal to turn on the air conditioner, and the air conditioner starts operating.

[0926] Detection of dangerous situations and emergency notifications

[0927] The device is equipped with a fire alarm and carbon monoxide sensor, and if a dangerous situation is detected, it immediately notifies the server. The server quickly analyzes the situation and sends an emergency notification to the parent's or administrator's smartphone app. For example, if smoke is detected, it has a function to notify the parent that there is a possibility of fire.

[0928] Children's learning support

[0929] The server generates a new learning program based on the user's (child's) age data and past learning progress. This program is sent to the device, which then assists with learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[0930] Specific example

[0931] For example, if a worker in a factory says to a robot, "I'm tired today," the robot captures this voice and facial expression data and sends it to a server. The server uses a generative AI model to analyze the data, and if it determines that the worker's fatigue level is high, it notifies the administrator's tablet. Based on this information, the administrator can instruct the worker to take an appropriate break.

[0932] Example of a prompt

[0933] The prompt statement is designed as follows:

[0934] Create a robot app that uses a voice capture system to monitor the health of workers. The app will use voice recognition to analyze the worker's fatigue level and notify the manager. For example, if a worker says, "I'm tired today," the robot will capture this, use a generative AI model to estimate the fatigue level, and notify the manager's tablet.

[0935] Hardware and software

[0936] The hardware used includes microphones for voice capture, cameras for facial recognition, and temperature, humidity, and CO2 sensors, fire alarms, and carbon monoxide sensors for environmental monitoring. The software uses speech recognition libraries (e.g., speech_recognition), generative AI models (e.g., "GPT" and "BERT"), and web APIs for data analysis.

[0937] By combining these elements, it is possible to create a system that comprehensively supports children's development, ensures their safety, and monitors the health of factory workers.

[0938] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0939] Step 1:

[0940] Audio data capture

[0941] The device's microphone captures the user's (child or worker's) voice input. The captured audio data is temporarily stored on the device and then sent to the speech recognition engine.

[0942] Input: Audio data

[0943] Output: Audio data sent to the speech recognition engine

[0944] Step 2:

[0945] Converting audio data to text

[0946] The device converts the captured audio data into text using a speech recognition library. The converted text data is then sent to the server.

[0947] Input: Voice data sent to the speech recognition engine

[0948] Output: Text data sent to the server

[0949] Step 3:

[0950] Text data analysis and content generation

[0951] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., stories, learning questions, health assessments). GPT and BERT are used as generative AI models for this analysis.

[0952] Input: Text data sent to the server

[0953] Output: Content data generated using a generative AI model

[0954] Step 4:

[0955] Sending and playing content data

[0956] The generated content data is sent from the server to the terminal, and the terminal provides it to the user via audio using its audio playback function.

[0957] Input: Content data generated using a generation AI model.

[0958] Output: Audio content provided using the audio playback function.

[0959] Step 5:

[0960] Inference of mental state

[0961] The device uses cameras for voice and facial expression capture to collect user conversation and facial expression data. The collected data is sent to a server. The server uses a generative AI model to analyze voice tone, words used, and facial expression data to infer the user's mental state.

[0962] Input: Audio data and facial expression data collected by the camera and microphone.

[0963] Output: Estimated mental state data

[0964] Step 6:

[0965] Notification of mental state

[0966] The estimated mental state data is sent from the server to a smartphone app for parents and administrators. This allows parents and administrators to quickly understand the situation.

[0967] Input: Estimated mental state data

[0968] Output: Notification to the parent's or administrator's smartphone app

[0969] Step 7:

[0970] Environmental data monitoring

[0971] The device's temperature, humidity, and CO2 sensors periodically collect indoor environmental data and transmit it to the server.

[0972] Input: Environmental data from temperature sensor, humidity sensor, and CO2 sensor.

[0973] Output: Monitoring data sent to the server

[0974] Step 8:

[0975] Automatic environment adjustment

[0976] The server analyzes the monitoring data and sends instructions to the terminal to adjust the air conditioning and ventilation systems as needed. The terminal then automatically adjusts the environment according to these instructions.

[0977] Input: Monitoring data sent to the server

[0978] Output: Instructions for adjusting air conditioning and ventilation systems

[0979] Step 9:

[0980] Detection of dangerous situations and emergency notifications

[0981] The device is equipped with a fire alarm and a carbon monoxide sensor, and immediately sends data to a server if a dangerous situation is detected. The server quickly analyzes this data and sends an emergency notification to the parent's or administrator's smartphone app.

[0982] Input: Data from fire alarms and carbon monoxide sensors

[0983] Output: Sending emergency notification data to the parent's or administrator's smartphone app.

[0984] Step 10:

[0985] Health status monitoring and notification

[0986] When a worker in the factory speaks into a terminal, their voice and facial expression data are captured and sent to a server. The server uses a generative AI model to estimate the worker's health status and notifies the administrator of the results.

[0987] Input: Voice data and facial expression data of factory workers

[0988] Output: Estimated health status data and notification to administrator

[0989] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0990] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[0991] 1. Audio capture and content generation

[0992] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and understands the request. Next, a generative model generates appropriate content (stories or learning questions) and sends the audio data to the device. The device then plays the audio data to provide the user with storytelling or study assistance.

[0993] Specific example:

[0994] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[0995] 2. Understanding and sharing the child's mental state

[0996] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[0997] Specific example:

[0998] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[0999] 3. Monitoring and automatic adjustment of the indoor environment

[1000] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[1001] Specific example:

[1002] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1003] 4. Detection of dangerous situations and emergency notification

[1004] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1005] Specific example:

[1006] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1007] 5. Personalized learning tailored to the user's age.

[1008] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1009] Specific example:

[1010] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1011] 6. Use of the Emotion Engine

[1012] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[1013] Specific example:

[1014] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[1015] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[1016] The following describes the processing flow.

[1017] 1. Audio capture and content generation

[1018] Step 1:

[1019] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[1020] Step 2:

[1021] The device's microphone captures the user's voice.

[1022] Step 3:

[1023] The device converts the captured audio into text data and sends that data to the server.

[1024] Step 4:

[1025] The server analyzes the text data it receives to understand the content of the request.

[1026] Step 5:

[1027] The server's generative model generates appropriate content (stories and learning questions).

[1028] Step 6:

[1029] The server generates audio data and sends it to the terminal.

[1030] Step 7:

[1031] The device plays back received audio data to provide the user with read-alouds and study support.

[1032] 2. Understanding and sharing the child's mental state

[1033] Step 1:

[1034] The device's microphone and camera collect data on the user's conversation and facial expressions.

[1035] Step 2:

[1036] The device sends the collected data to the server.

[1037] Step 3:

[1038] The server's generative model analyzes the received data and infers the user's mental state based on voice tone and the words used.

[1039] Step 4:

[1040] The server's emotion engine analyzes voice and facial expression data to recognize the user's emotions.

[1041] Step 5:

[1042] The server notifies the parent's smartphone app of their estimated mental state and emotions.

[1043] 3. Monitoring and automatic adjustment of the indoor environment

[1044] Step 1:

[1045] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[1046] Step 2:

[1047] The device sends the collected data to the server.

[1048] Step 3:

[1049] The server analyzes the environmental data it receives and compares it to the set baseline values.

[1050] Step 4:

[1051] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[1052] Step 5:

[1053] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[1054] 4. Detection of dangerous situations and emergency notification

[1055] Step 1:

[1056] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[1057] Step 2:

[1058] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[1059] Step 3:

[1060] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[1061] 5. Personalized learning tailored to the user's age.

[1062] Step 1:

[1063] The server retrieves the user's lunar age data and past learning progress data.

[1064] Step 2:

[1065] The server's generation model generates a new learning program based on the lunar phase and past data.

[1066] Step 3:

[1067] The server sends the generated learning program to the terminal.

[1068] Step 4:

[1069] The device supports the user's learning through voice guidance based on the learning program.

[1070] Step 5:

[1071] The device periodically sends learning progress data to the server.

[1072] Step 6:

[1073] The server analyzes the received progress data and adjusts the learning program as needed.

[1074] 6. Use of the Emotion Engine

[1075] Step 1:

[1076] The device sends the collected audio and facial expression data to the server.

[1077] Step 2:

[1078] The server's emotion engine analyzes the data to recognize the user's emotions.

[1079] Step 3:

[1080] Based on the sentiment data recognized by the server, the generative model individually adjusts the content.

[1081] Step 4:

[1082] Based on the data in which the server recognizes emotions, it fine-tunes the content of the notifications sent to the parents.

[1083] The above outlines the specific processing steps for each function.

[1084] (Example 2)

[1085] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[1086] To effectively promote and monitor children's development and provide a safe and comfortable environment, a system is needed that collects and analyzes various information in real time and optimizes operations based on the results. Conventional systems often perform tasks such as voice capture, content generation, mental state assessment, emergency response, and environmental monitoring and adjustment individually, and few systems integrate these functions, making it difficult for parents to watch over their children with peace of mind. The present invention aims to provide a comprehensive system to solve these problems.

[1087] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data and facial expression data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for monitoring the indoor environment using temperature, humidity, and CO2 sensors; means for automatically adjusting the environment based on the monitoring data; and means for detecting dangerous situations using fire alarms and carbon monoxide sensors and notifying the parent of an emergency. This enables support for the child's growth, emotionally responsive care, environment optimization, and rapid response in emergencies.

[1088] "Voice capture" refers to recording the audio when a user speaks.

[1089] A "generative model" is an algorithm that generates text or content based on received data.

[1090] "Content generation" refers to creating information using a predefined algorithm.

[1091] "Content playback" refers to the process by which a device delivers generated information to the user.

[1092] "Conversation data" refers to audio data generated when a user speaks to the system.

[1093] "Facial expression data" refers to data recorded by a camera or other device capturing a user's facial expressions.

[1094] "Inferring mental state" refers to judging a user's emotions and psychological state based on collected data.

[1095] A "notification" is a message used to convey specific information to an administrator, such as a parent.

[1096] A "temperature sensor" is a device used to measure the temperature of a room.

[1097] A "humidity sensor" is a device used to measure the humidity in a room.

[1098] A "CO2 sensor" is a device used to measure the concentration of carbon dioxide in a room.

[1099] "Monitoring" refers to the continuous observation of specific environmental conditions and the collection of data.

[1100] "Automatic adjustment" refers to the automatic modification of the device's operation based on the received data.

[1101] "Detecting dangerous situations" means sensing emergencies such as fires or carbon monoxide poisoning.

[1102] An "emergency notification" is a system that sends out an immediate warning when danger is detected.

[1103] A "personalized learning program" refers to learning content tailored to each individual user.

[1104] "Progress data" refers to data that records the progress of learning or activities.

[1105] An "emotion engine" is a system that analyzes voice and facial expression data to recognize emotions.

[1106] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[1107] Audio capture and content generation

[1108] The user requests:

[1109] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice, converts the data into text, and sends it to the server. The server uses a generative AI model to analyze the voice data, understand the request, and generate appropriate content. This generated content is saved as text data, converted back into audio data, and sent back to the device. The device then plays the audio data, providing storytelling or study assistance.

[1110] Specific example:

[1111] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[1112] Examples of prompts to input into a generative AI model:

[1113] The user requested, "Tell me a story." Please generate a Japanese children's fairy tale.

[1114] Understanding and sharing a child's mental state

[1115] Device features:

[1116] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[1117] Specific example:

[1118] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[1119] Examples of prompts to input into a generative AI model:

[1120] The user said, "Something unpleasant happened at school today." Based on this statement, identify their emotions and determine whether they are feeling anxious.

[1121] Indoor environment monitoring and automatic adjustment

[1122] Device environment monitoring:

[1123] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[1124] Specific example:

[1125] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1126] Examples of prompts to input into a generative AI model:

[1127] The indoor temperature is over 30°C. Please generate a command to turn on the air conditioner.

[1128] Detection of dangerous situations and emergency notifications

[1129] Device risk detection function:

[1130] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1131] Specific example:

[1132] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1133] Examples of prompts to input into a generative AI model:

[1134] The device has detected smoke. Please generate an emergency notification message for your parents.

[1135] Personalized learning tailored to the child's age.

[1136] Server's learning program generation function:

[1137] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1138] Specific example:

[1139] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1140] Examples of prompts to input into a generative AI model:

[1141] Please generate a program that teaches the hiragana character "い" based on the user's age in months and past learning data.

[1142] Use of the emotion engine

[1143] Emotion recognition feature of the device:

[1144] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[1145] Specific example:

[1146] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[1147] Examples of prompts to input into a generative AI model:

[1148] The user said, "I had fun today!" Use the emotion engine to generate content that the child will find enjoyable, once the emotion of joy is recognized.

[1149] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[1150] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1151] Step 1:

[1152] The user requests:

[1153] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice and collects that audio data. Specifically, when the user says "Tell me a story," the microphone acquires audio data, and that data is processed in the next step.

[1154] Input: User's voice

[1155] Output: Collected audio data

[1156] Step 2:

[1157] Converting audio data to text:

[1158] The device uses speech recognition software to convert the captured audio data into text data. Specifically, the speech recognition engine analyzes the audio and generates text data such as "Tell me your story." This text data is then sent to the server in the next step.

[1159] Input: Audio data

[1160] Output: Text data

[1161] Step 3:

[1162] Sending text data:

[1163] The terminal sends the converted text data to the server. The server receives the text data and prepares it for analysis. Specifically, the text data is sent to the server via the internet.

[1164] Input: Text data

[1165] Output: Text data sent to the server

[1166] Step 4:

[1167] Content generation:

[1168] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., fairy tales or learning questions). Specifically, the generative AI model processes a request like "Tell me a story" and generates a Japanese fairy tale. The generated text data is then converted into audio data.

[1169] Input: Text data

[1170] Output: Generated content (text data)

[1171] Step 5:

[1172] Generating audio data:

[1173] The generated text data is converted into speech data by the server's text-to-speech engine. Specifically, the text-to-speech engine generates speech data such as "Once upon a time, in a certain place..." This speech data is then sent to the terminal in the next step.

[1174] Input: Generated content (text data)

[1175] Output: Audio data

[1176] Step 6:

[1177] Sending audio data:

[1178] The server sends audio data to the terminal. The terminal prepares to play the received audio data. Specifically, the audio data is sent to the terminal via the internet.

[1179] Input: Audio data

[1180] Output: Audio data sent to the terminal

[1181] Step 7:

[1182] Play content:

[1183] The device plays the audio data it receives and provides it to the user. Specifically, the device's speaker plays the audio "Once upon a time, in a certain place..." This allows the user to listen to a fairy tale.

[1184] Input: Audio data sent to the device

[1185] Output: Audio data that the user will listen to.

[1186] ---

[1187] The above outlines the specific processing flow for voice capture and content generation in this system. This allows users to seamlessly receive content tailored to their requests.

[1188] (Application Example 2)

[1189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1190] In modern society, ensuring family safety and a comfortable shopping experience in physical stores is a crucial challenge. However, existing systems struggle to comprehensively address issues such as providing engaging content for children, understanding their emotional state, ensuring a comfortable environment, and responding quickly to emergencies. Furthermore, there is a lack of means for parents to know their child's current location and status in real time. This leads to problems where parental peace of mind and child safety are not adequately ensured.

[1191] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for emergency notification to the parent when a dangerous situation is detected; means for understanding the location information of the parent and child within the physical store and ensuring the child's safety; and means for providing various information within the store. This makes it possible for families to enjoy shopping safely and comfortably in physical stores.

[1192] "Voice capture" is a technology that captures the voice of a child when they speak and converts it into data that can be analyzed.

[1193] A "generative model" is an algorithm used to generate content based on input data.

[1194] "Read-aloud content" refers to stories and narratives intended to be read aloud to children.

[1195] "Study content" refers to problems and learning materials created to support children's learning.

[1196] "User conversation data" refers to audio and text data recorded during interactions with children.

[1197] "Inferring mental state" refers to judging a user's emotions and mood from their voice and facial expressions.

[1198] "Notification to parents" refers to a means of communication to inform parents about their child's presumed mental state or emergency situation.

[1199] "Indoor environment monitoring" refers to the act of measuring indoor environmental information such as temperature, humidity, and carbon dioxide concentration using sensors.

[1200] "Automatic environmental adjustment" refers to the automatic operation of devices such as air conditioners and humidifiers based on monitoring data to maintain an appropriate indoor environment.

[1201] "Detection of dangerous situations" refers to the act of using sensors to detect emergencies such as fires or carbon monoxide emissions.

[1202] "Emergency notification" refers to a notification function that quickly informs parents when a dangerous situation occurs.

[1203] "Location tracking" refers to technology that tracks and records a child's current location within a physical store.

[1204] "Means for ensuring safety" refers to systems that use location information to prevent children from getting lost and to respond to emergencies.

[1205] "In-store information guidance" refers to a function that informs users of information provided within the store, such as the location of restrooms and the location of sales areas.

[1206] In order to implement this invention, the following system configuration and processing are necessary.

[1207] System configuration:

[1208] 1. Voice capture device: This includes a microphone for capturing the child's voice and converting it into digital data.

[1209] 2. Generative Model Server: Receives audio data as text data, analyzes it, and generates appropriate content (stories and learning questions). For this purpose, it uses a database and AI algorithms (e.g., generative AI models such as GPT).

[1210] 3. Content playback device: Equipped with speakers for playing the generated content as audio.

[1211] 4. Emotion Recognition Engine: A device that analyzes conversation data and facial expression data with a user to infer their mental state, and includes a camera and emotion recognition software (e.g., emotion recognition API).

[1212] 5. Notification system: Equipped with a communication module to notify the parent's smartphone of the suspected mental state.

[1213] 6. Environmental monitoring sensors: Equipped with sensors (e.g., DHT22 or MQ-135 sensors) for measuring indoor temperature, humidity, CO2 concentration, etc.

[1214] 7. Automatic environmental control device: A device for controlling environmental control devices such as air conditioners and humidifiers.

[1215] 8. Emergency notification system: A device that detects fires, carbon monoxide emissions, and other emergencies and sends emergency notifications to the parent's smartphone.

[1216] 9. Location tracking system: GPS device and tracking software for determining a child's current location in real time.

[1217] Processing procedure:

[1218] This system supports children's safety and development through a multi-layered process.

[1219] 1. Audio capture and content generation:

[1220] When the user (child) says something like "Tell me a story" or "Help me with my studies," the device's microphone captures the audio and sends it to the generative model server.

[1221] The server analyzes the audio data as text data and generates appropriate content using a generative model.

[1222] For example, if a user says, "Tell me a story," the server generates a fairy tale, and the device plays the audio data.

[1223] As an example of a prompt message using a generative AI model, you can input the instruction, "Generate a fairy tale suitable for children."

[1224] 2. Emotion recognition and notification to parents:

[1225] The system collects conversation data and facial expression data from users and sends them to the server.

[1226] The server uses an emotion recognition engine to analyze the data and infer the user's mental state.

[1227] If the estimated mental state is "anxiety" or "joy," a notification will be sent to the parent's smartphone stating, "Your child is feeling anxious" or "Your child is very happy."

[1228] For example, if a user says, "Something unpleasant happened at school today...", the device sends the audio and facial expression data to the server, infers the user's anxiety, and sends a notification.

[1229] 3. Monitoring and automatic adjustment of the indoor environment:

[1230] The device measures temperature, humidity, and CO2 concentration at regular intervals and sends the data to the server.

[1231] The server analyzes the data and sends adjustment instructions to the terminal to maintain a comfortable environment.

[1232] For example, when the server detects a room temperature of 30°C, it sends a command to the terminal to turn on the air conditioner, and the terminal follows that command to lower the room temperature.

[1233] 4. Detection of dangerous situations and emergency notification:

[1234] If the terminal detects a fire or an increase in carbon monoxide concentration, it will immediately send data to the server.

[1235] The server analyzes the information and sends a notification to the parent's smartphone saying, "Emergency: Potential fire."

[1236] For example, one possible implementation is to immediately send an emergency notification to the parent when smoke is detected.

[1237] 5. Location information tracking and guidance within physical stores:

[1238] The system tracks the child's location in real time and displays it on the parent's smartphone.

[1239] Parents can know their child's current location, allowing them to enjoy shopping with peace of mind.

[1240] For example, if a child says, "I need to go to the toilet," the system can guide them to the location of the toilet within the store.

[1241] In summary, this invention comprehensively supports the safety and development of children and provides parents with an environment in which they can feel at ease.

[1242] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1243] Step 1:

[1244] Voice capture and text conversion:

[1245] The user (child) speaks aloud, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures this audio, and voice capture software converts the audio data into text data. The input is audio data, and the output is text data.

[1246] Step 2:

[1247] Sending text data to the server:

[1248] The terminal sends the captured text data to the server. The input is text data, and the output is the data sent to the server. Specifically, the data is sent to the server using an HTTP request.

[1249] Step 3:

[1250] Content generation:

[1251] The server analyzes the received text data. Using a generative AI model (e.g., GPT-3), it generates appropriate content (fairy tales or learning questions) in response to the request. During this process, prompts are input to the generative AI model, and the generated stories or questions are output as prompt results. For example, the prompt "Generate a fairy tale suitable for children" might be used.

[1252] Step 4:

[1253] Converting generated content into audio data:

[1254] The generated content is converted into audio data. The server uses a text-to-speech (TTS) engine to convert the generated text content into audio data. The input is the generated text data, and the output is audio data.

[1255] Step 5:

[1256] Sending audio data to the terminal:

[1257] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the data sent to the terminal. Specifically, the data is sent to the terminal using an HTTP response.

[1258] Step 6:

[1259] Play content:

[1260] The device plays the received audio data. It uses a speaker to let the user (child) listen to the content aloud. The input is audio data, and the output is audio playback. For example, the device starts reading a fairy tale, such as "Once upon a time, in a certain place..."

[1261] Step 7:

[1262] Collection and inference of emotional data:

[1263] The system collects conversation data and facial expression data from the user using the device's microphone and camera. This data is sent to a server, which analyzes it using an emotion recognition engine. The input is voice and facial expression data, and the output is estimated emotion data. For example, from a conversation like, "Something unpleasant happened at school today...", the system might infer "anxiety".

[1264] Step 8:

[1265] Notification to parents regarding their child's mental state:

[1266] The server sends a notification to the parent based on the inferred emotion data. The notification is sent via a smartphone application. The input is emotion data, and the output is a notification message. For example, a notification saying "Your child is feeling anxious" is sent.

[1267] Step 9:

[1268] Environmental data monitoring:

[1269] The terminal is equipped with sensors to measure temperature, humidity, and CO2 concentration. This environmental data is transmitted to the server at regular intervals. The input is sensor data, and the output is data transmitted to the server.

[1270] Step 10:

[1271] Automatic environment adjustment:

[1272] The server analyzes the received environmental data and issues instructions for necessary adjustments. It remotely controls devices such as air conditioners and humidifiers to automatically adjust the environment. The input is environmental data, and the output is instructions for operating the devices. For example, if it detects a room temperature of 30°C, it will issue an instruction to turn on the air conditioner.

[1273] Step 11:

[1274] Detection of dangerous situations and emergency notifications:

[1275] The device has sensors to detect fire and elevated carbon monoxide levels. When these dangerous situations are detected, it immediately sends data to a server. The server then sends an emergency notification to the parent's smartphone. The input is sensor data, and the output is an emergency notification message. For example, if smoke is detected, a notification such as "Emergency: Potential fire" is sent.

[1276] Step 12:

[1277] Location tracking and guidance:

[1278] The device tracks the child's current location in real time and displays the location information on the parent's smartphone. The input is GPS data, and the output is a display of location information. For example, if the child says, "I need to go to the toilet," the system will guide them to the location of the toilet.

[1279] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1280] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1281] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1282] [Third Embodiment]

[1283] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1284] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1285] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1286] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1287] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1288] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1289] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1290] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1291] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1292] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1293] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1294] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1295] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications. The embodiments of each function in this system are described below.

[1296] 1. Audio capture and content generation

[1297] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and uses a generative model to generate appropriate content (e.g., a story or study question). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[1298] Specific example:

[1299] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[1300] 2. Understanding and sharing the child's mental state

[1301] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative model to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety and notifies the parent.

[1302] Specific example:

[1303] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[1304] 3. Monitoring and automatic adjustment of the indoor environment

[1305] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[1306] Specific example:

[1307] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1308] 4. Detection of dangerous situations and emergency notification

[1309] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1310] Specific example:

[1311] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1312] 5. Personalized learning tailored to the user's age.

[1313] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1314] Specific example:

[1315] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1316] The above describes the embodiments of the present invention. By using this system, support for children's growth and ensuring their safety can be effectively achieved.

[1317] The following describes the processing flow.

[1318] 1. Audio capture and content generation

[1319] Step 1:

[1320] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[1321] Step 2:

[1322] The device's microphone captures the user's voice.

[1323] Step 3:

[1324] The device converts the captured audio into text data and sends that data to the server.

[1325] Step 4:

[1326] The server analyzes the text data it receives to understand the content of the request.

[1327] Step 5:

[1328] The server's generative model generates appropriate content (stories and learning questions).

[1329] Step 6:

[1330] The server generates audio data and sends it to the terminal.

[1331] Step 7:

[1332] The device plays back received audio data to provide the user with read-alouds and study support.

[1333] 2. Understanding and sharing the child's mental state

[1334] Step 1:

[1335] The device's microphone and camera collect data on the user's conversation and facial expressions.

[1336] Step 2:

[1337] The device sends the collected data to the server.

[1338] Step 3:

[1339] The server analyzes the received data and infers the user's mental state based on their tone of voice, the words they use, and their facial expressions.

[1340] Step 4:

[1341] The server notifies the parent's smartphone app of the estimated mental state.

[1342] 3. Monitoring and automatic adjustment of the indoor environment

[1343] Step 1:

[1344] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[1345] Step 2:

[1346] The device sends the collected data to the server.

[1347] Step 3:

[1348] The server analyzes the environmental data it receives and compares it to the set baseline values.

[1349] Step 4:

[1350] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[1351] Step 5:

[1352] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[1353] 4. Detection of dangerous situations and emergency notification

[1354] Step 1:

[1355] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[1356] Step 2:

[1357] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[1358] Step 3:

[1359] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[1360] 5. Personalized learning tailored to the user's age.

[1361] Step 1:

[1362] The server retrieves the user's lunar age data and past learning progress data.

[1363] Step 2:

[1364] The server's generation model generates a new learning program based on the lunar phase and past data.

[1365] Step 3:

[1366] The server sends the generated learning program to the terminal.

[1367] Step 4:

[1368] The device supports the user's learning through voice guidance based on the learning program.

[1369] Step 5:

[1370] The device periodically sends learning progress data to the server.

[1371] Step 6:

[1372] The server analyzes the received progress data and adjusts the learning program as needed.

[1373] The above outlines the specific processing steps for each function.

[1374] (Example 1)

[1375] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1376] In modern families, there is a need for systems that effectively support children's development and allow them to grow up in a safe and secure environment. In particular, support for children's learning, monitoring of their mental state, monitoring of the indoor environment, and rapid response in emergencies are crucial. However, existing systems struggle to provide these functions in a unified and effective manner.

[1377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1378] In this invention, the server includes means for capturing the user's voice and performing speech recognition, means for converting the recognized voice into text data, means for inputting the text data into a generative AI model and generating appropriate content, means for synthesizing the generated voice content, means for transmitting the synthesized content to a terminal, means for collecting conversation data and facial expression data with the user and inferring their mental state, means for notifying the parent of the inferred mental state, means for using sensors to monitor the indoor environment, means for automatically adjusting the environment based on the monitoring data, means for sending an emergency notification to the parent when a dangerous situation is detected, and means for playing the content regenerated by the generative model on the terminal. This enables seamless support for the child's development, safety, and learning.

[1379] "Children" refers to the youngest members of a family and users of this system.

[1380] "Methods for capturing audio" refer to methods of collecting user-generated audio as digital data using microphones, speech recognition devices, etc.

[1381] A "generative model" refers to an algorithm or system that automatically generates new content using machine learning or artificial intelligence technologies.

[1382] "Reading aloud" refers to the act of providing stories or educational content to users in audio format.

[1383] "Study content" refers to educational materials such as problems and explanations provided to support user learning.

[1384] "Means for playing generated content" refers to methods of providing users with viewable digital content generated by a generative model.

[1385] "User conversation data" refers to linguistic information that users communicate with the system, and is collected by a speech recognition system.

[1386] "Facial expression data" refers to data about a user's facial expressions collected using cameras and sensors.

[1387] "Means of inferring mental state" refers to algorithms and systems that analyze collected voice data and facial expression data to estimate the user's emotions and mental state.

[1388] "Methods for notifying parents" refers to methods of sending notifications to parents' smartphones or tablets to inform them of the user's mental state or emergency situation.

[1389] A "sensor for monitoring the indoor environment" refers to a sensor device used to measure temperature, humidity, CO2 concentration, etc.

[1390] "Means of automatically adjusting the environment" refers to methods of controlling environmental devices such as air conditioners and humidifiers based on monitoring data.

[1391] A "dangerous situation" refers to any condition that could potentially threaten the safety of the user or those around them, such as a fire or a carbon monoxide leak.

[1392] "Means of emergency notification" refers to a method of promptly sending notifications to parents or appropriate personnel when a dangerous situation is detected.

[1393] "Speech recognition" refers to a technology that analyzes a user's voice as digital data and converts it into text data.

[1394] "Text data" refers to data that stores speech information analyzed by a speech recognition system as text.

[1395] "Speech synthesis" refers to the technology that generates natural-sounding speech based on text data.

[1396] A "generative AI model" refers to a model that uses artificial intelligence algorithms to automatically generate new content such as speech and text.

[1397] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications.

[1398] Audio capture and content generation

[1399] When a user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies," the device's microphone captures the user's voice and sends that voice as text data to the server. The server analyzes the received text data and uses a generative AI model to generate appropriate content (for example, a story or study problem). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[1400] Specific example:

[1401] When the user says, "Tell me a story," the device captures the voice and sends it to the server. The server uses a generative AI model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, "Once upon a time, in a certain place..."

[1402] Example of a prompt:

[1403] "Please create a short children's story that children will enjoy listening to. The theme is animal friendship."

[1404] Understanding and sharing a child's mental state

[1405] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative AI model (e.g., an emotion analysis model) to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety.

[1406] Specific example:

[1407] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[1408] Indoor environment monitoring and automatic adjustment

[1409] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[1410] Specific example:

[1411] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1412] Detection of dangerous situations and emergency notifications

[1413] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1414] Specific example:

[1415] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1416] Personalized learning tailored to the child's age.

[1417] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1418] Specific example:

[1419] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1420] The above describes the embodiments of the present invention. By using this system, support for child development and safety can be achieved effectively and in a unified manner.

[1421] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1422] Audio capture and content generation

[1423] Step 1:

[1424] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[1425] Input: User's voice

[1426] Output: Captured audio data

[1427] Step 2:

[1428] The device's microphone captures the user's voice and saves that audio as digital data.

[1429] Input: Captured audio data

[1430] Output: Digital audio data

[1431] Step 3:

[1432] The device sends the captured digital audio data to the server.

[1433] Input: Digital audio data

[1434] Output: Data to send to the server

[1435] Step 4:

[1436] The server uses speech recognition software (e.g., Google Speech-to-Text) to convert the audio data into text data.

[1437] Input: Audio data

[1438] Output: Text data

[1439] Step 5:

[1440] The server inputs text data into an AI model (e.g., GPT-4) and generates appropriate content as a prompt.

[1441] Input: Text data

[1442] Output: Generated text content

[1443] Step 6:

[1444] The server converts the generated text content into speech data using speech synthesis software (e.g., Amazon Polly).

[1445] Input: Generated text content

[1446] Output: Generated audio data

[1447] Step 7:

[1448] The server sends the generated audio data to the terminal.

[1449] Input: Generated audio data

[1450] Output: Data to send to the terminal

[1451] Step 8:

[1452] The device plays back the received audio data, providing the user with read-alouds and study support.

[1453] Input: Generated audio data

[1454] Output: Audio to be played back to the user

[1455] Understanding and sharing a child's mental state

[1456] Step 1:

[1457] The device's microphone and camera capture the user's voice and facial expression data.

[1458] Input: User's voice and facial expressions

[1459] Output: Audio data and facial expression data

[1460] Step 2:

[1461] The device sends the captured audio and facial expression data to the server.

[1462] Input: Voice data and facial expression data

[1463] Output: Data to send to the server

[1464] Step 3:

[1465] The server uses voice analysis software and emotion analysis models (e.g., IBM Watson's emotion analysis) to analyze voice data and facial expression data and infer the user's mental state.

[1466] Input: Voice data and facial expression data

[1467] Output: Inferred mental state

[1468] Step 4:

[1469] The server notifies the parent's smartphone app of the prediction result.

[1470] Input: Estimated mental state

[1471] Output: Notification data sent to the parent's smartphone app

[1472] Indoor environment monitoring and automatic adjustment

[1473] Step 1:

[1474] The device acquires indoor environmental data using temperature, humidity, and CO2 sensors installed on the terminal.

[1475] Input: Indoor environment

[1476] Output: Environmental sensor data

[1477] Step 2:

[1478] The device sends the environmental sensor data it acquires to the server.

[1479] Input: Environmental sensor data

[1480] Output: Data to send to the server

[1481] Step 3:

[1482] The server analyzes environmental sensor data and generates necessary adjustment instructions.

[1483] Input: Environmental sensor data

[1484] Output: Environmental adjustment instructions

[1485] Step 4:

[1486] The server sends the generated adjustment instructions to the terminal.

[1487] Input: Environmental adjustment instructions

[1488] Output: Data to send to the terminal

[1489] Step 5:

[1490] The terminal follows the adjustment instructions and operates devices such as air conditioners and humidifiers to adjust the indoor environment.

[1491] Input: Environmental adjustment instructions

[1492] Output: Adjusted indoor environment

[1493] Detection of dangerous situations and emergency notifications

[1494] Step 1:

[1495] The device's built-in fire alarm and carbon monoxide sensor detect dangerous situations.

[1496] Input: Indoor environment

[1497] Output: Hazard detection data

[1498] Step 2:

[1499] The device sends the detected data to the server.

[1500] Input: Hazard detection data

[1501] Output: Data to send to the server

[1502] Step 3:

[1503] The server quickly analyzes the data and generates an emergency notification.

[1504] Input: Hazard detection data

[1505] Output: Emergency notification data

[1506] Step 4:

[1507] The server sends an emergency notification to the parent's smartphone app, prompting them to take immediate action.

[1508] Input: Emergency notification data

[1509] Output: Notification data sent to the parent's smartphone app

[1510] Personalized learning tailored to the child's age.

[1511] Step 1:

[1512] The server retrieves the user's lunar age data and past learning progress data.

[1513] Input: Lunar age data and learning progress data

[1514] Output: Analysis data

[1515] Step 2:

[1516] The server uses the generated AI model to create a new learning program.

[1517] Input: Analysis data

[1518] Output: New learning program

[1519] Step 3:

[1520] The server sends the generated learning program to the terminal.

[1521] Input: New learning program

[1522] Output: Data to send to the terminal

[1523] Step 4:

[1524] The device provides users with learning support content through voice guidance.

[1525] Input: New learning program

[1526] Output: Learning content for the user

[1527] Step 5:

[1528] The device feeds back user learning progress data to the server, which is then used to inform the next learning session.

[1529] Input: Learning progress data

[1530] Output: Feedback data to the server

[1531] (Application Example 1)

[1532] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1533] Supporting the development and ensuring the safety of children in modern society are crucial challenges. Existing systems do not adequately address the detailed understanding of children's mental states, provide appropriate learning support, or manage their environment. Furthermore, monitoring the health of factory workers is insufficient, making it difficult to provide appropriate breaks and preventative measures in a timely manner. To solve these problems, this invention aims to provide a system that effectively supports the development and safety of children, as well as monitors the health of factory workers.

[1534] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1535] In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or educational content using a generative model; means for playing the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for making an emergency notification to the parent when a dangerous situation is detected; and means for capturing the voice and facial expression data of a worker when they speak, inferring the worker's health status using a generative model, and notifying the manager. This enables support for the development and safety of children, as well as monitoring the health status of workers in the factory and taking appropriate action.

[1536] "Voice capture" is the process of collecting the voice of a user when they speak using an input device such as a microphone.

[1537] A "generative model" is a system that uses machine learning or artificial intelligence algorithms to generate new content or information based on input data.

[1538] "Content generation" refers to the process of creating new read-aloud or educational content based on collected data.

[1539] "Means of playback" refers to audio output devices or programs that allow users to listen to the generated content.

[1540] "Mental state estimation" is a process that analyzes voice and facial expression data to determine the user's emotions and psychological state.

[1541] "Means of notifying parents" refers to communication systems and devices used to inform parents or guardians of suspected mental states or emergencies.

[1542] "Indoor environment monitoring" is the process of measuring and collecting environmental data such as temperature, humidity, and CO2 using sensors.

[1543] "Means for automatically adjusting the environment" refers to a system that automatically operates air conditioners and ventilation devices based on data about the indoor environment.

[1544] "Means of emergency notification" refers to means of detecting dangerous situations such as fires or abnormal carbon monoxide levels and quickly informing relevant parties.

[1545] "Health status estimation" is a process that analyzes the worker's voice and facial expression data to determine their health status, such as fatigue and stress.

[1546] "Means of notifying the administrator" refers to a communication system device for informing the administrator of the estimated health status of the worker.

[1547] This invention provides a system that effectively supports the growth and safety of children, and monitors the health status of workers in a factory. The system consists of the following main components:

[1548] Audio capture and content generation

[1549] The system's terminal captures the voice spoken by the user (child or worker) using a microphone. This voice data is converted to text and sent to the server. The server uses a generative AI model to generate appropriate content (e.g., read-aloud stories, learning questions, or health assessments for workers) based on the input text data. For example, if a child says, "Tell me a story," the generated story is sent to the terminal and played back as audio.

[1550] Mental state assessment and notification

[1551] The device is equipped with voice capture capabilities and a camera to collect user conversation and facial expression data. This data is sent to a server, where a generative AI model is used to infer the user's mental state. For example, if a child says, "I had a bad day," the server analyzes the voice and facial expression to infer stress and anxiety, and notifies the parent's smartphone app. Similarly, in a factory, the voice and facial expressions of workers are captured to infer their health status and notify the manager.

[1552] Indoor environment monitoring and automatic adjustment

[1553] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and issues environmental adjustment instructions to the terminal as needed. For example, if the temperature is too high, it sends an instruction to the terminal to turn on the air conditioner, and the air conditioner starts operating.

[1554] Detection of dangerous situations and emergency notifications

[1555] The device is equipped with a fire alarm and carbon monoxide sensor, and if a dangerous situation is detected, it immediately notifies the server. The server quickly analyzes the situation and sends an emergency notification to the parent's or administrator's smartphone app. For example, if smoke is detected, it has a function to notify the parent that there is a possibility of fire.

[1556] Children's learning support

[1557] The server generates a new learning program based on the user's (child's) age data and past learning progress. This program is sent to the device, which then assists with learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1558] Specific example

[1559] For example, if a worker in a factory says to a robot, "I'm tired today," the robot captures this voice and facial expression data and sends it to a server. The server uses a generative AI model to analyze the data, and if it determines that the worker's fatigue level is high, it notifies the administrator's tablet. Based on this information, the administrator can instruct the worker to take an appropriate break.

[1560] Example of a prompt

[1561] The prompt statement is designed as follows:

[1562] Create a robot app that uses a voice capture system to monitor the health of workers. The app will use voice recognition to analyze the worker's fatigue level and notify the manager. For example, if a worker says, "I'm tired today," the robot will capture this, use a generative AI model to estimate the fatigue level, and notify the manager's tablet.

[1563] Hardware and software

[1564] The hardware used includes microphones for voice capture, cameras for facial recognition, and temperature, humidity, and CO2 sensors, fire alarms, and carbon monoxide sensors for environmental monitoring. The software uses speech recognition libraries (e.g., speech_recognition), generative AI models (e.g., "GPT" and "BERT"), and web APIs for data analysis.

[1565] By combining these elements, it is possible to create a system that comprehensively supports children's development, ensures their safety, and monitors the health of factory workers.

[1566] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1567] Step 1:

[1568] Audio data capture

[1569] The device's microphone captures the user's (child or worker's) voice input. The captured audio data is temporarily stored on the device and then sent to the speech recognition engine.

[1570] Input: Audio data

[1571] Output: Audio data sent to the speech recognition engine

[1572] Step 2:

[1573] Converting audio data to text

[1574] The device converts the captured audio data into text using a speech recognition library. The converted text data is then sent to the server.

[1575] Input: Voice data sent to the speech recognition engine

[1576] Output: Text data sent to the server

[1577] Step 3:

[1578] Text data analysis and content generation

[1579] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., stories, learning questions, health assessments). GPT and BERT are used as generative AI models for this analysis.

[1580] Input: Text data sent to the server

[1581] Output: Content data generated using a generative AI model

[1582] Step 4:

[1583] Sending and playing content data

[1584] The generated content data is sent from the server to the terminal, and the terminal provides it to the user via audio using its audio playback function.

[1585] Input: Content data generated using a generation AI model.

[1586] Output: Audio content provided using the audio playback function.

[1587] Step 5:

[1588] Inference of mental state

[1589] The device uses cameras for voice and facial expression capture to collect user conversation and facial expression data. The collected data is sent to a server. The server uses a generative AI model to analyze voice tone, words used, and facial expression data to infer the user's mental state.

[1590] Input: Audio data and facial expression data collected by the camera and microphone.

[1591] Output: Estimated mental state data

[1592] Step 6:

[1593] Notification of mental state

[1594] The estimated mental state data is sent from the server to a smartphone app for parents and administrators. This allows parents and administrators to quickly understand the situation.

[1595] Input: Estimated mental state data

[1596] Output: Notification to the parent's or administrator's smartphone app

[1597] Step 7:

[1598] Environmental data monitoring

[1599] The device's temperature, humidity, and CO2 sensors periodically collect indoor environmental data and transmit it to the server.

[1600] Input: Environmental data from temperature sensor, humidity sensor, and CO2 sensor.

[1601] Output: Monitoring data sent to the server

[1602] Step 8:

[1603] Automatic environment adjustment

[1604] The server analyzes the monitoring data and sends instructions to the terminal to adjust the air conditioning and ventilation systems as needed. The terminal then automatically adjusts the environment according to these instructions.

[1605] Input: Monitoring data sent to the server

[1606] Output: Instructions for adjusting air conditioning and ventilation systems

[1607] Step 9:

[1608] Detection of dangerous situations and emergency notifications

[1609] The device is equipped with a fire alarm and a carbon monoxide sensor, and immediately sends data to a server if a dangerous situation is detected. The server quickly analyzes this data and sends an emergency notification to the parent's or administrator's smartphone app.

[1610] Input: Data from fire alarms and carbon monoxide sensors

[1611] Output: Sending emergency notification data to the parent's or administrator's smartphone app.

[1612] Step 10:

[1613] Health status monitoring and notification

[1614] When a worker in the factory speaks into a terminal, their voice and facial expression data are captured and sent to a server. The server uses a generative AI model to estimate the worker's health status and notifies the administrator of the results.

[1615] Input: Voice data and facial expression data of factory workers

[1616] Output: Estimated health status data and notification to administrator

[1617] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1618] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[1619] 1. Audio capture and content generation

[1620] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and understands the request. Next, a generative model generates appropriate content (stories or learning questions) and sends the audio data to the device. The device then plays the audio data to provide the user with storytelling or study assistance.

[1621] Specific example:

[1622] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[1623] 2. Understanding and sharing the child's mental state

[1624] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[1625] Specific example:

[1626] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[1627] 3. Monitoring and automatic adjustment of the indoor environment

[1628] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[1629] Specific example:

[1630] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1631] 4. Detection of dangerous situations and emergency notification

[1632] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1633] Specific example:

[1634] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1635] 5. Personalized learning tailored to the user's age.

[1636] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1637] Specific example:

[1638] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1639] 6. Use of the Emotion Engine

[1640] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[1641] Specific example:

[1642] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[1643] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[1644] The following describes the processing flow.

[1645] 1. Audio capture and content generation

[1646] Step 1:

[1647] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[1648] Step 2:

[1649] The device's microphone captures the user's voice.

[1650] Step 3:

[1651] The device converts the captured audio into text data and sends that data to the server.

[1652] Step 4:

[1653] The server analyzes the text data it receives to understand the content of the request.

[1654] Step 5:

[1655] The server's generative model generates appropriate content (stories and learning questions).

[1656] Step 6:

[1657] The server generates audio data and sends it to the terminal.

[1658] Step 7:

[1659] The device plays back received audio data to provide the user with read-alouds and study support.

[1660] 2. Understanding and sharing the child's mental state

[1661] Step 1:

[1662] The device's microphone and camera collect data on the user's conversation and facial expressions.

[1663] Step 2:

[1664] The device sends the collected data to the server.

[1665] Step 3:

[1666] The server's generative model analyzes the received data and infers the user's mental state based on voice tone and the words used.

[1667] Step 4:

[1668] The server's emotion engine analyzes voice and facial expression data to recognize the user's emotions.

[1669] Step 5:

[1670] The server notifies the parent's smartphone app of their estimated mental state and emotions.

[1671] 3. Monitoring and automatic adjustment of the indoor environment

[1672] Step 1:

[1673] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[1674] Step 2:

[1675] The device sends the collected data to the server.

[1676] Step 3:

[1677] The server analyzes the environmental data it receives and compares it to the set baseline values.

[1678] Step 4:

[1679] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[1680] Step 5:

[1681] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[1682] 4. Detection of dangerous situations and emergency notification

[1683] Step 1:

[1684] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[1685] Step 2:

[1686] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[1687] Step 3:

[1688] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[1689] 5. Personalized learning tailored to the user's age.

[1690] Step 1:

[1691] The server retrieves the user's lunar age data and past learning progress data.

[1692] Step 2:

[1693] The server's generation model generates a new learning program based on the lunar phase and past data.

[1694] Step 3:

[1695] The server sends the generated learning program to the terminal.

[1696] Step 4:

[1697] The device supports the user's learning through voice guidance based on the learning program.

[1698] Step 5:

[1699] The device periodically sends learning progress data to the server.

[1700] Step 6:

[1701] The server analyzes the received progress data and adjusts the learning program as needed.

[1702] 6. Use of the Emotion Engine

[1703] Step 1:

[1704] The device sends the collected audio and facial expression data to the server.

[1705] Step 2:

[1706] The server's emotion engine analyzes the data to recognize the user's emotions.

[1707] Step 3:

[1708] Based on the sentiment data recognized by the server, the generative model individually adjusts the content.

[1709] Step 4:

[1710] Based on the data in which the server recognizes emotions, it fine-tunes the content of the notifications sent to the parents.

[1711] The above outlines the specific processing steps for each function.

[1712] (Example 2)

[1713] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1714] To effectively promote and monitor children's development and provide a safe and comfortable environment, a system is needed that collects and analyzes various information in real time and optimizes operations based on the results. Conventional systems often perform tasks such as voice capture, content generation, mental state assessment, emergency response, and environmental monitoring and adjustment individually, and few systems integrate these functions, making it difficult for parents to watch over their children with peace of mind. The present invention aims to provide a comprehensive system to solve these problems.

[1715] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data and facial expression data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for monitoring the indoor environment using temperature, humidity, and CO2 sensors; means for automatically adjusting the environment based on the monitoring data; and means for detecting dangerous situations using fire alarms and carbon monoxide sensors and notifying the parent of an emergency. This enables support for the child's growth, emotionally responsive care, environment optimization, and rapid response in emergencies.

[1716] "Voice capture" refers to recording the audio when a user speaks.

[1717] A "generative model" is an algorithm that generates text or content based on received data.

[1718] "Content generation" refers to creating information using a predefined algorithm.

[1719] "Content playback" refers to the process by which a device delivers generated information to the user.

[1720] "Conversation data" refers to audio data generated when a user speaks to the system.

[1721] "Facial expression data" refers to data recorded by a camera or other device capturing a user's facial expressions.

[1722] "Inferring mental state" refers to judging a user's emotions and psychological state based on collected data.

[1723] A "notification" is a message used to convey specific information to an administrator, such as a parent.

[1724] A "temperature sensor" is a device used to measure the temperature of a room.

[1725] A "humidity sensor" is a device used to measure the humidity in a room.

[1726] A "CO2 sensor" is a device used to measure the concentration of carbon dioxide in a room.

[1727] "Monitoring" refers to the continuous observation of specific environmental conditions and the collection of data.

[1728] "Automatic adjustment" refers to the automatic modification of the device's operation based on the received data.

[1729] "Detecting dangerous situations" means sensing emergencies such as fires or carbon monoxide poisoning.

[1730] An "emergency notification" is a system that sends out an immediate warning when danger is detected.

[1731] A "personalized learning program" refers to learning content tailored to each individual user.

[1732] "Progress data" refers to data that records the progress of learning or activities.

[1733] An "emotion engine" is a system that analyzes voice and facial expression data to recognize emotions.

[1734] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[1735] Audio capture and content generation

[1736] The user requests:

[1737] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice, converts the data into text, and sends it to the server. The server uses a generative AI model to analyze the voice data, understand the request, and generate appropriate content. This generated content is saved as text data, converted back into audio data, and sent back to the device. The device then plays the audio data, providing storytelling or study assistance.

[1738] Specific example:

[1739] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[1740] Examples of prompts to input into a generative AI model:

[1741] The user requested, "Tell me a story." Please generate a Japanese children's fairy tale.

[1742] Understanding and sharing a child's mental state

[1743] Device features:

[1744] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[1745] Specific example:

[1746] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[1747] Examples of prompts to input into a generative AI model:

[1748] The user said, "Something unpleasant happened at school today." Based on this statement, identify their emotions and determine whether they are feeling anxious.

[1749] Indoor environment monitoring and automatic adjustment

[1750] Device environment monitoring:

[1751] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[1752] Specific example:

[1753] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1754] Examples of prompts to input into a generative AI model:

[1755] The indoor temperature is over 30°C. Please generate a command to turn on the air conditioner.

[1756] Detection of dangerous situations and emergency notifications

[1757] Device risk detection function:

[1758] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1759] Specific example:

[1760] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1761] Examples of prompts to input into a generative AI model:

[1762] The device has detected smoke. Please generate an emergency notification message for your parents.

[1763] Personalized learning tailored to the child's age.

[1764] Server's learning program generation function:

[1765] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1766] Specific example:

[1767] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1768] Examples of prompts to input into a generative AI model:

[1769] Please generate a program that teaches the hiragana character "い" based on the user's age in months and past learning data.

[1770] Use of the emotion engine

[1771] Emotion recognition feature of the device:

[1772] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[1773] Specific example:

[1774] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[1775] Examples of prompts to input into a generative AI model:

[1776] The user said, "I had fun today!" Use the emotion engine to generate content that the child will find enjoyable, once the emotion of joy is recognized.

[1777] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[1778] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1779] Step 1:

[1780] The user requests:

[1781] The user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures the user's voice and collects that audio data. Specifically, when the user says "Tell me a story," the microphone acquires audio data, and that data is processed in the next step.

[1782] Input: User's voice

[1783] Output: Collected audio data

[1784] Step 2:

[1785] Converting audio data to text:

[1786] The device uses speech recognition software to convert the captured audio data into text data. Specifically, the speech recognition engine analyzes the audio and generates text data such as "Tell me your story." This text data is then sent to the server in the next step.

[1787] Input: Audio data

[1788] Output: Text data

[1789] Step 3:

[1790] Sending text data:

[1791] The terminal sends the converted text data to the server. The server receives the text data and prepares it for analysis. Specifically, the text data is sent to the server via the internet.

[1792] Input: Text data

[1793] Output: Text data sent to the server

[1794] Step 4:

[1795] Content generation:

[1796] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., fairy tales or learning questions). Specifically, the generative AI model processes a request like "Tell me a story" and generates a Japanese fairy tale. The generated text data is then converted into audio data.

[1797] Input: Text data

[1798] Output: Generated content (text data)

[1799] Step 5:

[1800] Generating audio data:

[1801] The generated text data is converted into speech data by the server's text-to-speech engine. Specifically, the text-to-speech engine generates speech data such as "Once upon a time, in a certain place..." This speech data is then sent to the terminal in the next step.

[1802] Input: Generated content (text data)

[1803] Output: Audio data

[1804] Step 6:

[1805] Sending audio data:

[1806] The server sends audio data to the terminal. The terminal prepares to play the received audio data. Specifically, the audio data is sent to the terminal via the internet.

[1807] Input: Audio data

[1808] Output: Audio data sent to the terminal

[1809] Step 7:

[1810] Play content:

[1811] The device plays the audio data it receives and provides it to the user. Specifically, the device's speaker plays the audio "Once upon a time, in a certain place..." This allows the user to listen to a fairy tale.

[1812] Input: Audio data sent to the device

[1813] Output: Audio data that the user will listen to.

[1814] ---

[1815] The above outlines the specific processing flow for voice capture and content generation in this system. This allows users to seamlessly receive content tailored to their requests.

[1816] (Application Example 2)

[1817] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1818] In modern society, ensuring family safety and a comfortable shopping experience in physical stores is a crucial challenge. However, existing systems struggle to comprehensively address issues such as providing engaging content for children, understanding their emotional state, ensuring a comfortable environment, and responding quickly to emergencies. Furthermore, there is a lack of means for parents to know their child's current location and status in real time. This leads to problems where parental peace of mind and child safety are not adequately ensured.

[1819] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or study content using a generative model; means for playing back the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for emergency notification to the parent when a dangerous situation is detected; means for understanding the location information of the parent and child within the physical store and ensuring the child's safety; and means for providing various information within the store. This makes it possible for families to enjoy shopping safely and comfortably in physical stores.

[1820] "Voice capture" is a technology that captures the voice of a child when they speak and converts it into data that can be analyzed.

[1821] A "generative model" is an algorithm used to generate content based on input data.

[1822] "Read-aloud content" refers to stories and narratives intended to be read aloud to children.

[1823] "Study content" refers to problems and learning materials created to support children's learning.

[1824] "User conversation data" refers to audio and text data recorded during interactions with children.

[1825] "Inferring mental state" refers to judging a user's emotions and mood from their voice and facial expressions.

[1826] "Notification to parents" refers to a means of communication to inform parents about their child's presumed mental state or emergency situation.

[1827] "Indoor environment monitoring" refers to the act of measuring indoor environmental information such as temperature, humidity, and carbon dioxide concentration using sensors.

[1828] "Automatic environmental adjustment" refers to the automatic operation of devices such as air conditioners and humidifiers based on monitoring data to maintain an appropriate indoor environment.

[1829] "Detection of dangerous situations" refers to the act of using sensors to detect emergencies such as fires or carbon monoxide emissions.

[1830] "Emergency notification" refers to a notification function that quickly informs parents when a dangerous situation occurs.

[1831] "Location tracking" refers to technology that tracks and records a child's current location within a physical store.

[1832] "Means for ensuring safety" refers to systems that use location information to prevent children from getting lost and to respond to emergencies.

[1833] "In-store information guidance" refers to a function that informs users of information provided within the store, such as the location of restrooms and the location of sales areas.

[1834] In order to implement this invention, the following system configuration and processing are necessary.

[1835] System configuration:

[1836] 1. Voice capture device: This includes a microphone for capturing the child's voice and converting it into digital data.

[1837] 2. Generative Model Server: Receives audio data as text data, analyzes it, and generates appropriate content (stories and learning questions). For this purpose, it uses a database and AI algorithms (e.g., generative AI models such as GPT).

[1838] 3. Content playback device: Equipped with speakers for playing the generated content as audio.

[1839] 4. Emotion Recognition Engine: A device that analyzes conversation data and facial expression data with a user to infer their mental state, and includes a camera and emotion recognition software (e.g., emotion recognition API).

[1840] 5. Notification system: Equipped with a communication module to notify the parent's smartphone of the suspected mental state.

[1841] 6. Environmental monitoring sensors: Equipped with sensors (e.g., DHT22 or MQ-135 sensors) for measuring indoor temperature, humidity, CO2 concentration, etc.

[1842] 7. Automatic environmental control device: A device for controlling environmental control devices such as air conditioners and humidifiers.

[1843] 8. Emergency notification system: A device that detects fires, carbon monoxide emissions, and other emergencies and sends emergency notifications to the parent's smartphone.

[1844] 9. Location tracking system: GPS device and tracking software for determining a child's current location in real time.

[1845] Processing procedure:

[1846] This system supports children's safety and development through a multi-layered process.

[1847] 1. Audio capture and content generation:

[1848] When the user (child) says something like "Tell me a story" or "Help me with my studies," the device's microphone captures the audio and sends it to the generative model server.

[1849] The server analyzes the audio data as text data and generates appropriate content using a generative model.

[1850] For example, if a user says, "Tell me a story," the server generates a fairy tale, and the device plays the audio data.

[1851] As an example of a prompt message using a generative AI model, you can input the instruction, "Generate a fairy tale suitable for children."

[1852] 2. Emotion recognition and notification to parents:

[1853] The system collects conversation data and facial expression data from users and sends them to the server.

[1854] The server uses an emotion recognition engine to analyze the data and infer the user's mental state.

[1855] If the estimated mental state is "anxiety" or "joy," a notification will be sent to the parent's smartphone stating, "Your child is feeling anxious" or "Your child is very happy."

[1856] For example, if a user says, "Something unpleasant happened at school today...", the device sends the audio and facial expression data to the server, infers the user's anxiety, and sends a notification.

[1857] 3. Monitoring and automatic adjustment of the indoor environment:

[1858] The device measures temperature, humidity, and CO2 concentration at regular intervals and sends the data to the server.

[1859] The server analyzes the data and sends adjustment instructions to the terminal to maintain a comfortable environment.

[1860] For example, when the server detects a room temperature of 30°C, it sends a command to the terminal to turn on the air conditioner, and the terminal follows that command to lower the room temperature.

[1861] 4. Detection of dangerous situations and emergency notification:

[1862] If the terminal detects a fire or an increase in carbon monoxide concentration, it will immediately send data to the server.

[1863] The server analyzes the information and sends a notification to the parent's smartphone saying, "Emergency: Potential fire."

[1864] For example, one possible implementation is to immediately send an emergency notification to the parent when smoke is detected.

[1865] 5. Location information tracking and guidance within physical stores:

[1866] The system tracks the child's location in real time and displays it on the parent's smartphone.

[1867] Parents can know their child's current location, allowing them to enjoy shopping with peace of mind.

[1868] For example, if a child says, "I need to go to the toilet," the system can guide them to the location of the toilet within the store.

[1869] In summary, this invention comprehensively supports the safety and development of children and provides parents with an environment in which they can feel at ease.

[1870] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1871] Step 1:

[1872] Voice capture and text conversion:

[1873] The user (child) speaks aloud, saying things like "Tell me a story" or "Help me with my studies." The device's microphone captures this audio, and voice capture software converts the audio data into text data. The input is audio data, and the output is text data.

[1874] Step 2:

[1875] Sending text data to the server:

[1876] The terminal sends the captured text data to the server. The input is text data, and the output is the data sent to the server. Specifically, the data is sent to the server using an HTTP request.

[1877] Step 3:

[1878] Content generation:

[1879] The server analyzes the received text data. Using a generative AI model (e.g., GPT-3), it generates appropriate content (fairy tales or learning questions) in response to the request. During this process, prompts are input to the generative AI model, and the generated stories or questions are output as prompt results. For example, the prompt "Generate a fairy tale suitable for children" might be used.

[1880] Step 4:

[1881] Converting generated content into audio data:

[1882] The generated content is converted into audio data. The server uses a text-to-speech (TTS) engine to convert the generated text content into audio data. The input is the generated text data, and the output is audio data.

[1883] Step 5:

[1884] Sending audio data to the terminal:

[1885] The server sends the generated audio data to the terminal. The input is the audio data, and the output is the data sent to the terminal. Specifically, the data is sent to the terminal using an HTTP response.

[1886] Step 6:

[1887] Play content:

[1888] The device plays the received audio data. It uses a speaker to let the user (child) listen to the content aloud. The input is audio data, and the output is audio playback. For example, the device starts reading a fairy tale, such as "Once upon a time, in a certain place..."

[1889] Step 7:

[1890] Collection and inference of emotional data:

[1891] The system collects conversation data and facial expression data from the user using the device's microphone and camera. This data is sent to a server, which analyzes it using an emotion recognition engine. The input is voice and facial expression data, and the output is estimated emotion data. For example, from a conversation like, "Something unpleasant happened at school today...", the system might infer "anxiety".

[1892] Step 8:

[1893] Notification to parents regarding their child's mental state:

[1894] The server sends a notification to the parent based on the inferred emotion data. The notification is sent via a smartphone application. The input is emotion data, and the output is a notification message. For example, a notification saying "Your child is feeling anxious" is sent.

[1895] Step 9:

[1896] Environmental data monitoring:

[1897] The terminal is equipped with sensors to measure temperature, humidity, and CO2 concentration. This environmental data is transmitted to the server at regular intervals. The input is sensor data, and the output is data transmitted to the server.

[1898] Step 10:

[1899] Automatic environment adjustment:

[1900] The server analyzes the received environmental data and issues instructions for necessary adjustments. It remotely controls devices such as air conditioners and humidifiers to automatically adjust the environment. The input is environmental data, and the output is instructions for operating the devices. For example, if it detects a room temperature of 30°C, it will issue an instruction to turn on the air conditioner.

[1901] Step 11:

[1902] Detection of dangerous situations and emergency notifications:

[1903] The device has sensors to detect fire and elevated carbon monoxide levels. When these dangerous situations are detected, it immediately sends data to a server. The server then sends an emergency notification to the parent's smartphone. The input is sensor data, and the output is an emergency notification message. For example, if smoke is detected, a notification such as "Emergency: Potential fire" is sent.

[1904] Step 12:

[1905] Location tracking and guidance:

[1906] The device tracks the child's current location in real time and displays the location information on the parent's smartphone. The input is GPS data, and the output is a display of location information. For example, if the child says, "I need to go to the toilet," the system will guide them to the location of the toilet.

[1907] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1908] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1909] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1910] [Fourth Embodiment]

[1911] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1912] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1913] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1914] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1915] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1916] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1917] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1918] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1919] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1920] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1921] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1922] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1923] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1924] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications. The embodiments of each function in this system are described below.

[1925] 1. Audio capture and content generation

[1926] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and uses a generative model to generate appropriate content (e.g., a story or study question). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[1927] Specific example:

[1928] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[1929] 2. Understanding and sharing the child's mental state

[1930] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative model to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety and notifies the parent.

[1931] Specific example:

[1932] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[1933] 3. Monitoring and automatic adjustment of the indoor environment

[1934] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[1935] Specific example:

[1936] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[1937] 4. Detection of dangerous situations and emergency notification

[1938] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[1939] Specific example:

[1940] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[1941] 5. Personalized learning tailored to the user's age.

[1942] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[1943] Specific example:

[1944] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[1945] The above describes the embodiments of the present invention. By using this system, support for children's growth and ensuring their safety can be effectively achieved.

[1946] The following describes the processing flow.

[1947] 1. Audio capture and content generation

[1948] Step 1:

[1949] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[1950] Step 2:

[1951] The device's microphone captures the user's voice.

[1952] Step 3:

[1953] The device converts the captured audio into text data and sends that data to the server.

[1954] Step 4:

[1955] The server analyzes the text data it receives to understand the content of the request.

[1956] Step 5:

[1957] The server's generative model generates appropriate content (stories and learning questions).

[1958] Step 6:

[1959] The server generates audio data and sends it to the terminal.

[1960] Step 7:

[1961] The device plays back received audio data to provide the user with read-alouds and study support.

[1962] 2. Understanding and sharing the child's mental state

[1963] Step 1:

[1964] The device's microphone and camera collect data on the user's conversation and facial expressions.

[1965] Step 2:

[1966] The device sends the collected data to the server.

[1967] Step 3:

[1968] The server analyzes the received data and infers the user's mental state based on their tone of voice, the words they use, and their facial expressions.

[1969] Step 4:

[1970] The server notifies the parent's smartphone app of the estimated mental state.

[1971] 3. Monitoring and automatic adjustment of the indoor environment

[1972] Step 1:

[1973] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[1974] Step 2:

[1975] The device sends the collected data to the server.

[1976] Step 3:

[1977] The server analyzes the environmental data it receives and compares it to the set baseline values.

[1978] Step 4:

[1979] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[1980] Step 5:

[1981] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[1982] 4. Detection of dangerous situations and emergency notification

[1983] Step 1:

[1984] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[1985] Step 2:

[1986] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[1987] Step 3:

[1988] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[1989] 5. Personalized learning tailored to the user's age.

[1990] Step 1:

[1991] The server retrieves the user's lunar age data and past learning progress data.

[1992] Step 2:

[1993] The server's generation model generates a new learning program based on the lunar phase and past data.

[1994] Step 3:

[1995] The server sends the generated learning program to the terminal.

[1996] Step 4:

[1997] The device supports the user's learning through voice guidance based on the learning program.

[1998] Step 5:

[1999] The device periodically sends learning progress data to the server.

[2000] Step 6:

[2001] The server analyzes the received progress data and adjusts the learning program as needed.

[2002] The above outlines the specific processing steps for each function.

[2003] (Example 1)

[2004] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2005] In modern families, there is a need for systems that effectively support children's development and allow them to grow up in a safe and secure environment. In particular, support for children's learning, monitoring of their mental state, monitoring of the indoor environment, and rapid response in emergencies are crucial. However, existing systems struggle to provide these functions in a unified and effective manner.

[2006] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[2007] In this invention, the server includes means for capturing the user's voice and performing speech recognition, means for converting the recognized voice into text data, means for inputting the text data into a generative AI model and generating appropriate content, means for synthesizing the generated voice content, means for transmitting the synthesized content to a terminal, means for collecting conversation data and facial expression data with the user and inferring their mental state, means for notifying the parent of the inferred mental state, means for using sensors to monitor the indoor environment, means for automatically adjusting the environment based on the monitoring data, means for sending an emergency notification to the parent when a dangerous situation is detected, and means for playing the content regenerated by the generative model on the terminal. This enables seamless support for the child's development, safety, and learning.

[2008] "Children" refers to the youngest members of a family and users of this system.

[2009] "Methods for capturing audio" refer to methods of collecting user-generated audio as digital data using microphones, speech recognition devices, etc.

[2010] A "generative model" refers to an algorithm or system that automatically generates new content using machine learning or artificial intelligence technologies.

[2011] "Reading aloud" refers to the act of providing stories or educational content to users in audio format.

[2012] "Study content" refers to educational materials such as problems and explanations provided to support user learning.

[2013] "Means for playing generated content" refers to methods of providing users with viewable digital content generated by a generative model.

[2014] "User conversation data" refers to linguistic information that users communicate with the system, and is collected by a speech recognition system.

[2015] "Facial expression data" refers to data about a user's facial expressions collected using cameras and sensors.

[2016] "Means of inferring mental state" refers to algorithms and systems that analyze collected voice data and facial expression data to estimate the user's emotions and mental state.

[2017] "Methods for notifying parents" refers to methods of sending notifications to parents' smartphones or tablets to inform them of the user's mental state or emergency situation.

[2018] A "sensor for monitoring the indoor environment" refers to a sensor device used to measure temperature, humidity, CO2 concentration, etc.

[2019] "Means of automatically adjusting the environment" refers to methods of controlling environmental devices such as air conditioners and humidifiers based on monitoring data.

[2020] A "dangerous situation" refers to any condition that could potentially threaten the safety of the user or those around them, such as a fire or a carbon monoxide leak.

[2021] "Means of emergency notification" refers to a method of promptly sending notifications to parents or appropriate personnel when a dangerous situation is detected.

[2022] "Speech recognition" refers to a technology that analyzes a user's voice as digital data and converts it into text data.

[2023] "Text data" refers to data that stores speech information analyzed by a speech recognition system as text.

[2024] "Speech synthesis" refers to the technology that generates natural-sounding speech based on text data.

[2025] A "generative AI model" refers to a model that uses artificial intelligence algorithms to automatically generate new content such as speech and text.

[2026] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, and emergency notifications.

[2027] Audio capture and content generation

[2028] When a user (child) speaks to the device, saying things like "Tell me a story" or "Help me with my studies," the device's microphone captures the user's voice and sends that voice as text data to the server. The server analyzes the received text data and uses a generative AI model to generate appropriate content (for example, a story or study problem). The generated audio data is then sent to the device, which plays the audio to provide storytelling or study assistance to the user.

[2029] Specific example:

[2030] When the user says, "Tell me a story," the device captures the voice and sends it to the server. The server uses a generative AI model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, "Once upon a time, in a certain place..."

[2031] Example of a prompt:

[2032] "Please create a short children's story that children will enjoy listening to. The theme is animal friendship."

[2033] Understanding and sharing a child's mental state

[2034] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, which uses a generative AI model (e.g., an emotion analysis model) to estimate the user's mental state. The estimation results are notified to the parent's smartphone app. For example, if a user says, "Something bad happened at school today...", the server analyzes the voice and facial expressions to estimate stress and anxiety.

[2035] Specific example:

[2036] When a user speaks to the device saying, "Something bad happened at school today," the device sends audio and facial expression data to the server. The server analyzes the audio and facial expressions and infers that the user is feeling "anxious." This information is then sent to the parent's smartphone app as a notification stating, "Your child is feeling anxious today."

[2037] Indoor environment monitoring and automatic adjustment

[2038] The terminal is equipped with temperature, humidity, and CO2 sensors, and sends this data to the server at regular intervals. The server analyzes the data and sends instructions to the terminal to make appropriate environmental adjustments. For example, if the temperature is too high, the server will instruct the terminal to turn on the air conditioner, and the terminal will operate the air conditioner.

[2039] Specific example:

[2040] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[2041] Detection of dangerous situations and emergency notifications

[2042] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[2043] Specific example:

[2044] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[2045] Personalized learning tailored to the child's age.

[2046] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[2047] Specific example:

[2048] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[2049] The above describes the embodiments of the present invention. By using this system, support for child development and safety can be achieved effectively and in a unified manner.

[2050] The flow of the specific processing in Example 1 will be explained using Figure 11.

[2051] Audio capture and content generation

[2052] Step 1:

[2053] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[2054] Input: User's voice

[2055] Output: Captured audio data

[2056] Step 2:

[2057] The device's microphone captures the user's voice and saves that audio as digital data.

[2058] Input: Captured audio data

[2059] Output: Digital audio data

[2060] Step 3:

[2061] The device sends the captured digital audio data to the server.

[2062] Input: Digital audio data

[2063] Output: Data to send to the server

[2064] Step 4:

[2065] The server uses speech recognition software (e.g., Google Speech-to-Text) to convert the audio data into text data.

[2066] Input: Audio data

[2067] Output: Text data

[2068] Step 5:

[2069] The server inputs text data into an AI model (e.g., GPT-4) and generates appropriate content as a prompt.

[2070] Input: Text data

[2071] Output: Generated text content

[2072] Step 6:

[2073] The server converts the generated text content into speech data using speech synthesis software (e.g., Amazon Polly).

[2074] Input: Generated text content

[2075] Output: Generated audio data

[2076] Step 7:

[2077] The server sends the generated audio data to the terminal.

[2078] Input: Generated audio data

[2079] Output: Data to send to the terminal

[2080] Step 8:

[2081] The device plays back the received audio data, providing the user with read-alouds and study support.

[2082] Input: Generated audio data

[2083] Output: Audio to be played back to the user

[2084] Understanding and sharing a child's mental state

[2085] Step 1:

[2086] The device's microphone and camera capture the user's voice and facial expression data.

[2087] Input: User's voice and facial expressions

[2088] Output: Audio data and facial expression data

[2089] Step 2:

[2090] The device sends the captured audio and facial expression data to the server.

[2091] Input: Voice data and facial expression data

[2092] Output: Data to send to the server

[2093] Step 3:

[2094] The server uses voice analysis software and emotion analysis models (e.g., IBM Watson's emotion analysis) to analyze voice data and facial expression data and infer the user's mental state.

[2095] Input: Voice data and facial expression data

[2096] Output: Inferred mental state

[2097] Step 4:

[2098] The server notifies the parent's smartphone app of the prediction result.

[2099] Input: Estimated mental state

[2100] Output: Notification data sent to the parent's smartphone app

[2101] Indoor environment monitoring and automatic adjustment

[2102] Step 1:

[2103] The device acquires indoor environmental data using temperature, humidity, and CO2 sensors installed on the terminal.

[2104] Input: Indoor environment

[2105] Output: Environmental sensor data

[2106] Step 2:

[2107] The device sends the environmental sensor data it acquires to the server.

[2108] Input: Environmental sensor data

[2109] Output: Data to send to the server

[2110] Step 3:

[2111] The server analyzes environmental sensor data and generates necessary adjustment instructions.

[2112] Input: Environmental sensor data

[2113] Output: Environmental adjustment instructions

[2114] Step 4:

[2115] The server sends the generated adjustment instructions to the terminal.

[2116] Input: Environmental adjustment instructions

[2117] Output: Data to send to the terminal

[2118] Step 5:

[2119] The terminal follows the adjustment instructions and operates devices such as air conditioners and humidifiers to adjust the indoor environment.

[2120] Input: Environmental adjustment instructions

[2121] Output: Adjusted indoor environment

[2122] Detection of dangerous situations and emergency notifications

[2123] Step 1:

[2124] The device's built-in fire alarm and carbon monoxide sensor detect dangerous situations.

[2125] Input: Indoor environment

[2126] Output: Hazard detection data

[2127] Step 2:

[2128] The device sends the detected data to the server.

[2129] Input: Hazard detection data

[2130] Output: Data to send to the server

[2131] Step 3:

[2132] The server quickly analyzes the data and generates an emergency notification.

[2133] Input: Hazard detection data

[2134] Output: Emergency notification data

[2135] Step 4:

[2136] The server sends an emergency notification to the parent's smartphone app, prompting them to take immediate action.

[2137] Input: Emergency notification data

[2138] Output: Notification data sent to the parent's smartphone app

[2139] Personalized learning tailored to the child's age.

[2140] Step 1:

[2141] The server retrieves the user's lunar age data and past learning progress data.

[2142] Input: Lunar age data and learning progress data

[2143] Output: Analysis data

[2144] Step 2:

[2145] The server uses the generated AI model to create a new learning program.

[2146] Input: Analysis data

[2147] Output: New learning program

[2148] Step 3:

[2149] The server sends the generated learning program to the terminal.

[2150] Input: New learning program

[2151] Output: Data to send to the terminal

[2152] Step 4:

[2153] The device provides users with learning support content through voice guidance.

[2154] Input: New learning program

[2155] Output: Learning content for the user

[2156] Step 5:

[2157] The device feeds back user learning progress data to the server, which is then used to inform the next learning session.

[2158] Input: Learning progress data

[2159] Output: Feedback data to the server

[2160] (Application Example 1)

[2161] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2162] Supporting the development and ensuring the safety of children in modern society are crucial challenges. Existing systems do not adequately address the detailed understanding of children's mental states, provide appropriate learning support, or manage their environment. Furthermore, monitoring the health of factory workers is insufficient, making it difficult to provide appropriate breaks and preventative measures in a timely manner. To solve these problems, this invention aims to provide a system that effectively supports the development and safety of children, as well as monitors the health of factory workers.

[2163] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[2164] In this invention, the server includes means for capturing the voice of a child when they speak and generating read-aloud or educational content using a generative model; means for playing the generated content; means for collecting conversation data with the user and inferring their mental state; means for notifying the parent of the inferred mental state; means for using sensors to monitor the indoor environment; means for automatically adjusting the environment based on monitoring data; means for making an emergency notification to the parent when a dangerous situation is detected; and means for capturing the voice and facial expression data of a worker when they speak, inferring the worker's health status using a generative model, and notifying the manager. This enables support for the development and safety of children, as well as monitoring the health status of workers in the factory and taking appropriate action.

[2165] "Voice capture" is the process of collecting the voice of a user when they speak using an input device such as a microphone.

[2166] A "generative model" is a system that uses machine learning or artificial intelligence algorithms to generate new content or information based on input data.

[2167] "Content generation" refers to the process of creating new read-aloud or educational content based on collected data.

[2168] "Means of playback" refers to audio output devices or programs that allow users to listen to the generated content.

[2169] "Mental state estimation" is a process that analyzes voice and facial expression data to determine the user's emotions and psychological state.

[2170] "Means of notifying parents" refers to communication systems and devices used to inform parents or guardians of suspected mental states or emergencies.

[2171] "Indoor environment monitoring" is the process of measuring and collecting environmental data such as temperature, humidity, and CO2 using sensors.

[2172] "Means for automatically adjusting the environment" refers to a system that automatically operates air conditioners and ventilation devices based on data about the indoor environment.

[2173] "Means of emergency notification" refers to means of detecting dangerous situations such as fires or abnormal carbon monoxide levels and quickly informing relevant parties.

[2174] "Health status estimation" is a process that analyzes the worker's voice and facial expression data to determine their health status, such as fatigue and stress.

[2175] "Means of notifying the administrator" refers to a communication system device for informing the administrator of the estimated health status of the worker.

[2176] This invention provides a system that effectively supports the growth and safety of children, and monitors the health status of workers in a factory. The system consists of the following main components:

[2177] Audio capture and content generation

[2178] The system's terminal captures the voice spoken by the user (child or worker) using a microphone. This voice data is converted to text and sent to the server. The server uses a generative AI model to generate appropriate content (e.g., read-aloud stories, learning questions, or health assessments for workers) based on the input text data. For example, if a child says, "Tell me a story," the generated story is sent to the terminal and played back as audio.

[2179] Mental state assessment and notification

[2180] The device is equipped with voice capture capabilities and a camera to collect user conversation and facial expression data. This data is sent to a server, where a generative AI model is used to infer the user's mental state. For example, if a child says, "I had a bad day," the server analyzes the voice and facial expression to infer stress and anxiety, and notifies the parent's smartphone app. Similarly, in a factory, the voice and facial expressions of workers are captured to infer their health status and notify the manager.

[2181] Indoor environment monitoring and automatic adjustment

[2182] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and issues environmental adjustment instructions to the terminal as needed. For example, if the temperature is too high, it sends an instruction to the terminal to turn on the air conditioner, and the air conditioner starts operating.

[2183] Detection of dangerous situations and emergency notifications

[2184] The device is equipped with a fire alarm and carbon monoxide sensor, and if a dangerous situation is detected, it immediately notifies the server. The server quickly analyzes the situation and sends an emergency notification to the parent's or administrator's smartphone app. For example, if smoke is detected, it has a function to notify the parent that there is a possibility of fire.

[2185] Children's learning support

[2186] The server generates a new learning program based on the user's (child's) age data and past learning progress. This program is sent to the device, which then assists with learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[2187] Specific example

[2188] For example, if a worker in a factory says to a robot, "I'm tired today," the robot captures this voice and facial expression data and sends it to a server. The server uses a generative AI model to analyze the data, and if it determines that the worker's fatigue level is high, it notifies the administrator's tablet. Based on this information, the administrator can instruct the worker to take an appropriate break.

[2189] Example of a prompt

[2190] The prompt statement is designed as follows:

[2191] Create a robot app that uses a voice capture system to monitor the health of workers. The app will use voice recognition to analyze the worker's fatigue level and notify the manager. For example, if a worker says, "I'm tired today," the robot will capture this, use a generative AI model to estimate the fatigue level, and notify the manager's tablet.

[2192] Hardware and software

[2193] The hardware used includes microphones for voice capture, cameras for facial recognition, and temperature, humidity, and CO2 sensors, fire alarms, and carbon monoxide sensors for environmental monitoring. The software uses speech recognition libraries (e.g., speech_recognition), generative AI models (e.g., "GPT" and "BERT"), and web APIs for data analysis.

[2194] By combining these elements, it is possible to create a system that comprehensively supports children's development, ensures their safety, and monitors the health of factory workers.

[2195] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[2196] Step 1:

[2197] Audio data capture

[2198] The device's microphone captures the user's (child or worker's) voice input. The captured audio data is temporarily stored on the device and then sent to the speech recognition engine.

[2199] Input: Audio data

[2200] Output: Audio data sent to the speech recognition engine

[2201] Step 2:

[2202] Converting audio data to text

[2203] The device converts the captured audio data into text using a speech recognition library. The converted text data is then sent to the server.

[2204] Input: Voice data sent to the speech recognition engine

[2205] Output: Text data sent to the server

[2206] Step 3:

[2207] Text data analysis and content generation

[2208] The server analyzes the received text data and uses a generative AI model to generate appropriate content (e.g., stories, learning questions, health assessments). GPT and BERT are used as generative AI models for this analysis.

[2209] Input: Text data sent to the server

[2210] Output: Content data generated using a generative AI model

[2211] Step 4:

[2212] Sending and playing content data

[2213] The generated content data is sent from the server to the terminal, and the terminal provides it to the user via audio using its audio playback function.

[2214] Input: Content data generated using a generation AI model.

[2215] Output: Audio content provided using the audio playback function.

[2216] Step 5:

[2217] Inference of mental state

[2218] The device uses cameras for voice and facial expression capture to collect user conversation and facial expression data. The collected data is sent to a server. The server uses a generative AI model to analyze voice tone, words used, and facial expression data to infer the user's mental state.

[2219] Input: Audio data and facial expression data collected by the camera and microphone.

[2220] Output: Estimated mental state data

[2221] Step 6:

[2222] Notification of mental state

[2223] The estimated mental state data is sent from the server to a smartphone app for parents and administrators. This allows parents and administrators to quickly understand the situation.

[2224] Input: Estimated mental state data

[2225] Output: Notification to the parent's or administrator's smartphone app

[2226] Step 7:

[2227] Environmental data monitoring

[2228] The device's temperature, humidity, and CO2 sensors periodically collect indoor environmental data and transmit it to the server.

[2229] Input: Environmental data from temperature sensor, humidity sensor, and CO2 sensor.

[2230] Output: Monitoring data sent to the server

[2231] Step 8:

[2232] Automatic environment adjustment

[2233] The server analyzes the monitoring data and sends instructions to the terminal to adjust the air conditioning and ventilation systems as needed. The terminal then automatically adjusts the environment according to these instructions.

[2234] Input: Monitoring data sent to the server

[2235] Output: Instructions for adjusting air conditioning and ventilation systems

[2236] Step 9:

[2237] Detection of dangerous situations and emergency notifications

[2238] The device is equipped with a fire alarm and a carbon monoxide sensor, and immediately sends data to a server if a dangerous situation is detected. The server quickly analyzes this data and sends an emergency notification to the parent's or administrator's smartphone app.

[2239] Input: Data from fire alarms and carbon monoxide sensors

[2240] Output: Sending emergency notification data to the parent's or administrator's smartphone app.

[2241] Step 10:

[2242] Health status monitoring and notification

[2243] When a worker in the factory speaks into a terminal, their voice and facial expression data are captured and sent to a server. The server uses a generative AI model to estimate the worker's health status and notifies the administrator of the results.

[2244] Input: Voice data and facial expression data of factory workers

[2245] Output: Estimated health status data and notification to administrator

[2246] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[2247] This invention is a system aimed at promoting and monitoring child development, and it has a wide range of functions, including voice capture, use of generative models, indoor environment management, emotion recognition by an emotion engine, and emergency notifications. The embodiments of each function in this system are described below.

[2248] 1. Audio capture and content generation

[2249] The user (child) speaks to the device, saying things like, "Tell me a story," or "Help me with my studies." The device's microphone captures the user's voice and sends it to the server as text data. The server analyzes the received text data and understands the request. Next, a generative model generates appropriate content (stories or learning questions) and sends the audio data to the device. The device then plays the audio data to provide the user with storytelling or study assistance.

[2250] Specific example:

[2251] When the user says "Tell me a story," the device captures the audio and sends it to the server. The server uses a generative model to generate a fairy tale and sends the audio data back to the device. The device then begins reading aloud, starting with "Once upon a time, in a certain place..."

[2252] 2. Understanding and sharing the child's mental state

[2253] The device's microphone and camera collect data on the user's conversation and facial expressions. This data is sent to a server, where a generative model analyzes the received data. Furthermore, an emotion engine recognizes the user's emotions based on their voice and facial expressions and infers their mental state. The inference results are notified to the parent's smartphone app.

[2254] Specific example:

[2255] When a user says, "Something bad happened at school today...", the device sends voice and facial expression data to the server. The server uses a generative model and emotion engine to infer that the user is "anxious" and notifies the parent of this information. A notification appears on the smartphone app stating, "Your child is feeling anxious today."

[2256] 3. Monitoring and automatic adjustment of the indoor environment

[2257] The terminal is equipped with temperature, humidity, and CO2 sensors, and transmits this data to the server at regular intervals. The server analyzes the data and compares it to set baseline values. The server then sends adjustment instructions to the terminal to maintain a comfortable environment. Based on the instructions from the server, the terminal operates devices such as air conditioners and humidifiers.

[2258] Specific example:

[2259] When the terminal detects a room temperature of 30°C, it sends that data to the server. The server determines that the temperature is high and sends a command to the terminal to turn on the air conditioner. The terminal follows the command and starts the air conditioner to lower the room temperature.

[2260] 4. Detection of dangerous situations and emergency notification

[2261] The device is equipped with a fire alarm and carbon monoxide sensor, and immediately notifies the server if a dangerous situation is detected. The server quickly analyzes the information and sends an emergency notification to the parent's smartphone app. This enables a rapid response.

[2262] Specific example:

[2263] When the device detects smoke, it immediately sends the data to the server. The server determines the possibility of a fire and sends a notification to the parent's smartphone app saying, "Emergency: Potential fire."

[2264] 5. Personalized learning tailored to the user's age.

[2265] The server generates a new learning program based on the user's lunar age data and past learning progress. This program is sent to the device, which then assists the user in learning through voice guidance. Progress data is periodically fed back to the server, and the program is adjusted as needed.

[2266] Specific example:

[2267] The server analyzes the user's age in months and past learning data to generate a program for learning the hiragana character "i". The terminal provides voice guidance such as, "Today let's try writing the hiragana character 'i'," and the user learns. After completion, progress data is sent to the server, and the next learning content is adjusted.

[2268] 6. Use of the Emotion Engine

[2269] The device sends collected voice and facial expression data to the server, where the server's emotion engine analyzes the data to recognize the user's emotions. The recognized emotion data is integrated with a generative model to individually adjust the content provided to the user according to their emotions. In addition, the content of notifications to parents is also finely adjusted based on the recognized emotion data.

[2270] Specific example:

[2271] When a user says, "I had fun today!", the device sends data of their cheerful voice and smiling face to the server. The server's emotion engine recognizes the emotion of "joy," and a generative model generates content that the user will find enjoyable (for example, a new game or quiz). The device plays the content, and the parent receives a notification that "the child is very happy."

[2272] The above describes the embodiments of the present invention. By using this system, support for children's growth, ensuring their safety, and providing care that is tailored to their emotions can be effectively achieved.

[2273] The following describes the processing flow.

[2274] 1. Audio capture and content generation

[2275] Step 1:

[2276] The user speaks to the device, saying things like, "Tell me a story," or "Help me with my studies."

[2277] Step 2:

[2278] The device's microphone captures the user's voice.

[2279] Step 3:

[2280] The device converts the captured audio into text data and sends that data to the server.

[2281] Step 4:

[2282] The server analyzes the text data it receives to understand the content of the request.

[2283] Step 5:

[2284] The server's generative model generates appropriate content (stories and learning questions).

[2285] Step 6:

[2286] The server generates audio data and sends it to the terminal.

[2287] Step 7:

[2288] The device plays back received audio data to provide the user with read-alouds and study support.

[2289] 2. Understanding and sharing the child's mental state

[2290] Step 1:

[2291] The device's microphone and camera collect data on the user's conversation and facial expressions.

[2292] Step 2:

[2293] The device sends the collected data to the server.

[2294] Step 3:

[2295] The server's generative model analyzes the received data and infers the user's mental state based on voice tone and the words used.

[2296] Step 4:

[2297] The server's emotion engine analyzes voice and facial expression data to recognize the user's emotions.

[2298] Step 5:

[2299] The server notifies the parent's smartphone app of their estimated mental state and emotions.

[2300] 3. Monitoring and automatic adjustment of the indoor environment

[2301] Step 1:

[2302] Temperature, humidity, and CO2 sensors built into the device collect environmental data at regular intervals.

[2303] Step 2:

[2304] The device sends the collected data to the server.

[2305] Step 3:

[2306] The server analyzes the environmental data it receives and compares it to the set baseline values.

[2307] Step 4:

[2308] The server sends adjustment instructions to the terminal to maintain a comfortable environment.

[2309] Step 5:

[2310] The terminal operates devices such as air conditioners and humidifiers based on instructions from the server.

[2311] 4. Detection of dangerous situations and emergency notification

[2312] Step 1:

[2313] The terminal's built-in fire alarm and carbon monoxide sensor constantly monitor the environment.

[2314] Step 2:

[2315] If the device detects a dangerous situation (for example, smoke or carbon monoxide), it immediately notifies the server.

[2316] Step 3:

[2317] The server analyzes the received danger data and sends an emergency notification to the parent's smartphone app.

[2318] 5. Personalized learning tailored to the user's age.

[2319] Step 1:

[2320] The server retrieves the user's lunar age data and past learning progress data.

[2321] Step 2:

[2322] The server's generation model generates a new learning program based on the lunar phase and past data.

[2323] Step 3:

[2324] The server sends the generated learning program to the terminal.

[2325] Step 4:

[2326] The device supports the user's learning through voice guidance based on the learning program.

[2327] Step 5:

[2328] The device periodically sends learning progress data to the server.

[2329] Step 6:

[2330] The server analyzes the received progress data and adjusts the learning program as needed.

[2331] 6. Use of the Emotion Engine

[2332] Step 1:

[2333] The device sends the collected audio and facial expression data to the server.

[2334] Step 2:

[2335] The server's emotion engine analyzes the data to recognize the user's emotions.

[2336] Step 3:

[2337] Based on the sentiment data recognized by the server, the generative model individually adjusts the content.

[2338] Step 4:

[2339] Based on the data in which the server recognizes emotions, it fine-tunes the content of the notifications sent to the parents.

[2340] The above outlines the specific processing steps for each function.

[2341] (Example 2)

[2342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2343] To effectively promote and monitor children's development and provide a safe and comfortable environment, a system is needed that collects and analyzes various information in real time and ...

Claims

1. A method for capturing a child's voice when they speak and using a generative model to generate read-aloud or educational content, A means of playing the generated content, A means of collecting conversation data with users and inferring their mental state, Means of notifying parents of the suspected mental state, A means of using sensors to monitor the indoor environment, A means of automatically adjusting the environment based on monitoring data, A system that includes a means to send an emergency notification to parents when a dangerous situation is detected.

2. A means of generating a personalized learning program tailored to a child's development, A means of guiding children through the generated learning program via voice and collecting progress data, The system according to claim 1, characterized in that it has means for adjusting the learning content based on progress data.

3. The system according to claim 1, characterized in that the means for estimating the user's mental state includes means for analyzing voice tone, words used, and facial expression data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A