System

A system for monitoring and interacting with children using sensors and AI-driven dialogue ensures their safety and educational engagement, addressing the challenges of dual-income households by providing real-time parental involvement.

JP2026022511APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124028
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

The increase in dual-income households has led to parents spending less time with their children, making it difficult to ensure their safety, provide educational opportunities, and engage in play, resulting in stress and feelings of guilt due to the inability to monitor their children's activities in real time.

Method used

A system comprising a monitoring means for tracking children's activities, a transmission means for real-time data transfer to parental devices, an interaction means for dialogue and play based on parental settings, and an instruction means for sending commands, along with customization and feedback features to tailor the system to household preferences.

Benefits of technology

Ensures the safety and educational engagement of children even when parents are not present, providing real-time monitoring and interaction, and allowing parents to respond promptly to abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022511000001_ABST
    Figure 2026022511000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system includes monitoring means for monitoring an activity state of a child in a home, transmission means for transmitting data acquired by the monitoring means to a terminal of a parent, interaction means for performing interaction and play on the basis of setting data received from the terminal of the parent, and instruction means for sending an instruction to the child on the basis of a command received from the terminal of the parent.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, the increase in dual-income households has led to parents spending less time with their children, leaving many unable to ensure sufficient time for monitoring their children's safety, providing educational opportunities, and playing. This makes it difficult for parents to balance work and family life, often resulting in stress and feelings of guilt. Another problem is the inability to ensure the safety of children when parents are not around and to keep track of their children's situation in real time. [Means for solving the problem]

[0005] To solve these problems, the present invention provides the following means: a system comprising a monitoring means for monitoring a child's activities within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's device, an interaction means for engaging in dialogue and play based on setting data received from the parent's device, and an instruction means for sending instructions to the child based on commands received from the parent's device. The system further includes a customization means for customizing the interaction means with content based on the family's specific rules and preferences. It also includes a feedback means for transmitting feedback data on the child's activities to the parent's device in real time. This ensures the safety of the child even when the parent is not at home and supports the child's development through educational dialogue and play.

[0006] "Inside the home" refers to the interior of a home or residential facility, including living spaces.

[0007] "Child" refers to a minor who is under care and grows up under the care of a parent or guardian.

[0008] "Activity status" refers to a child's daily behavior and state, including playing, learning, eating, sleeping, etc.

[0009] "Monitoring methods" refers to technologies that use devices such as sensors, cameras, and microphones to observe children's activities and collect data.

[0010] "Transmission means" refers to communication technology for transferring acquired data to other devices or systems.

[0011] "Parental devices" refer to electronic devices such as computers and smartphones used by parents or guardians.

[0012] "Configuration data" refers to information entered by parents about household-specific rules and preferences that customize the system's behavior.

[0013] "Dialogue means" refers to the system's functions for dialogue with children using voice recognition and voice synthesis technology.

[0014] "Play" refers to activities and games that promote children's growth and learning.

[0015] "Instruction means" refers to the function for transmitting commands received from the parent's device to the child, and refers to the technology for providing instructions through voice, screen display, etc.

[0016] "Customization options" refers to the ability to tailor the system's behavior to suit a household's specific rules and preferences.

[0017] "Feedback measures" refers to the function of reporting and providing information to parents on their children's activities in real time. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention provides a system that monitors the activities of children at home, transmits necessary information to a parent's terminal, and provides dialogue and play based on settings data from the parent. Specific embodiments for carrying out the present invention will be described below.

[0040] overview

[0041] This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The terminal monitors the child's activities and sends this information to the parent's terminal via the server. It also provides dialogue and play based on the setting data received from the parent, and conveys instructions from the parent to the child. Furthermore, a feedback function reports the child's activities to the parent in real time.

[0042] Entering User Data

[0043] The user (parent) launches the application on their smartphone or computer and inputs their household's specific rules, preferences, schedules, etc. This setting data is sent to the server and synchronized with the device.

[0044] Saving and Syncing Settings

[0045] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[0046] Real-time dialogue generation

[0047] When the user (child) speaks to the device, the device uses speech recognition technology to convert the speech into text and sends the data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[0048] Send Feedback

[0049] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[0050] Specific examples

[0051] Example 1: A child plays with a "cat sitter"

[0052] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0053] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0054] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[0055] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0056] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0057] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0058] Example 2: Detecting abnormal behavior in children

[0059] 1. The device monitors the child's movements using an accelerometer and camera.

[0060] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0061] 3. The server notifies the parent device of this information.

[0062] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[0063] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0064] The above is a specific embodiment for carrying out the present invention. This system ensures the safety of children even when their parents are away, and can provide educational conversations and play.

[0065] The processing flow will be explained below.

[0066] Entering User Data

[0067] Step 1:

[0068] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[0069] Step 2:

[0070] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[0071] Step 3:

[0072] The server receives the authentication information and checks it against information in a database.

[0073] Step 4:

[0074] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[0075] Step 5:

[0076] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[0077] Step 6:

[0078] The terminal transmits the input setting data to the server.

[0079] Saving and Syncing Settings

[0080] Step 1:

[0081] The server receives the setting data sent from the terminal and temporarily stores it.

[0082] Step 2:

[0083] The server checks the integrity of the received data and stores it in the database once the check is complete.

[0084] Step 3:

[0085] The server updates the version of the new setting data and notifies the terminal of the information.

[0086] Step 4:

[0087] The device receives the notification from the server and downloads the new configuration data.

[0088] Step 5:

[0089] The device analyzes the received data and updates its internal settings.

[0090] Real-time dialogue generation

[0091] Step 1:

[0092] The user (child) talks to "Nyanko Sitter."

[0093] Step 2:

[0094] The device converts voice input into text data using voice recognition technology.

[0095] Step 3:

[0096] The terminal transmits the converted text data to the server.

[0097] Step 4:

[0098] The server receives the text data and uses a generation AI to generate an appropriate response.

[0099] Step 5:

[0100] The server transmits the generated response text data to the terminal.

[0101] Step 6:

[0102] The terminal converts the received response text data into voice using voice synthesis technology.

[0103] Step 7:

[0104] The device plays the converted audio to the child.

[0105] Send Feedback

[0106] Step 1:

[0107] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[0108] Step 2:

[0109] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[0110] Step 3:

[0111] The server receives the abnormality detection information and notifies the parent terminal.

[0112] Step 4:

[0113] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[0114] Step 5:

[0115] The parent terminal transmits the input instruction data to the server.

[0116] Step 6:

[0117] The server receives the instruction data and transmits the data to the terminal.

[0118] Step 7:

[0119] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[0120] Example 1

[0121] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0122] There is a need to provide a means to effectively monitor children's activities at home and to provide a safe environment for parents to watch over them even when they are not around. It is also important to provide appropriate dialogue and play to facilitate two-way communication between children and parents. Without such a system, it is difficult to ensure children's safety and promote educational dialogue and play.

[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0124] In this invention, the server includes: a user input means for a parent to input family rules, preferences, schedules, etc. on a smartphone or computer; a storage / synchronization means for saving the setting data entered by the user input means in a database and synchronizing it with the device; a voice recognition means for converting the child's speech into text using voice recognition technology and sending the data to the server; a response generation means for the server to generate an appropriate response using a generative AI model and send it to the device; a voice synthesis means for the device to convert the response text into speech using voice synthesis technology and return it to the child; a monitoring means for the device to monitor the child's activity status and abnormalities using various sensors and send the data to the server; a notification means for the server to notify the parent's device of the monitored data in real time; an instruction sending means for the parent to check the child's status through an application and send necessary instructions; and an instruction execution means for the device to execute the parent's instructions using voice synthesis technology and convey them to the child. This enables monitoring of the child's activity status and smooth communication with the parent.

[0125] "User input means" refers to the means by which parents input their household's specific rules, preferences, schedules, etc. on their smartphone or computer.

[0126] The "storage and synchronization means" is a means for storing the setting data input by the user input means in a database and synchronizing it with the terminal.

[0127] The "voice recognition means" is a means of converting the voice spoken by the child into text using voice recognition technology and sending that data to a server.

[0128] The "response generation means" is a means by which the server uses a generative AI model to generate an appropriate response and transmits it to the terminal.

[0129] "Speech synthesis means" refers to the means by which the terminal uses speech synthesis technology to convert the response text into speech and return it to the child.

[0130] The "monitoring means" is a means by which the device uses various sensors to monitor the child's activity status and any abnormalities, and sends that data to the server.

[0131] The "notification means" is a means by which the server notifies the parent terminal of the monitoring data in real time.

[0132] The "means for sending instructions" is a means by which parents can check their child's status through the application and send necessary instructions.

[0133] The "means for executing instructions" refers to the means by which the device executes instructions from the parent using voice synthesis technology and conveys them to the child.

[0134] The present invention provides a system that monitors the activities of children at home, allowing parents to watch over them with peace of mind. This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The specific hardware and software configurations and their operation are described below.

[0135] Entering User Data

[0136] The user (parent) launches a dedicated application on a smartphone or computer and inputs setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack at 3 pm every day" or "Play with building blocks." This setting data is then sent to a server via the Internet.

[0137] Saving and Syncing Settings

[0138] The server receives the configuration data sent from the parent and stores it in a database. This data is managed in real time. The server synchronizes the saved data with the device. For example, if a network error occurs, the server will repeatedly retry, and if synchronization is successful, the device will receive the new configuration data.

[0139] Real-time dialogue generation

[0140] When a user (child) speaks to the device, the device converts the speech into text using voice recognition technology (e.g., Google Cloud Speech-to-Text), and this text data is sent to the server.

[0141] The server uses a generative AI model (e.g., OpenAI GPT-3.5) to generate an appropriate response. For example, in response to the question, "What do you want to do?", it generates a response like, "Let's play with building blocks!" The generated text data is sent to the device.

[0142] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the text into speech and then returns that speech to the child.

[0143] Send Feedback

[0144] The device uses an accelerometer and camera to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to a server.

[0145] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send instructions via the application if necessary.

[0146] Carrying out parental instructions

[0147] The user (parent) can check the child's status through the application and input and send necessary instructions, such as "Are you OK?"

[0148] The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and convey it to the child, for example, "Mommy is asking if you're OK."

[0149] Specific examples

[0150] Example 1: A child plays with a "cat sitter"

[0151] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0152] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0153] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[0154] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0155] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0156] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0157] Example 2: Detecting abnormal behavior in children

[0158] 1. The device monitors the child's movements using an accelerometer and camera.

[0159] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0160] 3. The server notifies the parent device of this information.

[0161] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[0162] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0163] Examples of prompt statements

[0164] Enter the following prompts into the generative AI model:

[0165] "Situation: Child has fallen. Parents are concerned. Generate an appropriate response message for the child."

[0166] Example of the resulting result:

[0167] "Are you okay? Can you get up now? Is anything hurting? Be careful."

[0168] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0169] Step 1: Enter and submit user data

[0170] The user (parent) launches a dedicated application on their smartphone or computer. They input setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack every day at 3 p.m." or "Play with building blocks." They confirm the input setting data and press the "Submit" button.

[0171] Input: Configuration data such as household-specific rules, preferences, and schedules

[0172] Output: The configuration data sent

[0173] Specific operation: The user operates the application, inputs and submits setting data.

[0174] Step 2: Save and sync configuration data

[0175] The server receives the setting data sent by the parent. It stores the received data in a database and manages it in real time. The stored data is time-stamped. The server then starts the process of synchronizing the stored data with the device. For example, if the setting data is "Eat a snack at 3 pm," it is synchronized with the device.

[0176] Input: Submitted configuration data

[0177] Output: Settings data synced to the device

[0178] Specific behavior: The server stores the received data and starts the synchronization process.

[0179] Step 3: Handling real-time interactions with your child

[0180] The user (child) speaks to the device, asking questions such as "What are you doing today?", and the device uses voice recognition technology (e.g., Google Cloud Speech-to-Text) to convert the speech into text. This text data is then sent to the server.

[0181] Input: Child's voice data

[0182] Output: Audio data converted to text

[0183] Specific operation: The device uses voice recognition technology to convert the child's voice into text and send it to the server.

[0184] Step 4: Server Response Generation

[0185] The server generates an appropriate response using a generative AI model (e.g., OpenAI GPT-3.5) and sends the generated response text to the device. For example, in response to the question "What do you want to do?", the server generates the response "Let's play with building blocks!"

[0186] Input: Text-converted audio data

[0187] Output: The generated response text

[0188] Specific operation: The server uses the generative AI model to generate an appropriate response and sends it to the device.

[0189] Step 5: Reply with a text-to-speech response

[0190] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the response text received from the server into speech, which is then returned to the child. For example, the device might respond with a voice saying, "Let's play with building blocks!"

[0191] Input: Generated response text

[0192] Output: Audio data

[0193] Specific operation: The device uses speech synthesis technology to convert the response text into speech and return it to the child.

[0194] Step 6: Monitor your child's activity

[0195] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to the server.

[0196] Input: Child's activity data

[0197] Output: Anomaly detection data

[0198] Specific operation: The device uses various sensors to monitor the child's activities, and if it detects any abnormalities, it sends the information to the server.

[0199] Step 7: Server notification

[0200] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send any necessary instructions via the application.

[0201] Input: Anomaly detection data

[0202] Output: Notification data to parent device

[0203] Specific operation: The server notifies the parent device of abnormal data in real time.

[0204] Step 8: Send and Execute Parental Instructions

[0205] The user (parent) checks the child's status through the application and inputs and sends the necessary instructions. For example, a user can send the instruction "Are you OK?" The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and conveys it to the child. For example, the user can tell the child "Mommy is asking if you're OK."

[0206] Input: Parent instruction data

[0207] Output: Voice instructions to the child

[0208] Specific operation: The parent inputs instructions into the application, the server sends the instructions to the device, and the device uses speech synthesis technology to convert the instructions into voice and convey them to the child.

[0209] (Application example 1)

[0210] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0211] In modern families, it is extremely important for parents to monitor their children's activities in real time and respond quickly when necessary. However, conventional technologies have struggled to ensure children's safety while parents are away and to provide educational conversations and play. Furthermore, there have been insufficient means for immediately notifying parents if their children exhibit abnormal behavior. The present invention addresses these issues.

[0212] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0213] In this invention, the server includes a monitoring means for monitoring the child's activities within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's terminal, a dialogue means for dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, and a notification means for detecting abnormal behavior and notifying the parent's terminal of that information in real time. This ensures the safety of the child even when the parent is away, allows the child to receive necessary feedback immediately, and takes appropriate action.

[0214] "Within the home" refers to the environment that forms the basis of life, such as a home or residence.

[0215] "Children's activity status" refers to the series of actions and behaviors that children perform at home, including play, learning, and daily activities.

[0216] "Monitoring means" refers to devices and systems such as sensors and cameras that are used to observe and record children's activities in real time.

[0217] "Data" refers to information or records obtained by monitoring means, including operating conditions, location information, audio data, etc.

[0218] "Parent's device" refers to an information processing device such as a smartphone, tablet, or computer used by the parent.

[0219] "Transmission means" refers to a communication device or communication protocol for transmitting the acquired data to the parent terminal.

[0220] "Settings data" refers to data such as instructions, wishes, rules, etc. regarding a child's activities that a parent sends from their device.

[0221] "Interaction means" refers to a device or system that provides the function of interacting with or playing with a child based on the received setting data.

[0222] "Instruction means" refers to a device or function that transmits commands received from the parent's device to the child.

[0223] "Abnormal behavior" refers to behavior that deviates from a child's normal activity pattern or is risky.

[0224] "Notification means" refers to a communication device or communication protocol that immediately notifies the parent device of abnormal behavior when it is detected.

[0225] "Customization means" refers to devices or systems that provide the ability to tailor the content of the interaction means based on household-specific rules and preferences.

[0226] "Feedback means" refers to a device or system for transmitting feedback data about a child's activities to a parent's device in real time.

[0227] "Voice recognition technology" refers to the technology that converts voice into text data.

[0228] A "generative AI model" refers to an artificial intelligence function that generates appropriate responses or instructions based on input data.

[0229] This system monitors children's activities at home, transmits necessary information to a parent's device, and provides dialogue and play based on the parent's settings. This system consists of a stuffed toy robot (hereinafter referred to as the "device"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's device"), and a server.

[0230] System configuration

[0231] monitoring means

[0232] The device monitors the child's activities in real time using monitoring means such as cameras and acceleration sensors. For example, it can detect whether the child is playing normally or making abnormal movements.

[0233] Transmission method

[0234] The data collected by the device is sent to a server via Wi-Fi or Bluetooth, and the necessary information is then sent to the parent device. This process uses a communication protocol to provide stable data communication.

[0235] Interaction methods

[0236] The device provides dialogue and play based on the configuration data received from the parent's device. For example, if a child says, "Let's build a castle today!", the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The server then uses a generative AI model to generate an appropriate response and sends it back to the device.

[0237] means of instruction

[0238] The device sends instructions to the child based on the commands received from the parent's device. For example, if the parent sends the command "Clean up," the device will use voice synthesis technology to convey the command to the child aloud.

[0239] Notification means

[0240] If abnormal behavior is detected, the device sends the information to the server, which then notifies the parent in real time, allowing the parent to immediately understand the child's situation and take appropriate action.

[0241] Hardware and software used

[0242] Hardware:

[0243] Stuffed toy robot (with camera and accelerometer)

[0244] Smartphone

[0245] software:

[0246] Mobile application (iOS / Android)

[0247] Server (Cloud server: AWS, GCP)

[0248] Speech recognition / speech synthesis engine (Google Cloud Speech-to-Text, Text-to-Speech API, etc.)

[0249] Generative AI models (e.g., OpenAI's GPT-3)

[0250] Data processing and calculation

[0251] Monitoring Data Collection

[0252] The server collects real-time data from the robot's camera and accelerometer to monitor the child's activity.

[0253] Data transmission and processing

[0254] The collected data is sent to a server using a communication protocol such as Google Cloud Pub / Sub, and the data is processed using AWS Lambda, etc. If abnormal behavior is detected, the information is immediately sent as a push notification to the parent's smartphone.

[0255] Dialogue Generation

[0256] The voice input from the child is converted into text data using a speech recognition engine, and an appropriate response is generated by a generative AI model (e.g., GPT-3). The generated text is sent from the server to the device, where it is converted into speech by a speech synthesis engine and responded to the child.

[0257] Specific examples

[0258] Example 1: When a child says to the robot, "I'm looking for my ball."

[0259] The robot uses a voice recognition engine to convert "I'm looking for my ball" into text.

[0260] This text is sent to a server, which uses a generative AI model to generate a response.

[0261] The server sends a response to the terminal saying, "Which color ball do you like? Shall we look for it together?"

[0262] The device uses a speech synthesis engine to convert the response into voice and convey it to the child.

[0263] Prompt Sentence Examples

[0264] User: I'm looking for a ball

[0265] AI: What color ball do you like? Shall we find it together?

[0266] The above is a specific embodiment of the present invention. This system ensures the safety of children even when parents are away, and provides appropriate feedback and dialogue.

[0267] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0268] Step 1:

[0269] The stuffed toy robot (terminal) uses a camera and an accelerometer to collect data to monitor children's activities at home. This data includes the child's movements, location, and speech. The input is raw data from the sensors, and the output is organized activity data.

[0270] Step 2:

[0271] The device transmits the acquired data to the server via Wi-Fi or Bluetooth for transmission in real time to the parent device. The input is the activity data and the output is the data sent to the server. The server receives this data using an appropriate connection protocol (e.g. HTTPS).

[0272] Step 3:

[0273] The server synchronizes the configuration data (e.g., child schedules and household rules) received from the parent's device with the monitoring data. The input is the configuration data from the parent, and the output is the configuration data updated by the server.

[0274] Step 4:

[0275] When a child speaks to the device, the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The input is voice data and the output is text data.

[0276] Step 5:

[0277] The server uses a generative AI model to generate an appropriate response based on the text data. For example, if a child says, "I'm looking for my ball," the generative AI model generates the response, "What color ball do you like? Want to look for it together?" The input is text data, and the output is a response text.

[0278] Step 6:

[0279] The server sends the generated response text to the device, which then converts the text into speech using speech synthesis technology. The input is the response text and the output is audio data. The device then transmits the audio to the child.

[0280] Step 7:

[0281] The device constantly monitors the child's activity and abnormal behavior, and if an abnormality is detected, it sends the information to the server in real time. The input is the monitoring data, and the output is the abnormality detection information.

[0282] Step 8:

[0283] The server immediately notifies the parent device of the anomaly detection information and provides a means for the parent to check the situation through the application.The input is the anomaly detection information and the output is a notification to the parent device.

[0284] Step 9:

[0285] Parents can check their child's situation through the application and send instructions (e.g., "Clean up") as needed. The input is instruction data from the parent, and the output is instructions to the device.

[0286] Step 10:

[0287] The device receives instructions from the parent and uses voice synthesis technology to convey those instructions to the child. The input is instruction data from the parent, and the output is voice instructions to the child.

[0288] These are the program processing steps for the system that realizes this application example. This ensures the safety of children and provides appropriate feedback and dialogue even when parents are not present.

[0289] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0290] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[0291] overview

[0292] The system consists of the following main components:

[0293] 1. A device (stuffed toy robot) that monitors children's activities.

[0294] 2. Smartphones and computers used by parents (parental devices).

[0295] 3. A central management server.

[0296] 4. Emotion engine that recognizes user emotions.

[0297] Entering User Data

[0298] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[0299] Saving and Syncing Settings

[0300] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[0301] Real-time dialogue generation

[0302] When a user (child) speaks to the device, the device uses voice recognition technology to convert the speech into text data and sends that data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[0303] Recognizing emotions and regulating responses

[0304] The device is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words will be generated.

[0305] Send Feedback

[0306] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[0307] Specific examples

[0308] Example 1: A child plays with a "cat sitter"

[0309] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0310] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0311] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[0312] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0313] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0314] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0315] Example 2: When a child is sad

[0316] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[0317] 2. The device sends the emotion data to the server.

[0318] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[0319] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[0320] Example 3: Detecting abnormal behavior in children

[0321] 1. The device monitors the child's movements using an accelerometer and camera.

[0322] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0323] 3. The server notifies the parent device of this information.

[0324] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[0325] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0326] The above is a concrete example of how to implement the present invention. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[0327] The processing flow will be explained below.

[0328] Entering User Data

[0329] Step 1:

[0330] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[0331] Step 2:

[0332] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[0333] Step 3:

[0334] The server receives the authentication information and checks it against information in a database.

[0335] Step 4:

[0336] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[0337] Step 5:

[0338] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[0339] Step 6:

[0340] The terminal transmits the input setting data to the server.

[0341] Saving and Syncing Settings

[0342] Step 1:

[0343] The server receives the setting data sent from the terminal and temporarily stores it.

[0344] Step 2:

[0345] The server checks the integrity of the received data and stores it in a database after checking.

[0346] Step 3:

[0347] The server updates the version of the new setting data and notifies the terminal of the information.

[0348] Step 4:

[0349] The device receives the notification from the server and downloads the new configuration data.

[0350] Step 5:

[0351] The device analyzes the received data and updates its internal settings.

[0352] Real-time dialogue generation

[0353] Step 1:

[0354] The user (child) talks to the device.

[0355] Step 2:

[0356] The device converts voice input into text data using voice recognition technology.

[0357] Step 3:

[0358] The terminal transmits the converted text data to the server.

[0359] Step 4:

[0360] The server receives the text data and uses a generation AI to generate an appropriate response.

[0361] Step 5:

[0362] The server transmits the generated response text data to the terminal.

[0363] Step 6:

[0364] The terminal converts the received response text data into voice using voice synthesis technology.

[0365] Step 7:

[0366] The device plays the converted audio to the child.

[0367] Recognizing emotions and regulating responses

[0368] Step 1:

[0369] When a user (child) speaks to the device, the device's emotion engine analyzes the voice tone and facial expressions to recognize the emotion.

[0370] Step 2:

[0371] The device transmits the analyzed emotion data to the server.

[0372] Step 3:

[0373] The server uses a generative AI to generate an appropriate response based on the received emotional data.

[0374] Step 4:

[0375] The server transmits the generated response text data to the terminal.

[0376] Step 5:

[0377] The terminal converts the received response text data into voice using voice synthesis technology.

[0378] Step 6:

[0379] The device plays appropriate voice responses based on the child's emotions.

[0380] Send Feedback

[0381] Step 1:

[0382] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[0383] Step 2:

[0384] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[0385] Step 3:

[0386] The server receives the abnormality detection information and notifies the parent terminal.

[0387] Step 4:

[0388] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[0389] Step 5:

[0390] The parent terminal transmits the input instruction data to the server.

[0391] Step 6:

[0392] The server receives the instruction data and transmits the data to the terminal.

[0393] Step 7:

[0394] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[0395] Example 2

[0396] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0397] Conventional home child activity monitoring systems primarily detect children's movements and location information, but lack appropriate dialogue or feedback regarding children's emotional state or specific behavior. This can result in parents being unable to respond appropriately to their children's emotional state, potentially having a negative impact on their children's mental health and education. Furthermore, conventional systems have limited functionality for customizing dialogue content based on parental rules and preferences, making it difficult to fully meet the needs of individual families.

[0398] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0399] In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's device, a dialogue means for conducting dialogue and play based on setting data received from the parent's device, an instruction means for sending instructions to the child based on instructions received from the parent's device, an emotion recognition means for analyzing the user's emotions and acquiring emotion data, and an emotion transmission means for transmitting the acquired emotion data to the parent's device. This makes it possible to grasp not only the child's activity status but also their emotional state, enabling customizable dialogue and feedback based on individual family rules and preferences.

[0400] "Monitoring means" refers to a system that includes devices and sensors for monitoring children's activities within the home.

[0401] The "transmission means" is a communication function for transferring data acquired by the monitoring means to the parent terminal.

[0402] "Interaction means" refers to technology or devices for interacting with or playing with a child based on the setting data received from the parent's terminal.

[0403] The "instruction means" is a function for sending instructions to the child based on commands received from the parent's terminal.

[0404] "Emotion recognition means" refers to technology or devices for analyzing the user's tone of voice and facial expressions to obtain emotional data.

[0405] "Emotion transmission means" is a communication function for transmitting the acquired emotion data to the parent terminal.

[0406] "Customization tools" are features that allow users to individually adjust interaction tools based on their household's specific rules and preferences.

[0407] The "feedback means" is a function for sending data about a child's activity status to a parent's device in real time.

[0408] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[0409] The system consists of the following main components:

[0410] 1. A stuffed toy robot (device) that monitors children's activities

[0411] 2. Smartphones and computers used by parents (parental devices)

[0412] 3. Central management server

[0413] 4. Emotion engine that recognizes user emotions

[0414] Entering User Data

[0415] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[0416] Saving and Syncing Settings

[0417] The server stores the received setting data in a database and synchronizes that data with the device. Specifically, a database such as MySQL is used. The device analyzes the received setting data and updates its internal settings. For example, based on the setting data for a game, the server might suggest to the child, "Let's build a castle today!"

[0418] Real-time dialogue generation

[0419] When a user (child) speaks to the device, the device uses speech recognition technology (e.g., Google Speech-to-Text API) to convert the speech into text data and sends that data to the server. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response and sends the response text to the device. The device then uses speech synthesis technology (e.g., Amazon Polly) to convert the response text into speech and returns it to the child.

[0420] Recognizing emotions and regulating responses

[0421] The device is equipped with an emotion engine (e.g., Affectiva's SDK) that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words is generated.

[0422] Send Feedback

[0423] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity and any abnormalities, and sends the data to a server. The server then notifies the parent's device of this data in real time. The parent can then check the child's status through an application and send instructions as necessary.

[0424] Specific examples

[0425] Example 1: A child plays with a "cat sitter"

[0426] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0427] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0428] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[0429] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0430] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0431] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0432] Example 2: When a child is sad

[0433] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[0434] 2. The device sends the emotion data to the server.

[0435] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[0436] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[0437] Example 3: Detecting abnormal behavior in children

[0438] 1. The device monitors the child's movements using an accelerometer and camera.

[0439] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0440] 3. The server notifies the parent device of this information.

[0441] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[0442] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0443] Example prompts for generative AI models

[0444] "Please enter the setting data required to do ____"

[0445] "Generate a response if a child is crying"

[0446] The above is a concrete example of how the present invention can be implemented. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[0447] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0448] Step 1:

[0449] A user (parent) launches an application on a smartphone or computer and enters configuration data such as household rules, preferences, schedules, etc. This configuration data includes information related to children's activities within the home, and specific actions involve entering data into fields on the application's settings screen.

[0450] Input: Household-specific rules and schedule information

[0451] Output: Setting data

[0452] Step 2:

[0453] The server receives the setting data sent by the user, encrypts it, and stores it in a database, such as MySQL. The received data is sent via the HTTP protocol.

[0454] Input: Configuration data submitted by the user

[0455] Output: Encrypted configuration data stored in the database

[0456] Step 3:

[0457] The device periodically sends a synchronization request to the server to obtain the latest configuration data. At this time, the device analyzes the data received from the server and updates its internal settings. Specifically, the device sends an API request to the server and receives a response.

[0458] Input: Configuration data stored on the server

[0459] Output: Updated internal settings

[0460] Step 4:

[0461] When a user (child) speaks to the device, the device captures the voice using a built-in microphone and converts it into text data using speech recognition technology (such as the Google Speech-to-Text API). Specifically, utterances such as "I'll bring the block!" are converted into text.

[0462] Input: Voice input from child

[0463] Output: Text data

[0464] Step 5:

[0465] The device sends the converted text data to the server, which uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the text data. The generated response is then sent back to the device. Specifically, the server generates the response using an NLP engine.

[0466] Input: Text data sent from the terminal

[0467] Output: The generated response text

[0468] Step 6:

[0469] The device converts the response text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and responds to the child through the speaker. Specifically, the device generates synthesized speech and plays it back on the speaker.

[0470] Input: Response text sent by the server

[0471] Output: Synthesized speech

[0472] Step 7:

[0473] The device uses an emotion engine (such as Affectiva's SDK) to analyze the child's emotions by analyzing voice tones and facial expressions in real time. The analysis results are sent to the server as emotion data. For example, if a child is crying, emotion data is generated.

[0474] Input: Child's vocal tone and facial expression data

[0475] Output: Generated emotion data

[0476] Step 8:

[0477] The server generates an appropriate response based on the emotion data. For example, if the child is sad, it generates an encouraging response. The generated response text is then sent back to the device.

[0478] Input: Emotion data sent from the device

[0479] Output: Sentiment-based response text

[0480] Step 9:

[0481] The device uses speech synthesis technology to convert the generated response text and respond to the child, playing the synthesized voice from the speaker.

[0482] Input: Response text sent by the server

[0483] Output: Synthesized speech

[0484] Step 10:

[0485] The device monitors the child's movements in real time using an accelerometer and camera. If it detects any abnormal movements, it sends the data to a server. For example, if a child falls, that data is generated.

[0486] Input: Data on children's movements

[0487] Output: Abnormal behavior data

[0488] Step 11:

[0489] The server receives the abnormal behavior data and notifies the parent device in real time via push notifications, emails, etc.

[0490] Input: Abnormal behavior data sent from the device

[0491] Output: Notification sent to the parent device

[0492] Step 12:

[0493] The user (parent) checks the child's status through the application and sends appropriate instructions as necessary. The instructions are transmitted to the device via the server. Specifically, the parent enters instructions on the application screen and presses the send button.

[0494] Input: Parental instruction data

[0495] Output: Instruction data sent to the terminal

[0496] Step 13:

[0497] The device receives instruction data sent by the parent and uses voice synthesis technology to convey the message to the child, for example, by playing the message "Are you OK?" over the speaker.

[0498] Input: Instruction data sent from the parent

[0499] Output: Synthesized speech

[0500] (Application example 2)

[0501] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0502] There is a need for a system that monitors children's activities at home and allows parents to respond appropriately based on their children's emotional state. Current systems lack the ability to recognize children's emotions, making it difficult for parents to respond based on their children's emotions. There is also a need for a system that allows parents to understand their children's situation in real time and send appropriate instructions, even when they are at work or out.

[0503] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to the parent's terminal, a dialogue means for conducting dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, an emotion recognition means for recognizing the user's emotion, and a response generation means for generating an appropriate response based on the recognized emotion data. This enables detailed feedback according to the child's emotional state, supports educational dialogue and appropriate play, and allows parents to watch over their children with peace of mind even when they are not at home.

[0504] "Monitoring means for monitoring children's activities within the home" refers to devices and systems such as sensors and cameras that monitor children's movements and behavior within the home in real time and obtain the necessary data.

[0505] The "transmission means for transmitting data acquired by the monitoring means to the parent's terminal" refers to a communication device and its system for transmitting data acquired by the monitoring means to a smartphone or computer used by the parent.

[0506] "An interaction means for interacting and playing based on setting data received from the parent's terminal" refers to a device and system for interacting with and playing with a child in accordance with setting data sent from the parent's terminal.

[0507] An "instruction means for sending instructions to a child based on an instruction received from a parent's terminal" is a device and system for receiving an instruction sent from a parent's terminal and encouraging a child to take a specific action based on that instruction.

[0508] The "emotion recognition means for recognizing the user's emotions" refers to a device and system for recognizing the user's emotional state in real time by analyzing the user's tone of voice and facial expressions.

[0509] The "response generating means for generating an appropriate response based on the recognized emotion data" refers to a device and system for generating an appropriate response to a user based on the recognized emotion data.

[0510] A system for implementing the present invention comprises the following major components:

[0511] 1. A stuffed toy robot equipped with a camera and sensors that can be used as a device (monitoring means) to monitor children's activities at home.

[0512] 2. Smartphones and computers used by parents (parental devices).

[0513] 3. Server.

[0514] 4. Emotion recognition means to recognize the user's emotions.

[0515] 5. A response generation means for generating an appropriate response based on the recognized emotion data.

[0516] Monitoring and Transmission Methods

[0517] The monitoring means uses cameras and sensors to monitor the child's activities in real time, making it possible to detect what the child is currently doing. The acquired data is then sent to the parent's device by the transmission means. Communication technologies such as Wi-Fi and Bluetooth can be used for transmission.

[0518] Interaction and instruction methods

[0519] The dialogue means operates based on the setting data received from the parent's device. For example, if the parent sets up a game such as "building a castle" through the application, the dialogue means will make suggestions to the child based on that information. The dialogue means also uses voice recognition technology to converse with the child. This voice data is converted into text data and sent to the server.

[0520] The instruction means prompts the child to perform a specific action based on a command received from the parent's device. For example, if the parent commands the child to "bring a block," the instruction means prompts the child to perform the action in accordance with the command.

[0521] Emotion recognition and response generation

[0522] The emotion recognition means analyzes the user's (child's) emotions in real time using voice tone and facial expression analysis technology. The analysis results are sent to the server, and the response generation means generates an appropriate response. For example, if a child has a sad expression, the response generation means generates encouraging words such as "What's wrong? It's okay."

[0523] Feedback and Notifications

[0524] Various sensors monitor the child's activity and any abnormalities, and send the data to a server. Feedback is used to notify the parent's device in real time. The parent can check the child's condition through the application and send instructions as necessary. For example, if a child falls, data stating "a fall has been detected" is sent to the server, and a notification is also sent to the parent's device.

[0525] Specific examples

[0526] 1. Example of a set interaction: A parent sets up a game called "Build a castle" in the application. The device suggests, "Let's build a castle today!", and the child responds, "I'll bring the blocks!". This interaction is coordinated via the server.

[0527] 2. Example of emotion recognition: If a child is crying, the emotion recognition means will analyze this and generate an encouraging message such as "What's wrong? It's okay" and return it to the child.

[0528] 3. Example of abnormal behavior detection: An accelerometer or camera detects abnormal behavior and sends a real-time notification to the parent's device saying, "A fall has been detected." The parent can then check the notification and send appropriate instructions.

[0529] Example prompts to input to the generative AI model:

[0530] "Design a system that recognizes the emotions of workers from video data and generates messages to help them work safely and efficiently. Please include specific example instructions."

[0531] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0532] Step 1:

[0533] The device uses cameras and sensors to monitor the child's activities in real time. The input is camera footage and sensor data, and the device acquires the child's movements and emotions based on this. The output is the acquired activity data.

[0534] Step 2:

[0535] The device sends the data acquired by the monitoring means to the server. The transmission method is Wi-Fi or Bluetooth, and the input is activity data, and the output is data sent to the server. The server receives this data and analyzes it.

[0536] Step 3:

[0537] Based on the configuration data received from the parent's device, the server generates a dialogue or play scenario and sends it to the device. The input is the parent's configuration data, and the output is the dialogue or play scenario. This allows the device to begin a dialogue with the child according to the settings.

[0538] Step 4:

[0539] As the device engages in a dialogue with the child, it uses voice recognition technology to convert the child's voice into text data. The input is the child's voice data, and the output is text data. The content of the dialogue is sent to the server.

[0540] Step 5:

[0541] The server analyzes the received text data and generates an appropriate response. The input is the child's text data, and the output is the response text. The response text is sent to the terminal and returned to the child via voice synthesis.

[0542] Step 6:

[0543] The device's emotion recognition means analyzes the user's (child's) voice tone and facial expression in real time and generates emotion data. The input is voice tone and facial expression data, and the output is emotion data. This data is sent to the server.

[0544] Step 7:

[0545] The server receives the emotion data and generates an appropriate response based on the emotion. The input is emotion data, and the output is a response text based on the emotion. For example, if the analysis finds that a child is sad, an encouraging message is generated.

[0546] Step 8:

[0547] The server sends the generated response text to the device, which then uses speech synthesis technology to communicate it to the child. The input is the response text, and the output is a voice message to the child.

[0548] Step 9:

[0549] The device uses various sensors to monitor children's abnormal behavior and notifies the server if an abnormality is detected. The input is sensor data, and the output is notification data of the abnormality detection.

[0550] Step 10:

[0551] The server sends a notification of anomaly detection to the parent's device in real time. The input is the anomaly detection data, and the output is a notification to the parent's device. When the parent receives the notification, they can check the child's status through the application and send necessary instructions to the device.

[0552] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0553] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0554] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0555] [Second embodiment]

[0556] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0557] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0558] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0559] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0560] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0561] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0562] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0563] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0564] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0565] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0566] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0567] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0568] The present invention provides a system that monitors the activities of children at home, transmits necessary information to a parent's terminal, and provides dialogue and play based on settings data from the parent. Specific embodiments for carrying out the present invention will be described below.

[0569] overview

[0570] This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The terminal monitors the child's activities and sends this information to the parent's terminal via the server. It also provides dialogue and play based on the setting data received from the parent, and conveys instructions from the parent to the child. Furthermore, a feedback function reports the child's activities to the parent in real time.

[0571] Entering User Data

[0572] The user (parent) launches the application on their smartphone or computer and inputs their household's specific rules, preferences, schedules, etc. This setting data is sent to the server and synchronized with the device.

[0573] Saving and Syncing Settings

[0574] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[0575] Real-time dialogue generation

[0576] When the user (child) speaks to the device, the device uses speech recognition technology to convert the speech into text and sends the data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[0577] Send Feedback

[0578] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[0579] Specific examples

[0580] Example 1: A child plays with a "cat sitter"

[0581] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0582] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0583] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[0584] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0585] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0586] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0587] Example 2: Detecting abnormal behavior in children

[0588] 1. The device monitors the child's movements using an accelerometer and camera.

[0589] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0590] 3. The server notifies the parent device of this information.

[0591] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[0592] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0593] The above is a specific embodiment for carrying out the present invention. This system ensures the safety of children even when their parents are away, and can provide educational conversations and play.

[0594] The processing flow will be explained below.

[0595] Entering User Data

[0596] Step 1:

[0597] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[0598] Step 2:

[0599] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[0600] Step 3:

[0601] The server receives the authentication information and checks it against information in a database.

[0602] Step 4:

[0603] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[0604] Step 5:

[0605] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[0606] Step 6:

[0607] The terminal transmits the input setting data to the server.

[0608] Saving and Syncing Settings

[0609] Step 1:

[0610] The server receives the setting data sent from the terminal and temporarily stores it.

[0611] Step 2:

[0612] The server checks the integrity of the received data and stores it in the database once the check is complete.

[0613] Step 3:

[0614] The server updates the version of the new setting data and notifies the terminal of the information.

[0615] Step 4:

[0616] The device receives the notification from the server and downloads the new configuration data.

[0617] Step 5:

[0618] The device analyzes the received data and updates its internal settings.

[0619] Real-time dialogue generation

[0620] Step 1:

[0621] The user (child) talks to "Nyanko Sitter."

[0622] Step 2:

[0623] The device converts voice input into text data using voice recognition technology.

[0624] Step 3:

[0625] The terminal transmits the converted text data to the server.

[0626] Step 4:

[0627] The server receives the text data and uses a generation AI to generate an appropriate response.

[0628] Step 5:

[0629] The server transmits the generated response text data to the terminal.

[0630] Step 6:

[0631] The terminal converts the received response text data into voice using voice synthesis technology.

[0632] Step 7:

[0633] The device plays the converted audio to the child.

[0634] Send Feedback

[0635] Step 1:

[0636] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[0637] Step 2:

[0638] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[0639] Step 3:

[0640] The server receives the abnormality detection information and notifies the parent terminal.

[0641] Step 4:

[0642] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[0643] Step 5:

[0644] The parent terminal transmits the input instruction data to the server.

[0645] Step 6:

[0646] The server receives the instruction data and transmits the data to the terminal.

[0647] Step 7:

[0648] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[0649] Example 1

[0650] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0651] There is a need to provide a means to effectively monitor children's activities at home and to provide a safe environment for parents to watch over them even when they are not around. It is also important to provide appropriate dialogue and play to facilitate two-way communication between children and parents. Without such a system, it is difficult to ensure children's safety and promote educational dialogue and play.

[0652] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0653] In this invention, the server includes: a user input means for a parent to input family rules, preferences, schedules, etc. on a smartphone or computer; a storage / synchronization means for saving the setting data entered by the user input means in a database and synchronizing it with the device; a voice recognition means for converting the child's speech into text using voice recognition technology and sending the data to the server; a response generation means for the server to generate an appropriate response using a generative AI model and send it to the device; a voice synthesis means for the device to convert the response text into speech using voice synthesis technology and return it to the child; a monitoring means for the device to monitor the child's activity status and abnormalities using various sensors and send the data to the server; a notification means for the server to notify the parent's device of the monitored data in real time; an instruction sending means for the parent to check the child's status through an application and send necessary instructions; and an instruction execution means for the device to execute the parent's instructions using voice synthesis technology and convey them to the child. This enables monitoring of the child's activity status and smooth communication with the parent.

[0654] "User input means" refers to the means by which parents input their household's specific rules, preferences, schedules, etc. on their smartphone or computer.

[0655] The "storage and synchronization means" is a means for storing the setting data input by the user input means in a database and synchronizing it with the terminal.

[0656] The "voice recognition means" is a means of converting the voice spoken by the child into text using voice recognition technology and sending that data to a server.

[0657] The "response generation means" is a means by which the server uses a generative AI model to generate an appropriate response and transmits it to the terminal.

[0658] "Speech synthesis means" refers to the means by which the terminal uses speech synthesis technology to convert the response text into speech and return it to the child.

[0659] The "monitoring means" is a means by which the device uses various sensors to monitor the child's activity status and any abnormalities, and sends that data to the server.

[0660] The "notification means" is a means by which the server notifies the parent terminal of the monitoring data in real time.

[0661] The "means for sending instructions" is a means by which parents can check their child's status through the application and send necessary instructions.

[0662] The "means for executing instructions" refers to the means by which the device executes instructions from the parent using voice synthesis technology and conveys them to the child.

[0663] The present invention provides a system that monitors the activities of children at home, allowing parents to watch over them with peace of mind. This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The specific hardware and software configurations and their operation are described below.

[0664] Entering User Data

[0665] The user (parent) launches a dedicated application on a smartphone or computer and inputs setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack at 3 pm every day" or "Play with building blocks." This setting data is then sent to a server via the Internet.

[0666] Saving and Syncing Settings

[0667] The server receives the configuration data sent from the parent and stores it in a database. This data is managed in real time. The server synchronizes the saved data with the device. For example, if a network error occurs, the server will repeatedly retry, and if synchronization is successful, the device will receive the new configuration data.

[0668] Real-time dialogue generation

[0669] When a user (child) speaks to the device, the device converts the speech into text using voice recognition technology (e.g., Google Cloud Speech-to-Text), and this text data is sent to the server.

[0670] The server uses a generative AI model (e.g., OpenAI GPT-3.5) to generate an appropriate response. For example, in response to the question, "What do you want to do?", it generates a response like, "Let's play with building blocks!" The generated text data is sent to the device.

[0671] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the text into speech and then returns that speech to the child.

[0672] Send Feedback

[0673] The device uses an accelerometer and camera to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to a server.

[0674] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send instructions via the application if necessary.

[0675] Carrying out parental instructions

[0676] The user (parent) can check the child's status through the application and input and send necessary instructions, such as "Are you OK?"

[0677] The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and convey it to the child, for example, "Mommy is asking if you're OK."

[0678] Specific examples

[0679] Example 1: A child plays with a "cat sitter"

[0680] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0681] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0682] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[0683] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0684] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0685] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0686] Example 2: Detecting abnormal behavior in children

[0687] 1. The device monitors the child's movements using an accelerometer and camera.

[0688] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0689] 3. The server notifies the parent device of this information.

[0690] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[0691] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0692] Examples of prompt statements

[0693] Enter the following prompts into the generative AI model:

[0694] "Situation: Child has fallen. Parents are concerned. Generate an appropriate response message for the child."

[0695] Example of the resulting result:

[0696] "Are you okay? Can you get up now? Is anything hurting? Be careful."

[0697] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0698] Step 1: Enter and submit user data

[0699] The user (parent) launches a dedicated application on their smartphone or computer. They input setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack every day at 3 p.m." or "Play with building blocks." They confirm the input setting data and press the "Submit" button.

[0700] Input: Configuration data such as household-specific rules, preferences, and schedules

[0701] Output: The configuration data sent

[0702] Specific operation: The user operates the application, inputs and submits setting data.

[0703] Step 2: Save and sync configuration data

[0704] The server receives the setting data sent by the parent. It stores the received data in a database and manages it in real time. The stored data is time-stamped. The server then starts the process of synchronizing the stored data with the device. For example, if the setting data is "Eat a snack at 3 pm," it is synchronized with the device.

[0705] Input: Submitted configuration data

[0706] Output: Settings data synced to the device

[0707] Specific behavior: The server stores the received data and starts the synchronization process.

[0708] Step 3: Handling real-time interactions with your child

[0709] The user (child) speaks to the device, asking questions such as "What are you doing today?", and the device uses voice recognition technology (e.g., Google Cloud Speech-to-Text) to convert the speech into text. This text data is then sent to the server.

[0710] Input: Child's voice data

[0711] Output: Audio data converted to text

[0712] Specific operation: The device uses voice recognition technology to convert the child's voice into text and send it to the server.

[0713] Step 4: Server Response Generation

[0714] The server generates an appropriate response using a generative AI model (e.g., OpenAI GPT-3.5) and sends the generated response text to the device. For example, in response to the question "What do you want to do?", the server generates the response "Let's play with building blocks!"

[0715] Input: Text-converted audio data

[0716] Output: The generated response text

[0717] Specific operation: The server uses the generative AI model to generate an appropriate response and sends it to the device.

[0718] Step 5: Reply with a text-to-speech response

[0719] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the response text received from the server into speech, which is then returned to the child. For example, the device might respond with a voice saying, "Let's play with building blocks!"

[0720] Input: Generated response text

[0721] Output: Audio data

[0722] Specific operation: The device uses speech synthesis technology to convert the response text into speech and return it to the child.

[0723] Step 6: Monitor your child's activity

[0724] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to the server.

[0725] Input: Child's activity data

[0726] Output: Anomaly detection data

[0727] Specific operation: The device uses various sensors to monitor the child's activities, and if it detects any abnormalities, it sends the information to the server.

[0728] Step 7: Server notification

[0729] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send any necessary instructions via the application.

[0730] Input: Anomaly detection data

[0731] Output: Notification data to parent device

[0732] Specific operation: The server notifies the parent device of abnormal data in real time.

[0733] Step 8: Send and Execute Parental Instructions

[0734] The user (parent) checks the child's status through the application and inputs and sends the necessary instructions. For example, a user can send the instruction "Are you OK?" The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and conveys it to the child. For example, the user can tell the child "Mommy is asking if you're OK."

[0735] Input: Parent instruction data

[0736] Output: Voice instructions to the child

[0737] Specific operation: The parent inputs instructions into the application, the server sends the instructions to the device, and the device uses speech synthesis technology to convert the instructions into voice and convey them to the child.

[0738] (Application example 1)

[0739] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0740] In modern families, it is extremely important for parents to monitor their children's activities in real time and respond quickly when necessary. However, conventional technologies have struggled to ensure children's safety while parents are away and to provide educational conversations and play. Furthermore, there have been insufficient means for immediately notifying parents if their children exhibit abnormal behavior. The present invention addresses these issues.

[0741] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0742] In this invention, the server includes a monitoring means for monitoring the child's activities within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's terminal, a dialogue means for dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, and a notification means for detecting abnormal behavior and notifying the parent's terminal of that information in real time. This ensures the safety of the child even when the parent is away, allows the child to receive necessary feedback immediately, and takes appropriate action.

[0743] "Within the home" refers to the environment that forms the basis of life, such as a home or residence.

[0744] "Children's activity status" refers to the series of actions and behaviors that children perform at home, including play, learning, and daily activities.

[0745] "Monitoring means" refers to devices and systems such as sensors and cameras that are used to observe and record children's activities in real time.

[0746] "Data" refers to information or records obtained by monitoring means, including operating conditions, location information, audio data, etc.

[0747] "Parent's device" refers to an information processing device such as a smartphone, tablet, or computer used by the parent.

[0748] "Transmission means" refers to a communication device or communication protocol for transmitting the acquired data to the parent terminal.

[0749] "Settings data" refers to data such as instructions, wishes, rules, etc. regarding a child's activities that a parent sends from their device.

[0750] "Interaction means" refers to a device or system that provides the function of interacting with or playing with a child based on the received setting data.

[0751] "Instruction means" refers to a device or function that transmits commands received from the parent's device to the child.

[0752] "Abnormal behavior" refers to behavior that deviates from a child's normal activity pattern or is risky.

[0753] "Notification means" refers to a communication device or communication protocol that immediately notifies the parent device of abnormal behavior when it is detected.

[0754] "Customization means" refers to devices or systems that provide the ability to tailor the content of the interaction means based on household-specific rules and preferences.

[0755] "Feedback means" refers to a device or system for transmitting feedback data about a child's activities to a parent's device in real time.

[0756] "Voice recognition technology" refers to the technology that converts voice into text data.

[0757] A "generative AI model" refers to an artificial intelligence function that generates appropriate responses or instructions based on input data.

[0758] This system monitors children's activities at home, transmits necessary information to a parent's device, and provides dialogue and play based on the parent's settings. This system consists of a stuffed toy robot (hereinafter referred to as the "device"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's device"), and a server.

[0759] System configuration

[0760] monitoring means

[0761] The device monitors the child's activities in real time using monitoring means such as cameras and acceleration sensors. For example, it can detect whether the child is playing normally or making abnormal movements.

[0762] Transmission method

[0763] The data collected by the device is sent to a server via Wi-Fi or Bluetooth, and the necessary information is then sent to the parent device. This process uses a communication protocol to provide stable data communication.

[0764] Interaction methods

[0765] The device provides dialogue and play based on the configuration data received from the parent's device. For example, if a child says, "Let's build a castle today!", the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The server then uses a generative AI model to generate an appropriate response and sends it back to the device.

[0766] means of instruction

[0767] The device sends instructions to the child based on the commands received from the parent's device. For example, if the parent sends the command "Clean up," the device will use voice synthesis technology to convey the command to the child aloud.

[0768] Notification means

[0769] If abnormal behavior is detected, the device sends the information to the server, which then notifies the parent in real time, allowing the parent to immediately understand the child's situation and take appropriate action.

[0770] Hardware and software used

[0771] Hardware:

[0772] Stuffed toy robot (with camera and accelerometer)

[0773] Smartphone

[0774] software:

[0775] Mobile application (iOS / Android)

[0776] Server (Cloud server: AWS, GCP)

[0777] Speech recognition / speech synthesis engine (Google Cloud Speech-to-Text, Text-to-Speech API, etc.)

[0778] Generative AI models (e.g., OpenAI's GPT-3)

[0779] Data processing and calculation

[0780] Monitoring Data Collection

[0781] The server collects real-time data from the robot's camera and accelerometer to monitor the child's activity.

[0782] Data transmission and processing

[0783] The collected data is sent to a server using a communication protocol such as Google Cloud Pub / Sub, and the data is processed using AWS Lambda, etc. If abnormal behavior is detected, the information is immediately sent as a push notification to the parent's smartphone.

[0784] Dialogue Generation

[0785] The voice input from the child is converted into text data using a speech recognition engine, and an appropriate response is generated by a generative AI model (e.g., GPT-3). The generated text is sent from the server to the device, where it is converted into speech by a speech synthesis engine and responded to the child.

[0786] Specific examples

[0787] Example 1: When a child says to the robot, "I'm looking for my ball."

[0788] The robot uses a voice recognition engine to convert "I'm looking for my ball" into text.

[0789] This text is sent to a server, which uses a generative AI model to generate a response.

[0790] The server sends a response to the terminal saying, "Which color ball do you like? Shall we look for it together?"

[0791] The device uses a speech synthesis engine to convert the response into voice and convey it to the child.

[0792] Prompt Sentence Examples

[0793] User: I'm looking for a ball

[0794] AI: What color ball do you like? Shall we find it together?

[0795] The above is a specific embodiment of the present invention. This system ensures the safety of children even when parents are away, and provides appropriate feedback and dialogue.

[0796] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0797] Step 1:

[0798] The stuffed toy robot (terminal) uses a camera and an accelerometer to collect data to monitor children's activities at home. This data includes the child's movements, location, and speech. The input is raw data from the sensors, and the output is organized activity data.

[0799] Step 2:

[0800] The device transmits the acquired data to the server via Wi-Fi or Bluetooth for transmission in real time to the parent device. The input is the activity data and the output is the data sent to the server. The server receives this data using an appropriate connection protocol (e.g. HTTPS).

[0801] Step 3:

[0802] The server synchronizes the configuration data (e.g., child schedules and household rules) received from the parent's device with the monitoring data. The input is the configuration data from the parent, and the output is the configuration data updated by the server.

[0803] Step 4:

[0804] When a child speaks to the device, the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The input is voice data and the output is text data.

[0805] Step 5:

[0806] The server uses a generative AI model to generate an appropriate response based on the text data. For example, if a child says, "I'm looking for my ball," the generative AI model generates the response, "What color ball do you like? Want to look for it together?" The input is text data, and the output is a response text.

[0807] Step 6:

[0808] The server sends the generated response text to the device, which then converts the text into speech using speech synthesis technology. The input is the response text and the output is audio data. The device then transmits the audio to the child.

[0809] Step 7:

[0810] The device constantly monitors the child's activity and abnormal behavior, and if an abnormality is detected, it sends the information to the server in real time. The input is the monitoring data, and the output is the abnormality detection information.

[0811] Step 8:

[0812] The server immediately notifies the parent device of the anomaly detection information and provides a means for the parent to check the situation through the application.The input is the anomaly detection information and the output is a notification to the parent device.

[0813] Step 9:

[0814] Parents can check their child's situation through the application and send instructions (e.g., "Clean up") as needed. The input is instruction data from the parent, and the output is instructions to the device.

[0815] Step 10:

[0816] The device receives instructions from the parent and uses voice synthesis technology to convey those instructions to the child. The input is instruction data from the parent, and the output is voice instructions to the child.

[0817] These are the program processing steps for the system that realizes this application example. This ensures the safety of children and provides appropriate feedback and dialogue even when parents are not present.

[0818] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0819] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[0820] overview

[0821] The system consists of the following main components:

[0822] 1. A device (stuffed toy robot) that monitors children's activities.

[0823] 2. Smartphones and computers used by parents (parental devices).

[0824] 3. A central management server.

[0825] 4. Emotion engine that recognizes user emotions.

[0826] Entering User Data

[0827] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[0828] Saving and Syncing Settings

[0829] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[0830] Real-time dialogue generation

[0831] When a user (child) speaks to the device, the device uses voice recognition technology to convert the speech into text data and sends that data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[0832] Recognizing emotions and regulating responses

[0833] The device is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words will be generated.

[0834] Send Feedback

[0835] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[0836] Specific examples

[0837] Example 1: A child plays with a "cat sitter"

[0838] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0839] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0840] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[0841] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0842] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0843] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0844] Example 2: When a child is sad

[0845] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[0846] 2. The device sends the emotion data to the server.

[0847] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[0848] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[0849] Example 3: Detecting abnormal behavior in children

[0850] 1. The device monitors the child's movements using an accelerometer and camera.

[0851] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0852] 3. The server notifies the parent device of this information.

[0853] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[0854] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0855] The above is a concrete example of how to implement the present invention. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[0856] The processing flow will be explained below.

[0857] Entering User Data

[0858] Step 1:

[0859] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[0860] Step 2:

[0861] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[0862] Step 3:

[0863] The server receives the authentication information and checks it against information in a database.

[0864] Step 4:

[0865] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[0866] Step 5:

[0867] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[0868] Step 6:

[0869] The terminal transmits the input setting data to the server.

[0870] Saving and Syncing Settings

[0871] Step 1:

[0872] The server receives the setting data sent from the terminal and temporarily stores it.

[0873] Step 2:

[0874] The server checks the integrity of the received data and stores it in a database after checking.

[0875] Step 3:

[0876] The server updates the version of the new setting data and notifies the terminal of the information.

[0877] Step 4:

[0878] The device receives the notification from the server and downloads the new configuration data.

[0879] Step 5:

[0880] The device analyzes the received data and updates its internal settings.

[0881] Real-time dialogue generation

[0882] Step 1:

[0883] The user (child) talks to the device.

[0884] Step 2:

[0885] The device converts voice input into text data using voice recognition technology.

[0886] Step 3:

[0887] The terminal transmits the converted text data to the server.

[0888] Step 4:

[0889] The server receives the text data and uses a generation AI to generate an appropriate response.

[0890] Step 5:

[0891] The server transmits the generated response text data to the terminal.

[0892] Step 6:

[0893] The terminal converts the received response text data into voice using voice synthesis technology.

[0894] Step 7:

[0895] The device plays the converted audio to the child.

[0896] Recognizing emotions and regulating responses

[0897] Step 1:

[0898] When a user (child) speaks to the device, the device's emotion engine analyzes the voice tone and facial expressions to recognize the emotion.

[0899] Step 2:

[0900] The device transmits the analyzed emotion data to the server.

[0901] Step 3:

[0902] The server uses a generative AI to generate an appropriate response based on the received emotional data.

[0903] Step 4:

[0904] The server transmits the generated response text data to the terminal.

[0905] Step 5:

[0906] The terminal converts the received response text data into voice using voice synthesis technology.

[0907] Step 6:

[0908] The device plays appropriate voice responses based on the child's emotions.

[0909] Send Feedback

[0910] Step 1:

[0911] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[0912] Step 2:

[0913] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[0914] Step 3:

[0915] The server receives the abnormality detection information and notifies the parent terminal.

[0916] Step 4:

[0917] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[0918] Step 5:

[0919] The parent terminal transmits the input instruction data to the server.

[0920] Step 6:

[0921] The server receives the instruction data and transmits the data to the terminal.

[0922] Step 7:

[0923] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[0924] Example 2

[0925] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0926] Conventional home child activity monitoring systems primarily detect children's movements and location information, but lack appropriate dialogue or feedback regarding children's emotional state or specific behavior. This can result in parents being unable to respond appropriately to their children's emotional state, potentially having a negative impact on their children's mental health and education. Furthermore, conventional systems have limited functionality for customizing dialogue content based on parental rules and preferences, making it difficult to fully meet the needs of individual families.

[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0928] In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's device, a dialogue means for conducting dialogue and play based on setting data received from the parent's device, an instruction means for sending instructions to the child based on instructions received from the parent's device, an emotion recognition means for analyzing the user's emotions and acquiring emotion data, and an emotion transmission means for transmitting the acquired emotion data to the parent's device. This makes it possible to grasp not only the child's activity status but also their emotional state, enabling customizable dialogue and feedback based on individual family rules and preferences.

[0929] "Monitoring means" refers to a system that includes devices and sensors for monitoring children's activities within the home.

[0930] The "transmission means" is a communication function for transferring data acquired by the monitoring means to the parent terminal.

[0931] "Interaction means" refers to technology or devices for interacting with or playing with a child based on the setting data received from the parent's terminal.

[0932] The "instruction means" is a function for sending instructions to the child based on commands received from the parent's terminal.

[0933] "Emotion recognition means" refers to technology or devices for analyzing the user's tone of voice and facial expressions to obtain emotional data.

[0934] "Emotion transmission means" is a communication function for transmitting the acquired emotion data to the parent terminal.

[0935] "Customization tools" are features that allow users to individually adjust interaction tools based on their household's specific rules and preferences.

[0936] The "feedback means" is a function for sending data about a child's activity status to a parent's device in real time.

[0937] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[0938] The system consists of the following main components:

[0939] 1. A stuffed toy robot (device) that monitors children's activities

[0940] 2. Smartphones and computers used by parents (parental devices)

[0941] 3. Central management server

[0942] 4. Emotion engine that recognizes user emotions

[0943] Entering User Data

[0944] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[0945] Saving and Syncing Settings

[0946] The server stores the received setting data in a database and synchronizes that data with the device. Specifically, a database such as MySQL is used. The device analyzes the received setting data and updates its internal settings. For example, based on the setting data for a game, the server might suggest to the child, "Let's build a castle today!"

[0947] Real-time dialogue generation

[0948] When a user (child) speaks to the device, the device uses speech recognition technology (e.g., Google Speech-to-Text API) to convert the speech into text data and sends that data to the server. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response and sends the response text to the device. The device then uses speech synthesis technology (e.g., Amazon Polly) to convert the response text into speech and returns it to the child.

[0949] Recognizing emotions and regulating responses

[0950] The device is equipped with an emotion engine (e.g., Affectiva's SDK) that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words is generated.

[0951] Send Feedback

[0952] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity and any abnormalities, and sends the data to a server. The server then notifies the parent's device of this data in real time. The parent can then check the child's status through an application and send instructions as necessary.

[0953] Specific examples

[0954] Example 1: A child plays with a "cat sitter"

[0955] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[0956] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[0957] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[0958] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[0959] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[0960] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[0961] Example 2: When a child is sad

[0962] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[0963] 2. The device sends the emotion data to the server.

[0964] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[0965] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[0966] Example 3: Detecting abnormal behavior in children

[0967] 1. The device monitors the child's movements using an accelerometer and camera.

[0968] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[0969] 3. The server notifies the parent device of this information.

[0970] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[0971] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[0972] Example prompts for generative AI models

[0973] "Please enter the setting data required to do ____"

[0974] "Generate a response if a child is crying"

[0975] The above is a concrete example of how the present invention can be implemented. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[0976] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0977] Step 1:

[0978] A user (parent) launches an application on a smartphone or computer and enters configuration data such as household rules, preferences, schedules, etc. This configuration data includes information related to children's activities within the home, and specific actions involve entering data into fields on the application's settings screen.

[0979] Input: Household-specific rules and schedule information

[0980] Output: Setting data

[0981] Step 2:

[0982] The server receives the setting data sent by the user, encrypts it, and stores it in a database, such as MySQL. The received data is sent via the HTTP protocol.

[0983] Input: Configuration data submitted by the user

[0984] Output: Encrypted configuration data stored in the database

[0985] Step 3:

[0986] The device periodically sends a synchronization request to the server to obtain the latest configuration data. At this time, the device analyzes the data received from the server and updates its internal settings. Specifically, the device sends an API request to the server and receives a response.

[0987] Input: Configuration data stored on the server

[0988] Output: Updated internal settings

[0989] Step 4:

[0990] When a user (child) speaks to the device, the device captures the voice using a built-in microphone and converts it into text data using speech recognition technology (such as the Google Speech-to-Text API). Specifically, utterances such as "I'll bring the block!" are converted into text.

[0991] Input: Voice input from child

[0992] Output: Text data

[0993] Step 5:

[0994] The device sends the converted text data to the server, which uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the text data. The generated response is then sent back to the device. Specifically, the server generates the response using an NLP engine.

[0995] Input: Text data sent from the terminal

[0996] Output: The generated response text

[0997] Step 6:

[0998] The device converts the response text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and responds to the child through the speaker. Specifically, the device generates synthesized speech and plays it back on the speaker.

[0999] Input: Response text sent by the server

[1000] Output: Synthesized speech

[1001] Step 7:

[1002] The device uses an emotion engine (such as Affectiva's SDK) to analyze the child's emotions by analyzing voice tones and facial expressions in real time. The analysis results are sent to the server as emotion data. For example, if a child is crying, emotion data is generated.

[1003] Input: Child's vocal tone and facial expression data

[1004] Output: Generated emotion data

[1005] Step 8:

[1006] The server generates an appropriate response based on the emotion data. For example, if the child is sad, it generates an encouraging response. The generated response text is then sent back to the device.

[1007] Input: Emotion data sent from the device

[1008] Output: Sentiment-based response text

[1009] Step 9:

[1010] The device uses speech synthesis technology to convert the generated response text and respond to the child, playing the synthesized voice from the speaker.

[1011] Input: Response text sent by the server

[1012] Output: Synthesized speech

[1013] Step 10:

[1014] The device monitors the child's movements in real time using an accelerometer and camera. If it detects any abnormal movements, it sends the data to a server. For example, if a child falls, that data is generated.

[1015] Input: Data on children's movements

[1016] Output: Abnormal behavior data

[1017] Step 11:

[1018] The server receives the abnormal behavior data and notifies the parent device in real time via push notifications, emails, etc.

[1019] Input: Abnormal behavior data sent from the device

[1020] Output: Notification sent to the parent device

[1021] Step 12:

[1022] The user (parent) checks the child's status through the application and sends appropriate instructions as necessary. The instructions are transmitted to the device via the server. Specifically, the parent enters instructions on the application screen and presses the send button.

[1023] Input: Parental instruction data

[1024] Output: Instruction data sent to the terminal

[1025] Step 13:

[1026] The device receives instruction data sent by the parent and uses voice synthesis technology to convey the message to the child, for example, by playing the message "Are you OK?" over the speaker.

[1027] Input: Instruction data sent from the parent

[1028] Output: Synthesized speech

[1029] (Application example 2)

[1030] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1031] There is a need for a system that monitors children's activities at home and allows parents to respond appropriately based on their children's emotional state. Current systems lack the ability to recognize children's emotions, making it difficult for parents to respond based on their children's emotions. There is also a need for a system that allows parents to understand their children's situation in real time and send appropriate instructions, even when they are at work or out.

[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to the parent's terminal, a dialogue means for conducting dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, an emotion recognition means for recognizing the user's emotion, and a response generation means for generating an appropriate response based on the recognized emotion data. This enables detailed feedback according to the child's emotional state, supports educational dialogue and appropriate play, and allows parents to watch over their children with peace of mind even when they are not at home.

[1033] "Monitoring means for monitoring children's activities within the home" refers to devices and systems such as sensors and cameras that monitor children's movements and behavior within the home in real time and obtain the necessary data.

[1034] The "transmission means for transmitting data acquired by the monitoring means to the parent's terminal" refers to a communication device and its system for transmitting data acquired by the monitoring means to a smartphone or computer used by the parent.

[1035] "An interaction means for interacting and playing based on setting data received from the parent's terminal" refers to a device and system for interacting with and playing with a child in accordance with setting data sent from the parent's terminal.

[1036] An "instruction means for sending instructions to a child based on an instruction received from a parent's terminal" is a device and system for receiving an instruction sent from a parent's terminal and encouraging a child to take a specific action based on that instruction.

[1037] The "emotion recognition means for recognizing the user's emotions" refers to a device and system for recognizing the user's emotional state in real time by analyzing the user's tone of voice and facial expressions.

[1038] The "response generating means for generating an appropriate response based on the recognized emotion data" refers to a device and system for generating an appropriate response to a user based on the recognized emotion data.

[1039] A system for implementing the present invention comprises the following major components:

[1040] 1. A stuffed toy robot equipped with a camera and sensors that can be used as a device (monitoring means) to monitor children's activities at home.

[1041] 2. Smartphones and computers used by parents (parental devices).

[1042] 3. Server.

[1043] 4. Emotion recognition means to recognize the user's emotions.

[1044] 5. A response generation means for generating an appropriate response based on the recognized emotion data.

[1045] Monitoring and Transmission Methods

[1046] The monitoring means uses cameras and sensors to monitor the child's activities in real time, making it possible to detect what the child is currently doing. The acquired data is then sent to the parent's device by the transmission means. Communication technologies such as Wi-Fi and Bluetooth can be used for transmission.

[1047] Interaction and instruction methods

[1048] The dialogue means operates based on the setting data received from the parent's device. For example, if the parent sets up a game such as "building a castle" through the application, the dialogue means will make suggestions to the child based on that information. The dialogue means also uses voice recognition technology to converse with the child. This voice data is converted into text data and sent to the server.

[1049] The instruction means prompts the child to perform a specific action based on a command received from the parent's device. For example, if the parent commands the child to "bring a block," the instruction means prompts the child to perform the action in accordance with the command.

[1050] Emotion recognition and response generation

[1051] The emotion recognition means analyzes the user's (child's) emotions in real time using voice tone and facial expression analysis technology. The analysis results are sent to the server, and the response generation means generates an appropriate response. For example, if a child has a sad expression, the response generation means generates encouraging words such as "What's wrong? It's okay."

[1052] Feedback and Notifications

[1053] Various sensors monitor the child's activity and any abnormalities, and send the data to a server. Feedback is used to notify the parent's device in real time. The parent can check the child's condition through the application and send instructions as necessary. For example, if a child falls, data stating "a fall has been detected" is sent to the server, and a notification is also sent to the parent's device.

[1054] Specific examples

[1055] 1. Example of a set interaction: A parent sets up a game called "Build a castle" in the application. The device suggests, "Let's build a castle today!", and the child responds, "I'll bring the blocks!". This interaction is coordinated via the server.

[1056] 2. Example of emotion recognition: If a child is crying, the emotion recognition means will analyze this and generate an encouraging message such as "What's wrong? It's okay" and return it to the child.

[1057] 3. Example of abnormal behavior detection: An accelerometer or camera detects abnormal behavior and sends a real-time notification to the parent's device saying, "A fall has been detected." The parent can then check the notification and send appropriate instructions.

[1058] Example prompts to input to the generative AI model:

[1059] "Design a system that recognizes the emotions of workers from video data and generates messages to help them work safely and efficiently. Please include specific example instructions."

[1060] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1061] Step 1:

[1062] The device uses cameras and sensors to monitor the child's activities in real time. The input is camera footage and sensor data, and the device acquires the child's movements and emotions based on this. The output is the acquired activity data.

[1063] Step 2:

[1064] The device sends the data acquired by the monitoring means to the server. The transmission method is Wi-Fi or Bluetooth, and the input is activity data, and the output is data sent to the server. The server receives this data and analyzes it.

[1065] Step 3:

[1066] Based on the configuration data received from the parent's device, the server generates a dialogue or play scenario and sends it to the device. The input is the parent's configuration data, and the output is the dialogue or play scenario. This allows the device to begin a dialogue with the child according to the settings.

[1067] Step 4:

[1068] As the device engages in a dialogue with the child, it uses voice recognition technology to convert the child's voice into text data. The input is the child's voice data, and the output is text data. The content of the dialogue is sent to the server.

[1069] Step 5:

[1070] The server analyzes the received text data and generates an appropriate response. The input is the child's text data, and the output is the response text. The response text is sent to the terminal and returned to the child via voice synthesis.

[1071] Step 6:

[1072] The device's emotion recognition means analyzes the user's (child's) voice tone and facial expression in real time and generates emotion data. The input is voice tone and facial expression data, and the output is emotion data. This data is sent to the server.

[1073] Step 7:

[1074] The server receives the emotion data and generates an appropriate response based on the emotion. The input is emotion data, and the output is a response text based on the emotion. For example, if the analysis finds that a child is sad, an encouraging message is generated.

[1075] Step 8:

[1076] The server sends the generated response text to the device, which then uses speech synthesis technology to communicate it to the child. The input is the response text, and the output is a voice message to the child.

[1077] Step 9:

[1078] The device uses various sensors to monitor children's abnormal behavior and notifies the server if an abnormality is detected. The input is sensor data, and the output is notification data of the abnormality detection.

[1079] Step 10:

[1080] The server sends a notification of anomaly detection to the parent's device in real time. The input is the anomaly detection data, and the output is a notification to the parent's device. When the parent receives the notification, they can check the child's status through the application and send necessary instructions to the device.

[1081] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1082] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1083] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1084] [Third embodiment]

[1085] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1086] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1087] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1088] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1089] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1090] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1091] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1092] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1093] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1094] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1095] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1096] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1097] The present invention provides a system that monitors the activities of children at home, transmits necessary information to a parent's terminal, and provides dialogue and play based on settings data from the parent. Specific embodiments for carrying out the present invention will be described below.

[1098] overview

[1099] This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The terminal monitors the child's activities and sends this information to the parent's terminal via the server. It also provides dialogue and play based on the setting data received from the parent, and conveys instructions from the parent to the child. Furthermore, a feedback function reports the child's activities to the parent in real time.

[1100] Entering User Data

[1101] The user (parent) launches the application on their smartphone or computer and inputs their household's specific rules, preferences, schedules, etc. This setting data is sent to the server and synchronized with the device.

[1102] Saving and Syncing Settings

[1103] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[1104] Real-time dialogue generation

[1105] When the user (child) speaks to the device, the device uses speech recognition technology to convert the speech into text and sends the data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[1106] Send Feedback

[1107] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[1108] Specific examples

[1109] Example 1: A child plays with a "cat sitter"

[1110] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1111] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1112] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[1113] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1114] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1115] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1116] Example 2: Detecting abnormal behavior in children

[1117] 1. The device monitors the child's movements using an accelerometer and camera.

[1118] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1119] 3. The server notifies the parent device of this information.

[1120] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[1121] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1122] The above is a specific embodiment for carrying out the present invention. This system ensures the safety of children even when their parents are away, and can provide educational conversations and play.

[1123] The processing flow will be explained below.

[1124] Entering User Data

[1125] Step 1:

[1126] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[1127] Step 2:

[1128] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[1129] Step 3:

[1130] The server receives the authentication information and checks it against information in a database.

[1131] Step 4:

[1132] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[1133] Step 5:

[1134] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[1135] Step 6:

[1136] The terminal transmits the input setting data to the server.

[1137] Saving and Syncing Settings

[1138] Step 1:

[1139] The server receives the setting data sent from the terminal and temporarily stores it.

[1140] Step 2:

[1141] The server checks the integrity of the received data and stores it in the database once the check is complete.

[1142] Step 3:

[1143] The server updates the version of the new setting data and notifies the terminal of the information.

[1144] Step 4:

[1145] The device receives the notification from the server and downloads the new configuration data.

[1146] Step 5:

[1147] The device analyzes the received data and updates its internal settings.

[1148] Real-time dialogue generation

[1149] Step 1:

[1150] The user (child) talks to "Nyanko Sitter."

[1151] Step 2:

[1152] The device converts voice input into text data using voice recognition technology.

[1153] Step 3:

[1154] The terminal transmits the converted text data to the server.

[1155] Step 4:

[1156] The server receives the text data and uses a generation AI to generate an appropriate response.

[1157] Step 5:

[1158] The server transmits the generated response text data to the terminal.

[1159] Step 6:

[1160] The terminal converts the received response text data into voice using voice synthesis technology.

[1161] Step 7:

[1162] The device plays the converted audio to the child.

[1163] Send Feedback

[1164] Step 1:

[1165] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[1166] Step 2:

[1167] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[1168] Step 3:

[1169] The server receives the abnormality detection information and notifies the parent terminal.

[1170] Step 4:

[1171] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[1172] Step 5:

[1173] The parent terminal transmits the input instruction data to the server.

[1174] Step 6:

[1175] The server receives the instruction data and transmits the data to the terminal.

[1176] Step 7:

[1177] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[1178] Example 1

[1179] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1180] There is a need to provide a means to effectively monitor children's activities at home and to provide a safe environment for parents to watch over them even when they are not around. It is also important to provide appropriate dialogue and play to facilitate two-way communication between children and parents. Without such a system, it is difficult to ensure children's safety and promote educational dialogue and play.

[1181] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1182] In this invention, the server includes: a user input means for a parent to input family rules, preferences, schedules, etc. on a smartphone or computer; a storage / synchronization means for saving the setting data entered by the user input means in a database and synchronizing it with the device; a voice recognition means for converting the child's speech into text using voice recognition technology and sending the data to the server; a response generation means for the server to generate an appropriate response using a generative AI model and send it to the device; a voice synthesis means for the device to convert the response text into speech using voice synthesis technology and return it to the child; a monitoring means for the device to monitor the child's activity status and abnormalities using various sensors and send the data to the server; a notification means for the server to notify the parent's device of the monitored data in real time; an instruction sending means for the parent to check the child's status through an application and send necessary instructions; and an instruction execution means for the device to execute the parent's instructions using voice synthesis technology and convey them to the child. This enables monitoring of the child's activity status and smooth communication with the parent.

[1183] "User input means" refers to the means by which parents input their household's specific rules, preferences, schedules, etc. on their smartphone or computer.

[1184] The "storage and synchronization means" is a means for storing the setting data input by the user input means in a database and synchronizing it with the terminal.

[1185] The "voice recognition means" is a means of converting the voice spoken by the child into text using voice recognition technology and sending that data to a server.

[1186] The "response generation means" is a means by which the server uses a generative AI model to generate an appropriate response and transmits it to the terminal.

[1187] "Speech synthesis means" refers to the means by which the terminal uses speech synthesis technology to convert the response text into speech and return it to the child.

[1188] The "monitoring means" is a means by which the device uses various sensors to monitor the child's activity status and any abnormalities, and sends that data to the server.

[1189] The "notification means" is a means by which the server notifies the parent terminal of the monitoring data in real time.

[1190] The "means for sending instructions" is a means by which parents can check their child's status through the application and send necessary instructions.

[1191] The "means for executing instructions" refers to the means by which the device executes instructions from the parent using voice synthesis technology and conveys them to the child.

[1192] The present invention provides a system that monitors the activities of children at home, allowing parents to watch over them with peace of mind. This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The specific hardware and software configurations and their operation are described below.

[1193] Entering User Data

[1194] The user (parent) launches a dedicated application on a smartphone or computer and inputs setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack at 3 pm every day" or "Play with building blocks." This setting data is then sent to a server via the Internet.

[1195] Saving and Syncing Settings

[1196] The server receives the configuration data sent from the parent and stores it in a database. This data is managed in real time. The server synchronizes the saved data with the device. For example, if a network error occurs, the server will repeatedly retry, and if synchronization is successful, the device will receive the new configuration data.

[1197] Real-time dialogue generation

[1198] When a user (child) speaks to the device, the device converts the speech into text using voice recognition technology (e.g., Google Cloud Speech-to-Text), and this text data is sent to the server.

[1199] The server uses a generative AI model (e.g., OpenAI GPT-3.5) to generate an appropriate response. For example, in response to the question "What do you want to do?", it generates a response such as "Let's play with building blocks!". The generated text data is sent to the device.

[1200] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the text into speech and then returns that speech to the child.

[1201] Send Feedback

[1202] The device uses an accelerometer and camera to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to a server.

[1203] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send instructions via the application if necessary.

[1204] Carrying out parental instructions

[1205] The user (parent) can check the child's status through the application and input and send necessary instructions, such as "Are you OK?"

[1206] The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and convey it to the child, for example, "Mommy is asking if you're OK."

[1207] Specific examples

[1208] Example 1: A child plays with a "cat sitter"

[1209] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1210] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1211] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[1212] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1213] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1214] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1215] Example 2: Detecting abnormal behavior in children

[1216] 1. The device monitors the child's movements using an accelerometer and camera.

[1217] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1218] 3. The server notifies the parent device of this information.

[1219] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[1220] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1221] Examples of prompt statements

[1222] Enter the following prompts into the generative AI model:

[1223] "Situation: Child has fallen. Parents are concerned. Generate an appropriate response message for the child."

[1224] Example of the resulting result:

[1225] "Are you okay? Can you get up now? Is anything hurting? Be careful."

[1226] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1227] Step 1: Enter and submit user data

[1228] The user (parent) launches a dedicated application on their smartphone or computer. They input setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack every day at 3 p.m." or "Play with building blocks." They confirm the input setting data and press the "Submit" button.

[1229] Input: Configuration data such as household-specific rules, preferences, and schedules

[1230] Output: The configuration data sent

[1231] Specific operation: The user operates the application, inputs and submits setting data.

[1232] Step 2: Save and sync configuration data

[1233] The server receives the setting data sent by the parent. It stores the received data in a database and manages it in real time. The stored data is time-stamped. The server then starts the process of synchronizing the stored data with the device. For example, if the setting data is "Eat a snack at 3 pm," it is synchronized with the device.

[1234] Input: Submitted configuration data

[1235] Output: Settings data synced to the device

[1236] Specific behavior: The server stores the received data and starts the synchronization process.

[1237] Step 3: Handling real-time interactions with your child

[1238] The user (child) speaks to the device, asking questions such as "What are you doing today?", and the device uses voice recognition technology (e.g., Google Cloud Speech-to-Text) to convert the speech into text. This text data is then sent to the server.

[1239] Input: Child's voice data

[1240] Output: Audio data converted to text

[1241] Specific operation: The device uses voice recognition technology to convert the child's voice into text and send it to the server.

[1242] Step 4: Server Response Generation

[1243] The server generates an appropriate response using a generative AI model (e.g., OpenAI GPT-3.5) and sends the generated response text to the device. For example, in response to the question "What do you want to do?", the server generates the response "Let's play with building blocks!"

[1244] Input: Text-converted audio data

[1245] Output: The generated response text

[1246] Specific operation: The server uses the generative AI model to generate an appropriate response and sends it to the device.

[1247] Step 5: Reply with a text-to-speech response

[1248] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the response text received from the server into speech, which is then returned to the child. For example, the device might respond with a voice saying, "Let's play with building blocks!"

[1249] Input: Generated response text

[1250] Output: Audio data

[1251] Specific operation: The device uses speech synthesis technology to convert the response text into speech and return it to the child.

[1252] Step 6: Monitor your child's activity

[1253] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to the server.

[1254] Input: Child's activity data

[1255] Output: Anomaly detection data

[1256] Specific operation: The device uses various sensors to monitor the child's activities, and if it detects any abnormalities, it sends the information to the server.

[1257] Step 7: Server notification

[1258] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send any necessary instructions via the application.

[1259] Input: Anomaly detection data

[1260] Output: Notification data to parent device

[1261] Specific operation: The server notifies the parent device of abnormal data in real time.

[1262] Step 8: Send and Execute Parental Instructions

[1263] The user (parent) checks the child's status through the application and inputs and sends the necessary instructions. For example, sending the instruction "Are you OK?" The server sends the instruction to the device, which uses speech synthesis technology to convert the instruction into voice and conveys it to the child. For example, the device might tell the child, "Mommy is asking if you're OK."

[1264] Input: Parent instruction data

[1265] Output: Voice instructions to the child

[1266] Specific operation: The parent inputs instructions into the application, the server sends the instructions to the device, and the device uses speech synthesis technology to convert the instructions into voice and convey them to the child.

[1267] (Application example 1)

[1268] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1269] In modern families, it is extremely important for parents to monitor their children's activities in real time and respond quickly when necessary. However, conventional technologies have struggled to ensure children's safety while parents are away and to provide educational conversations and play. Furthermore, there have been insufficient means for immediately notifying parents if their children exhibit abnormal behavior. The present invention addresses these issues.

[1270] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1271] In this invention, the server includes a monitoring means for monitoring the child's activities within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's terminal, a dialogue means for dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, and a notification means for detecting abnormal behavior and notifying the parent's terminal of that information in real time. This ensures the safety of the child even when the parent is away, allows the child to receive necessary feedback immediately, and takes appropriate action.

[1272] "Within the home" refers to the environment that forms the basis of life, such as a home or residence.

[1273] "Children's activity status" refers to the series of actions and behaviors that children perform at home, including play, learning, and daily activities.

[1274] "Monitoring means" refers to devices and systems such as sensors and cameras that are used to observe and record children's activities in real time.

[1275] "Data" refers to information or records obtained by monitoring means, including operating conditions, location information, audio data, etc.

[1276] "Parent's device" refers to an information processing device such as a smartphone, tablet, or computer used by the parent.

[1277] "Transmission means" refers to a communication device or communication protocol for transmitting the acquired data to the parent terminal.

[1278] "Settings data" refers to data such as instructions, wishes, rules, etc. regarding a child's activities that a parent sends from their device.

[1279] "Interaction means" refers to a device or system that provides the function of interacting with or playing with a child based on the received setting data.

[1280] "Instruction means" refers to a device or function that transmits commands received from the parent's device to the child.

[1281] "Abnormal behavior" refers to behavior that deviates from a child's normal activity pattern or is risky.

[1282] "Notification means" refers to a communication device or communication protocol that immediately notifies the parent device of abnormal behavior when it is detected.

[1283] "Customization means" refers to devices or systems that provide the ability to tailor the content of the interaction means based on household-specific rules and preferences.

[1284] "Feedback means" refers to a device or system for transmitting feedback data about a child's activities to a parent's device in real time.

[1285] "Voice recognition technology" refers to the technology that converts voice into text data.

[1286] A "generative AI model" refers to an artificial intelligence function that generates appropriate responses or instructions based on input data.

[1287] This system monitors children's activities at home, transmits necessary information to a parent's device, and provides dialogue and play based on the parent's settings. This system consists of a stuffed toy robot (hereinafter referred to as the "device"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's device"), and a server.

[1288] System configuration

[1289] monitoring means

[1290] The device monitors the child's activities in real time using monitoring means such as cameras and acceleration sensors. For example, it can detect whether the child is playing normally or making abnormal movements.

[1291] Transmission method

[1292] The data collected by the device is sent to a server via Wi-Fi or Bluetooth, and the necessary information is then sent to the parent device. This process uses a communication protocol to provide stable data communication.

[1293] Interaction methods

[1294] The device provides dialogue and play based on the configuration data received from the parent's device. For example, if a child says, "Let's build a castle today!", the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The server then uses a generative AI model to generate an appropriate response and sends it back to the device.

[1295] means of instruction

[1296] The device sends instructions to the child based on the commands received from the parent's device. For example, if the parent sends the command "Clean up," the device will use voice synthesis technology to convey the command to the child aloud.

[1297] Notification means

[1298] If abnormal behavior is detected, the device sends the information to the server, which then notifies the parent in real time, allowing the parent to immediately understand the child's situation and take appropriate action.

[1299] Hardware and software used

[1300] Hardware:

[1301] Stuffed toy robot (with camera and accelerometer)

[1302] Smartphone

[1303] software:

[1304] Mobile application (iOS / Android)

[1305] Server (Cloud server: AWS, GCP)

[1306] Speech recognition / speech synthesis engine (Google Cloud Speech-to-Text, Text-to-Speech API, etc.)

[1307] Generative AI models (e.g., OpenAI's GPT-3)

[1308] Data processing and calculation

[1309] Monitoring Data Collection

[1310] The server collects real-time data from the robot's camera and accelerometer to monitor the child's activity.

[1311] Data transmission and processing

[1312] The collected data is sent to a server using a communication protocol such as Google Cloud Pub / Sub, and the data is processed using AWS Lambda, etc. If abnormal behavior is detected, the information is immediately sent as a push notification to the parent's smartphone.

[1313] Dialogue Generation

[1314] Voice input from the child is converted into text data using a speech recognition engine, and an appropriate response is generated by a generative AI model (e.g., GPT-3). The generated text is sent from the server to the device, where it is converted into speech by a speech synthesis engine and responded to the child.

[1315] Specific examples

[1316] Example 1: When a child says to the robot, "I'm looking for my ball."

[1317] The robot uses a voice recognition engine to convert "I'm looking for my ball" into text.

[1318] This text is sent to a server, which uses a generative AI model to generate a response.

[1319] The server sends a response to the terminal saying, "Which color ball do you like? Shall we look for it together?"

[1320] The device uses a speech synthesis engine to convert the response into voice and convey it to the child.

[1321] Prompt Sentence Examples

[1322] User: I'm looking for a ball

[1323] AI: What color ball do you like? Shall we find it together?

[1324] The above is a specific embodiment of the present invention. This system ensures the safety of children even when parents are away, and provides appropriate feedback and dialogue.

[1325] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1326] Step 1:

[1327] The stuffed toy robot (terminal) uses a camera and an accelerometer to collect data to monitor children's activities at home. This data includes the child's movements, location, and speech. The input is raw data from the sensors, and the output is organized activity data.

[1328] Step 2:

[1329] The device transmits the acquired data to the server via Wi-Fi or Bluetooth for transmission to the parent device in real time. The input is the activity data and the output is the data sent to the server. The server receives this data using an appropriate connection protocol (e.g. HTTPS).

[1330] Step 3:

[1331] The server synchronizes the configuration data (e.g., child schedules and household rules) received from the parent's device with the monitoring data. The input is the configuration data from the parent, and the output is the configuration data updated by the server.

[1332] Step 4:

[1333] When a child speaks to the device, the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The input is voice data and the output is text data.

[1334] Step 5:

[1335] The server uses a generative AI model to generate an appropriate response based on the text data. For example, if a child says, "I'm looking for my ball," the generative AI model generates the response, "What color ball do you like? Want to look for it together?" The input is text data, and the output is a response text.

[1336] Step 6:

[1337] The server sends the generated response text to the device, which then converts the text into speech using speech synthesis technology. The input is the response text and the output is audio data. The device then transmits the audio to the child.

[1338] Step 7:

[1339] The device constantly monitors the child's activity and abnormal behavior, and if an abnormality is detected, it sends the information to the server in real time. The input is the monitoring data, and the output is the abnormality detection information.

[1340] Step 8:

[1341] The server immediately notifies the parent device of the anomaly detection information and provides a means for the parent to check the situation through the application.The input is the anomaly detection information, and the output is a notification to the parent device.

[1342] Step 9:

[1343] Parents can check their children's status through the application and send instructions (e.g., "Clean up") as needed. The input is instruction data from the parent, and the output is instructions to the device.

[1344] Step 10:

[1345] The device receives instructions from the parent and uses voice synthesis technology to communicate those instructions to the child. The input is instruction data from the parent, and the output is voice instructions to the child.

[1346] These are the program processing steps for the system that realizes this application example. This makes it possible to ensure the safety of children and provide appropriate feedback and dialogue even when parents are not present.

[1347] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1348] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[1349] overview

[1350] The system consists of the following main components:

[1351] 1. A device (stuffed toy robot) that monitors children's activities.

[1352] 2. Smartphones and computers used by parents (parental devices).

[1353] 3. A central management server.

[1354] 4. Emotion engine that recognizes user emotions.

[1355] Entering User Data

[1356] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[1357] Saving and Syncing Settings

[1358] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[1359] Real-time dialogue generation

[1360] When a user (child) speaks to the device, the device uses voice recognition technology to convert the speech into text data and sends that data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[1361] Recognizing emotions and regulating responses

[1362] The device is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words will be generated.

[1363] Send Feedback

[1364] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[1365] Specific examples

[1366] Example 1: A child plays with a "cat sitter"

[1367] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1368] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1369] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[1370] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1371] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1372] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1373] Example 2: When a child is sad

[1374] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[1375] 2. The device sends the emotion data to the server.

[1376] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[1377] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[1378] Example 3: Detecting abnormal behavior in children

[1379] 1. The device monitors the child's movements using an accelerometer and camera.

[1380] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1381] 3. The server notifies the parent device of this information.

[1382] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[1383] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1384] The above is a concrete example of how to implement the present invention. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[1385] The processing flow will be explained below.

[1386] Entering User Data

[1387] Step 1:

[1388] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[1389] Step 2:

[1390] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[1391] Step 3:

[1392] The server receives the authentication information and checks it against information in a database.

[1393] Step 4:

[1394] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[1395] Step 5:

[1396] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[1397] Step 6:

[1398] The terminal transmits the input setting data to the server.

[1399] Saving and Syncing Settings

[1400] Step 1:

[1401] The server receives the setting data sent from the terminal and temporarily stores it.

[1402] Step 2:

[1403] The server checks the integrity of the received data and stores it in a database after checking.

[1404] Step 3:

[1405] The server updates the version of the new setting data and notifies the terminal of the information.

[1406] Step 4:

[1407] The device receives the notification from the server and downloads the new configuration data.

[1408] Step 5:

[1409] The device analyzes the received data and updates its internal settings.

[1410] Real-time dialogue generation

[1411] Step 1:

[1412] The user (child) talks to the device.

[1413] Step 2:

[1414] The device converts voice input into text data using voice recognition technology.

[1415] Step 3:

[1416] The terminal transmits the converted text data to the server.

[1417] Step 4:

[1418] The server receives the text data and uses a generation AI to generate an appropriate response.

[1419] Step 5:

[1420] The server transmits the generated response text data to the terminal.

[1421] Step 6:

[1422] The terminal converts the received response text data into voice using voice synthesis technology.

[1423] Step 7:

[1424] The device plays the converted audio to the child.

[1425] Recognizing emotions and regulating responses

[1426] Step 1:

[1427] When a user (child) speaks to the device, the device's emotion engine analyzes the voice tone and facial expressions to recognize the emotion.

[1428] Step 2:

[1429] The device transmits the analyzed emotion data to the server.

[1430] Step 3:

[1431] The server uses a generative AI to generate an appropriate response based on the received emotional data.

[1432] Step 4:

[1433] The server transmits the generated response text data to the terminal.

[1434] Step 5:

[1435] The terminal converts the received response text data into voice using voice synthesis technology.

[1436] Step 6:

[1437] The device plays appropriate voice responses based on the child's emotions.

[1438] Send Feedback

[1439] Step 1:

[1440] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[1441] Step 2:

[1442] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[1443] Step 3:

[1444] The server receives the abnormality detection information and notifies the parent terminal.

[1445] Step 4:

[1446] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[1447] Step 5:

[1448] The parent terminal transmits the input instruction data to the server.

[1449] Step 6:

[1450] The server receives the instruction data and transmits the data to the terminal.

[1451] Step 7:

[1452] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[1453] Example 2

[1454] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1455] Conventional home child activity monitoring systems primarily detect children's movements and location information, but lack appropriate dialogue or feedback regarding children's emotional state or specific behavior. This can result in parents being unable to respond appropriately to their children's emotional state, potentially having a negative impact on their children's mental health and education. Furthermore, conventional systems have limited functionality for customizing dialogue content based on parental rules and preferences, making it difficult to fully meet the needs of individual families.

[1456] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1457] In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's device, a dialogue means for conducting dialogue and play based on setting data received from the parent's device, an instruction means for sending instructions to the child based on instructions received from the parent's device, an emotion recognition means for analyzing the user's emotions and acquiring emotion data, and an emotion transmission means for transmitting the acquired emotion data to the parent's device. This makes it possible to grasp not only the child's activity status but also their emotional state, enabling customizable dialogue and feedback based on individual family rules and preferences.

[1458] "Monitoring means" refers to a system that includes devices and sensors for monitoring children's activities within the home.

[1459] The "transmission means" is a communication function for transferring data acquired by the monitoring means to the parent terminal.

[1460] "Interaction means" refers to technology or devices for interacting with or playing with a child based on the setting data received from the parent's terminal.

[1461] The "instruction means" is a function for sending instructions to the child based on commands received from the parent's terminal.

[1462] "Emotion recognition means" refers to technology or devices for analyzing the user's tone of voice and facial expressions to obtain emotional data.

[1463] "Emotion transmission means" is a communication function for transmitting the acquired emotion data to the parent terminal.

[1464] "Customization tools" are features that allow users to individually adjust interaction tools based on their household's specific rules and preferences.

[1465] The "feedback means" is a function for sending data about a child's activity status to a parent's device in real time.

[1466] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[1467] The system consists of the following main components:

[1468] 1. A stuffed toy robot (device) that monitors children's activities

[1469] 2. Smartphones and computers used by parents (parental devices)

[1470] 3. Central management server

[1471] 4. Emotion engine that recognizes user emotions

[1472] Entering User Data

[1473] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[1474] Saving and Syncing Settings

[1475] The server stores the received setting data in a database and synchronizes that data with the device. Specifically, a database such as MySQL is used. The device analyzes the received setting data and updates its internal settings. For example, based on the play setting data, the server might suggest to the child, "Let's build a castle today!"

[1476] Real-time dialogue generation

[1477] When a user (child) speaks to the device, the device uses speech recognition technology (e.g., Google Speech-to-Text API) to convert the speech into text data and sends that data to the server. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response and sends the response text to the device. The device then uses speech synthesis technology (e.g., Amazon Polly) to convert the response text into speech and returns it to the child.

[1478] Recognizing emotions and regulating responses

[1479] The device is equipped with an emotion engine (e.g., Affectiva's SDK) that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words is generated.

[1480] Send Feedback

[1481] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity and any abnormalities, and sends the data to a server. The server then notifies the parent's device of this data in real time. The parent can then check the child's status through an application and send instructions as necessary.

[1482] Specific examples

[1483] Example 1: A child plays with a "cat sitter"

[1484] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1485] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1486] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[1487] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1488] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1489] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1490] Example 2: When a child is sad

[1491] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[1492] 2. The device sends the emotion data to the server.

[1493] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[1494] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[1495] Example 3: Detecting abnormal behavior in children

[1496] 1. The device monitors the child's movements using an accelerometer and camera.

[1497] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1498] 3. The server notifies the parent device of this information.

[1499] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[1500] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1501] Example prompts for generative AI models

[1502] "Please enter the setting data required to do ____"

[1503] "Generate a response if a child is crying"

[1504] The above is a concrete example of how the present invention can be implemented. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[1505] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1506] Step 1:

[1507] A user (parent) launches an application on a smartphone or computer and enters configuration data such as household rules, preferences, schedules, etc. This configuration data includes information related to children's activities within the home, and specific actions involve entering data into fields on the application's settings screen.

[1508] Input: Family-specific rules and schedule information

[1509] Output: Setting data

[1510] Step 2:

[1511] The server receives the setting data sent by the user, encrypts it, and stores it in a database, such as MySQL. The received data is sent via the HTTP protocol.

[1512] Input: Configuration data submitted by the user

[1513] Output: Encrypted configuration data stored in the database

[1514] Step 3:

[1515] The device periodically sends a synchronization request to the server to obtain the latest configuration data. At this time, the device analyzes the data received from the server and updates its internal settings. Specifically, the device sends an API request to the server and receives a response.

[1516] Input: Configuration data stored on the server

[1517] Output: Updated internal settings

[1518] Step 4:

[1519] When a user (child) speaks to the device, the device captures the voice using a built-in microphone and converts it into text data using speech recognition technology (such as the Google Speech-to-Text API). Specifically, utterances such as "I'll bring the block!" are converted into text.

[1520] Input: Voice input from children

[1521] Output: Text data

[1522] Step 5:

[1523] The device sends the converted text data to the server, which uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the text data. The generated response is then sent back to the device. Specifically, the server generates the response using an NLP engine.

[1524] Input: Text data sent from the terminal

[1525] Output: The generated response text

[1526] Step 6:

[1527] The device converts the response text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and responds to the child through the speaker. Specifically, the device generates synthesized speech and plays it back on the speaker.

[1528] Input: Response text sent by the server

[1529] Output: Synthesized speech

[1530] Step 7:

[1531] The device uses an emotion engine (such as Affectiva's SDK) to analyze the child's emotions by analyzing voice tones and facial expressions in real time. The analysis results are sent to the server as emotion data. For example, if a child is crying, emotion data is generated.

[1532] Input: Child's vocal tone and facial expression data

[1533] Output: Generated emotion data

[1534] Step 8:

[1535] The server generates an appropriate response based on the emotion data. For example, if the child is sad, it generates an encouraging response. The generated response text is then sent back to the device.

[1536] Input: Emotion data sent from the device

[1537] Output: Sentiment-based response text

[1538] Step 9:

[1539] The device uses speech synthesis technology to convert the generated response text and respond to the child, playing the synthesized voice from the speaker.

[1540] Input: Response text sent by the server

[1541] Output: Synthesized speech

[1542] Step 10:

[1543] The device monitors the child's movements in real time using an accelerometer and camera. If it detects any abnormal movements, it sends the data to a server. For example, if a child falls, that data is generated.

[1544] Input: Data on children's movements

[1545] Output: Abnormal behavior data

[1546] Step 11:

[1547] The server receives the abnormal behavior data and notifies the parent device in real time via push notifications, emails, etc.

[1548] Input: Abnormal behavior data sent from the device

[1549] Output: Notification sent to parent device

[1550] Step 12:

[1551] The user (parent) checks the child's status through the application and sends appropriate instructions as necessary. The instructions are transmitted to the device via the server. Specifically, the parent enters instructions on the application screen and presses the send button.

[1552] Input: Parental instruction data

[1553] Output: Instruction data sent to the terminal

[1554] Step 13:

[1555] The device receives instruction data sent by the parent and uses voice synthesis technology to convey it to the child, for example by playing the message "Are you OK?" over the speaker.

[1556] Input: Instruction data sent from the parent

[1557] Output: Synthesized speech

[1558] (Application example 2)

[1559] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1560] There is a need for a system that monitors children's activities at home and allows parents to respond appropriately based on their children's emotional state. Current systems lack the ability to recognize children's emotions, making it difficult for parents to respond based on their children's emotions. There is also a need for a system that allows parents to understand their children's situation in real time and send appropriate instructions, even when they are at work or out.

[1561] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to the parent's terminal, a dialogue means for conducting dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, an emotion recognition means for recognizing the user's emotion, and a response generation means for generating an appropriate response based on the recognized emotion data. This enables detailed feedback according to the child's emotional state, supports educational dialogue and appropriate play, and allows parents to watch over their children with peace of mind even when they are not at home.

[1562] "Monitoring means for monitoring children's activities within the home" refers to devices and systems such as sensors and cameras that monitor children's movements and behavior within the home in real time and obtain the necessary data.

[1563] The "transmission means for transmitting data acquired by the monitoring means to the parent's terminal" refers to a communication device and its system for transmitting data acquired by the monitoring means to a smartphone or computer used by the parent.

[1564] "An interaction means for interacting and playing based on setting data received from the parent's terminal" refers to a device and system for interacting with and playing with a child in accordance with setting data sent from the parent's terminal.

[1565] An "instruction means for sending instructions to a child based on an instruction received from a parent's terminal" is a device and system for receiving an instruction sent from a parent's terminal and encouraging a child to take a specific action based on that instruction.

[1566] The "emotion recognition means for recognizing the user's emotions" refers to a device and system for recognizing the user's emotional state in real time by analyzing the user's tone of voice and facial expressions.

[1567] The "response generating means for generating an appropriate response based on the recognized emotion data" refers to a device and system for generating an appropriate response to a user based on the recognized emotion data.

[1568] A system for implementing the present invention comprises the following major components:

[1569] 1. A stuffed toy robot equipped with a camera and sensors that can be used as a device (monitoring means) to monitor children's activities at home.

[1570] 2. Smartphones and computers used by parents (parental devices).

[1571] 3. Server.

[1572] 4. Emotion recognition means to recognize the user's emotions.

[1573] 5. A response generation means for generating an appropriate response based on the recognized emotion data.

[1574] Monitoring and Transmission Methods

[1575] The monitoring means uses cameras and sensors to monitor the child's activities in real time, making it possible to detect what the child is currently doing. The acquired data is then sent to the parent's device by the transmission means. Communication technologies such as Wi-Fi and Bluetooth can be used for transmission.

[1576] Interaction and instruction methods

[1577] The dialogue means operates based on the setting data received from the parent's device. For example, if the parent sets up a game such as "building a castle" through the application, the dialogue means will make suggestions to the child based on that information. The dialogue means also uses voice recognition technology to converse with the child. This voice data is converted into text data and sent to the server.

[1578] The instruction means prompts the child to perform a specific action based on a command received from the parent's device. For example, if the parent instructs the child to "bring a block," the instruction means prompts the child to perform the action in accordance with the command.

[1579] Emotion recognition and response generation

[1580] The emotion recognition means analyzes the user's (child's) emotions in real time using voice tone and facial expression analysis technology. The analysis results are sent to the server, and the response generation means generates an appropriate response. For example, if a child has a sad expression, the response generation means generates encouraging words such as "What's wrong? It's okay."

[1581] Feedback and Notifications

[1582] Various sensors monitor the child's activity and any abnormalities, and send the data to a server. Feedback is used to notify the parent's device in real time. The parent can check the child's condition through the application and send instructions as necessary. For example, if a child falls, data stating "a fall has been detected" is sent to the server, and a notification is also sent to the parent's device.

[1583] Specific examples

[1584] 1. Example of a set interaction: A parent sets up a game called "Build a castle" in the application. The device suggests, "Let's build a castle today!", and the child responds, "I'll bring the blocks!". This interaction is coordinated via the server.

[1585] 2. Example of emotion recognition: If a child is crying, the emotion recognition means will analyze this and generate an encouraging message such as "What's wrong? It's okay" and return it to the child.

[1586] 3. Example of abnormal behavior detection: An accelerometer or camera detects abnormal behavior and sends a real-time notification to the parent's device saying, "A fall has been detected." The parent can then check the notification and send appropriate instructions.

[1587] Example prompts to input to a generative AI model:

[1588] "Design a system that recognizes the emotions of workers from video data and generates messages to help them work safely and efficiently. Please include specific example instructions."

[1589] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1590] Step 1:

[1591] The device uses cameras and sensors to monitor the child's activities in real time. The input is camera footage and sensor data, and the device acquires the child's movements and emotions based on this. The output is the acquired activity data.

[1592] Step 2:

[1593] The device sends the data acquired by the monitoring means to the server. The transmission method is Wi-Fi or Bluetooth, and the input is activity data, and the output is data sent to the server. The server receives this data and analyzes it.

[1594] Step 3:

[1595] Based on the configuration data received from the parent's device, the server generates a dialogue or play scenario and sends it to the device. The input is the parent's configuration data, and the output is the dialogue or play scenario. This allows the device to begin a dialogue with the child according to the settings.

[1596] Step 4:

[1597] As the device engages in a dialogue with the child, it uses voice recognition technology to convert the child's voice into text data. The input is the child's voice data, and the output is text data. The content of the dialogue is sent to the server.

[1598] Step 5:

[1599] The server analyzes the received text data and generates an appropriate response. The input is the child's text data, and the output is the response text. The response text is sent to the terminal and returned to the child via voice synthesis.

[1600] Step 6:

[1601] The device's emotion recognition means analyzes the user's (child's) voice tone and facial expression in real time and generates emotion data. The input is voice tone and facial expression data, and the output is emotion data. This data is sent to the server.

[1602] Step 7:

[1603] The server receives the emotion data and generates an appropriate response based on the emotion. The input is emotion data, and the output is a response text based on the emotion. For example, if the analysis finds that a child is sad, an encouraging message is generated.

[1604] Step 8:

[1605] The server sends the generated response text to the device, which then uses speech synthesis technology to communicate it to the child. The input is the response text, and the output is a voice message to the child.

[1606] Step 9:

[1607] The device uses various sensors to monitor children's abnormal behavior and notifies the server if an abnormality is detected. The input is sensor data, and the output is notification data of the abnormality detection.

[1608] Step 10:

[1609] The server sends a notification of anomaly detection to the parent's device in real time. The input is the anomaly detection data, and the output is a notification to the parent's device. When the parent receives the notification, they can check the child's status through the application and send necessary instructions to the device.

[1610] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1611] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1612] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1613] [Fourth embodiment]

[1614] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1615] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1616] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1617] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1618] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1619] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1620] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1621] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1622] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1623] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1624] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1625] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1626] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1627] The present invention provides a system that monitors the activities of children at home, transmits necessary information to a parent's terminal, and provides dialogue and play based on settings data from the parent. Specific embodiments for carrying out the present invention will be described below.

[1628] overview

[1629] This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The terminal monitors the child's activities and sends this information to the parent's terminal via the server. It also provides dialogue and play based on the setting data received from the parent, and conveys instructions from the parent to the child. Furthermore, a feedback function reports the child's activities to the parent in real time.

[1630] Entering User Data

[1631] The user (parent) launches the application on their smartphone or computer and inputs their household's specific rules, preferences, schedules, etc. This setting data is sent to the server and synchronized with the device.

[1632] Saving and Syncing Settings

[1633] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[1634] Real-time dialogue generation

[1635] When the user (child) speaks to the device, the device uses speech recognition technology to convert the speech into text and sends the data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[1636] Send Feedback

[1637] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[1638] Specific examples

[1639] Example 1: A child plays with a "cat sitter"

[1640] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1641] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1642] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[1643] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1644] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1645] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1646] Example 2: Detecting abnormal behavior in children

[1647] 1. The device monitors the child's movements using an accelerometer and camera.

[1648] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1649] 3. The server notifies the parent device of this information.

[1650] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[1651] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1652] The above is a specific embodiment for carrying out the present invention. This system ensures the safety of children even when their parents are away, and can provide educational conversations and play.

[1653] The processing flow will be explained below.

[1654] Entering User Data

[1655] Step 1:

[1656] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[1657] Step 2:

[1658] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[1659] Step 3:

[1660] The server receives the authentication information and checks it against information in a database.

[1661] Step 4:

[1662] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[1663] Step 5:

[1664] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[1665] Step 6:

[1666] The terminal transmits the input setting data to the server.

[1667] Saving and Syncing Settings

[1668] Step 1:

[1669] The server receives the setting data sent from the terminal and temporarily stores it.

[1670] Step 2:

[1671] The server checks the integrity of the received data and stores it in the database once the check is complete.

[1672] Step 3:

[1673] The server updates the version of the new setting data and notifies the terminal of the information.

[1674] Step 4:

[1675] The device receives the notification from the server and downloads the new configuration data.

[1676] Step 5:

[1677] The device analyzes the received data and updates its internal settings.

[1678] Real-time dialogue generation

[1679] Step 1:

[1680] The user (child) talks to "Nyanko Sitter."

[1681] Step 2:

[1682] The device converts voice input into text data using voice recognition technology.

[1683] Step 3:

[1684] The terminal transmits the converted text data to the server.

[1685] Step 4:

[1686] The server receives the text data and uses a generation AI to generate an appropriate response.

[1687] Step 5:

[1688] The server transmits the generated response text data to the terminal.

[1689] Step 6:

[1690] The terminal converts the received response text data into voice using voice synthesis technology.

[1691] Step 7:

[1692] The device plays the converted audio to the child.

[1693] Send Feedback

[1694] Step 1:

[1695] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[1696] Step 2:

[1697] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[1698] Step 3:

[1699] The server receives the abnormality detection information and notifies the parent terminal.

[1700] Step 4:

[1701] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[1702] Step 5:

[1703] The parent terminal transmits the input instruction data to the server.

[1704] Step 6:

[1705] The server receives the instruction data and transmits the data to the terminal.

[1706] Step 7:

[1707] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[1708] Example 1

[1709] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1710] There is a need to provide a means to effectively monitor children's activities at home and to provide a safe environment for parents to watch over them even when they are not around. It is also important to provide appropriate dialogue and play to facilitate two-way communication between children and parents. Without such a system, it is difficult to ensure children's safety and promote educational dialogue and play.

[1711] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1712] In this invention, the server includes: a user input means for a parent to input family rules, preferences, schedules, etc. on a smartphone or computer; a storage / synchronization means for saving the setting data entered by the user input means in a database and synchronizing it with the device; a voice recognition means for converting the child's speech into text using voice recognition technology and sending the data to the server; a response generation means for the server to generate an appropriate response using a generative AI model and send it to the device; a voice synthesis means for the device to convert the response text into speech using voice synthesis technology and return it to the child; a monitoring means for the device to monitor the child's activity status and abnormalities using various sensors and send the data to the server; a notification means for the server to notify the parent's device of the monitored data in real time; an instruction sending means for the parent to check the child's status through an application and send necessary instructions; and an instruction execution means for the device to execute the parent's instructions using voice synthesis technology and convey them to the child. This enables monitoring of the child's activity status and smooth communication with the parent.

[1713] "User input means" refers to the means by which parents input their household's specific rules, preferences, schedules, etc. on their smartphone or computer.

[1714] The "storage and synchronization means" is a means for storing the setting data input by the user input means in a database and synchronizing it with the terminal.

[1715] The "voice recognition means" is a means of converting the voice spoken by the child into text using voice recognition technology and sending that data to a server.

[1716] The "response generation means" is a means by which the server uses a generative AI model to generate an appropriate response and transmits it to the terminal.

[1717] "Speech synthesis means" refers to the means by which the terminal uses speech synthesis technology to convert the response text into speech and return it to the child.

[1718] The "monitoring means" is a means by which the device uses various sensors to monitor the child's activity status and any abnormalities, and sends that data to the server.

[1719] The "notification means" is a means by which the server notifies the parent terminal of the monitoring data in real time.

[1720] The "means for sending instructions" is a means by which parents can check their child's status through the application and send necessary instructions.

[1721] The "means for executing instructions" refers to the means by which the device executes instructions from the parent using voice synthesis technology and conveys them to the child.

[1722] The present invention provides a system that monitors the activities of children at home, allowing parents to watch over them with peace of mind. This system consists of a stuffed toy robot (hereinafter referred to as the "terminal"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's terminal"), and a server. The specific hardware and software configurations and their operation are described below.

[1723] Entering User Data

[1724] The user (parent) launches a dedicated application on a smartphone or computer and inputs setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack at 3 pm every day" or "Play with building blocks." This setting data is then sent to a server via the Internet.

[1725] Saving and Syncing Settings

[1726] The server receives the configuration data sent from the parent and stores it in a database. This data is managed in real time. The server synchronizes the saved data with the device. For example, if a network error occurs, the server will repeatedly retry, and if synchronization is successful, the device will receive the new configuration data.

[1727] Real-time dialogue generation

[1728] When a user (child) speaks to the device, the device converts the speech into text using voice recognition technology (e.g., Google Cloud Speech-to-Text), and this text data is sent to the server.

[1729] The server uses a generative AI model (e.g., OpenAI GPT-3.5) to generate an appropriate response. For example, in response to the question "What do you want to do?", it generates a response such as "Let's play with building blocks!". The generated text data is sent to the device.

[1730] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the text into speech and then returns that speech to the child.

[1731] Send Feedback

[1732] The device uses an accelerometer and camera to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to a server.

[1733] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send instructions via the application if necessary.

[1734] Carrying out parental instructions

[1735] The user (parent) can check the child's status through the application and input and send necessary instructions, such as "Are you OK?"

[1736] The server sends the instruction to the device, which then uses speech synthesis technology to convert the instruction into voice and convey it to the child, for example, "Mommy is asking if you're OK."

[1737] Specific examples

[1738] Example 1: A child plays with a "cat sitter"

[1739] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1740] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1741] 3. The child speaks to the device, saying, "I'll bring you the blocks!"

[1742] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1743] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1744] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1745] Example 2: Detecting abnormal behavior in children

[1746] 1. The device monitors the child's movements using an accelerometer and camera.

[1747] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1748] 3. The server notifies the parent device of this information.

[1749] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[1750] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1751] Examples of prompt statements

[1752] Enter the following prompts into the generative AI model:

[1753] "Situation: Child has fallen. Parents are concerned. Generate an appropriate response message for the child."

[1754] Example of the resulting result:

[1755] "Are you okay? Can you get up now? Is anything hurting? Be careful."

[1756] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1757] Step 1: Enter and submit user data

[1758] The user (parent) launches a dedicated application on their smartphone or computer. They input setting data such as household rules, preferences, and schedules. For example, they input information such as "Eat a snack every day at 3 p.m." or "Play with building blocks." They confirm the input setting data and press the "Submit" button.

[1759] Input: Configuration data such as household-specific rules, preferences, and schedules

[1760] Output: The configuration data sent

[1761] Specific operation: The user operates the application, inputs and submits setting data.

[1762] Step 2: Save and sync configuration data

[1763] The server receives the setting data sent by the parent. It stores the received data in a database and manages it in real time. The stored data is time-stamped. The server then starts the process of synchronizing the stored data with the device. For example, if the setting data is "Eat a snack at 3 pm," it is synchronized with the device.

[1764] Input: Submitted configuration data

[1765] Output: Settings data synced to the device

[1766] Specific behavior: The server stores the received data and starts the synchronization process.

[1767] Step 3: Handling real-time interactions with your child

[1768] The user (child) speaks to the device, asking questions such as "What are you doing today?", and the device uses voice recognition technology (e.g., Google Cloud Speech-to-Text) to convert the speech into text. This text data is then sent to the server.

[1769] Input: Child's voice data

[1770] Output: Audio data converted to text

[1771] Specific operation: The device uses voice recognition technology to convert the child's voice into text and send it to the server.

[1772] Step 4: Server Response Generation

[1773] The server generates an appropriate response using a generative AI model (e.g., OpenAI GPT-3.5) and sends the generated response text to the device. For example, in response to the question "What do you want to do?", the server generates the response "Let's play with building blocks!"

[1774] Input: Text-converted audio data

[1775] Output: The generated response text

[1776] Specific operation: The server uses the generative AI model to generate an appropriate response and sends it to the device.

[1777] Step 5: Reply with a text-to-speech response

[1778] The device uses speech synthesis technology (e.g., Amazon Polly) to convert the response text received from the server into speech, which is then returned to the child. For example, the device might respond with a voice saying, "Let's play with building blocks!"

[1779] Input: Generated response text

[1780] Output: Audio data

[1781] Specific operation: The device uses speech synthesis technology to convert the response text into speech and return it to the child.

[1782] Step 6: Monitor your child's activity

[1783] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity in real time. For example, if it detects a fall or abnormal movement, it sends that information to the server.

[1784] Input: Child's activity data

[1785] Output: Anomaly detection data

[1786] Specific operation: The device uses various sensors to monitor the child's activities, and if it detects any abnormalities, it sends the information to the server.

[1787] Step 7: Server notification

[1788] The server then sends the received abnormal data to the parent's device in real time. For example, a notification saying "A fall has been detected" is displayed. The parent can check the situation and send any necessary instructions via the application.

[1789] Input: Anomaly detection data

[1790] Output: Notification data to parent device

[1791] Specific operation: The server notifies the parent device of abnormal data in real time.

[1792] Step 8: Send and Execute Parental Instructions

[1793] The user (parent) checks the child's status through the application and inputs and sends the necessary instructions. For example, sending the instruction "Are you OK?" The server sends the instruction to the device, which uses speech synthesis technology to convert the instruction into voice and conveys it to the child. For example, the device might tell the child, "Mommy is asking if you're OK."

[1794] Input: Parent instruction data

[1795] Output: Voice instructions to the child

[1796] Specific operation: The parent inputs instructions into the application, the server sends the instructions to the device, and the device uses speech synthesis technology to convert the instructions into voice and convey them to the child.

[1797] (Application example 1)

[1798] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1799] In modern families, it is extremely important for parents to monitor their children's activities in real time and respond quickly when necessary. However, conventional technologies have struggled to ensure children's safety while parents are away and to provide educational conversations and play. Furthermore, there have been insufficient means for immediately notifying parents if their children exhibit abnormal behavior. The present invention addresses these issues.

[1800] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1801] In this invention, the server includes a monitoring means for monitoring the child's activities within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's terminal, a dialogue means for dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, and a notification means for detecting abnormal behavior and notifying the parent's terminal of that information in real time. This ensures the safety of the child even when the parent is away, allows the child to receive necessary feedback immediately, and takes appropriate action.

[1802] "Within the home" refers to the environment that forms the basis of life, such as a home or residence.

[1803] "Children's activity status" refers to the series of actions and behaviors that children perform at home, including play, learning, and daily activities.

[1804] "Monitoring means" refers to devices and systems such as sensors and cameras that are used to observe and record children's activities in real time.

[1805] "Data" refers to information or records obtained by monitoring means, including operating conditions, location information, audio data, etc.

[1806] "Parent's device" refers to an information processing device such as a smartphone, tablet, or computer used by the parent.

[1807] "Transmission means" refers to a communication device or communication protocol for transmitting the acquired data to the parent terminal.

[1808] "Settings data" refers to data such as instructions, wishes, rules, etc. regarding a child's activities that a parent sends from their device.

[1809] "Interaction means" refers to a device or system that provides the function of interacting with or playing with a child based on the received setting data.

[1810] "Instruction means" refers to a device or function that transmits commands received from the parent's device to the child.

[1811] "Abnormal behavior" refers to behavior that deviates from a child's normal activity pattern or is risky.

[1812] "Notification means" refers to a communication device or communication protocol that immediately notifies the parent device of abnormal behavior when it is detected.

[1813] "Customization means" refers to devices or systems that provide the ability to tailor the content of the interaction means based on household-specific rules and preferences.

[1814] "Feedback means" refers to a device or system for transmitting feedback data about a child's activities to a parent's device in real time.

[1815] "Voice recognition technology" refers to the technology that converts voice into text data.

[1816] A "generative AI model" refers to an artificial intelligence function that generates appropriate responses or instructions based on input data.

[1817] This system monitors children's activities at home, transmits necessary information to a parent's device, and provides dialogue and play based on the parent's settings. This system consists of a stuffed toy robot (hereinafter referred to as the "device"), a smartphone or computer used by the parent (hereinafter referred to as the "parent's device"), and a server.

[1818] System configuration

[1819] monitoring means

[1820] The device monitors the child's activities in real time using monitoring means such as cameras and acceleration sensors. For example, it can detect whether the child is playing normally or making abnormal movements.

[1821] Transmission method

[1822] The data collected by the device is sent to a server via Wi-Fi or Bluetooth, and the necessary information is then sent to the parent device. This process uses a communication protocol to provide stable data communication.

[1823] Interaction methods

[1824] The device provides dialogue and play based on the configuration data received from the parent's device. For example, if a child says, "Let's build a castle today!", the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The server then uses a generative AI model to generate an appropriate response and sends it back to the device.

[1825] means of instruction

[1826] The device sends instructions to the child based on the commands received from the parent's device. For example, if the parent sends the command "Clean up," the device will use voice synthesis technology to convey the command to the child aloud.

[1827] Notification means

[1828] If abnormal behavior is detected, the device sends the information to the server, which then notifies the parent in real time, allowing the parent to immediately understand the child's situation and take appropriate action.

[1829] Hardware and software used

[1830] Hardware:

[1831] Stuffed toy robot (with camera and accelerometer)

[1832] Smartphone

[1833] software:

[1834] Mobile application (iOS / Android)

[1835] Server (Cloud server: AWS, GCP)

[1836] Speech recognition / speech synthesis engine (Google Cloud Speech-to-Text, Text-to-Speech API, etc.)

[1837] Generative AI models (e.g., OpenAI's GPT-3)

[1838] Data processing and calculation

[1839] Monitoring Data Collection

[1840] The server collects real-time data from the robot's camera and accelerometer to monitor the child's activity.

[1841] Data transmission and processing

[1842] The collected data is sent to a server using a communication protocol such as Google Cloud Pub / Sub, and the data is processed using AWS Lambda, etc. If abnormal behavior is detected, the information is immediately sent as a push notification to the parent's smartphone.

[1843] Dialogue Generation

[1844] Voice input from the child is converted into text data using a speech recognition engine, and an appropriate response is generated by a generative AI model (e.g., GPT-3). The generated text is sent from the server to the device, where it is converted into speech by a speech synthesis engine and responded to the child.

[1845] Specific examples

[1846] Example 1: When a child says to the robot, "I'm looking for my ball."

[1847] The robot uses a voice recognition engine to convert "I'm looking for my ball" into text.

[1848] This text is sent to a server, which uses a generative AI model to generate a response.

[1849] The server sends a response to the terminal saying, "Which color ball do you like? Shall we look for it together?"

[1850] The device uses a speech synthesis engine to convert the response into voice and convey it to the child.

[1851] Prompt Sentence Examples

[1852] User: I'm looking for a ball

[1853] AI: What color ball do you like? Shall we find it together?

[1854] The above is a specific embodiment of the present invention. This system ensures the safety of children even when parents are away, and provides appropriate feedback and dialogue.

[1855] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1856] Step 1:

[1857] The stuffed toy robot (terminal) uses a camera and an accelerometer to collect data to monitor children's activities at home. This data includes the child's movements, location, and speech. The input is raw data from the sensors, and the output is organized activity data.

[1858] Step 2:

[1859] The device transmits the acquired data to the server via Wi-Fi or Bluetooth for transmission to the parent device in real time. The input is the activity data and the output is the data sent to the server. The server receives this data using an appropriate connection protocol (e.g. HTTPS).

[1860] Step 3:

[1861] The server synchronizes the configuration data (e.g., child schedules and household rules) received from the parent's device with the monitoring data. The input is the configuration data from the parent, and the output is the configuration data updated by the server.

[1862] Step 4:

[1863] When a child speaks to the device, the device uses voice recognition technology to convert the speech into text and sends the text data to the server. The input is voice data and the output is text data.

[1864] Step 5:

[1865] The server uses a generative AI model to generate an appropriate response based on the text data. For example, if a child says, "I'm looking for my ball," the generative AI model generates the response, "What color ball do you like? Want to look for it together?" The input is text data, and the output is a response text.

[1866] Step 6:

[1867] The server sends the generated response text to the device, which then converts the text into speech using speech synthesis technology. The input is the response text and the output is audio data. The device then transmits the audio to the child.

[1868] Step 7:

[1869] The device constantly monitors the child's activity and abnormal behavior, and if an abnormality is detected, it sends the information to the server in real time. The input is the monitoring data, and the output is the abnormality detection information.

[1870] Step 8:

[1871] The server immediately notifies the parent device of the anomaly detection information and provides a means for the parent to check the situation through the application.The input is the anomaly detection information, and the output is a notification to the parent device.

[1872] Step 9:

[1873] Parents can check their children's status through the application and send instructions (e.g., "Clean up") as needed. The input is instruction data from the parent, and the output is instructions to the device.

[1874] Step 10:

[1875] The device receives instructions from the parent and uses voice synthesis technology to communicate those instructions to the child. The input is instruction data from the parent, and the output is voice instructions to the child.

[1876] These are the program processing steps for the system that realizes this application example. This makes it possible to ensure the safety of children and provide appropriate feedback and dialogue even when parents are not present.

[1877] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1878] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[1879] overview

[1880] The system consists of the following main components:

[1881] 1. A device (stuffed toy robot) that monitors children's activities.

[1882] 2. Smartphones and computers used by parents (parental devices).

[1883] 3. A central management server.

[1884] 4. Emotion engine that recognizes user emotions.

[1885] Entering User Data

[1886] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[1887] Saving and Syncing Settings

[1888] The server stores the received setting data in a database and synchronizes the data with the device, which then analyzes the received setting data and updates its internal settings.

[1889] Real-time dialogue generation

[1890] When a user (child) speaks to the device, the device uses voice recognition technology to convert the speech into text data and sends that data to the server. The server generates an appropriate response and sends the response text to the device. The device then uses speech synthesis technology to convert the response text into speech and returns it to the child.

[1891] Recognizing emotions and regulating responses

[1892] The device is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words will be generated.

[1893] Send Feedback

[1894] The device uses various sensors to monitor the child's activity and any abnormalities, and sends the data to a server. The server then sends this data to the parent's device in real time. The parent can then check the child's status through an application and send instructions as needed.

[1895] Specific examples

[1896] Example 1: A child plays with a "cat sitter"

[1897] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[1898] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[1899] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[1900] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[1901] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[1902] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[1903] Example 2: When a child is sad

[1904] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[1905] 2. The device sends the emotion data to the server.

[1906] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[1907] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[1908] Example 3: Detecting abnormal behavior in children

[1909] 1. The device monitors the child's movements using an accelerometer and camera.

[1910] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[1911] 3. The server notifies the parent device of this information.

[1912] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking the child, "Are you okay?").

[1913] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[1914] The above is a concrete example of how to implement the present invention. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[1915] The processing flow will be explained below.

[1916] Entering User Data

[1917] Step 1:

[1918] The user (parent) launches the application on their smartphone or computer and the login screen appears.

[1919] Step 2:

[1920] The user logs in by entering their authentication information (e.g., user ID and password), and then sends the authentication information to the server.

[1921] Step 3:

[1922] The server receives the authentication information and checks it against information in a database.

[1923] Step 4:

[1924] If the authentication is successful, the server sends a settings menu screen to the terminal through the application.

[1925] Step 5:

[1926] The user enters household rules and preferences on the settings menu screen and saves the settings data.

[1927] Step 6:

[1928] The terminal transmits the input setting data to the server.

[1929] Saving and Syncing Settings

[1930] Step 1:

[1931] The server receives the setting data sent from the terminal and temporarily stores it.

[1932] Step 2:

[1933] The server checks the integrity of the received data and stores it in a database after checking.

[1934] Step 3:

[1935] The server updates the version of the new setting data and notifies the terminal of the information.

[1936] Step 4:

[1937] The device receives the notification from the server and downloads the new configuration data.

[1938] Step 5:

[1939] The device analyzes the received data and updates its internal settings.

[1940] Real-time dialogue generation

[1941] Step 1:

[1942] The user (child) talks to the device.

[1943] Step 2:

[1944] The device converts voice input into text data using voice recognition technology.

[1945] Step 3:

[1946] The terminal transmits the converted text data to the server.

[1947] Step 4:

[1948] The server receives the text data and uses a generation AI to generate an appropriate response.

[1949] Step 5:

[1950] The server transmits the generated response text data to the terminal.

[1951] Step 6:

[1952] The terminal converts the received response text data into voice using voice synthesis technology.

[1953] Step 7:

[1954] The device plays the converted audio to the child.

[1955] Recognizing emotions and regulating responses

[1956] Step 1:

[1957] When a user (child) speaks to the device, the device's emotion engine analyzes the voice tone and facial expressions to recognize the emotion.

[1958] Step 2:

[1959] The device transmits the analyzed emotion data to the server.

[1960] Step 3:

[1961] The server uses a generative AI to generate an appropriate response based on the received emotional data.

[1962] Step 4:

[1963] The server transmits the generated response text data to the terminal.

[1964] Step 5:

[1965] The terminal converts the received response text data into voice using voice synthesis technology.

[1966] Step 6:

[1967] The device plays appropriate voice responses based on the child's emotions.

[1968] Send Feedback

[1969] Step 1:

[1970] The device monitors children's activities using a camera, microphone, accelerometer, etc.

[1971] Step 2:

[1972] The terminal analyzes the monitoring data and, if it detects an abnormal situation, sends the information to the server.

[1973] Step 3:

[1974] The server receives the abnormality detection information and notifies the parent terminal.

[1975] Step 4:

[1976] The user (parent) checks the child's status through the application and inputs instructions as necessary.

[1977] Step 5:

[1978] The parent terminal transmits the input instruction data to the server.

[1979] Step 6:

[1980] The server receives the instruction data and transmits the data to the terminal.

[1981] Step 7:

[1982] Based on the received instruction data, the device uses voice synthesis technology to communicate the instructions to the child.

[1983] Example 2

[1984] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1985] Conventional home child activity monitoring systems primarily detect children's movements and location information, but lack appropriate dialogue or feedback regarding children's emotional state or specific behavior. This can result in parents being unable to respond appropriately to their children's emotional state, potentially having a negative impact on their children's mental health and education. Furthermore, conventional systems have limited functionality for customizing dialogue content based on parental rules and preferences, making it difficult to fully meet the needs of individual families.

[1986] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1987] In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to a parent's device, a dialogue means for conducting dialogue and play based on setting data received from the parent's device, an instruction means for sending instructions to the child based on instructions received from the parent's device, an emotion recognition means for analyzing the user's emotions and acquiring emotion data, and an emotion transmission means for transmitting the acquired emotion data to the parent's device. This makes it possible to grasp not only the child's activity status but also their emotional state, enabling customizable dialogue and feedback based on individual family rules and preferences.

[1988] "Monitoring means" refers to a system that includes devices and sensors for monitoring children's activities within the home.

[1989] The "transmission means" is a communication function for transferring data acquired by the monitoring means to the parent terminal.

[1990] "Interaction means" refers to technology or devices for interacting with or playing with a child based on the setting data received from the parent's terminal.

[1991] The "instruction means" is a function for sending instructions to the child based on commands received from the parent's terminal.

[1992] "Emotion recognition means" refers to technology or devices for analyzing the user's tone of voice and facial expressions to obtain emotional data.

[1993] "Emotion transmission means" is a communication function for transmitting the acquired emotion data to the parent terminal.

[1994] "Customization tools" are features that allow users to individually adjust interaction tools based on their household's specific rules and preferences.

[1995] The "feedback means" is a function for sending data about a child's activity status to a parent's device in real time.

[1996] The present invention relates to a system that monitors children's activities at home and provides appropriate information to a parent's terminal, and in particular, is combined with an emotion engine that recognizes the user's emotions and adjusts responses.

[1997] The system consists of the following main components:

[1998] 1. A stuffed toy robot (device) that monitors children's activities

[1999] 2. Smartphones and computers used by parents (parental devices)

[2000] 3. Central management server

[2001] 4. Emotion engine that recognizes user emotions

[2002] Entering User Data

[2003] The user (parent) starts the application on their smartphone or computer, logs in, and then enters configuration data such as household rules, preferences, and schedules. This configuration data is sent to the server and synchronized with the device. Interactions and play are carried out based on the configured data.

[2004] Saving and Syncing Settings

[2005] The server stores the received setting data in a database and synchronizes that data with the device. Specifically, a database such as MySQL is used. The device analyzes the received setting data and updates its internal settings. For example, based on the play setting data, the server might suggest to the child, "Let's build a castle today!"

[2006] Real-time dialogue generation

[2007] When a user (child) speaks to the device, the device uses speech recognition technology (e.g., Google Speech-to-Text API) to convert the speech into text data and sends that data to the server. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate response and sends the response text to the device. The device then uses speech synthesis technology (e.g., Amazon Polly) to convert the response text into speech and returns it to the child.

[2008] Recognizing emotions and regulating responses

[2009] The device is equipped with an emotion engine (e.g., Affectiva's SDK) that recognizes the user's emotions. The emotion engine analyzes the user's emotions in real time using voice tone and facial expression analysis technology. The analyzed emotion data is sent to the server and reflected in the generation of a response. For example, if the user is sad, a response containing encouraging words is generated.

[2010] Send Feedback

[2011] The device uses various sensors (such as an accelerometer and camera) to monitor the child's activity and any abnormalities, and sends the data to a server. The server then notifies the parent's device of this data in real time. The parent can then check the child's status through an application and send instructions as necessary.

[2012] Specific examples

[2013] Example 1: A child plays with a "cat sitter"

[2014] 1. The user (parent) sets up a game called "Build a Castle" in the application.

[2015] 2. Based on the configuration data, the device will suggest to the child, "Let's build a castle today!"

[2016] 3. The user (child) speaks to the device, saying, "I'll bring you the block!"

[2017] 4. The device converts the voice into text and sends the data "I'll bring the block!" to the server.

[2018] 5. The server generates an appropriate response and sends the text "Yup, that's fun!" to the device.

[2019] 6. The device uses voice synthesis technology to respond to the child, "Yes, that sounds fun!"

[2020] Example 2: When a child is sad

[2021] 1. The emotion engine analyzes the situation in which the user (child) is crying.

[2022] 2. The device sends the emotion data to the server.

[2023] 3. The server generates an encouraging response based on the emotional data and sends the text "What's wrong? I'm okay" to the device.

[2024] 4. The device uses voice synthesis technology to tell the child, "What's wrong? It's okay."

[2025] Example 3: Detecting abnormal behavior in children

[2026] 1. The device monitors the child's movements using an accelerometer and camera.

[2027] 2. The device detects abnormal movement and sends data to the server stating "A fall has been detected."

[2028] 3. The server notifies the parent device of this information.

[2029] 4. The user (parent) checks the situation through the application and sends instructions (e.g., asking "Are you okay?").

[2030] 5. The device uses voice synthesis technology to carry out the parent's instructions and convey them to the child.

[2031] Example prompts for generative AI models

[2032] "Please enter the setting data required to do ____"

[2033] "Generate a response if a child is crying"

[2034] The above is a concrete example of how the present invention can be implemented. This system ensures the safety of children even when their parents are away, provides educational dialogue and play, and even recognizes emotions to provide appropriate responses.

[2035] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2036] Step 1:

[2037] A user (parent) launches an application on a smartphone or computer and enters configuration data such as household rules, preferences, schedules, etc. This configuration data includes information related to children's activities within the home, and specific actions involve entering data into fields on the application's settings screen.

[2038] Input: Family-specific rules and schedule information

[2039] Output: Setting data

[2040] Step 2:

[2041] The server receives the setting data sent by the user, encrypts it, and stores it in a database, such as MySQL. The received data is sent via the HTTP protocol.

[2042] Input: Configuration data submitted by the user

[2043] Output: Encrypted configuration data stored in the database

[2044] Step 3:

[2045] The device periodically sends a synchronization request to the server to obtain the latest configuration data. At this time, the device analyzes the data received from the server and updates its internal settings. Specifically, the device sends an API request to the server and receives a response.

[2046] Input: Configuration data stored on the server

[2047] Output: Updated internal settings

[2048] Step 4:

[2049] When a user (child) speaks to the device, the device captures the voice using a built-in microphone and converts it into text data using speech recognition technology (such as the Google Speech-to-Text API). Specifically, utterances such as "I'll bring the block!" are converted into text.

[2050] Input: Voice input from children

[2051] Output: Text data

[2052] Step 5:

[2053] The device sends the converted text data to the server, which uses a generative AI model (e.g., GPT-3) to generate an appropriate response based on the text data. The generated response is then sent back to the device. Specifically, the server generates the response using an NLP engine.

[2054] Input: Text data sent from the terminal

[2055] Output: The generated response text

[2056] Step 6:

[2057] The device converts the response text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and responds to the child through the speaker. Specifically, the device generates synthesized speech and plays it back on the speaker.

[2058] Input: Response text sent by the server

[2059] Output: Synthesized speech

[2060] Step 7:

[2061] The device uses an emotion engine (such as Affectiva's SDK) to analyze the child's emotions by analyzing voice tones and facial expressions in real time. The analysis results are sent to the server as emotion data. For example, if a child is crying, emotion data is generated.

[2062] Input: Child's vocal tone and facial expression data

[2063] Output: Generated emotion data

[2064] Step 8:

[2065] The server generates an appropriate response based on the emotion data. For example, if the child is sad, it generates an encouraging response. The generated response text is then sent back to the device.

[2066] Input: Emotion data sent from the device

[2067] Output: Sentiment-based response text

[2068] Step 9:

[2069] The device uses speech synthesis technology to convert the generated response text and respond to the child, playing the synthesized voice from the speaker.

[2070] Input: Response text sent by the server

[2071] Output: Synthesized speech

[2072] Step 10:

[2073] The device monitors the child's movements in real time using an accelerometer and camera. If it detects any abnormal movements, it sends the data to a server. For example, if a child falls, that data is generated.

[2074] Input: Data on children's movements

[2075] Output: Abnormal behavior data

[2076] Step 11:

[2077] The server receives the abnormal behavior data and notifies the parent device in real time via push notifications, emails, etc.

[2078] Input: Abnormal behavior data sent from the device

[2079] Output: Notification sent to parent device

[2080] Step 12:

[2081] The user (parent) checks the child's status through the application and sends appropriate instructions as necessary. The instructions are transmitted to the device via the server. Specifically, the parent enters instructions on the application screen and presses the send button.

[2082] Input: Parental instruction data

[2083] Output: Instruction data sent to the terminal

[2084] Step 13:

[2085] The device receives instruction data sent by the parent and uses voice synthesis technology to convey it to the child, for example by playing the message "Are you OK?" over the speaker.

[2086] Input: Instruction data sent from the parent

[2087] Output: Synthesized speech

[2088] (Application example 2)

[2089] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2090] There is a need for a system that monitors children's activities at home and allows parents to respond appropriately based on their children's emotional state. Current systems lack the ability to recognize children's emotions, making it difficult for parents to respond based on their children's emotions. There is also a need for a system that allows parents to understand their children's situation in real time and send appropriate instructions, even when they are at work or out.

[2091] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means for monitoring the child's activity status within the home, a transmission means for transmitting data acquired by the monitoring means to the parent's terminal, a dialogue means for conducting dialogue and play based on setting data received from the parent's terminal, an instruction means for sending instructions to the child based on commands received from the parent's terminal, an emotion recognition means for recognizing the user's emotion, and a response generation means for generating an appropriate response based on the recognized emotion data. This enables detailed feedback according to the child's emotional state, supports educational dialogue and appropriate play, and allows parents to watch over their children with peace of mind even when they are not at home.

[2092] "Monitoring means for monitoring children's activities within the home" refers to devices and systems such as sensors and cameras that monitor children's movements and behavior within the home in real time and obtain the necessary data.

[2093] The "transmission means for transmitting data acquired by the monitoring means to the parent's terminal" refers to a communication device and its system for transmitting data acquired by the monitoring means to a smartphone or computer used by the parent.

[2094] "An interaction means for interacting and playing based on setting data received from the parent's terminal" refers to a device and system for interacting with and playing with a child in accordance with setting data sent from the parent's terminal.

[2095] An "instruction means for sending instructions to a child based on an instruction received from a parent's terminal" is a device and system for receiving an instruction sent from a parent's terminal and encouraging a child to take a specific action based on that instruction.

[2096] The "emotion recognition means for recognizing the user's emotions" refers to a device and system for recognizing the user's emotional state in real time by analyzing the user's tone of voice and facial expressions.

[2097] The "response generating means for generating an appropriate response based on the recognized emotion data" refers to a device and system for generating an appropriate response to a user based on the recognized emotion data.

[2098] A system for implementing the present invention comprises the following major components:

[2099] 1. A stuffed toy robot equipped with a camera and sensors that can be used as a device (monitoring means) to monitor children's activities at home.

[2100] 2. Smartphones and computers used by parents (parental devices).

[2101] 3. Server.

[2102] 4. Emotion recognition means to recognize the user's emotions.

[2103] 5. A response generation means for generating an appropriate response based on the recognized emotion data.

[2104] Monitoring and Transmission Methods

[2105] The monitoring means uses cameras and sensors to monitor the child's activities in real time, making it possible to detect what the child is currently doing. The acquired data is then sent to the parent's device by the transmission means. Communication technologies such as Wi-Fi and Bluetooth can be used for transmission.

[2106] Interaction and instruction methods

[2107] The dialogue means operates based on the setting data received from the parent's device. For example, if the parent sets up a game such as "building a castle" through the application, the dialogue means will make suggestions to the child based on that information. The dialogue means also uses voice recognition technology to converse with the child. This voice data is converted into text data and sent to the server.

[2108] The instruction means prompts the child to perform a specific action based on a command received from the parent's device. For example, if the parent instructs the child to "bring a block," the instruction means prompts the child to perform the action in accordance with the command.

[2109] Emotion recognition and response generation

[2110] The emotion recognition means analyzes the user's (child's) emotions in real time using voice tone and facial expression analysis technology. The analysis results are sent to the server, and the response generation means generates an appropriate response. For example, if a child has a sad expression, the response generation means generates encouraging words such as "What's wrong? It's okay."

[2111] Feedback and Notifications

[2112] Various sensors monitor the child's activity and any abnormalities, and send the data to a server. Feedback is used to notify the parent's device in real time. The parent can check the child's condition through the application and send instructions as necessary. For example, if a child falls, data stating "a fall has been detected" is sent to the server, and a notification is also sent to the parent's device.

[2113] Specific examples

[2114] 1. Example of a set interaction: A parent sets up a game called "Build a castle" in the application. The device suggests, "Let's build a castle today!", and the child responds, "I'll bring the blocks!". This interaction is coordinated via the server.

[2115] 2. Example of emotion recognition: If a child is crying, the emotion recognition means will analyze this and generate an encouraging message such as "What's wrong? It's okay" and return it to the child.

[2116] 3. Example of abnormal behavior detection: An accelerometer or camera detects abnormal behavior and sends a real-time notification to the parent's device saying, "A fall has been detected." The parent can then check the notification and send appropriate instructions.

[2117] Example prompts to input to a generative AI model:

[2118] "Design a system that recognizes the emotions of workers from video data and generates messages to help them work safely and efficiently. Please include specific example instructions."

[2119] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2120] Step 1:

[2121] The device uses cameras and sensors to monitor the child's activities in real time. The input is camera footage and sensor data, and the device acquires the child's movements and emotions based on this. The output is the acquired activity data.

[2122] Step 2:

[2123] The device sends the data acquired by the monitoring means to the server. The transmission method is Wi-Fi or Bluetooth, and the input is activity data, and the output is data sent to the server. The server receives this data and analyzes it.

[2124] Step 3:

[2125] Based on the configuration data received from the parent's device, the server generates a dialogue or play scenario and sends it to the device. The input is the parent's configuration data, and the output is the dialogue or play scenario. This allows the device to begin a dialogue with the child according to the settings.

[2126] Step 4:

[2127] As the device engages in a dialogue with the child, it uses voice recognition technology to convert the child's voice into text data. The input is the child's voice data, and the output is text data. The content of the dialogue is sent to the server.

[2128] Step 5:

[2129] The server analyzes the received text data and generates an appropriate response. The input is the child's text data, and the output is the response text. The response text is sent to the terminal and returned to the child via voice synthesis.

[2130] Step 6:

[2131] The device's emotion recognition means analyzes the user's (child's) voice tone and facial expression in real time and generates emotion data. The input is voice tone and facial expression data, and the output is emotion data. This data is sent to the server.

[2132] Step 7:

[2133] The server receives the emotion data and generates an appropriate response based on the emotion. The input is emotion data, and the output is a response text based on the emotion. For example, if the analysis finds that a child is sad, an encouraging message is generated.

[2134] Step 8:

[2135] The server sends the generated response text to the device, which then uses speech synthesis technology to communicate it to the child. The input is the response text, and the output is a voice message to the child.

[2136] Step 9:

[2137] The device uses various sensors to monitor children's abnormal behavior and notifies the server if an abnormality is detected. The input is sensor data, and the output is notification data of the abnormality detection.

[2138] Step 10:

[2139] The server sends a notification of anomaly detection to the parent's device in real time. The input is the anomaly detection data, and the output is a notification to the parent's device. When the parent receives the notification, they can check the child's status through the application and send necessary instructions to the device.

[2140] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2141] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2142] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2143] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2144] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2145] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2146] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2147] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2148] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2149] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2150] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2151] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2152] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2153] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2154] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2155] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2156] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2157] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2158] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2159] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effec...

Claims

1. a monitoring means for monitoring the activities of children in the home; a transmitting means for transmitting the data acquired by the monitoring means to the parent terminal; an interaction means for interacting and playing with the child based on the setting data received from the parent's terminal; instruction means for sending instructions to the child based on a command received from the parent's terminal; A system including:

2. 10. The system of claim 1, further comprising a customization means for customizing the interaction means with content based on household specific rules and preferences.

3. The system of claim 1 , further comprising a feedback means for transmitting feedback data on the child's activities to the parent's terminal in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A