system

A system using a camera, temperature sensor, and audio device with AI-generated lullabies and environmental control addresses the challenge of putting babies to sleep, enhancing parental well-being by providing real-time monitoring and safety.

JP2026036095APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

Smart Images

  • Figure 2026036095000001_ABST
    Figure 2026036095000001_ABST
Patent Text Reader

Abstract

To provide a smart sleep-training system that appropriately monitors a baby's condition, generates and plays unique lullabies, and adjusts the environment and detects abnormalities as needed. [Solution] A system including: means for acquiring baby's movements from a camera device; means for acquiring environmental and baby's body temperature data from a temperature sensor device; means for acquiring baby's voice from an audio device; means for analyzing the acquired data and determining the baby's condition; means for generating a lullaby according to the baby's condition using music generation AI; means for playing the generated lullaby; means for controlling the air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment; means for sending a notification to the parent's communication device when the baby falls asleep; and means for immediately sending an alert to the parent's communication device if something abnormal occurs with the baby.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] For many parents, putting a baby to sleep is a very time-consuming task, and the effectiveness of lullabies can diminish if the same lullaby is played over and over again. Furthermore, parents often struggle to get enough rest at night because it is difficult to quickly respond to changes in the baby's physical condition or environment. This can lead to stress and fatigue for parents, affecting the quality of life of the entire family. To address these issues, a smart sleep-training system is needed that can appropriately monitor the baby's condition, generate and play unique lullabies, and adjust the environment or detect abnormalities as needed. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means.

[0006] 1. A means of capturing baby movements from a camera device.

[0007] 2. A means of obtaining environmental and baby temperature data from a temperature sensor device.

[0008] 3. A means of capturing baby audio from an audio device.

[0009] 4. A means of analyzing the acquired data and determining the baby's condition.

[0010] 5. A means of using music generation AI to generate lullabies according to the baby's condition.

[0011] 6. A means of playing the generated lullaby.

[0012] 7. A means to control the air conditioning as needed and adjust the temperature to keep your baby in a comfortable environment.

[0013] 8. A way to send a notification to the parent's communication device when the baby falls asleep.

[0014] 9. A means of instantly sending alerts to parents' communication devices if something abnormal happens to the baby.

[0015] These measures provide an intelligent system that consistently supports putting babies to sleep, reduces the burden on parents, and improves the quality of life for the entire family.

[0016] A "camera device" is a photographic device used to monitor and record a baby's movements and status in real time.

[0017] A "temperature sensor device" is a sensor device for measuring a baby's body temperature and the ambient temperature of a room.

[0018] An "audio device" is a sound collection device for capturing a baby's cry and other sounds.

[0019] The "means for analyzing data" is a data processing system for determining the baby's condition based on the acquired video, audio, and temperature data.

[0020] "Music generation AI" is an artificial intelligence technology that generates the optimal lullaby depending on the baby's condition.

[0021] A "means for generating lullabies" is a device or system that automatically creates lullabies based on the baby's condition.

[0022] The "means for playing a lullaby" is an audio playback device for playing the generated lullaby to the baby.

[0023] The "means for controlling the air conditioner" is a system that adjusts the operation of the air conditioner to maintain a comfortable environment for the baby.

[0024] The "means for sending notifications" is a system that sends information to the parent's communication device when the baby is sleeping stably.

[0025] The "means for sending alerts" is a system that immediately sends a warning to the parent's communication device if something abnormal happens to the baby. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0027] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0028] First, the terms used in the following description will be explained.

[0029] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0030] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0031] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0032] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0034] [First embodiment]

[0035] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0036] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0037] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0038] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0039] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0041] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0042] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0043] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0044] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0045] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0046] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0047] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0048] System Configuration

[0049] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0050] What the program does

[0051] 1. Data Collection

[0052] The server collects real-time baby movement data from the camera device, which can determine whether the baby is moving or quiet.

[0053] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[0054] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0055] 2. Data analysis and feedback

[0056] The server analyzes the collected data and evaluates the baby's condition (e.g., excited, relaxed, normal body temperature, high temperature). For example, if the baby continues to cry, it is determined to be excited.

[0057] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0058] 3. Lullaby Generation and Playback

[0059] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited, a lullaby with a slow rhythm will be generated.

[0060] The device (e.g., a smartphone or tablet) plays the generated lullaby through a speaker in the baby's room, providing a sound environment that suits the baby's condition.

[0061] 4. Environmental adjustment

[0062] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the room.

[0063] It receives feedback from the temperature sensor and continues to adjust until it reaches the target temperature.

[0064] 5. Notifications and Anomaly Detection

[0065] The server checks the baby's sleep status and, if it is confirmed that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, allowing the parent to know that the baby has fallen asleep safely.

[0066] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends an alert to the parent's communication device, allowing the parent to respond promptly.

[0067] Specific examples

[0068] Example 1: If a baby continues to cry, the camera and audio devices will be used to check the baby's condition, and the music generation AI will create a soothing lullaby. The lullaby will be played through the speakers, and the air conditioner will start cooling the room. As a result, the baby will gradually calm down and fall asleep.

[0069] Example 2: If the baby's body temperature is high and the room temperature is also high, the air conditioner will be controlled based on data from the temperature sensor device and the cooling will start. At the same time, a gentle lullaby will be generated and played to help the baby relax. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0070] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. It also provides a safe and secure environment for parents, reducing the burden of childcare.

[0071] The processing flow will be explained below.

[0072] Step 1:

[0073] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[0074] Step 2:

[0075] The server receives the room temperature data and the baby's temperature data from the temperature sensor device. Specifically, it records the baby's temperature and the room's ambient temperature based on the numerical data obtained from the temperature sensor.

[0076] Step 3:

[0077] The server receives the baby's crying and other sounds in real time from the audio device, and analyzes the audio input to determine whether the baby is crying or quiet.

[0078] Step 4:

[0079] The server integrates and analyzes the data from steps 1 to 3 to evaluate the baby's current state (excited, relaxed, normal body temperature, high temperature, etc.) For example, if the baby is crying continuously, has a high body temperature, and is active, it is determined to be in an excited state.

[0080] Step 5:

[0081] The server uses music generation AI to generate the optimal lullaby based on the analysis results. Specifically, it creates lullabies with slow rhythms or gentle melodies depending on the results of data analysis.

[0082] Step 6:

[0083] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data is transmitted to the speaker using Wi-Fi, and the audio is played back.

[0084] Step 7:

[0085] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it automatically adjusts the air conditioner's cooling or heating mode based on the temperature data and sets the target temperature.

[0086] Step 8:

[0087] The server receives feedback from the temperature sensor and continues to appropriately control the air conditioner until the target temperature is reached. For example, it may continue cooling until the target temperature is reached, and then stop the air conditioner when the temperature is appropriate.

[0088] Step 9:

[0089] The server confirms that the baby is asleep, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[0090] Step 10:

[0091] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[0092] This allows the system to monitor the baby's condition in real time and respond appropriately to the situation, reducing the burden on parents and providing a comfortable sleeping environment for the baby.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] Conventional baby monitoring systems tend to rely on specific devices, making it difficult to comprehensively monitor a baby's condition and provide appropriate feedback. As a result, parents are unable to consistently monitor their baby's condition, placing a heavy burden on childcare. Furthermore, due to insufficient environmental adjustment and anomaly detection functions, it is difficult to provide a sustainable, comfortable environment for babies. To solve these problems, a system is needed that can comprehensively monitor a baby's condition in real time and automatically take appropriate action.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for analyzing the acquired data and evaluating the baby's condition, means for generating a lullaby according to the baby's condition using a music generation algorithm, means for playing the generated lullaby through a playback device, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if something abnormal occurs with the baby. This makes it possible to monitor the baby's condition from various angles and automatically take appropriate measures.

[0098] A "camera device" is a device that captures a baby's movements and acquires the video data in real time.

[0099] A "temperature sensor device" is a device that measures the baby's surrounding environment and the baby's own body temperature and provides that data to a server.

[0100] An "audio device" is a device that collects the baby's crying and other surrounding sounds and sends the data to a server.

[0101] The "server" is a central computer system that integrates and analyzes data collected from various devices and generates appropriate feedback.

[0102] The "means for analyzing the acquired data and assessing the baby's condition" refers to algorithms or software for estimating the baby's current condition based on the collected movement data, temperature data, and audio data.

[0103] A "music generation algorithm" is an AI (artificial intelligence) model or program that generates the optimal lullaby depending on the baby's condition.

[0104] A "playback device" is a speaker or audio device for playing back the generated lullaby sent from the server.

[0105] "Air conditioning equipment" refers to equipment with heating and cooling functions for adjusting the temperature in a room, and includes air conditioners.

[0106] A "parent communication device" is a mobile device, such as a smartphone or tablet, used by a parent to receive notifications and alerts.

[0107] "Means for sending an alert in the event of an abnormality" refers to software or hardware configurations that immediately send an alert message to the parent's communication device when an abnormality in the baby is detected from the collected and analyzed data.

[0108] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0109] System Configuration

[0110] This system is composed of multiple hardware and software devices, including the following:

[0111] Camera device: A device for capturing baby's movements in real time, such as an infrared camera or a high-resolution camera.

[0112] Temperature sensor device: A sensor device that measures the temperature of the room and the baby. A high-precision digital thermometer is used.

[0113] Audio device: A microphone device that collects the baby's cries and surrounding sounds.

[0114] Server: A central computer system that analyzes data collected from each device and generates the necessary feedback. A server equipped with a high-performance processor is used here.

[0115] Device: A communication device used by the end user (parent). For example, a smartphone or tablet.

[0116] Data collection

[0117] The server collects the baby's movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0118] The server receives real-time room temperature and baby temperature data from the temperature sensor device, which can determine whether the baby is within a comfortable temperature range.

[0119] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0120] Data analysis and feedback

[0121] The server then aggregates and analyzes the collected data, for example analyzing the audio data to determine whether the baby is crying, using a voice analysis algorithm.

[0122] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0123] Lullaby generation and playback

[0124] The server uses a music generation algorithm to generate an appropriate lullaby depending on the baby's state. For example, if the baby is excited, a lullaby with a calming rhythm will be generated.

[0125] The device receives the generated lullaby and plays it through a speaker in the room, which may be connected to the smartphone via Wi-Fi.

[0126] environmental adjustment

[0127] The server controls the air conditioning unit as needed to adjust the temperature to keep the baby in a comfortable environment. For example, if the room temperature is high, the server activates the air conditioning's cooling function.

[0128] It receives feedback from temperature sensor devices and adjusts the air conditioner until the target temperature is reached.

[0129] Notifications and Anomaly Detection

[0130] If the server determines that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, for example, a message saying "Baby has fallen asleep."

[0131] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends a warning to the parent's communication device, such as a message saying, "Your baby's temperature is high. Please check it."

[0132] Specific examples

[0133] Example 1:

[0134] If the baby continues to cry, the server will check the baby's condition from the camera and audio data and use a music generation algorithm to generate a soothing lullaby. The lullaby will be played through the speaker, and the air conditioner will start cooling to adjust the room temperature. As a result, the baby will gradually calm down and fall asleep.

[0135] Example 2:

[0136] If the baby's body temperature is high and the room temperature is also high, the server will control the air conditioning based on the data from the temperature sensor device and turn on the air conditioner. At the same time, a gentle lullaby to help the baby relax will be generated and played. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0137] Prompt Sentence Examples

[0138] "Monitor your baby's crying and play a soothing lullaby."

[0139] "While controlling the air conditioning unit if the room temperature is high, it also checks the baby's temperature and generates an appropriate lullaby."

[0140] In this way, the present invention monitors the baby's condition from multiple angles and responds appropriately to provide a comfortable sleeping environment for the baby. It also provides a safe and secure childcare environment for parents, reducing the burden of childcare.

[0141] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0142] Step 1:

[0143] Data collection

[0144] The server collects baby's movement data in real time from the camera device. The input is the video data captured by the camera. The server analyzes this video data frame by frame and extracts the baby's movements (e.g., whether or not the baby is moving and the speed of the movements). The output is the baby's movement pattern data.

[0145] The server obtains the room temperature and the baby's body temperature data from the temperature sensor device. The temperature data measured by the sensor at regular intervals is the input. The server receives and analyzes this data to determine whether the room is within the appropriate temperature range or whether the baby's body temperature is abnormal. The output is the environmental temperature data and the baby's body temperature data.

[0146] The server collects the baby's crying and other ambient sounds from the audio device. The audio data captured by the audio device's microphone is the input. The server uses an audio analysis algorithm to analyze the characteristics of the audio (e.g., volume, frequency) and determine whether the baby is crying. The output is the audio analysis result.

[0147] Step 2:

[0148] Data analysis and feedback

[0149] The server integrates and analyzes all the data collected in step 1 (movement pattern data, ambient temperature data, body temperature data, and voice analysis results). Based on these analysis results, the baby's condition is evaluated. For example, if the baby continues to cry, it is judged to be in an "excited state" based on the voice analysis results. The inputs are each sensor data and its analysis results. The output is the evaluation result of the baby's condition.

[0150] Step 3:

[0151] Lullaby generation and playback

[0152] The server uses a music generation algorithm to generate an appropriate lullaby based on the baby's state evaluated in step 2. For example, the prompt sentence is "Generate a calming song." The music generation AI model generates a lullaby based on this prompt sentence, and the generated lullaby data is obtained as the output.

[0153] The device receives the lullaby data sent from the server and plays it through the room's speakers. The input is the lullaby data from the server. The device sends it to a playback device, and the sound is played as the output.

[0154] Step 4:

[0155] environmental adjustment

[0156] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. If the temperature sensor device data indicates that the room is hot, the server issues a command to the air conditioner to operate the cooling function. The operating status of the air conditioner is obtained as an output.

[0157] The server continues to receive feedback from the temperature sensor and continues to control the air conditioning until the target temperature is reached. The input is the continuous temperature feedback data. The output is the final adjusted temperature data.

[0158] Step 5:

[0159] Notifications and Anomaly Detection

[0160] When the server confirms that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." The input includes the analysis results of motion data and voice data. The output is a notification message sent to the parent's communication device.

[0161] The server immediately sends an alert to the parent's communication device if the baby experiences any abnormalities (e.g., high temperature, respiratory failure, etc.). The input is temperature data and other sensor data indicating an abnormality. The output is a warning message sent to the parent's communication device.

[0162] (Application example 1)

[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0164] In conventional factory environments, monitoring of workers' movements, physical condition, and working environment is insufficient, resulting in reduced work efficiency and health risks for workers. In particular, working in the same position for long periods of time and working in high-temperature environments can cause fatigue and health problems for workers. There is also a need for a system that can respond immediately when a worker becomes ill.

[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0166] In this invention, the server includes means for acquiring the target person's movements from a camera device, means for acquiring environmental and target person's body temperature data from a temperature sensor device, means for acquiring the target person's voice from an audio device, means for analyzing the acquired data and determining the target person's condition, means for generating voice instructions according to the target person's condition using a music generation AI, means for playing the generated voice instructions, means for controlling the air conditioner as necessary to adjust the temperature so that the target person is in a comfortable environment, means for sending a notification to a manager's communication terminal when the target person enters a specific state, and means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person. This enables real-time monitoring of worker movements and physical condition, enabling optimization of the work environment and health management.

[0167] A "camera device" is a photographing device for capturing the movements of a target person in real time.

[0168] A "temperature sensor device" is a measuring device for acquiring body temperature data of the environment and a target person.

[0169] An "audio device" is a sound collection device for acquiring the voice of a target person.

[0170] "Means for analyzing acquired data and determining the status of the target person" refers to a method or apparatus for processing information acquired from the camera device, temperature sensor device, and audio device and assessing the current status of the target person.

[0171] "Means for generating voice instructions according to the state of a target person using music generation AI" refers to a method or device that uses artificial intelligence technology to generate voice instructions that are adapted to the state of a target person.

[0172] "Means for playing generated voice instructions" refers to a device or method for playing voice instructions generated by the music generation AI.

[0173] "Means for controlling an air conditioner and adjusting the temperature so that the target person is in a comfortable environment" refers to a method or device for operating an air conditioner to appropriately adjust the environmental temperature.

[0174] "Means for sending a notification to the manager's communication terminal when the target person enters a specific state" refers to a method or device for notifying the manager of this information when the target person reaches a specific state, such as a situation where a break is required.

[0175] "Means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person" refers to a method or device for immediately sending a warning to the manager when an abnormality occurs in the target person's health.

[0176] The present invention aims to provide a work support robot system that monitors the movements and environment of a target person (worker) and provides necessary instructions and adjusts the environment. Specific embodiments of the present invention will be described below.

[0177] System Configuration

[0178] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0179] What the program does

[0180] 1. Data Collection

[0181] The server collects real-time motion data of the target person from the camera device, which allows it to determine whether the target person is moving or stationary.

[0182] The temperature sensor device acquires the ambient temperature and the target person's body temperature data, which can then be used to determine whether the target person is within a comfortable temperature range.

[0183] Collecting the subject's voice from a voice device, which allows us to determine whether the subject understands the instructions.

[0184] 2. Data analysis and feedback

[0185] The server analyzes the collected data and evaluates the target person's status (e.g., working, needing a break, abnormal state). For example, if the target person remains motionless for a long time, it determines that the person needs a break.

[0186] If each piece of data is outside the specified range (e.g., room temperature is high, lighting is low), the system determines how to respond accordingly.

[0187] 3. Generation and playback of voice instructions

[0188] The server uses music generation AI to generate optimal voice instructions based on the analysis results. For example, if the person is tired, it will generate a voice instruction to encourage them to take a break.

[0189] The terminal (e.g., a smartphone) plays the generated voice instructions through an audio device that transmits the voice instructions to the target person, thereby providing appropriate instructions according to the target person's state.

[0190] 4. Environmental adjustment

[0191] The server controls the air conditioner as needed to keep the temperature in the factory within an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the factory.

[0192] 5. Notifications and Anomaly Detection

[0193] When the target person enters a specific state (e.g., needs a break), the server sends a notification to the administrator's communication terminal, allowing the administrator to grasp the target person's state.

[0194] Additionally, if an abnormality occurs with the target person (e.g., high fever, immobility, etc.), an alert is immediately sent to the administrator's communication terminal, allowing the administrator to respond promptly.

[0195] Hardware and software used

[0196] Camera devices: Used to monitor the behavior of subjects.

[0197] Temperature sensor device: Used to monitor environmental and body temperatures.

[0198] Audio device: Used to communicate instructions to the target person.

[0199] Server: Data processing and analysis.

[0200] Communication device: Used to receive notifications and alerts.

[0201] OpenCV: Software used to analyze camera footage.

[0202] pyttsx3: Software used to generate and play audio instructions.

[0203] GPIO Zero: A library used to acquire data from the temperature sensor.

[0204] Specific examples

[0205] Example 1: If a worker remains motionless for more than 30 minutes, the system will issue a voice message saying "Please take a break" and will also send a work status notification to the manager.

[0206] Example 2: If the temperature inside the factory goes outside the set range (e.g., above 28 degrees), the system automatically switches the air conditioner to cooling mode.

[0207] Example prompt for a generative AI model:

[0208] Similar to the system that provides voice alerts and adjusts the temperature based on a baby's movements and temperature data, create a program that monitors the movements and temperature data of factory workers, issues voice alerts when they become fatigued, and controls the air conditioner as necessary. Specific examples include notifications if a worker has not moved for 30 minutes and switching on the air conditioner if the temperature exceeds 28 degrees.

[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0210] Step 1:

[0211] The server acquires the motion data of the target person from the camera device in real time. Specifically, the camera device captures the video of the target person and sends the video data to the server. The server analyzes the video data to determine whether the target person is moving. The input is the video data from the camera device, and the output is the motion status of the target person (e.g., moving, stationary).

[0212] Step 2:

[0213] The server receives the ambient temperature and the target person's body temperature data from the temperature sensor device. Specifically, the temperature sensor device measures the temperature data in real time and sends it to the server. The server analyzes the measured data to determine whether the ambient temperature and the target person's body temperature are within an appropriate range. The input is the temperature data from the temperature sensor device, and the output is the temperature status (e.g., normal, high temperature, low temperature).

[0214] Step 3:

[0215] The server acquires the target person's voice data from the audio device. Specifically, it picks up the target person's voice and other sounds and sends them to the server. The server analyzes the voice data to determine what the target person is saying or what sounds are being made. The input is the voice data from the audio device, and the output is the content of the voice (e.g., response to commands, background sounds).

[0216] Step 4:

[0217] The server comprehensively analyzes the collected data and determines the target person's condition. Specifically, it integrates motion data, temperature data, and voice data, and evaluates whether the target person is working, needs a break, or is in an abnormal state based on each piece of data. The input is motion data, temperature data, and voice data, and the output is the target person's overall condition (e.g., working, fatigue, abnormal).

[0218] Step 5:

[0219] The server uses music generation AI to generate optimal voice instructions based on the analysis results. Specifically, it generates appropriate instructions (e.g., "Please take a break" or "Please continue working") depending on the target person's condition. The input is the target person's overall condition, and the output is the generated voice instructions.

[0220] Step 6:

[0221] The server transmits the generated voice instructions to the target person through a terminal (e.g., a smartphone) or an audio device. Specifically, the generated voice instructions are played back as audio so that the target person can confirm the instructions. The input is the generated voice instructions, and the output is the played back audio.

[0222] Step 7:

[0223] The server controls the air conditioner as needed to adjust the temperature in the factory to an appropriate range. Specifically, if the temperature data exceeds the set range, the air conditioner will operate to cool or heat the factory. The input is the temperature state, and the output is the appropriate adjustment of the environmental temperature.

[0224] Step 8:

[0225] The server sends a notification to the manager's communication terminal when the target person enters a specific state. Specifically, for example, if it determines that the target person needs a break, it notifies the manager of that information. The input is the target person's overall state, and the output is a notification to the manager.

[0226] Step 9:

[0227] The server immediately sends an alert to the administrator's communication terminal if an abnormality occurs in the target person. Specifically, if it determines that the target person is in an abnormal condition, such as having a high fever or not moving, it sends an emergency alert to the administrator. The input is the target person's abnormal condition, and the output is an emergency alert.

[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0229] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[0230] System Configuration

[0231] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[0232] What the program does

[0233] 1. Data Collection

[0234] The server receives real-time baby movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0235] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[0236] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0237] 2. Emotion recognition by emotion engine

[0238] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. For example, if the user is tired or stressed, the server obtains that information.

[0239] 3. Data analysis and feedback

[0240] The server integrates and analyzes the baby's data and the user's emotional data to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state.

[0241] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, the crying continues, or the user is feeling stressed), the system determines how to respond accordingly.

[0242] 4. Lullaby Generation and Playback

[0243] The server uses music generation AI to generate the optimal lullaby based on the baby's state and the user's emotions. For example, if the baby is excited and the user is tired, a particularly relaxing lullaby will be generated.

[0244] The device (e.g., a smartphone or tablet) then transmits the generated lullaby audio data to a speaker in the baby's room and plays it back, providing a sound environment tailored to the baby's condition.

[0245] 5. Environmental adjustment

[0246] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. For example, if the room temperature is high, the server will operate the air conditioner to cool it down.

[0247] The lighting in a room can also be adjusted based on the user's emotional data, for example, by adjusting the lighting in the room to a softer light to help the user relax.

[0248] 6. Notifications and Anomaly Detection

[0249] The server checks the baby's sleeping state, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[0250] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[0251] Specific examples

[0252] Example 1: If a baby continues to cry and the user is feeling stressed, the camera and audio devices will be used to check the baby's condition. The music generation AI will generate a relaxing lullaby and play it through the speaker. Furthermore, the air conditioner will start cooling and the room lighting will be adjusted to a softer light, creating a comfortable environment for the baby and the user.

[0253] Example 2: If the baby's body temperature is high and the user is tired at the same time, the air conditioner is controlled based on the data from the temperature sensor device and the analysis results of the emotion engine. At the same time, a gentle lullaby to help the baby relax is generated and played. When the baby falls asleep, a notification "The baby has fallen asleep" is sent to the parent's communication device.

[0254] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. In addition, by taking the user's emotions into consideration, it also provides a safe and secure environment for parents and reduces the burden of childcare.

[0255] The processing flow will be explained below.

[0256] Step 1:

[0257] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[0258] Step 2:

[0259] The server acquires the room temperature data and the baby's body temperature data from the temperature sensor device. Specifically, it records the baby's body temperature and the room ambient temperature based on the numerical data obtained from the temperature sensor.

[0260] Step 3:

[0261] The server receives the baby's crying and other sounds in real time from the audio device, analyzes the audio data, and determines whether the baby is crying or quiet.

[0262] Step 4:

[0263] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. Specifically, it obtains data on the user's emotional state, such as whether they are tired, stressed, or relaxed.

[0264] Step 5:

[0265] The server integrates and analyzes the data from steps 1 to 4 to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state. For example, if the baby continues to cry and the user is tired, the server determines that the baby is excited and the user is tired.

[0266] Step 6:

[0267] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited and the user is tired, it will generate a lullaby that has a particularly relaxing effect.

[0268] Step 7:

[0269] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data can be transmitted to the speaker using Wi-Fi, and the relaxing lullaby will be played.

[0270] Step 8:

[0271] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it sets the air conditioner to cooling or heating mode based on the temperature data and adjusts it to the target temperature.

[0272] Step 9:

[0273] The server then adjusts the lighting in the room appropriately based on the user's emotional data, for example, changing the lighting to softer light to help the user relax.

[0274] Step 10:

[0275] The server checks whether the baby is asleep, and if it is confirmed that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying "The baby has fallen asleep."

[0276] Step 11:

[0277] If something unusual happens to the baby, the server will immediately send an alert to the parent's communication device. For example, if the baby's temperature exceeds a predetermined range or breathing stops, an alert will be sent immediately saying "There is something wrong with the baby."

[0278] This allows the system to monitor the baby's condition from multiple angles and take optimal action while taking into consideration the user's emotions, thereby providing a comfortable sleeping environment for the baby and reducing the burden of childcare on the parents.

[0279] Example 2

[0280] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0281] In modern childcare settings, it is necessary to constantly monitor a baby's health and comfort and respond appropriately. However, parents are not always close to their babies, which increases the burden of childcare. In addition, it is necessary to respond by taking into account not only the baby's condition but also the parent's emotions and stress, but conventional systems do not adequately address this. Therefore, the challenge is to provide a comfortable environment for both babies and parents and reduce the burden of childcare on parents.

[0282] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0283] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for recognizing a user's emotion using an emotion engine, means for integrating and analyzing the baby's condition data and the user's emotion data, means for generating a lullaby according to the baby's condition using a music generation AI, means for playing the generated lullaby, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for adjusting the environment based on the user's emotion data, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if an abnormality occurs in the baby. This makes it possible to comprehensively monitor and analyze the conditions of both the baby and the parent and automatically take appropriate measures to provide an optimal environment for the baby and the parent and reduce the burden of childcare.

[0284] The "camera device" is a video capture device for monitoring and capturing the baby's movements in real time.

[0285] A "temperature sensor device" is a device for measuring and acquiring the temperature of the environment and the baby's body temperature.

[0286] "Audio Device" refers to a recording device for collecting and capturing a baby's cry or other sounds.

[0287] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize the user's emotional state.

[0288] "Music Generation AI" is an artificial intelligence system that generates optimal lullabies based on the baby's condition and the user's emotions.

[0289] An "air conditioner" is an air conditioning device that adjusts and maintains the temperature in a room.

[0290] "Communication devices" are communication devices such as smartphones and tablets used by parents.

[0291] "Lullaby generation" is the process in which the music generation AI creates a lullaby taking into account the baby's condition and the user's emotions.

[0292] "Playback" is the process of outputting the generated lullaby as sound through a speaker.

[0293] "Notifications" are messages sent to parents' communication devices to inform them of changes in the baby's condition or environment.

[0294] An "alert" is a warning message that sends an emergency notification to parents when there is an abnormality in the baby's health or safety.

[0295] "User emotion data" is data that indicates the user's emotional state analyzed by the emotion engine.

[0296] "Environmental adjustment" is the process of changing settings such as temperature and lighting to provide a comfortable environment for the baby and the user.

[0297] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[0298] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[0299] Data collection

[0300] The server receives real-time motion data from the camera device, including the baby's limb movements, rolling over, etc. For example, if the baby is waving its arms, the data is sent to the server.

[0301] The server receives the room temperature data and the baby's temperature data from the temperature sensor device, and the specific values ​​that the room temperature is 25 degrees and the baby's temperature is 37 degrees are sent to the server.

[0302] The server collects the baby's cry and other sounds from the audio device. For example, if the baby is crying loudly, the decibel level of the cry is sent to the server.

[0303] Recognizing user emotions with an emotion engine

[0304] The server uses the camera device to analyze the user's (parent's) facial expressions. For example, when the user's face is facing the camera, the server uses a facial expression recognition algorithm to identify emotions such as smile, anger, sadness, etc.

[0305] The server uses the audio device to analyze the tone of the user's voice. For example, if the user's voice is high-pitched and fast, it determines that the user is stressed. By combining these data, it determines whether the user is tired or not.

[0306] Data analysis and feedback

[0307] The server analyzes the collected data on the baby's movements, temperature, voice, and user's emotions. For example, if a baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, the server will determine that the baby may be crying because of the heat.

[0308] Lullaby generation and playback

[0309] The server uses a music generation AI model to generate the optimal lullaby based on the baby's state and the user's emotions. For example, the prompt text is "My baby is excited. Please generate a lullaby that has a relaxing effect."

[0310] The device receives the generated lullaby audio data and transmits it to a speaker in the baby's room, for example, a smartphone connected to the speaker via Bluetooth, which plays the lullaby.

[0311] environmental adjustment

[0312] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, the server sends a command to the air conditioner to set the temperature to 24 degrees.

[0313] The server can also adjust the lighting in the room based on the user's emotional data, for example adjusting the lighting to 30% brightness to help the user relax.

[0314] Notifications and Anomaly Detection

[0315] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's device saying "The baby has fallen asleep." Specifically, it sends a push notification to the parent's smartphone.

[0316] If something abnormal happens to the baby, the server will quickly send an alert to the parent's device. For example, if the baby's temperature exceeds 39 degrees, an alert will be sent saying "There is something abnormal with the baby."

[0317] Example prompts for generative AI models

[0318] "Generate a relaxing lullaby to play when your baby is crying."

[0319] "Right now, your baby is excited and the room temperature is high. Please generate a lullaby that will help your baby relax."

[0320] "When parents are tired, they need a lullaby to calm their baby. Please generate an appropriate lullaby."

[0321] As described above, this system comprehensively monitors the baby's condition and the user's emotions and takes appropriate action to support the baby's comfortable sleep and reduce the burden of childcare.

[0322] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0323] Step 1:

[0324] Data collection

[0325] The server obtains the baby's motion data in real time from the camera device. As input, it receives the video stream from the camera device and processes it with an image analysis algorithm. As output, it obtains motion data such as the baby's limb movements and rolling over. Specifically, the camera captures the baby's movements and sends the data to the server.

[0326] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device. As input, it receives and analyzes the temperature data from the temperature sensor device. As output, it obtains specific numerical data about the room temperature and the baby's body temperature. For example, data that the room temperature is 25 degrees and the baby's body temperature is 37 degrees is sent to the server.

[0327] The server collects the baby's cry and other sounds from the audio device. As input, it receives the audio stream from the audio device and processes it with an audio analysis algorithm. As output, it obtains data about the decibel level and type of sound of the baby's cry. Specifically, if the baby is crying hard, the decibel level of the cry is sent to the server.

[0328] Step 2:

[0329] Recognizing user emotions with an emotion engine

[0330] The server analyzes the user's facial expressions using a camera device. As input, it receives the video stream from the camera device and processes it with a facial expression recognition algorithm. As output, it obtains data about the user's emotional state (e.g., smiling, angry, sad, etc.). Specifically, when the user's face is facing the camera, the server analyzes their facial expressions.

[0331] The server analyzes the tone of the user's voice using the audio device. As input, it receives the audio stream from the audio device and processes it with an audio tone analysis algorithm. As output, it obtains data about the user's emotional state (e.g., stress, fatigue, etc.). Specifically, if the user's voice is high-pitched and fast, it is determined that the user is stressed.

[0332] Step 3:

[0333] Data analysis and feedback

[0334] The server integrates and analyzes the collected baby's movement data, temperature data, voice data, and user's emotional data. As input, it receives the data obtained in each step above and processes it with a data integration algorithm. As output, it obtains judgment data regarding the baby's current state (e.g., crying, warm, etc.) and the user's emotional state. In a specific operation, assuming that the baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, it will make the judgment that "the baby may be crying because of the heat."

[0335] Step 4:

[0336] Lullaby generation and playback

[0337] The server uses a music generation AI model to generate an optimal lullaby based on the baby's state and the user's emotions. The AI ​​model receives the prompt "My baby is excited. Please generate a lullaby that has a relaxing effect." The output is audio data of a lullaby with a relaxing effect. In concrete terms, the AI ​​generates a lullaby based on the prompt, and the server receives the data.

[0338] The device receives the generated lullaby audio data and sends it to the speaker in the baby's room. As input, it receives the audio data sent from the server and sends it to the speaker. As output, the lullaby is played from the speaker. Specifically, the smartphone is connected to the speaker via Bluetooth and the lullaby is played.

[0339] Step 5:

[0340] environmental adjustment

[0341] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. As input, it receives room temperature data and processes it with the air conditioner control algorithm. As output, it obtains commands to send to the air conditioner. In concrete terms, the server sends a command to the air conditioner to "set the temperature to 24 degrees."

[0342] The server can also adjust the lighting in a room based on the user's emotional data. It receives the user's emotional data as input and processes it with a lighting control algorithm. The output is a command to send to the lighting. Specifically, the server sends a command to adjust the lighting to 30% brightness to help the user relax.

[0343] Step 6:

[0344] Notifications and Anomaly Detection

[0345] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." As input, it receives and analyzes the baby's movement data. As output, it generates a notification message and sends it to the parent's device. Specifically, a push notification is sent to the smartphone, displaying the message, "The baby has fallen asleep."

[0346] If something abnormal occurs with the baby, the server quickly sends an alert to the parent's communication device. As input, it receives and analyzes the baby's temperature and breathing data. As output, it generates an alert message informing the parent of the abnormality and sends it to the parent's device. For example, if the body temperature exceeds 39 degrees, an alert saying "There is something abnormal with the baby" is sent to the smartphone.

[0347] (Application example 2)

[0348] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0349] In today's physical stores, customers often feel stressed, which causes a poor customer experience. It is also difficult for store staff to grasp customers' emotional state in real time and respond appropriately, making it a challenge to improve customer satisfaction. Furthermore, some stores do not adequately adjust the environment to maintain customer comfort.

[0350] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0351] In this invention, the server includes means for acquiring motion from a camera device, means for acquiring environmental and body temperature data from a temperature sensor device, means for acquiring audio from an audio device, means for analyzing the acquired data and determining the state, means for generating music according to the state using a music generation AI, means for playing the generated music, means for controlling an air conditioner as needed to adjust the temperature to maintain a comfortable environment, means for sending a notification to a notification terminal when the state changes, means for immediately sending an alert to the notification terminal when an abnormality occurs, means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize emotions, means for adjusting the environment according to the user's emotional state based on data from the audio device, means for the music generation AI to generate music with a high relaxing effect based on the emotional data, and means for controlling lighting to adjust the ambient light, thereby enabling real-time environmental adjustment according to the customer's emotional state.

[0352] A "camera device" is an electronic device for capturing video and image data.

[0353] "Movement" refers to physical movement or action.

[0354] A "temperature sensor device" is an electronic device for measuring the temperature of an environment or object.

[0355] "Environmental data" refers to data that represents the state or condition of the surroundings, such as temperature and light intensity.

[0356] "Body temperature data" refers to data obtained by measuring and recording the temperature of a living body.

[0357] An "audio device" is an electronic device for capturing sound, such as a microphone.

[0358] A "means for determining the status" is a method or system that analyzes and evaluates the current situation or conditions based on acquired data.

[0359] "Music generation AI" is a system that automatically generates music using artificial intelligence technology.

[0360] An "air conditioning device" is a device that adjusts temperature and humidity. Air conditioners are examples of this.

[0361] A "notification terminal" is a device that receives signals and messages, such as a smartphone.

[0362] An "alert" is a warning message or notification of a particular condition.

[0363] The "emotion engine" is a system that analyzes emotions from facial expressions and tone of voice.

[0364] "User" means any person who uses a system or service.

[0365] "Environmental control" is the act of controlling environmental conditions such as temperature and light to maintain or improve comfort.

[0366] "Music with a high relaxing effect" is music that is intended to relax the listener.

[0367] A "means for controlling lighting" is a system or method for adjusting the intensity or color temperature of light.

[0368] This invention is a system that monitors the emotional state of customers in real time in physical stores and adjusts the environment accordingly to improve customer experience. This system consists of a camera device, a temperature sensor device, a voice device, an emotion engine, a music generation AI, an air conditioning system, a lighting system, and a notification terminal.

[0369] System Configuration

[0370] The system includes the following major components:

[0371] 1. Camera Device

[0372] The camera device is used to capture customer behavior in real time, specifically, a common webcam such as the Logitech C920 can be used.

[0373] 2. Temperature sensor device

[0374] Temperature sensor devices are used to obtain environmental and customer body temperature data. For example, DHT22 sensors are available.

[0375] 3. Audio Devices

[0376] The audio device is used to capture the customer's voice. A high-sensitivity microphone such as a Blue Yeti can be used.

[0377] 4. Emotion Engine

[0378] The emotion engine uses data from cameras and audio devices to analyze customers' facial expressions and tone of voice to recognize emotions, using an AI model powered by OpenCV and TENSORFLOW®.

[0379] 5. Music Generation AI

[0380] The music generation AI generates relaxing music according to the customer's emotional state, using AI frameworks such as Google Magenta.

[0381] 6. Air conditioner

[0382] An air conditioning device, such as a Nest Thermostat, to adjust the ambient temperature as needed.

[0383] 7. Lighting System

[0384] A lighting system for adjusting ambient light, compatible with smart lights such as Phillips Hue.

[0385] 8. Notification terminal

[0386] A device that sends notifications when a customer's condition changes or an abnormality occurs. A standard smartphone can be used.

[0387] Program processing overview

[0388] The server acquires real-time customer behavior data from the camera device, environmental and customer body temperature data from the temperature sensor device, and customer voice data from the audio device. Using these data, the emotion engine analyzes the customer's emotional state.

[0389] The acquired data is analyzed by a server to determine the current situation. The music generation AI generates relaxing music according to the customer's emotional state and plays it through the store's speakers. In addition, the air conditioning and lighting systems adjust the environment as needed.

[0390] Specific examples

[0391] For example, if a customer is feeling stressed, that information is collected from the camera and audio device. The emotion engine determines the emotional state as "stressed," and the music generation AI generates relaxing music. This music is played through speakers in the store, and the lighting is adjusted to a softer light. At the same time, the air conditioning is adjusted to an appropriate temperature.

[0392] Prompt Sentence Examples

[0393] "Generate appropriate relaxing music when the customer is feeling stressed. The text data corresponding to the customer's emotional state is as follows: 'Stress'"

[0394] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0395] Step 1:

[0396] The server acquires customer behavior data in real time from the camera device. The acquired video data (input) is used to analyze the customer's location, behavior patterns, etc., and the results (output) become preprocessed data for emotion recognition.

[0397] Step 2:

[0398] The server acquires temperature data of the environment and the customer from the temperature sensor device. The acquired temperature data (input) is analyzed by the server to determine the appropriate range of the environment temperature and the customer's body temperature (output).

[0399] Step 3:

[0400] The server captures the customer's voice from the audio device, and after noise filtering and voice analysis, the captured voice data (input) is used to detect the customer's tone of voice and emotional state (output).

[0401] Step 4:

[0402] The server uses an emotion engine to analyze data acquired from the camera device and audio device to determine the customer's emotional state. Based on the input data (video and audio data), the emotional state (e.g., stress, relaxation) is detected from the customer's facial expression, posture, and tone of voice, and the emotional state (output) is obtained.

[0403] Step 5:

[0404] Based on the output of the emotion engine, the server sends a prompt to the music generation AI. An example of a prompt is, "Please generate appropriate relaxing music for when the customer is feeling stressed." Based on the input prompt, the music generation AI generates (outputs) music with a high relaxing effect, and the music data is obtained.

[0405] Step 6:

[0406] The server then transmits the generated music data to speakers in the store, where it is played, providing an acoustic environment that responds to the customer's emotional state.

[0407] Step 7:

[0408] The server controls the air conditioner based on the environmental data. Based on the acquired temperature data (input), it sets the appropriate room temperature (output) and sends temperature adjustment instructions to the air conditioner. This maintains the optimal ambient temperature.

[0409] Step 8:

[0410] The server controls the lighting system to adjust the ambient light. Based on the output data (input) of the emotion engine, it sets the appropriate lighting settings (output) and sends dimming instructions to the lighting system. This provides a comfortable lighting environment.

[0411] Step 9:

[0412] If the server detects an abnormality, it immediately sends an alert to the notification terminal. Based on the abnormal data (input), it generates an appropriate warning message (output) and sends it to the notification terminal. This enables a quick response.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0416] [Second embodiment]

[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0429] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0430] System Configuration

[0431] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0432] What the program does

[0433] 1. Data Collection

[0434] The server collects real-time baby movement data from the camera device, which can determine whether the baby is moving or quiet.

[0435] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[0436] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0437] 2. Data analysis and feedback

[0438] The server analyzes the collected data and evaluates the baby's condition (e.g., excited, relaxed, normal body temperature, high temperature). For example, if the baby continues to cry, it is determined to be excited.

[0439] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0440] 3. Lullaby Generation and Playback

[0441] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited, a lullaby with a slow rhythm will be generated.

[0442] The device (e.g., a smartphone or tablet) plays the generated lullaby through a speaker in the baby's room, providing a sound environment that suits the baby's condition.

[0443] 4. Environmental adjustment

[0444] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the room.

[0445] It receives feedback from the temperature sensor and continues to adjust until it reaches the target temperature.

[0446] 5. Notifications and Anomaly Detection

[0447] The server checks the baby's sleep status and, if it is confirmed that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, allowing the parent to know that the baby has fallen asleep safely.

[0448] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends an alert to the parent's communication device, allowing the parent to respond promptly.

[0449] Specific examples

[0450] Example 1: If a baby continues to cry, the camera and audio devices will be used to check the baby's condition, and the music generation AI will create a soothing lullaby. The lullaby will be played through the speakers, and the air conditioner will start cooling the room. As a result, the baby will gradually calm down and fall asleep.

[0451] Example 2: If the baby's body temperature is high and the room temperature is also high, the air conditioner will be controlled based on data from the temperature sensor device and the cooling will start. At the same time, a gentle lullaby will be generated and played to help the baby relax. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0452] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. It also provides a safe and secure environment for parents, reducing the burden of childcare.

[0453] The processing flow will be explained below.

[0454] Step 1:

[0455] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[0456] Step 2:

[0457] The server receives the room temperature data and the baby's temperature data from the temperature sensor device. Specifically, it records the baby's temperature and the room's ambient temperature based on the numerical data obtained from the temperature sensor.

[0458] Step 3:

[0459] The server receives the baby's crying and other sounds in real time from the audio device, and analyzes the audio input to determine whether the baby is crying or quiet.

[0460] Step 4:

[0461] The server integrates and analyzes the data from steps 1 to 3 to evaluate the baby's current state (excited, relaxed, normal body temperature, high temperature, etc.) For example, if the baby is crying continuously, has a high body temperature, and is active, it is determined to be in an excited state.

[0462] Step 5:

[0463] The server uses music generation AI to generate the optimal lullaby based on the analysis results. Specifically, it creates lullabies with slow rhythms or gentle melodies depending on the results of data analysis.

[0464] Step 6:

[0465] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data is transmitted to the speaker using Wi-Fi, and the audio is played back.

[0466] Step 7:

[0467] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it automatically adjusts the air conditioner's cooling or heating mode based on the temperature data and sets the target temperature.

[0468] Step 8:

[0469] The server receives feedback from the temperature sensor and continues to appropriately control the air conditioner until the target temperature is reached. For example, it may continue cooling until the target temperature is reached, and then stop the air conditioner when the temperature is appropriate.

[0470] Step 9:

[0471] The server confirms that the baby is asleep, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[0472] Step 10:

[0473] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[0474] This allows the system to monitor the baby's condition in real time and respond appropriately to the situation, reducing the burden on parents and providing a comfortable sleeping environment for the baby.

[0475] Example 1

[0476] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0477] Conventional baby monitoring systems tend to rely on specific devices, making it difficult to comprehensively monitor a baby's condition and provide appropriate feedback. As a result, parents are unable to consistently monitor their baby's condition, placing a heavy burden on childcare. Furthermore, due to insufficient environmental adjustment and anomaly detection functions, it is difficult to provide a sustainable, comfortable environment for babies. To solve these problems, a system is needed that can comprehensively monitor a baby's condition in real time and automatically take appropriate action.

[0478] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0479] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for analyzing the acquired data and evaluating the baby's condition, means for generating a lullaby according to the baby's condition using a music generation algorithm, means for playing the generated lullaby through a playback device, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if something abnormal occurs with the baby. This makes it possible to monitor the baby's condition from various angles and automatically take appropriate measures.

[0480] A "camera device" is a device that captures a baby's movements and acquires the video data in real time.

[0481] A "temperature sensor device" is a device that measures the baby's surrounding environment and the baby's own body temperature and provides that data to a server.

[0482] An "audio device" is a device that collects the baby's crying and other surrounding sounds and sends the data to a server.

[0483] The "server" is a central computer system that integrates and analyzes data collected from various devices and generates appropriate feedback.

[0484] The "means for analyzing the acquired data and assessing the baby's condition" refers to algorithms or software for estimating the baby's current condition based on the collected movement data, temperature data, and audio data.

[0485] A "music generation algorithm" is an AI (artificial intelligence) model or program that generates the optimal lullaby depending on the baby's condition.

[0486] A "playback device" is a speaker or audio device for playing back the generated lullaby sent from the server.

[0487] "Air conditioning equipment" refers to equipment with heating and cooling functions for adjusting the temperature in a room, and includes air conditioners.

[0488] A "parent communication device" is a mobile device, such as a smartphone or tablet, used by a parent to receive notifications and alerts.

[0489] "Means for sending an alert in the event of an abnormality" refers to software or hardware configurations that immediately send an alert message to the parent's communication device when an abnormality in the baby is detected from the collected and analyzed data.

[0490] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0491] System Configuration

[0492] This system is composed of multiple hardware and software devices, including the following:

[0493] Camera device: A device for capturing baby's movements in real time, such as an infrared camera or a high-resolution camera.

[0494] Temperature sensor device: A sensor device that measures the temperature of the room and the baby. A high-precision digital thermometer is used.

[0495] Audio device: A microphone device that collects the baby's cries and surrounding sounds.

[0496] Server: A central computer system that analyzes data collected from each device and generates the necessary feedback. A server equipped with a high-performance processor is used here.

[0497] Device: A communication device used by the end user (parent). For example, a smartphone or tablet.

[0498] Data collection

[0499] The server collects the baby's movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0500] The server receives real-time room temperature and baby temperature data from the temperature sensor device, which can determine whether the baby is within a comfortable temperature range.

[0501] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0502] Data analysis and feedback

[0503] The server then aggregates and analyzes the collected data, for example analyzing the audio data to determine whether the baby is crying, using a voice analysis algorithm.

[0504] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0505] Lullaby generation and playback

[0506] The server uses a music generation algorithm to generate an appropriate lullaby depending on the baby's state. For example, if the baby is excited, a lullaby with a calming rhythm will be generated.

[0507] The device receives the generated lullaby and plays it through a speaker in the room, which may be connected to the smartphone via Wi-Fi.

[0508] environmental adjustment

[0509] The server controls the air conditioning unit as needed to adjust the temperature to keep the baby in a comfortable environment. For example, if the room temperature is high, the server activates the air conditioning's cooling function.

[0510] It receives feedback from temperature sensor devices and adjusts the air conditioner until the target temperature is reached.

[0511] Notifications and Anomaly Detection

[0512] If the server determines that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, for example, a message saying "Baby has fallen asleep."

[0513] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends a warning to the parent's communication device, such as a message saying, "Your baby's temperature is high. Please check it."

[0514] Specific examples

[0515] Example 1:

[0516] If the baby continues to cry, the server will check the baby's condition from the camera and audio data and use a music generation algorithm to generate a soothing lullaby. The lullaby will be played through the speaker, and the air conditioner will start cooling to adjust the room temperature. As a result, the baby will gradually calm down and fall asleep.

[0517] Example 2:

[0518] If the baby's body temperature is high and the room temperature is also high, the server will control the air conditioning based on the data from the temperature sensor device and turn on the air conditioner. At the same time, a gentle lullaby to help the baby relax will be generated and played. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0519] Prompt Sentence Examples

[0520] "Monitor your baby's crying and play a soothing lullaby."

[0521] "While controlling the air conditioning unit if the room temperature is high, it also checks the baby's temperature and generates an appropriate lullaby."

[0522] In this way, the present invention monitors the baby's condition from multiple angles and responds appropriately to provide a comfortable sleeping environment for the baby. It also provides a safe and secure childcare environment for parents, reducing the burden of childcare.

[0523] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0524] Step 1:

[0525] Data collection

[0526] The server collects baby's movement data in real time from the camera device. The input is the video data captured by the camera. The server analyzes this video data frame by frame and extracts the baby's movements (e.g., whether or not the baby is moving and the speed of the movements). The output is the baby's movement pattern data.

[0527] The server obtains the room temperature and the baby's body temperature data from the temperature sensor device. The temperature data measured by the sensor at regular intervals is the input. The server receives and analyzes this data to determine whether the room is within the appropriate temperature range or whether the baby's body temperature is abnormal. The output is the environmental temperature data and the baby's body temperature data.

[0528] The server collects the baby's crying and other ambient sounds from the audio device. The audio data captured by the audio device's microphone is the input. The server uses an audio analysis algorithm to analyze the characteristics of the audio (e.g., volume, frequency) and determine whether the baby is crying. The output is the audio analysis result.

[0529] Step 2:

[0530] Data analysis and feedback

[0531] The server integrates and analyzes all the data collected in step 1 (movement pattern data, ambient temperature data, body temperature data, and voice analysis results). Based on these analysis results, the baby's condition is evaluated. For example, if the baby continues to cry, it is judged to be in an "excited state" based on the voice analysis results. The inputs are each sensor data and its analysis results. The output is the evaluation result of the baby's condition.

[0532] Step 3:

[0533] Lullaby generation and playback

[0534] The server uses a music generation algorithm to generate an appropriate lullaby based on the baby's state evaluated in step 2. For example, the prompt sentence is "Generate a calming song." The music generation AI model generates a lullaby based on this prompt sentence, and the generated lullaby data is obtained as the output.

[0535] The device receives the lullaby data sent from the server and plays it through the room's speakers. The input is the lullaby data from the server. The device sends it to a playback device, and the sound is played as the output.

[0536] Step 4:

[0537] environmental adjustment

[0538] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. If the temperature sensor device data indicates that the room is hot, the server issues a command to the air conditioner to operate the cooling function. The operating status of the air conditioner is obtained as an output.

[0539] The server continues to receive feedback from the temperature sensor and continues to control the air conditioning until the target temperature is reached. The input is the continuous temperature feedback data. The output is the final adjusted temperature data.

[0540] Step 5:

[0541] Notifications and Anomaly Detection

[0542] When the server confirms that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." The input includes the analysis results of motion data and voice data. The output is a notification message sent to the parent's communication device.

[0543] The server immediately sends an alert to the parent's communication device if the baby experiences any abnormalities (e.g., high temperature, respiratory failure, etc.). The input is temperature data and other sensor data indicating an abnormality. The output is a warning message sent to the parent's communication device.

[0544] (Application example 1)

[0545] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0546] In conventional factory environments, monitoring of workers' movements, physical condition, and working environment is insufficient, resulting in reduced work efficiency and health risks for workers. In particular, working in the same position for long periods of time and working in high-temperature environments can cause fatigue and health problems for workers. There is also a need for a system that can respond immediately when a worker becomes ill.

[0547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0548] In this invention, the server includes means for acquiring the target person's movements from a camera device, means for acquiring environmental and target person's body temperature data from a temperature sensor device, means for acquiring the target person's voice from an audio device, means for analyzing the acquired data and determining the target person's condition, means for generating voice instructions according to the target person's condition using a music generation AI, means for playing the generated voice instructions, means for controlling the air conditioner as necessary to adjust the temperature so that the target person is in a comfortable environment, means for sending a notification to a manager's communication terminal when the target person enters a specific state, and means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person. This enables real-time monitoring of worker movements and physical condition, enabling optimization of the work environment and health management.

[0549] A "camera device" is a photographing device for capturing the movements of a target person in real time.

[0550] A "temperature sensor device" is a measuring device for acquiring body temperature data of the environment and a target person.

[0551] An "audio device" is a sound collection device for acquiring the voice of a target person.

[0552] "Means for analyzing acquired data and determining the status of the target person" refers to a method or apparatus for processing information acquired from the camera device, temperature sensor device, and audio device and assessing the current status of the target person.

[0553] "Means for generating voice instructions according to the state of a target person using music generation AI" refers to a method or device that uses artificial intelligence technology to generate voice instructions that are adapted to the state of a target person.

[0554] "Means for playing generated voice instructions" refers to a device or method for playing voice instructions generated by the music generation AI.

[0555] "Means for controlling an air conditioner and adjusting the temperature so that the target person is in a comfortable environment" refers to a method or device for operating an air conditioner to appropriately adjust the environmental temperature.

[0556] "Means for sending a notification to the manager's communication terminal when the target person enters a specific state" refers to a method or device for notifying the manager of this information when the target person reaches a specific state, such as a situation where a break is required.

[0557] "Means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person" refers to a method or device for immediately sending a warning to the manager when an abnormality occurs in the target person's health.

[0558] The present invention aims to provide a work support robot system that monitors the movements and environment of a target person (worker) and provides necessary instructions and adjusts the environment. Specific embodiments of the present invention will be described below.

[0559] System Configuration

[0560] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0561] What the program does

[0562] 1. Data Collection

[0563] The server collects real-time motion data of the target person from the camera device, which allows it to determine whether the target person is moving or stationary.

[0564] The temperature sensor device acquires the ambient temperature and the target person's body temperature data, which can then be used to determine whether the target person is within a comfortable temperature range.

[0565] Collecting the subject's voice from a voice device, which allows us to determine whether the subject understands the instructions.

[0566] 2. Data analysis and feedback

[0567] The server analyzes the collected data and evaluates the target person's status (e.g., working, needing a break, abnormal state). For example, if the target person remains motionless for a long time, it determines that the person needs a break.

[0568] If each piece of data is outside the specified range (e.g., room temperature is high, lighting is low), the system determines how to respond accordingly.

[0569] 3. Generation and playback of voice instructions

[0570] The server uses music generation AI to generate optimal voice instructions based on the analysis results. For example, if the person is tired, it will generate a voice instruction to encourage them to take a break.

[0571] The terminal (e.g., a smartphone) plays the generated voice instructions through an audio device that transmits the voice instructions to the target person, thereby providing appropriate instructions according to the target person's state.

[0572] 4. Environmental adjustment

[0573] The server controls the air conditioner as needed to keep the temperature in the factory within an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the factory.

[0574] 5. Notifications and Anomaly Detection

[0575] When the target person enters a specific state (e.g., needs a break), the server sends a notification to the administrator's communication terminal, allowing the administrator to grasp the target person's state.

[0576] Additionally, if an abnormality occurs with the target person (e.g., high fever, immobility, etc.), an alert is immediately sent to the administrator's communication terminal, allowing the administrator to respond promptly.

[0577] Hardware and software used

[0578] Camera devices: Used to monitor the behavior of subjects.

[0579] Temperature sensor device: Used to monitor environmental and body temperatures.

[0580] Audio device: Used to communicate instructions to the target person.

[0581] Server: Data processing and analysis.

[0582] Communication device: Used to receive notifications and alerts.

[0583] OpenCV: Software used to analyze camera footage.

[0584] pyttsx3: Software used to generate and play audio instructions.

[0585] GPIO Zero: A library used to acquire data from the temperature sensor.

[0586] Specific examples

[0587] Example 1: If a worker remains motionless for more than 30 minutes, the system will issue a voice message saying "Please take a break" and will also send a work status notification to the manager.

[0588] Example 2: If the temperature inside the factory goes outside the set range (e.g., above 28 degrees), the system automatically switches the air conditioner to cooling mode.

[0589] Example prompt for a generative AI model:

[0590] Similar to the system that provides voice alerts and adjusts the temperature based on a baby's movements and temperature data, create a program that monitors the movements and temperature data of factory workers, issues voice alerts when they become fatigued, and controls the air conditioner as necessary. Specific examples include notifications if a worker has not moved for 30 minutes and switching on the air conditioner if the temperature exceeds 28 degrees.

[0591] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0592] Step 1:

[0593] The server acquires the motion data of the target person from the camera device in real time. Specifically, the camera device captures the video of the target person and sends the video data to the server. The server analyzes the video data to determine whether the target person is moving. The input is the video data from the camera device, and the output is the motion status of the target person (e.g., moving, stationary).

[0594] Step 2:

[0595] The server receives the ambient temperature and the target person's body temperature data from the temperature sensor device. Specifically, the temperature sensor device measures the temperature data in real time and sends it to the server. The server analyzes the measured data to determine whether the ambient temperature and the target person's body temperature are within an appropriate range. The input is the temperature data from the temperature sensor device, and the output is the temperature status (e.g., normal, high temperature, low temperature).

[0596] Step 3:

[0597] The server acquires the target person's voice data from the audio device. Specifically, it picks up the target person's voice and other sounds and sends them to the server. The server analyzes the voice data to determine what the target person is saying or what sounds are being made. The input is the voice data from the audio device, and the output is the content of the voice (e.g., response to commands, background sounds).

[0598] Step 4:

[0599] The server comprehensively analyzes the collected data and determines the target person's condition. Specifically, it integrates motion data, temperature data, and voice data, and evaluates whether the target person is working, needs a break, or is in an abnormal state based on each piece of data. The input is motion data, temperature data, and voice data, and the output is the target person's overall condition (e.g., working, fatigue, abnormal).

[0600] Step 5:

[0601] The server uses music generation AI to generate optimal voice instructions based on the analysis results. Specifically, it generates appropriate instructions (e.g., "Please take a break" or "Please continue working") depending on the target person's condition. The input is the target person's overall condition, and the output is the generated voice instructions.

[0602] Step 6:

[0603] The server transmits the generated voice instructions to the target person through a terminal (e.g., a smartphone) or an audio device. Specifically, the generated voice instructions are played back as audio so that the target person can confirm the instructions. The input is the generated voice instructions, and the output is the played back audio.

[0604] Step 7:

[0605] The server controls the air conditioner as needed to adjust the temperature in the factory to an appropriate range. Specifically, if the temperature data exceeds the set range, the air conditioner will operate to cool or heat the factory. The input is the temperature state, and the output is the appropriate adjustment of the environmental temperature.

[0606] Step 8:

[0607] The server sends a notification to the manager's communication terminal when the target person enters a specific state. Specifically, for example, if it determines that the target person needs a break, it notifies the manager of that information. The input is the target person's overall state, and the output is a notification to the manager.

[0608] Step 9:

[0609] The server immediately sends an alert to the administrator's communication terminal if an abnormality occurs in the target person. Specifically, if it determines that the target person is in an abnormal condition, such as having a high fever or not moving, it sends an emergency alert to the administrator. The input is the target person's abnormal condition, and the output is an emergency alert.

[0610] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0611] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[0612] System Configuration

[0613] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[0614] What the program does

[0615] 1. Data Collection

[0616] The server receives real-time baby movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0617] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[0618] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0619] 2. Emotion recognition by emotion engine

[0620] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. For example, if the user is tired or stressed, the server obtains that information.

[0621] 3. Data analysis and feedback

[0622] The server integrates and analyzes the baby's data and the user's emotional data to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state.

[0623] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, the crying continues, or the user is feeling stressed), the system determines how to respond accordingly.

[0624] 4. Lullaby Generation and Playback

[0625] The server uses music generation AI to generate the optimal lullaby based on the baby's state and the user's emotions. For example, if the baby is excited and the user is tired, a particularly relaxing lullaby will be generated.

[0626] The device (e.g., a smartphone or tablet) then transmits the generated lullaby audio data to a speaker in the baby's room and plays it back, providing a sound environment tailored to the baby's condition.

[0627] 5. Environmental adjustment

[0628] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. For example, if the room temperature is high, the server will operate the air conditioner to cool it down.

[0629] The lighting in a room can also be adjusted based on the user's emotional data, for example, by adjusting the lighting in the room to a softer light to help the user relax.

[0630] 6. Notifications and Anomaly Detection

[0631] The server checks the baby's sleeping state, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[0632] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[0633] Specific examples

[0634] Example 1: If a baby continues to cry and the user is feeling stressed, the camera and audio devices will be used to check the baby's condition. The music generation AI will generate a relaxing lullaby and play it through the speaker. Furthermore, the air conditioner will start cooling and the room lighting will be adjusted to a softer light, creating a comfortable environment for the baby and the user.

[0635] Example 2: If the baby's body temperature is high and the user is tired at the same time, the air conditioner is controlled based on the data from the temperature sensor device and the analysis results of the emotion engine. At the same time, a gentle lullaby to help the baby relax is generated and played. When the baby falls asleep, a notification "The baby has fallen asleep" is sent to the parent's communication device.

[0636] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. In addition, by taking the user's emotions into consideration, it also provides a safe and secure environment for parents and reduces the burden of childcare.

[0637] The processing flow will be explained below.

[0638] Step 1:

[0639] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[0640] Step 2:

[0641] The server acquires the room temperature data and the baby's body temperature data from the temperature sensor device. Specifically, it records the baby's body temperature and the room ambient temperature based on the numerical data obtained from the temperature sensor.

[0642] Step 3:

[0643] The server receives the baby's crying and other sounds in real time from the audio device, analyzes the audio data, and determines whether the baby is crying or quiet.

[0644] Step 4:

[0645] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. Specifically, it obtains data on the user's emotional state, such as whether they are tired, stressed, or relaxed.

[0646] Step 5:

[0647] The server integrates and analyzes the data from steps 1 to 4 to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state. For example, if the baby continues to cry and the user is tired, the server determines that the baby is excited and the user is tired.

[0648] Step 6:

[0649] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited and the user is tired, it will generate a lullaby that has a particularly relaxing effect.

[0650] Step 7:

[0651] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data can be transmitted to the speaker using Wi-Fi, and the relaxing lullaby will be played.

[0652] Step 8:

[0653] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it sets the air conditioner to cooling or heating mode based on the temperature data and adjusts it to the target temperature.

[0654] Step 9:

[0655] The server then adjusts the lighting in the room appropriately based on the user's emotional data, for example, changing the lighting to softer light to help the user relax.

[0656] Step 10:

[0657] The server checks whether the baby is asleep, and if it is confirmed that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying "The baby has fallen asleep."

[0658] Step 11:

[0659] If something unusual happens to the baby, the server will immediately send an alert to the parent's communication device. For example, if the baby's temperature exceeds a predetermined range or breathing stops, an alert will be sent immediately saying "There is something wrong with the baby."

[0660] This allows the system to monitor the baby's condition from multiple angles and take optimal action while taking into consideration the user's emotions, thereby providing a comfortable sleeping environment for the baby and reducing the burden of childcare on the parents.

[0661] Example 2

[0662] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0663] In modern childcare settings, it is necessary to constantly monitor a baby's health and comfort and respond appropriately. However, parents are not always close to their babies, which increases the burden of childcare. In addition, it is necessary to respond by taking into account not only the baby's condition but also the parent's emotions and stress, but conventional systems do not adequately address this. Therefore, the challenge is to provide a comfortable environment for both babies and parents and reduce the burden of childcare on parents.

[0664] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0665] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for recognizing a user's emotion using an emotion engine, means for integrating and analyzing the baby's condition data and the user's emotion data, means for generating a lullaby according to the baby's condition using a music generation AI, means for playing the generated lullaby, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for adjusting the environment based on the user's emotion data, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if an abnormality occurs in the baby. This makes it possible to comprehensively monitor and analyze the conditions of both the baby and the parent and automatically take appropriate measures to provide an optimal environment for the baby and the parent and reduce the burden of childcare.

[0666] The "camera device" is a video capture device for monitoring and capturing the baby's movements in real time.

[0667] A "temperature sensor device" is a device for measuring and acquiring the temperature of the environment and the baby's body temperature.

[0668] "Audio Device" refers to a recording device for collecting and capturing a baby's cry or other sounds.

[0669] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize the user's emotional state.

[0670] "Music Generation AI" is an artificial intelligence system that generates optimal lullabies based on the baby's condition and the user's emotions.

[0671] An "air conditioner" is an air conditioning device that adjusts and maintains the temperature in a room.

[0672] "Communication devices" are communication devices such as smartphones and tablets used by parents.

[0673] "Lullaby generation" is the process in which the music generation AI creates a lullaby taking into account the baby's condition and the user's emotions.

[0674] "Playback" is the process of outputting the generated lullaby as sound through a speaker.

[0675] "Notifications" are messages sent to parents' communication devices to inform them of changes in the baby's condition or environment.

[0676] An "alert" is a warning message that sends an emergency notification to parents when there is an abnormality in the baby's health or safety.

[0677] "User emotion data" is data that indicates the user's emotional state analyzed by the emotion engine.

[0678] "Environmental adjustment" is the process of changing settings such as temperature and lighting to provide a comfortable environment for the baby and the user.

[0679] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[0680] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[0681] Data collection

[0682] The server receives real-time motion data from the camera device, including the baby's limb movements, rolling over, etc. For example, if the baby is waving its arms, the data is sent to the server.

[0683] The server receives the room temperature data and the baby's temperature data from the temperature sensor device, and the specific values ​​that the room temperature is 25 degrees and the baby's temperature is 37 degrees are sent to the server.

[0684] The server collects the baby's cry and other sounds from the audio device. For example, if the baby is crying loudly, the decibel level of the cry is sent to the server.

[0685] Recognizing user emotions with an emotion engine

[0686] The server uses the camera device to analyze the user's (parent's) facial expressions. For example, when the user's face is facing the camera, the server uses a facial expression recognition algorithm to identify emotions such as smile, anger, sadness, etc.

[0687] The server uses the audio device to analyze the tone of the user's voice. For example, if the user's voice is high-pitched and fast, it determines that the user is stressed. By combining these data, it determines whether the user is tired or not.

[0688] Data analysis and feedback

[0689] The server analyzes the collected data on the baby's movements, temperature, voice, and user's emotions. For example, if a baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, the server will determine that the baby may be crying because of the heat.

[0690] Lullaby generation and playback

[0691] The server uses a music generation AI model to generate the optimal lullaby based on the baby's state and the user's emotions. For example, the prompt text is "My baby is excited. Please generate a lullaby that has a relaxing effect."

[0692] The device receives the generated lullaby audio data and transmits it to a speaker in the baby's room, for example, a smartphone connected to the speaker via Bluetooth, which plays the lullaby.

[0693] environmental adjustment

[0694] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, the server sends a command to the air conditioner to set the temperature to 24 degrees.

[0695] The server can also adjust the lighting in the room based on the user's emotional data, for example adjusting the lighting to 30% brightness to help the user relax.

[0696] Notifications and Anomaly Detection

[0697] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's device saying "The baby has fallen asleep." Specifically, it sends a push notification to the parent's smartphone.

[0698] If something abnormal happens to the baby, the server will quickly send an alert to the parent's device. For example, if the baby's temperature exceeds 39 degrees, an alert will be sent saying "There is something abnormal with the baby."

[0699] Example prompts for generative AI models

[0700] "Generate a relaxing lullaby to play when your baby is crying."

[0701] "Right now, your baby is excited and the room temperature is high. Please generate a lullaby that will help your baby relax."

[0702] "When parents are tired, they need a lullaby to calm their baby. Please generate an appropriate lullaby."

[0703] As described above, this system comprehensively monitors the baby's condition and the user's emotions and takes appropriate action to support the baby's comfortable sleep and reduce the burden of childcare.

[0704] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0705] Step 1:

[0706] Data collection

[0707] The server obtains the baby's motion data in real time from the camera device. As input, it receives the video stream from the camera device and processes it with an image analysis algorithm. As output, it obtains motion data such as the baby's limb movements and rolling over. Specifically, the camera captures the baby's movements and sends the data to the server.

[0708] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device. As input, it receives and analyzes the temperature data from the temperature sensor device. As output, it obtains specific numerical data about the room temperature and the baby's body temperature. For example, data that the room temperature is 25 degrees and the baby's body temperature is 37 degrees is sent to the server.

[0709] The server collects the baby's cry and other sounds from the audio device. As input, it receives the audio stream from the audio device and processes it with an audio analysis algorithm. As output, it obtains data about the decibel level and type of sound of the baby's cry. Specifically, if the baby is crying hard, the decibel level of the cry is sent to the server.

[0710] Step 2:

[0711] Recognizing user emotions with an emotion engine

[0712] The server analyzes the user's facial expressions using a camera device. As input, it receives the video stream from the camera device and processes it with a facial expression recognition algorithm. As output, it obtains data about the user's emotional state (e.g., smiling, angry, sad, etc.). Specifically, when the user's face is facing the camera, the server analyzes their facial expressions.

[0713] The server analyzes the tone of the user's voice using the audio device. As input, it receives the audio stream from the audio device and processes it with an audio tone analysis algorithm. As output, it obtains data about the user's emotional state (e.g., stress, fatigue, etc.). Specifically, if the user's voice is high-pitched and fast, it is determined that the user is stressed.

[0714] Step 3:

[0715] Data analysis and feedback

[0716] The server integrates and analyzes the collected baby's movement data, temperature data, voice data, and user's emotional data. As input, it receives the data obtained in each step above and processes it with a data integration algorithm. As output, it obtains judgment data regarding the baby's current state (e.g., crying, warm, etc.) and the user's emotional state. In a specific operation, assuming that the baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, it will make the judgment that "the baby may be crying because of the heat."

[0717] Step 4:

[0718] Lullaby generation and playback

[0719] The server uses a music generation AI model to generate an optimal lullaby based on the baby's state and the user's emotions. The AI ​​model receives the prompt "My baby is excited. Please generate a lullaby that has a relaxing effect." The output is audio data of a lullaby with a relaxing effect. In concrete terms, the AI ​​generates a lullaby based on the prompt, and the server receives the data.

[0720] The device receives the generated lullaby audio data and sends it to the speaker in the baby's room. As input, it receives the audio data sent from the server and sends it to the speaker. As output, the lullaby is played from the speaker. Specifically, the smartphone is connected to the speaker via Bluetooth and the lullaby is played.

[0721] Step 5:

[0722] environmental adjustment

[0723] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. As input, it receives room temperature data and processes it with the air conditioner control algorithm. As output, it obtains commands to send to the air conditioner. In concrete terms, the server sends a command to the air conditioner to "set the temperature to 24 degrees."

[0724] The server can also adjust the lighting in a room based on the user's emotional data. It receives the user's emotional data as input and processes it with a lighting control algorithm. The output is a command to send to the lighting. Specifically, the server sends a command to adjust the lighting to 30% brightness to help the user relax.

[0725] Step 6:

[0726] Notifications and Anomaly Detection

[0727] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." As input, it receives and analyzes the baby's movement data. As output, it generates a notification message and sends it to the parent's device. Specifically, a push notification is sent to the smartphone, displaying the message, "The baby has fallen asleep."

[0728] If something abnormal occurs with the baby, the server quickly sends an alert to the parent's communication device. As input, it receives and analyzes the baby's temperature and breathing data. As output, it generates an alert message informing the parent of the abnormality and sends it to the parent's device. For example, if the body temperature exceeds 39 degrees, an alert saying "There is something abnormal with the baby" is sent to the smartphone.

[0729] (Application example 2)

[0730] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] In today's physical stores, customers often feel stressed, which causes a poor customer experience. It is also difficult for store staff to grasp customers' emotional state in real time and respond appropriately, making it a challenge to improve customer satisfaction. Furthermore, some stores do not adequately adjust the environment to maintain customer comfort.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0733] In this invention, the server includes means for acquiring motion from a camera device, means for acquiring environmental and body temperature data from a temperature sensor device, means for acquiring audio from an audio device, means for analyzing the acquired data and determining the state, means for generating music according to the state using a music generation AI, means for playing the generated music, means for controlling an air conditioner as needed to adjust the temperature to maintain a comfortable environment, means for sending a notification to a notification terminal when the state changes, means for immediately sending an alert to the notification terminal when an abnormality occurs, means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize emotions, means for adjusting the environment according to the user's emotional state based on data from the audio device, means for the music generation AI to generate music with a high relaxing effect based on the emotional data, and means for controlling lighting to adjust the ambient light, thereby enabling real-time environmental adjustment according to the customer's emotional state.

[0734] A "camera device" is an electronic device for capturing video and image data.

[0735] "Movement" refers to physical movement or action.

[0736] A "temperature sensor device" is an electronic device for measuring the temperature of an environment or object.

[0737] "Environmental data" refers to data that represents the state or condition of the surroundings, such as temperature and light intensity.

[0738] "Body temperature data" refers to data obtained by measuring and recording the temperature of a living body.

[0739] An "audio device" is an electronic device for capturing sound, such as a microphone.

[0740] A "means for determining the status" is a method or system that analyzes and evaluates the current situation or conditions based on acquired data.

[0741] "Music generation AI" is a system that automatically generates music using artificial intelligence technology.

[0742] An "air conditioning device" is a device that adjusts temperature and humidity. Air conditioners are examples of this.

[0743] A "notification terminal" is a device that receives signals and messages, such as a smartphone.

[0744] An "alert" is a warning message or notification of a particular condition.

[0745] The "emotion engine" is a system that analyzes emotions from facial expressions and tone of voice.

[0746] "User" means any person who uses a system or service.

[0747] "Environmental control" is the act of controlling environmental conditions such as temperature and light to maintain or improve comfort.

[0748] "Music with a high relaxing effect" is music that is intended to relax the listener.

[0749] A "means for controlling lighting" is a system or method for adjusting the intensity or color temperature of light.

[0750] This invention is a system that monitors the emotional state of customers in real time in physical stores and adjusts the environment accordingly to improve customer experience. This system consists of a camera device, a temperature sensor device, a voice device, an emotion engine, a music generation AI, an air conditioning system, a lighting system, and a notification terminal.

[0751] System Configuration

[0752] The system includes the following major components:

[0753] 1. Camera Device

[0754] The camera device is used to capture customer behavior in real time, specifically, a common webcam such as the Logitech C920 can be used.

[0755] 2. Temperature sensor device

[0756] Temperature sensor devices are used to obtain environmental and customer body temperature data. For example, DHT22 sensors are available.

[0757] 3. Audio Devices

[0758] The audio device is used to capture the customer's voice. A high-sensitivity microphone such as a Blue Yeti can be used.

[0759] 4. Emotion Engine

[0760] The emotion engine uses data from cameras and audio devices to analyze customers' facial expressions and tone of voice to recognize emotions, using an AI model powered by OpenCV and TensorFlow.

[0761] 5. Music Generation AI

[0762] The music generation AI generates relaxing music based on the customer's emotional state, using AI frameworks such as Google Magenta.

[0763] 6. Air conditioner

[0764] An air conditioning device, such as a Nest Thermostat, to adjust the ambient temperature as needed.

[0765] 7. Lighting System

[0766] A lighting system for adjusting ambient light, compatible with smart lights such as Phillips Hue.

[0767] 8. Notification terminal

[0768] A device that sends notifications when a customer's condition changes or an abnormality occurs. A standard smartphone can be used.

[0769] Program processing overview

[0770] The server acquires real-time customer behavior data from the camera device, environmental and customer body temperature data from the temperature sensor device, and customer voice data from the audio device. Using these data, the emotion engine analyzes the customer's emotional state.

[0771] The acquired data is analyzed by a server to determine the current situation. The music generation AI generates relaxing music according to the customer's emotional state and plays it through the store's speakers. In addition, the air conditioning and lighting systems adjust the environment as needed.

[0772] Specific examples

[0773] For example, if a customer is feeling stressed, that information is collected from the camera and audio device. The emotion engine determines the emotional state as "stressed," and the music generation AI generates relaxing music. This music is played through speakers in the store, and the lighting is adjusted to a softer light. At the same time, the air conditioning is adjusted to an appropriate temperature.

[0774] Prompt Sentence Examples

[0775] "Generate appropriate relaxing music when the customer is feeling stressed. The text data corresponding to the customer's emotional state is as follows: 'Stress'"

[0776] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0777] Step 1:

[0778] The server acquires customer behavior data in real time from the camera device. The acquired video data (input) is used to analyze the customer's location, behavior patterns, etc., and the results (output) become preprocessed data for emotion recognition.

[0779] Step 2:

[0780] The server acquires temperature data of the environment and the customer from the temperature sensor device. The acquired temperature data (input) is analyzed by the server to determine the appropriate range of the environment temperature and the customer's body temperature (output).

[0781] Step 3:

[0782] The server captures the customer's voice from the audio device, and after noise filtering and voice analysis, the captured voice data (input) is used to detect the customer's tone of voice and emotional state (output).

[0783] Step 4:

[0784] The server uses an emotion engine to analyze data acquired from the camera device and audio device to determine the customer's emotional state. Based on the input data (video and audio data), the emotional state (e.g., stress, relaxation) is detected from the customer's facial expression, posture, and tone of voice, and the emotional state (output) is obtained.

[0785] Step 5:

[0786] Based on the output of the emotion engine, the server sends a prompt to the music generation AI. An example of a prompt is, "Please generate appropriate relaxing music for when the customer is feeling stressed." Based on the input prompt, the music generation AI generates (outputs) music with a high relaxing effect, and the music data is obtained.

[0787] Step 6:

[0788] The server then transmits the generated music data to speakers in the store, where it is played, providing an acoustic environment that responds to the customer's emotional state.

[0789] Step 7:

[0790] The server controls the air conditioner based on the environmental data. Based on the acquired temperature data (input), it sets the appropriate room temperature (output) and sends temperature adjustment instructions to the air conditioner. This maintains the optimal ambient temperature.

[0791] Step 8:

[0792] The server controls the lighting system to adjust the ambient light. Based on the output data (input) of the emotion engine, it sets the appropriate lighting settings (output) and sends dimming instructions to the lighting system. This provides a comfortable lighting environment.

[0793] Step 9:

[0794] If the server detects an abnormality, it immediately sends an alert to the notification terminal. Based on the abnormal data (input), it generates an appropriate warning message (output) and sends it to the notification terminal. This enables a quick response.

[0795] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0796] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0797] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0798] [Third embodiment]

[0799] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0800] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0801] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0802] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0803] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0804] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0805] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0806] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0807] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0808] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0809] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0810] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0811] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0812] System Configuration

[0813] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0814] What the program does

[0815] 1. Data Collection

[0816] The server collects real-time baby movement data from the camera device, which can determine whether the baby is moving or quiet.

[0817] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[0818] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0819] 2. Data analysis and feedback

[0820] The server analyzes the collected data and evaluates the baby's condition (e.g., excited, relaxed, normal body temperature, high temperature). For example, if the baby continues to cry, it is determined to be excited.

[0821] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0822] 3. Lullaby Generation and Playback

[0823] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited, a lullaby with a slow rhythm will be generated.

[0824] The device (e.g., a smartphone or tablet) plays the generated lullaby through a speaker in the baby's room, providing a sound environment that suits the baby's condition.

[0825] 4. Environmental adjustment

[0826] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the room.

[0827] It receives feedback from the temperature sensor and continues to adjust until it reaches the target temperature.

[0828] 5. Notifications and Anomaly Detection

[0829] The server checks the baby's sleep status and, if it is confirmed that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, allowing the parent to know that the baby has fallen asleep safely.

[0830] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends an alert to the parent's communication device, allowing the parent to respond promptly.

[0831] Specific examples

[0832] Example 1: If a baby continues to cry, the camera and audio devices will be used to check the baby's condition, and the music generation AI will create a soothing lullaby. The lullaby will be played through the speakers, and the air conditioner will start cooling the room. As a result, the baby will gradually calm down and fall asleep.

[0833] Example 2: If the baby's body temperature is high and the room temperature is also high, the air conditioner will be controlled based on data from the temperature sensor device and the cooling will start. At the same time, a gentle lullaby will be generated and played to help the baby relax. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0834] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. It also provides a safe and secure environment for parents, reducing the burden of childcare.

[0835] The processing flow will be explained below.

[0836] Step 1:

[0837] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[0838] Step 2:

[0839] The server receives the room temperature data and the baby's temperature data from the temperature sensor device. Specifically, it records the baby's temperature and the room's ambient temperature based on the numerical data obtained from the temperature sensor.

[0840] Step 3:

[0841] The server receives the baby's crying and other sounds in real time from the audio device, and analyzes the audio input to determine whether the baby is crying or quiet.

[0842] Step 4:

[0843] The server integrates and analyzes the data from steps 1 to 3 to evaluate the baby's current state (excited, relaxed, normal body temperature, high temperature, etc.) For example, if the baby is crying continuously, has a high body temperature, and is active, it is determined to be in an excited state.

[0844] Step 5:

[0845] The server uses music generation AI to generate the optimal lullaby based on the analysis results. Specifically, it creates lullabies with slow rhythms or gentle melodies depending on the results of data analysis.

[0846] Step 6:

[0847] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data is transmitted to the speaker using Wi-Fi, and the audio is played back.

[0848] Step 7:

[0849] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it automatically adjusts the air conditioner's cooling or heating mode based on the temperature data and sets the target temperature.

[0850] Step 8:

[0851] The server receives feedback from the temperature sensor and continues to appropriately control the air conditioner until the target temperature is reached. For example, it may continue cooling until the target temperature is reached, and then stop the air conditioner when the temperature is appropriate.

[0852] Step 9:

[0853] The server confirms that the baby is asleep, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[0854] Step 10:

[0855] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[0856] This allows the system to monitor the baby's condition in real time and respond appropriately to the situation, reducing the burden on parents and providing a comfortable sleeping environment for the baby.

[0857] Example 1

[0858] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0859] Conventional baby monitoring systems tend to rely on specific devices, making it difficult to comprehensively monitor a baby's condition and provide appropriate feedback. As a result, parents are unable to consistently monitor their baby's condition, placing a heavy burden on childcare. Furthermore, due to insufficient environmental adjustment and anomaly detection functions, it is difficult to provide a sustainable, comfortable environment for babies. To solve these problems, a system is needed that can comprehensively monitor a baby's condition in real time and automatically take appropriate action.

[0860] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0861] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for analyzing the acquired data and evaluating the baby's condition, means for generating a lullaby according to the baby's condition using a music generation algorithm, means for playing the generated lullaby through a playback device, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if something abnormal occurs with the baby. This makes it possible to monitor the baby's condition from various angles and automatically take appropriate measures.

[0862] A "camera device" is a device that captures a baby's movements and acquires the video data in real time.

[0863] A "temperature sensor device" is a device that measures the baby's surrounding environment and the baby's own body temperature and provides that data to a server.

[0864] An "audio device" is a device that collects the baby's crying and other surrounding sounds and sends the data to a server.

[0865] The "server" is a central computer system that integrates and analyzes data collected from various devices and generates appropriate feedback.

[0866] The "means for analyzing the acquired data and assessing the baby's condition" refers to algorithms or software for estimating the baby's current condition based on the collected movement data, temperature data, and audio data.

[0867] A "music generation algorithm" is an AI (artificial intelligence) model or program that generates the optimal lullaby depending on the baby's condition.

[0868] A "playback device" is a speaker or audio device for playing back the generated lullaby sent from the server.

[0869] "Air conditioning equipment" refers to equipment with heating and cooling functions for adjusting the temperature in a room, and includes air conditioners.

[0870] A "parent communication device" is a mobile device, such as a smartphone or tablet, used by a parent to receive notifications and alerts.

[0871] "Means for sending an alert in the event of an abnormality" refers to software or hardware configurations that immediately send an alert message to the parent's communication device when an abnormality in the baby is detected from the collected and analyzed data.

[0872] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[0873] System Configuration

[0874] This system is composed of multiple hardware and software devices, including the following:

[0875] Camera device: A device for capturing baby's movements in real time, such as an infrared camera or a high-resolution camera.

[0876] Temperature sensor device: A sensor device that measures the temperature of the room and the baby. A high-precision digital thermometer is used.

[0877] Audio device: A microphone device that collects the baby's cries and surrounding sounds.

[0878] Server: A central computer system that analyzes data collected from each device and generates the necessary feedback. A server equipped with a high-performance processor is used here.

[0879] Device: A communication device used by the end user (parent). For example, a smartphone or tablet.

[0880] Data collection

[0881] The server collects the baby's movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0882] The server receives real-time room temperature and baby temperature data from the temperature sensor device, which can determine whether the baby is within a comfortable temperature range.

[0883] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[0884] Data analysis and feedback

[0885] The server then aggregates and analyzes the collected data, for example analyzing the audio data to determine whether the baby is crying, using a voice analysis algorithm.

[0886] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[0887] Lullaby generation and playback

[0888] The server uses a music generation algorithm to generate an appropriate lullaby depending on the baby's state. For example, if the baby is excited, a lullaby with a calming rhythm will be generated.

[0889] The device receives the generated lullaby and plays it through a speaker in the room, which may be connected to the smartphone via Wi-Fi.

[0890] environmental adjustment

[0891] The server controls the air conditioning unit as needed to adjust the temperature to keep the baby in a comfortable environment. For example, if the room temperature is high, the server activates the air conditioning's cooling function.

[0892] It receives feedback from temperature sensor devices and adjusts the air conditioner until the target temperature is reached.

[0893] Notifications and Anomaly Detection

[0894] If the server determines that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, for example, a message saying "Baby has fallen asleep."

[0895] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends a warning to the parent's communication device, such as a message saying, "Your baby's temperature is high. Please check it."

[0896] Specific examples

[0897] Example 1:

[0898] If the baby continues to cry, the server will check the baby's condition from the camera and audio data and use a music generation algorithm to generate a soothing lullaby. The lullaby will be played through the speaker, and the air conditioner will start cooling to adjust the room temperature. As a result, the baby will gradually calm down and fall asleep.

[0899] Example 2:

[0900] If the baby's body temperature is high and the room temperature is also high, the server will control the air conditioning based on the data from the temperature sensor device and turn on the air conditioner. At the same time, a gentle lullaby to help the baby relax will be generated and played. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[0901] Prompt Sentence Examples

[0902] "Monitor your baby's crying and play a soothing lullaby."

[0903] "While controlling the air conditioning unit if the room temperature is high, it also checks the baby's temperature and generates an appropriate lullaby."

[0904] In this way, the present invention monitors the baby's condition from multiple angles and responds appropriately to provide a comfortable sleeping environment for the baby. It also provides a safe and secure childcare environment for parents, reducing the burden of childcare.

[0905] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0906] Step 1:

[0907] Data collection

[0908] The server collects baby's movement data in real time from the camera device. The input is the video data captured by the camera. The server analyzes this video data frame by frame and extracts the baby's movements (e.g., whether or not the baby is moving and the speed of the movements). The output is the baby's movement pattern data.

[0909] The server obtains the room temperature and the baby's body temperature data from the temperature sensor device. The temperature data measured by the sensor at regular intervals is the input. The server receives and analyzes this data to determine whether the room is within the appropriate temperature range or whether the baby's body temperature is abnormal. The output is the environmental temperature data and the baby's body temperature data.

[0910] The server collects the baby's crying and other ambient sounds from the audio device. The audio data captured by the audio device's microphone is the input. The server uses an audio analysis algorithm to analyze the characteristics of the audio (e.g., volume, frequency) and determine whether the baby is crying. The output is the audio analysis result.

[0911] Step 2:

[0912] Data analysis and feedback

[0913] The server integrates and analyzes all the data collected in step 1 (movement pattern data, ambient temperature data, body temperature data, and voice analysis results). Based on these analysis results, the baby's condition is evaluated. For example, if the baby continues to cry, it is judged to be in an "excited state" based on the voice analysis results. The inputs are each sensor data and its analysis results. The output is the evaluation result of the baby's condition.

[0914] Step 3:

[0915] Lullaby generation and playback

[0916] The server uses a music generation algorithm to generate an appropriate lullaby based on the baby's state evaluated in step 2. For example, the prompt sentence is "Generate a calming song." The music generation AI model generates a lullaby based on this prompt sentence, and the generated lullaby data is obtained as the output.

[0917] The device receives the lullaby data sent from the server and plays it through the room's speakers. The input is the lullaby data from the server. The device sends it to a playback device, and the sound is played as the output.

[0918] Step 4:

[0919] environmental adjustment

[0920] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. If the temperature sensor device data indicates that the room is hot, the server issues a command to the air conditioner to operate the cooling function. The operating status of the air conditioner is obtained as an output.

[0921] The server continues to receive feedback from the temperature sensor and continues to control the air conditioning until the target temperature is reached. The input is the continuous temperature feedback data. The output is the final adjusted temperature data.

[0922] Step 5:

[0923] Notifications and Anomaly Detection

[0924] When the server confirms that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." The input includes the analysis results of motion data and voice data. The output is a notification message sent to the parent's communication device.

[0925] The server immediately sends an alert to the parent's communication device if the baby experiences any abnormalities (e.g., high temperature, respiratory failure, etc.). The input is temperature data and other sensor data indicating an abnormality. The output is a warning message sent to the parent's communication device.

[0926] (Application example 1)

[0927] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0928] In conventional factory environments, monitoring of workers' movements, physical condition, and working environment is insufficient, resulting in reduced work efficiency and health risks for workers. In particular, working in the same position for long periods of time and working in high-temperature environments can cause fatigue and health problems for workers. There is also a need for a system that can respond immediately when a worker becomes ill.

[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0930] In this invention, the server includes means for acquiring the target person's movements from a camera device, means for acquiring environmental and target person's body temperature data from a temperature sensor device, means for acquiring the target person's voice from an audio device, means for analyzing the acquired data and determining the target person's condition, means for generating voice instructions according to the target person's condition using a music generation AI, means for playing the generated voice instructions, means for controlling the air conditioner as necessary to adjust the temperature so that the target person is in a comfortable environment, means for sending a notification to a manager's communication terminal when the target person enters a specific state, and means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person. This enables real-time monitoring of worker movements and physical condition, enabling optimization of the work environment and health management.

[0931] A "camera device" is a photographing device for capturing the movements of a target person in real time.

[0932] A "temperature sensor device" is a measuring device for acquiring body temperature data of the environment and a target person.

[0933] An "audio device" is a sound collection device for acquiring the voice of a target person.

[0934] "Means for analyzing acquired data and determining the status of the target person" refers to a method or apparatus for processing information acquired from the camera device, temperature sensor device, and audio device and assessing the current status of the target person.

[0935] "Means for generating voice instructions according to the state of a target person using music generation AI" refers to a method or device that uses artificial intelligence technology to generate voice instructions that are adapted to the state of a target person.

[0936] "Means for playing generated voice instructions" refers to a device or method for playing voice instructions generated by the music generation AI.

[0937] "Means for controlling an air conditioner and adjusting the temperature so that the target person is in a comfortable environment" refers to a method or device for operating an air conditioner to appropriately adjust the environmental temperature.

[0938] "Means for sending a notification to the manager's communication terminal when the target person enters a specific state" refers to a method or device for notifying the manager of this information when the target person reaches a specific state, such as a situation where a break is required.

[0939] "Means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person" refers to a method or device for immediately sending a warning to the manager when an abnormality occurs in the target person's health.

[0940] The present invention aims to provide a work support robot system that monitors the movements and environment of a target person (worker) and provides necessary instructions and adjusts the environment. Specific embodiments of the present invention will be described below.

[0941] System Configuration

[0942] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[0943] What the program does

[0944] 1. Data Collection

[0945] The server collects real-time motion data of the target person from the camera device, which allows it to determine whether the target person is moving or stationary.

[0946] The temperature sensor device acquires the ambient temperature and the target person's body temperature data, which can then be used to determine whether the target person is within a comfortable temperature range.

[0947] Collecting the subject's voice from a voice device, which allows us to determine whether the subject understands the instructions.

[0948] 2. Data analysis and feedback

[0949] The server analyzes the collected data and evaluates the target person's status (e.g., working, needing a break, abnormal state). For example, if the target person remains motionless for a long time, it determines that the person needs a break.

[0950] If each piece of data is outside the specified range (e.g., room temperature is high, lighting is low), the system determines how to respond accordingly.

[0951] 3. Generation and playback of voice instructions

[0952] The server uses music generation AI to generate optimal voice instructions based on the analysis results. For example, if the person is tired, it will generate a voice instruction to encourage them to take a break.

[0953] The terminal (e.g., a smartphone) plays the generated voice instructions through an audio device that transmits the voice instructions to the target person, thereby providing appropriate instructions according to the target person's state.

[0954] 4. Environmental adjustment

[0955] The server controls the air conditioner as needed to keep the temperature in the factory within an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the factory.

[0956] 5. Notifications and Anomaly Detection

[0957] When the target person enters a specific state (e.g., needs a break), the server sends a notification to the administrator's communication terminal, allowing the administrator to grasp the target person's state.

[0958] Additionally, if an abnormality occurs with the target person (e.g., high fever, immobility, etc.), an alert is immediately sent to the administrator's communication terminal, allowing the administrator to respond promptly.

[0959] Hardware and software used

[0960] Camera devices: Used to monitor the behavior of subjects.

[0961] Temperature sensor device: Used to monitor environmental and body temperatures.

[0962] Audio device: Used to communicate instructions to the target person.

[0963] Server: Data processing and analysis.

[0964] Communication device: Used to receive notifications and alerts.

[0965] OpenCV: Software used to analyze camera footage.

[0966] pyttsx3: Software used to generate and play audio instructions.

[0967] GPIO Zero: A library used to acquire data from the temperature sensor.

[0968] Specific examples

[0969] Example 1: If a worker remains motionless for more than 30 minutes, the system will issue a voice message saying "Please take a break" and will also send a work status notification to the manager.

[0970] Example 2: If the temperature inside the factory goes outside the set range (e.g., above 28 degrees), the system automatically switches the air conditioner to cooling mode.

[0971] Example prompt for a generative AI model:

[0972] Similar to the system that provides voice alerts and adjusts the temperature based on a baby's movements and temperature data, create a program that monitors the movements and temperature data of factory workers, issues voice alerts when they become fatigued, and controls the air conditioner as necessary. Specific examples include notifications if a worker has not moved for 30 minutes and switching on the air conditioner if the temperature exceeds 28 degrees.

[0973] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0974] Step 1:

[0975] The server acquires the motion data of the target person from the camera device in real time. Specifically, the camera device captures the video of the target person and sends the video data to the server. The server analyzes the video data to determine whether the target person is moving. The input is the video data from the camera device, and the output is the motion status of the target person (e.g., moving, stationary).

[0976] Step 2:

[0977] The server receives the ambient temperature and the target person's body temperature data from the temperature sensor device. Specifically, the temperature sensor device measures the temperature data in real time and sends it to the server. The server analyzes the measured data to determine whether the ambient temperature and the target person's body temperature are within an appropriate range. The input is the temperature data from the temperature sensor device, and the output is the temperature status (e.g., normal, high temperature, low temperature).

[0978] Step 3:

[0979] The server acquires the target person's voice data from the audio device. Specifically, it picks up the target person's voice and other sounds and sends them to the server. The server analyzes the voice data to determine what the target person is saying or what sounds are being made. The input is the voice data from the audio device, and the output is the content of the voice (e.g., response to commands, background sounds).

[0980] Step 4:

[0981] The server comprehensively analyzes the collected data and determines the target person's condition. Specifically, it integrates motion data, temperature data, and voice data, and evaluates whether the target person is working, needs a break, or is in an abnormal state based on each piece of data. The input is motion data, temperature data, and voice data, and the output is the target person's overall condition (e.g., working, fatigue, abnormal).

[0982] Step 5:

[0983] The server uses music generation AI to generate optimal voice instructions based on the analysis results. Specifically, it generates appropriate instructions (e.g., "Please take a break" or "Please continue working") depending on the target person's condition. The input is the target person's overall condition, and the output is the generated voice instructions.

[0984] Step 6:

[0985] The server transmits the generated voice instructions to the target person through a terminal (e.g., a smartphone) or an audio device. Specifically, the generated voice instructions are played back as audio so that the target person can confirm the instructions. The input is the generated voice instructions, and the output is the played back audio.

[0986] Step 7:

[0987] The server controls the air conditioner as needed to adjust the temperature in the factory to an appropriate range. Specifically, if the temperature data exceeds the set range, the air conditioner will operate to cool or heat the factory. The input is the temperature state, and the output is the appropriate adjustment of the environmental temperature.

[0988] Step 8:

[0989] The server sends a notification to the manager's communication terminal when the target person enters a specific state. Specifically, for example, if it determines that the target person needs a break, it notifies the manager of that information. The input is the target person's overall state, and the output is a notification to the manager.

[0990] Step 9:

[0991] The server immediately sends an alert to the administrator's communication terminal if an abnormality occurs in the target person. Specifically, if it determines that the target person is in an abnormal condition, such as having a high fever or not moving, it sends an emergency alert to the administrator. The input is the target person's abnormal condition, and the output is an emergency alert.

[0992] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0993] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[0994] System Configuration

[0995] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[0996] What the program does

[0997] 1. Data Collection

[0998] The server receives real-time baby movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[0999] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[1000] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[1001] 2. Emotion recognition by emotion engine

[1002] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. For example, if the user is tired or stressed, the server obtains that information.

[1003] 3. Data analysis and feedback

[1004] The server integrates and analyzes the baby's data and the user's emotional data to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state.

[1005] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, the crying continues, or the user is feeling stressed), the system determines how to respond accordingly.

[1006] 4. Lullaby Generation and Playback

[1007] The server uses music generation AI to generate the optimal lullaby based on the baby's state and the user's emotions. For example, if the baby is excited and the user is tired, a particularly relaxing lullaby will be generated.

[1008] The device (e.g., a smartphone or tablet) then transmits the generated lullaby audio data to a speaker in the baby's room and plays it back, providing a sound environment tailored to the baby's condition.

[1009] 5. Environmental adjustment

[1010] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. For example, if the room temperature is high, the server will operate the air conditioner to cool it down.

[1011] The lighting in a room can also be adjusted based on the user's emotional data, for example, by adjusting the lighting in the room to a softer light to help the user relax.

[1012] 6. Notifications and Anomaly Detection

[1013] The server checks the baby's sleeping state, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[1014] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[1015] Specific examples

[1016] Example 1: If a baby continues to cry and the user is feeling stressed, the camera and audio devices will be used to check the baby's condition. The music generation AI will generate a relaxing lullaby and play it through the speaker. Furthermore, the air conditioner will start cooling and the room lighting will be adjusted to a softer light, creating a comfortable environment for the baby and the user.

[1017] Example 2: If the baby's body temperature is high and the user is tired at the same time, the air conditioner is controlled based on the data from the temperature sensor device and the analysis results of the emotion engine. At the same time, a gentle lullaby to help the baby relax is generated and played. When the baby falls asleep, a notification "The baby has fallen asleep" is sent to the parent's communication device.

[1018] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. In addition, by taking the user's emotions into consideration, it also provides a safe and secure environment for parents and reduces the burden of childcare.

[1019] The processing flow will be explained below.

[1020] Step 1:

[1021] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[1022] Step 2:

[1023] The server acquires the room temperature data and the baby's body temperature data from the temperature sensor device. Specifically, it records the baby's body temperature and the room ambient temperature based on the numerical data obtained from the temperature sensor.

[1024] Step 3:

[1025] The server receives the baby's crying and other sounds in real time from the audio device, analyzes the audio data, and determines whether the baby is crying or quiet.

[1026] Step 4:

[1027] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. Specifically, it obtains data on the user's emotional state, such as whether they are tired, stressed, or relaxed.

[1028] Step 5:

[1029] The server integrates and analyzes the data from steps 1 to 4 to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state. For example, if the baby continues to cry and the user is tired, the server determines that the baby is excited and the user is tired.

[1030] Step 6:

[1031] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited and the user is tired, it will generate a lullaby that has a particularly relaxing effect.

[1032] Step 7:

[1033] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data can be transmitted to the speaker using Wi-Fi, and the relaxing lullaby will be played.

[1034] Step 8:

[1035] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it sets the air conditioner to cooling or heating mode based on the temperature data and adjusts it to the target temperature.

[1036] Step 9:

[1037] The server then adjusts the lighting in the room appropriately based on the user's emotional data, for example, changing the lighting to softer light to help the user relax.

[1038] Step 10:

[1039] The server checks whether the baby is asleep, and if it is confirmed that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying "The baby has fallen asleep."

[1040] Step 11:

[1041] If something unusual happens to the baby, the server will immediately send an alert to the parent's communication device. For example, if the baby's temperature exceeds a predetermined range or breathing stops, an alert will be sent immediately saying "There is something wrong with the baby."

[1042] This allows the system to monitor the baby's condition from multiple angles and take optimal action while taking into consideration the user's emotions, thereby providing a comfortable sleeping environment for the baby and reducing the burden of childcare on the parents.

[1043] Example 2

[1044] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1045] In modern childcare settings, it is necessary to constantly monitor a baby's health and comfort and respond appropriately. However, parents are not always close to their babies, which increases the burden of childcare. In addition, it is necessary to respond by taking into account not only the baby's condition but also the parent's emotions and stress, but conventional systems do not adequately address this. Therefore, the challenge is to provide a comfortable environment for both babies and parents and reduce the burden of childcare on parents.

[1046] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1047] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for recognizing a user's emotion using an emotion engine, means for integrating and analyzing the baby's condition data and the user's emotion data, means for generating a lullaby according to the baby's condition using a music generation AI, means for playing the generated lullaby, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for adjusting the environment based on the user's emotion data, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if an abnormality occurs in the baby. This makes it possible to comprehensively monitor and analyze the conditions of both the baby and the parent and automatically take appropriate measures to provide an optimal environment for the baby and the parent and reduce the burden of childcare.

[1048] The "camera device" is a video capture device for monitoring and capturing the baby's movements in real time.

[1049] A "temperature sensor device" is a device for measuring and acquiring the temperature of the environment and the baby's body temperature.

[1050] "Audio Device" refers to a recording device for collecting and capturing a baby's cry or other sounds.

[1051] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize the user's emotional state.

[1052] "Music Generation AI" is an artificial intelligence system that generates optimal lullabies based on the baby's condition and the user's emotions.

[1053] An "air conditioner" is an air conditioning device that adjusts and maintains the temperature in a room.

[1054] "Communication devices" are communication devices such as smartphones and tablets used by parents.

[1055] "Lullaby generation" is the process in which the music generation AI creates a lullaby taking into account the baby's condition and the user's emotions.

[1056] "Playback" is the process of outputting the generated lullaby as sound through a speaker.

[1057] "Notifications" are messages sent to parents' communication devices to inform them of changes in the baby's condition or environment.

[1058] An "alert" is a warning message that sends an emergency notification to parents when there is an abnormality in the baby's health or safety.

[1059] "User emotion data" is data that indicates the user's emotional state analyzed by the emotion engine.

[1060] "Environmental adjustment" is the process of changing settings such as temperature and lighting to provide a comfortable environment for the baby and the user.

[1061] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[1062] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[1063] Data collection

[1064] The server receives real-time motion data from the camera device, including the baby's limb movements, rolling over, etc. For example, if the baby is waving its arms, the data is sent to the server.

[1065] The server receives the room temperature data and the baby's temperature data from the temperature sensor device, and the specific values ​​that the room temperature is 25 degrees and the baby's temperature is 37 degrees are sent to the server.

[1066] The server collects the baby's cry and other sounds from the audio device. For example, if the baby is crying loudly, the decibel level of the cry is sent to the server.

[1067] Recognizing user emotions with an emotion engine

[1068] The server uses the camera device to analyze the user's (parent's) facial expressions. For example, when the user's face is facing the camera, the server uses a facial expression recognition algorithm to identify emotions such as smile, anger, sadness, etc.

[1069] The server uses the audio device to analyze the tone of the user's voice. For example, if the user's voice is high-pitched and fast, it determines that the user is stressed. By combining these data, it determines whether the user is tired or not.

[1070] Data analysis and feedback

[1071] The server analyzes the collected data on the baby's movements, temperature, voice, and user's emotions. For example, if a baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, the server will determine that the baby may be crying because of the heat.

[1072] Lullaby generation and playback

[1073] The server uses a music generation AI model to generate the optimal lullaby based on the baby's state and the user's emotions. For example, the prompt text is "My baby is excited. Please generate a lullaby that has a relaxing effect."

[1074] The device receives the generated lullaby audio data and transmits it to a speaker in the baby's room, for example, a smartphone connected to the speaker via Bluetooth, which plays the lullaby.

[1075] environmental adjustment

[1076] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, the server sends a command to the air conditioner to set the temperature to 24 degrees.

[1077] The server can also adjust the lighting in the room based on the user's emotional data, for example adjusting the lighting to 30% brightness to help the user relax.

[1078] Notifications and Anomaly Detection

[1079] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's device saying "The baby has fallen asleep." Specifically, it sends a push notification to the parent's smartphone.

[1080] If something abnormal happens to the baby, the server will quickly send an alert to the parent's device. For example, if the baby's temperature exceeds 39 degrees, an alert will be sent saying "There is something abnormal with the baby."

[1081] Example prompts for generative AI models

[1082] "Generate a relaxing lullaby to play when your baby is crying."

[1083] "Right now, your baby is excited and the room temperature is high. Please generate a lullaby that will help your baby relax."

[1084] "When parents are tired, they need a lullaby to calm their baby. Please generate an appropriate lullaby."

[1085] As described above, this system comprehensively monitors the baby's condition and the user's emotions and takes appropriate action to support the baby's comfortable sleep and reduce the burden of childcare.

[1086] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1087] Step 1:

[1088] Data collection

[1089] The server obtains the baby's motion data in real time from the camera device. As input, it receives the video stream from the camera device and processes it with an image analysis algorithm. As output, it obtains motion data such as the baby's limb movements and rolling over. Specifically, the camera captures the baby's movements and sends the data to the server.

[1090] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device. As input, it receives and analyzes the temperature data from the temperature sensor device. As output, it obtains specific numerical data about the room temperature and the baby's body temperature. For example, data that the room temperature is 25 degrees and the baby's body temperature is 37 degrees is sent to the server.

[1091] The server collects the baby's cry and other sounds from the audio device. As input, it receives the audio stream from the audio device and processes it with an audio analysis algorithm. As output, it obtains data about the decibel level and type of sound of the baby's cry. Specifically, if the baby is crying hard, the decibel level of the cry is sent to the server.

[1092] Step 2:

[1093] Recognizing user emotions with an emotion engine

[1094] The server analyzes the user's facial expressions using a camera device. As input, it receives the video stream from the camera device and processes it with a facial expression recognition algorithm. As output, it obtains data about the user's emotional state (e.g., smiling, angry, sad, etc.). Specifically, when the user's face is facing the camera, the server analyzes their facial expressions.

[1095] The server analyzes the tone of the user's voice using the audio device. As input, it receives the audio stream from the audio device and processes it with an audio tone analysis algorithm. As output, it obtains data about the user's emotional state (e.g., stress, fatigue, etc.). Specifically, if the user's voice is high-pitched and fast, it is determined that the user is stressed.

[1096] Step 3:

[1097] Data analysis and feedback

[1098] The server integrates and analyzes the collected baby's movement data, temperature data, voice data, and user's emotional data. As input, it receives the data obtained in each step above and processes it with a data integration algorithm. As output, it obtains judgment data regarding the baby's current state (e.g., crying, warm, etc.) and the user's emotional state. In a specific operation, assuming that the baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, it will make the judgment that "the baby may be crying because of the heat."

[1099] Step 4:

[1100] Lullaby generation and playback

[1101] The server uses a music generation AI model to generate an optimal lullaby based on the baby's state and the user's emotions. The AI ​​model receives the prompt "My baby is excited. Please generate a lullaby that has a relaxing effect." The output is audio data of a lullaby with a relaxing effect. In concrete terms, the AI ​​generates a lullaby based on the prompt, and the server receives the data.

[1102] The device receives the generated lullaby audio data and sends it to the speaker in the baby's room. As input, it receives the audio data sent from the server and sends it to the speaker. As output, the lullaby is played from the speaker. Specifically, the smartphone is connected to the speaker via Bluetooth and the lullaby is played.

[1103] Step 5:

[1104] environmental adjustment

[1105] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. As input, it receives room temperature data and processes it with the air conditioner control algorithm. As output, it obtains commands to send to the air conditioner. In concrete terms, the server sends a command to the air conditioner to "set the temperature to 24 degrees."

[1106] The server can also adjust the lighting in a room based on the user's emotional data. It receives the user's emotional data as input and processes it with a lighting control algorithm. The output is a command to send to the lighting. Specifically, the server sends a command to adjust the lighting to 30% brightness to help the user relax.

[1107] Step 6:

[1108] Notifications and Anomaly Detection

[1109] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." As input, it receives and analyzes the baby's movement data. As output, it generates a notification message and sends it to the parent's device. Specifically, a push notification is sent to the smartphone, displaying the message, "The baby has fallen asleep."

[1110] If something abnormal occurs with the baby, the server quickly sends an alert to the parent's communication device. As input, it receives and analyzes the baby's temperature and breathing data. As output, it generates an alert message informing the parent of the abnormality and sends it to the parent's device. For example, if the body temperature exceeds 39 degrees, an alert saying "There is something abnormal with the baby" is sent to the smartphone.

[1111] (Application example 2)

[1112] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1113] In today's physical stores, customers often feel stressed, which causes a poor customer experience. It is also difficult for store staff to grasp customers' emotional state in real time and respond appropriately, making it a challenge to improve customer satisfaction. Furthermore, some stores do not adequately adjust the environment to maintain customer comfort.

[1114] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1115] In this invention, the server includes means for acquiring motion from a camera device, means for acquiring environmental and body temperature data from a temperature sensor device, means for acquiring audio from an audio device, means for analyzing the acquired data and determining the state, means for generating music according to the state using a music generation AI, means for playing the generated music, means for controlling an air conditioner as needed to adjust the temperature to maintain a comfortable environment, means for sending a notification to a notification terminal when the state changes, means for immediately sending an alert to the notification terminal when an abnormality occurs, means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize emotions, means for adjusting the environment according to the user's emotional state based on data from the audio device, means for the music generation AI to generate music with a high relaxing effect based on the emotional data, and means for controlling lighting to adjust the ambient light, thereby enabling real-time environmental adjustment according to the customer's emotional state.

[1116] A "camera device" is an electronic device for capturing video and image data.

[1117] "Movement" refers to physical movement or action.

[1118] A "temperature sensor device" is an electronic device for measuring the temperature of an environment or object.

[1119] "Environmental data" refers to data that represents the state or condition of the surroundings, such as temperature and light intensity.

[1120] "Body temperature data" refers to data obtained by measuring and recording the temperature of a living body.

[1121] An "audio device" is an electronic device for capturing sound, such as a microphone.

[1122] A "means for determining the status" is a method or system that analyzes and evaluates the current situation or conditions based on acquired data.

[1123] "Music generation AI" is a system that automatically generates music using artificial intelligence technology.

[1124] An "air conditioning device" is a device that adjusts temperature and humidity. Air conditioners are examples of this.

[1125] A "notification terminal" is a device that receives signals and messages, such as a smartphone.

[1126] An "alert" is a warning message or notification of a particular condition.

[1127] The "emotion engine" is a system that analyzes emotions from facial expressions and tone of voice.

[1128] "User" means any person who uses a system or service.

[1129] "Environmental control" is the act of controlling environmental conditions such as temperature and light to maintain or improve comfort.

[1130] "Music with a high relaxing effect" is music that is intended to relax the listener.

[1131] A "means for controlling lighting" is a system or method for adjusting the intensity or color temperature of light.

[1132] This invention is a system that monitors the emotional state of customers in real time in physical stores and adjusts the environment accordingly to improve customer experience. This system consists of a camera device, a temperature sensor device, a voice device, an emotion engine, a music generation AI, an air conditioning system, a lighting system, and a notification terminal.

[1133] System Configuration

[1134] The system includes the following major components:

[1135] 1. Camera Device

[1136] The camera device is used to capture customer behavior in real time, specifically, a common webcam such as the Logitech C920 can be used.

[1137] 2. Temperature sensor device

[1138] Temperature sensor devices are used to obtain environmental and customer body temperature data. For example, DHT22 sensors are available.

[1139] 3. Audio Devices

[1140] The audio device is used to capture the customer's voice. A high-sensitivity microphone such as a Blue Yeti can be used.

[1141] 4. Emotion Engine

[1142] The emotion engine uses data from cameras and audio devices to analyze customers' facial expressions and tone of voice to recognize emotions, using an AI model powered by OpenCV and TensorFlow.

[1143] 5. Music Generation AI

[1144] The music generation AI generates relaxing music based on the customer's emotional state, using AI frameworks such as Google Magenta.

[1145] 6. Air conditioner

[1146] An air conditioning device, such as a Nest Thermostat, to adjust the ambient temperature as needed.

[1147] 7. Lighting System

[1148] A lighting system for adjusting ambient light, compatible with smart lights such as Phillips Hue.

[1149] 8. Notification terminal

[1150] A device that sends notifications when a customer's condition changes or an abnormality occurs. A standard smartphone can be used.

[1151] Program processing overview

[1152] The server acquires real-time customer behavior data from the camera device, environmental and customer body temperature data from the temperature sensor device, and customer voice data from the audio device. Using these data, the emotion engine analyzes the customer's emotional state.

[1153] The acquired data is analyzed by a server to determine the current situation. The music generation AI generates relaxing music according to the customer's emotional state and plays it through the store's speakers. In addition, the air conditioning and lighting systems adjust the environment as needed.

[1154] Specific examples

[1155] For example, if a customer is feeling stressed, that information is collected from the camera and audio device. The emotion engine determines the emotional state as "stressed," and the music generation AI generates relaxing music. This music is played through speakers in the store, and the lighting is adjusted to a softer light. At the same time, the air conditioning is adjusted to an appropriate temperature.

[1156] Prompt Sentence Examples

[1157] "Generate appropriate relaxing music when the customer is feeling stressed. The text data corresponding to the customer's emotional state is as follows: 'Stress'"

[1158] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1159] Step 1:

[1160] The server acquires customer behavior data in real time from the camera device. The acquired video data (input) is used to analyze the customer's location, behavior patterns, etc., and the results (output) become preprocessed data for emotion recognition.

[1161] Step 2:

[1162] The server acquires temperature data of the environment and the customer from the temperature sensor device. The acquired temperature data (input) is analyzed by the server to determine the appropriate range of the environment temperature and the customer's body temperature (output).

[1163] Step 3:

[1164] The server captures the customer's voice from the audio device, and after noise filtering and voice analysis, the captured voice data (input) is used to detect the customer's tone of voice and emotional state (output).

[1165] Step 4:

[1166] The server uses an emotion engine to analyze data acquired from the camera device and audio device to determine the customer's emotional state. Based on the input data (video and audio data), the emotional state (e.g., stress, relaxation) is detected from the customer's facial expression, posture, and tone of voice, and the emotional state (output) is obtained.

[1167] Step 5:

[1168] Based on the output of the emotion engine, the server sends a prompt to the music generation AI. An example of a prompt is, "Please generate appropriate relaxing music for when the customer is feeling stressed." Based on the input prompt, the music generation AI generates (outputs) music with a high relaxing effect, and the music data is obtained.

[1169] Step 6:

[1170] The server then transmits the generated music data to speakers in the store, where it is played, providing an acoustic environment that responds to the customer's emotional state.

[1171] Step 7:

[1172] The server controls the air conditioner based on the environmental data. Based on the acquired temperature data (input), it sets the appropriate room temperature (output) and sends temperature adjustment instructions to the air conditioner. This maintains the optimal ambient temperature.

[1173] Step 8:

[1174] The server controls the lighting system to adjust the ambient light. Based on the output data (input) of the emotion engine, it sets the appropriate lighting settings (output) and sends dimming instructions to the lighting system. This provides a comfortable lighting environment.

[1175] Step 9:

[1176] If the server detects an abnormality, it immediately sends an alert to the notification terminal. Based on the abnormal data (input), it generates an appropriate warning message (output) and sends it to the notification terminal. This enables a quick response.

[1177] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1178] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1179] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1180] [Fourth embodiment]

[1181] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1182] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1183] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1184] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1185] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1186] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1187] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1188] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1189] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1190] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1191] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1192] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1193] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1194] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[1195] System Configuration

[1196] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[1197] What the program does

[1198] 1. Data Collection

[1199] The server collects real-time baby movement data from the camera device, which can determine whether the baby is moving or quiet.

[1200] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[1201] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[1202] 2. Data analysis and feedback

[1203] The server analyzes the collected data and evaluates the baby's condition (e.g., excited, relaxed, normal body temperature, high temperature). For example, if the baby continues to cry, it is determined to be excited.

[1204] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[1205] 3. Lullaby Generation and Playback

[1206] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited, a lullaby with a slow rhythm will be generated.

[1207] The device (e.g., a smartphone or tablet) plays the generated lullaby through a speaker in the baby's room, providing a sound environment that suits the baby's condition.

[1208] 4. Environmental adjustment

[1209] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the room.

[1210] It receives feedback from the temperature sensor and continues to adjust until it reaches the target temperature.

[1211] 5. Notifications and Anomaly Detection

[1212] The server checks the baby's sleep status and, if it is confirmed that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, allowing the parent to know that the baby has fallen asleep safely.

[1213] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends an alert to the parent's communication device, allowing the parent to respond promptly.

[1214] Specific examples

[1215] Example 1: If a baby continues to cry, the camera and audio devices will be used to check the baby's condition, and the music generation AI will create a soothing lullaby. The lullaby will be played through the speakers, and the air conditioner will start cooling the room. As a result, the baby will gradually calm down and fall asleep.

[1216] Example 2: If the baby's body temperature is high and the room temperature is also high, the air conditioner will be controlled based on data from the temperature sensor device and the cooling will start. At the same time, a gentle lullaby will be generated and played to help the baby relax. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[1217] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. It also provides a safe and secure environment for parents, reducing the burden of childcare.

[1218] The processing flow will be explained below.

[1219] Step 1:

[1220] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[1221] Step 2:

[1222] The server receives the room temperature data and the baby's temperature data from the temperature sensor device. Specifically, it records the baby's temperature and the room's ambient temperature based on the numerical data obtained from the temperature sensor.

[1223] Step 3:

[1224] The server receives the baby's crying and other sounds in real time from the audio device, and analyzes the audio input to determine whether the baby is crying or quiet.

[1225] Step 4:

[1226] The server integrates and analyzes the data from steps 1 to 3 to evaluate the baby's current state (excited, relaxed, normal body temperature, high temperature, etc.) For example, if the baby is crying continuously, has a high body temperature, and is active, it is determined to be in an excited state.

[1227] Step 5:

[1228] The server uses music generation AI to generate the optimal lullaby based on the analysis results. Specifically, it creates lullabies with slow rhythms or gentle melodies depending on the results of data analysis.

[1229] Step 6:

[1230] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data is transmitted to the speaker using Wi-Fi, and the audio is played back.

[1231] Step 7:

[1232] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it automatically adjusts the air conditioner's cooling or heating mode based on the temperature data and sets the target temperature.

[1233] Step 8:

[1234] The server receives feedback from the temperature sensor and continues to appropriately control the air conditioner until the target temperature is reached. For example, it may continue cooling until the target temperature is reached, and then stop the air conditioner when the temperature is appropriate.

[1235] Step 9:

[1236] The server confirms that the baby is asleep, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[1237] Step 10:

[1238] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[1239] This allows the system to monitor the baby's condition in real time and respond appropriately to the situation, reducing the burden on parents and providing a comfortable sleeping environment for the baby.

[1240] Example 1

[1241] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1242] Conventional baby monitoring systems tend to rely on specific devices, making it difficult to comprehensively monitor a baby's condition and provide appropriate feedback. As a result, parents are unable to consistently monitor their baby's condition, placing a heavy burden on childcare. Furthermore, due to insufficient environmental adjustment and anomaly detection functions, it is difficult to provide a sustainable, comfortable environment for babies. To solve these problems, a system is needed that can comprehensively monitor a baby's condition in real time and automatically take appropriate action.

[1243] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1244] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for analyzing the acquired data and evaluating the baby's condition, means for generating a lullaby according to the baby's condition using a music generation algorithm, means for playing the generated lullaby through a playback device, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if something abnormal occurs with the baby. This makes it possible to monitor the baby's condition from various angles and automatically take appropriate measures.

[1245] A "camera device" is a device that captures a baby's movements and acquires the video data in real time.

[1246] A "temperature sensor device" is a device that measures the baby's surrounding environment and the baby's own body temperature and provides that data to a server.

[1247] An "audio device" is a device that collects the baby's crying and other surrounding sounds and sends the data to a server.

[1248] The "server" is a central computer system that integrates and analyzes data collected from various devices and generates appropriate feedback.

[1249] The "means for analyzing the acquired data and assessing the baby's condition" refers to algorithms or software for estimating the baby's current condition based on the collected movement data, temperature data, and audio data.

[1250] A "music generation algorithm" is an AI (artificial intelligence) model or program that generates the optimal lullaby depending on the baby's condition.

[1251] A "playback device" is a speaker or audio device for playing back the generated lullaby sent from the server.

[1252] "Air conditioning equipment" refers to equipment with heating and cooling functions for adjusting the temperature in a room, and includes air conditioners.

[1253] A "parent communication device" is a mobile device, such as a smartphone or tablet, used by a parent to receive notifications and alerts.

[1254] "Means for sending an alert in the event of an abnormality" refers to software or hardware configurations that immediately send an alert message to the parent's communication device when an abnormality in the baby is detected from the collected and analyzed data.

[1255] The present invention is a system that appropriately monitors the baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. The present invention can be implemented as follows.

[1256] System Configuration

[1257] This system is composed of multiple hardware and software devices, including the following:

[1258] Camera device: A device for capturing baby's movements in real time, such as an infrared camera or a high-resolution camera.

[1259] Temperature sensor device: A sensor device that measures the temperature of the room and the baby. A high-precision digital thermometer is used.

[1260] Audio device: A microphone device that collects the baby's cries and surrounding sounds.

[1261] Server: A central computer system that analyzes data collected from each device and generates the necessary feedback. A server equipped with a high-performance processor is used here.

[1262] Device: A communication device used by the end user (parent). For example, a smartphone or tablet.

[1263] Data collection

[1264] The server collects the baby's movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[1265] The server receives real-time room temperature and baby temperature data from the temperature sensor device, which can determine whether the baby is within a comfortable temperature range.

[1266] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[1267] Data analysis and feedback

[1268] The server then aggregates and analyzes the collected data, for example analyzing the audio data to determine whether the baby is crying, using a voice analysis algorithm.

[1269] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, or the crying continues), the system determines the appropriate response.

[1270] Lullaby generation and playback

[1271] The server uses a music generation algorithm to generate an appropriate lullaby depending on the baby's state. For example, if the baby is excited, a lullaby with a calming rhythm will be generated.

[1272] The device receives the generated lullaby and plays it through a speaker in the room, which may be connected to the smartphone via Wi-Fi.

[1273] environmental adjustment

[1274] The server controls the air conditioning unit as needed to adjust the temperature to keep the baby in a comfortable environment. For example, if the room temperature is high, the server activates the air conditioning's cooling function.

[1275] It receives feedback from temperature sensor devices and adjusts the air conditioner until the target temperature is reached.

[1276] Notifications and Anomaly Detection

[1277] If the server determines that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device, for example, a message saying "Baby has fallen asleep."

[1278] If the baby develops any abnormalities (e.g., high fever, cough, or breathing problems), the server immediately sends a warning to the parent's communication device, such as a message saying, "Your baby's temperature is high. Please check it."

[1279] Specific examples

[1280] Example 1:

[1281] If the baby continues to cry, the server will check the baby's condition from the camera and audio data and use a music generation algorithm to generate a soothing lullaby. The lullaby will be played through the speaker, and the air conditioner will start cooling to adjust the room temperature. As a result, the baby will gradually calm down and fall asleep.

[1282] Example 2:

[1283] If the baby's body temperature is high and the room temperature is also high, the server will control the air conditioning based on the data from the temperature sensor device and turn on the air conditioner. At the same time, a gentle lullaby to help the baby relax will be generated and played. After the baby falls asleep, a notification "The baby has fallen asleep" will be sent to the parent's communication device.

[1284] Prompt Sentence Examples

[1285] "Monitor your baby's crying and play a soothing lullaby."

[1286] "While controlling the air conditioning unit if the room temperature is high, it also checks the baby's temperature and generates an appropriate lullaby."

[1287] In this way, the present invention monitors the baby's condition from multiple angles and responds appropriately to provide a comfortable sleeping environment for the baby. It also provides a safe and secure childcare environment for parents, reducing the burden of childcare.

[1288] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1289] Step 1:

[1290] Data collection

[1291] The server collects baby's movement data in real time from the camera device. The input is the video data captured by the camera. The server analyzes this video data frame by frame and extracts the baby's movements (e.g., whether or not the baby is moving and the speed of the movements). The output is the baby's movement pattern data.

[1292] The server obtains the room temperature and the baby's body temperature data from the temperature sensor device. The temperature data measured by the sensor at regular intervals is the input. The server receives and analyzes this data to determine whether the room is within the appropriate temperature range or whether the baby's body temperature is abnormal. The output is the environmental temperature data and the baby's body temperature data.

[1293] The server collects the baby's crying and other ambient sounds from the audio device. The audio data captured by the audio device's microphone is the input. The server uses an audio analysis algorithm to analyze the characteristics of the audio (e.g., volume, frequency) and determine whether the baby is crying. The output is the audio analysis result.

[1294] Step 2:

[1295] Data analysis and feedback

[1296] The server integrates and analyzes all the data collected in step 1 (movement pattern data, ambient temperature data, body temperature data, and voice analysis results). Based on these analysis results, the baby's condition is evaluated. For example, if the baby continues to cry, it is judged to be in an "excited state" based on the voice analysis results. The inputs are each sensor data and its analysis results. The output is the evaluation result of the baby's condition.

[1297] Step 3:

[1298] Lullaby generation and playback

[1299] The server uses a music generation algorithm to generate an appropriate lullaby based on the baby's state evaluated in step 2. For example, the prompt sentence is "Generate a calming song." The music generation AI model generates a lullaby based on this prompt sentence, and the generated lullaby data is obtained as the output.

[1300] The device receives the lullaby data sent from the server and plays it through the room's speakers. The input is the lullaby data from the server. The device sends it to a playback device, and the sound is played as the output.

[1301] Step 4:

[1302] environmental adjustment

[1303] The server controls the air conditioner as needed to adjust the room temperature to an appropriate range. If the temperature sensor device data indicates that the room is hot, the server issues a command to the air conditioner to operate the cooling function. The operating status of the air conditioner is obtained as an output.

[1304] The server continues to receive feedback from the temperature sensor and continues to control the air conditioning until the target temperature is reached. The input is the continuous temperature feedback data. The output is the final adjusted temperature data.

[1305] Step 5:

[1306] Notifications and Anomaly Detection

[1307] When the server confirms that the baby has been quiet for a certain period of time (e.g., 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." The input includes the analysis results of motion data and voice data. The output is a notification message sent to the parent's communication device.

[1308] The server immediately sends an alert to the parent's communication device if the baby experiences any abnormalities (e.g., high temperature, respiratory failure, etc.). The input is temperature data and other sensor data indicating an abnormality. The output is a warning message sent to the parent's communication device.

[1309] (Application example 1)

[1310] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1311] In conventional factory environments, monitoring of workers' movements, physical condition, and working environment is insufficient, resulting in reduced work efficiency and health risks for workers. In particular, working in the same position for long periods of time and working in high-temperature environments can cause fatigue and health problems for workers. There is also a need for a system that can respond immediately when a worker becomes ill.

[1312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1313] In this invention, the server includes means for acquiring the target person's movements from a camera device, means for acquiring environmental and target person's body temperature data from a temperature sensor device, means for acquiring the target person's voice from an audio device, means for analyzing the acquired data and determining the target person's condition, means for generating voice instructions according to the target person's condition using a music generation AI, means for playing the generated voice instructions, means for controlling the air conditioner as necessary to adjust the temperature so that the target person is in a comfortable environment, means for sending a notification to a manager's communication terminal when the target person enters a specific state, and means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person. This enables real-time monitoring of worker movements and physical condition, enabling optimization of the work environment and health management.

[1314] A "camera device" is a photographing device for capturing the movements of a target person in real time.

[1315] A "temperature sensor device" is a measuring device for acquiring body temperature data of the environment and a target person.

[1316] An "audio device" is a sound collection device for acquiring the voice of a target person.

[1317] "Means for analyzing acquired data and determining the status of the target person" refers to a method or apparatus for processing information acquired from the camera device, temperature sensor device, and audio device and assessing the current status of the target person.

[1318] "Means for generating voice instructions according to the state of a target person using music generation AI" refers to a method or device that uses artificial intelligence technology to generate voice instructions that are adapted to the state of a target person.

[1319] "Means for playing generated voice instructions" refers to a device or method for playing voice instructions generated by the music generation AI.

[1320] "Means for controlling an air conditioner and adjusting the temperature so that the target person is in a comfortable environment" refers to a method or device for operating an air conditioner to appropriately adjust the environmental temperature.

[1321] "Means for sending a notification to the manager's communication terminal when the target person enters a specific state" refers to a method or device for notifying the manager of this information when the target person reaches a specific state, such as a situation where a break is required.

[1322] "Means for immediately sending an alert to the manager's communication terminal when an abnormality occurs in the target person" refers to a method or device for immediately sending a warning to the manager when an abnormality occurs in the target person's health.

[1323] The present invention aims to provide a work support robot system that monitors the movements and environment of a target person (worker) and provides necessary instructions and adjusts the environment. Specific embodiments of the present invention will be described below.

[1324] System Configuration

[1325] This system consists of a camera device, a temperature sensor device, an audio device, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and sends data. The server is also connected to the communication terminal and sends notifications and alerts.

[1326] What the program does

[1327] 1. Data Collection

[1328] The server collects real-time motion data of the target person from the camera device, which allows it to determine whether the target person is moving or stationary.

[1329] The temperature sensor device acquires the ambient temperature and the target person's body temperature data, which can then be used to determine whether the target person is within a comfortable temperature range.

[1330] Collecting the subject's voice from a voice device, which allows us to determine whether the subject understands the instructions.

[1331] 2. Data analysis and feedback

[1332] The server analyzes the collected data and evaluates the target person's status (e.g., working, needing a break, abnormal state). For example, if the target person remains motionless for a long time, it determines that the person needs a break.

[1333] If each piece of data is outside the specified range (e.g., room temperature is high, lighting is low), the system determines how to respond accordingly.

[1334] 3. Generation and playback of voice instructions

[1335] The server uses music generation AI to generate optimal voice instructions based on the analysis results. For example, if the person is tired, it will generate a voice instruction to encourage them to take a break.

[1336] The terminal (e.g., a smartphone) plays the generated voice instructions through an audio device that transmits the voice instructions to the target person, thereby providing appropriate instructions according to the target person's state.

[1337] 4. Environmental adjustment

[1338] The server controls the air conditioner as needed to keep the temperature in the factory within an appropriate range. For example, if the room temperature is high, the air conditioner will operate to cool the factory.

[1339] 5. Notifications and Anomaly Detection

[1340] When the target person enters a specific state (e.g., needs a break), the server sends a notification to the administrator's communication terminal, allowing the administrator to grasp the target person's state.

[1341] Additionally, if an abnormality occurs with the target person (e.g., high fever, immobility, etc.), an alert is immediately sent to the administrator's communication terminal, allowing the administrator to respond promptly.

[1342] Hardware and software used

[1343] Camera devices: Used to monitor the behavior of subjects.

[1344] Temperature sensor device: Used to monitor environmental and body temperatures.

[1345] Audio device: Used to communicate instructions to the target person.

[1346] Server: Data processing and analysis.

[1347] Communication device: Used to receive notifications and alerts.

[1348] OpenCV: Software used to analyze camera footage.

[1349] pyttsx3: Software used to generate and play audio instructions.

[1350] GPIO Zero: A library used to acquire data from the temperature sensor.

[1351] Specific examples

[1352] Example 1: If a worker remains motionless for more than 30 minutes, the system will issue a voice message saying "Please take a break" and will also send a work status notification to the manager.

[1353] Example 2: If the temperature inside the factory goes outside the set range (e.g., above 28 degrees), the system automatically switches the air conditioner to cooling mode.

[1354] Example prompt for a generative AI model:

[1355] Similar to the system that provides voice alerts and adjusts the temperature based on a baby's movements and temperature data, create a program that monitors the movements and temperature data of factory workers, issues voice alerts when they become fatigued, and controls the air conditioner as necessary. Specific examples include notifications if a worker has not moved for 30 minutes and switching on the air conditioner if the temperature exceeds 28 degrees.

[1356] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1357] Step 1:

[1358] The server acquires the motion data of the target person from the camera device in real time. Specifically, the camera device captures the video of the target person and sends the video data to the server. The server analyzes the video data to determine whether the target person is moving. The input is the video data from the camera device, and the output is the motion status of the target person (e.g., moving, stationary).

[1359] Step 2:

[1360] The server receives the ambient temperature and the target person's body temperature data from the temperature sensor device. Specifically, the temperature sensor device measures the temperature data in real time and sends it to the server. The server analyzes the measured data to determine whether the ambient temperature and the target person's body temperature are within an appropriate range. The input is the temperature data from the temperature sensor device, and the output is the temperature status (e.g., normal, high temperature, low temperature).

[1361] Step 3:

[1362] The server acquires the target person's voice data from the audio device. Specifically, it picks up the target person's voice and other sounds and sends them to the server. The server analyzes the voice data to determine what the target person is saying or what sounds are being made. The input is the voice data from the audio device, and the output is the content of the voice (e.g., response to commands, background sounds).

[1363] Step 4:

[1364] The server comprehensively analyzes the collected data and determines the target person's condition. Specifically, it integrates motion data, temperature data, and voice data, and evaluates whether the target person is working, needs a break, or is in an abnormal state based on each piece of data. The input is motion data, temperature data, and voice data, and the output is the target person's overall condition (e.g., working, fatigue, abnormal).

[1365] Step 5:

[1366] The server uses music generation AI to generate optimal voice instructions based on the analysis results. Specifically, it generates appropriate instructions (e.g., "Please take a break" or "Please continue working") depending on the target person's condition. The input is the target person's overall condition, and the output is the generated voice instructions.

[1367] Step 6:

[1368] The server transmits the generated voice instructions to the target person through a terminal (e.g., a smartphone) or an audio device. Specifically, the generated voice instructions are played back as audio so that the target person can confirm the instructions. The input is the generated voice instructions, and the output is the played back audio.

[1369] Step 7:

[1370] The server controls the air conditioner as needed to adjust the temperature in the factory to an appropriate range. Specifically, if the temperature data exceeds the set range, the air conditioner will operate to cool or heat the factory. The input is the temperature state, and the output is the appropriate adjustment of the environmental temperature.

[1371] Step 8:

[1372] The server sends a notification to the manager's communication terminal when the target person enters a specific state. Specifically, for example, if it determines that the target person needs a break, it notifies the manager of that information. The input is the target person's overall state, and the output is a notification to the manager.

[1373] Step 9:

[1374] The server immediately sends an alert to the administrator's communication terminal if an abnormality occurs in the target person. Specifically, if it determines that the target person is in an abnormal condition, such as having a high fever or not moving, it sends an emergency alert to the administrator. The input is the target person's abnormal condition, and the output is an emergency alert.

[1375] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1376] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[1377] System Configuration

[1378] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[1379] What the program does

[1380] 1. Data Collection

[1381] The server receives real-time baby movement data from the camera device, which allows it to determine whether the baby is moving or quiet.

[1382] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device, which allows it to determine whether the baby is within a comfortable temperature range.

[1383] The server collects the baby's cries and other sounds from the audio device, which allows it to determine whether the baby is crying or quiet.

[1384] 2. Emotion recognition by emotion engine

[1385] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. For example, if the user is tired or stressed, the server obtains that information.

[1386] 3. Data analysis and feedback

[1387] The server integrates and analyzes the baby's data and the user's emotional data to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state.

[1388] If each piece of data is outside a predetermined range (e.g., the room temperature exceeds the set range, the crying continues, or the user is feeling stressed), the system determines how to respond accordingly.

[1389] 4. Lullaby Generation and Playback

[1390] The server uses music generation AI to generate the optimal lullaby based on the baby's state and the user's emotions. For example, if the baby is excited and the user is tired, a particularly relaxing lullaby will be generated.

[1391] The device (e.g., a smartphone or tablet) then transmits the generated lullaby audio data to a speaker in the baby's room and plays it back, providing a sound environment tailored to the baby's condition.

[1392] 5. Environmental adjustment

[1393] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. For example, if the room temperature is high, the server will operate the air conditioner to cool it down.

[1394] The lighting in a room can also be adjusted based on the user's emotional data, for example, by adjusting the lighting in the room to a softer light to help the user relax.

[1395] 6. Notifications and Anomaly Detection

[1396] The server checks the baby's sleeping state, and if it determines that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep."

[1397] The server quickly sends an alert to the parent's communication device if something unusual happens to the baby, such as if the baby's temperature exceeds a certain range or breathing stops.

[1398] Specific examples

[1399] Example 1: If a baby continues to cry and the user is feeling stressed, the camera and audio devices will be used to check the baby's condition. The music generation AI will generate a relaxing lullaby and play it through the speaker. Furthermore, the air conditioner will start cooling and the room lighting will be adjusted to a softer light, creating a comfortable environment for the baby and the user.

[1400] Example 2: If the baby's body temperature is high and the user is tired at the same time, the air conditioner is controlled based on the data from the temperature sensor device and the analysis results of the emotion engine. At the same time, a gentle lullaby to help the baby relax is generated and played. When the baby falls asleep, a notification "The baby has fallen asleep" is sent to the parent's communication device.

[1401] As described above, the present invention supports a comfortable sleep for babies by monitoring the baby's condition from multiple angles and taking appropriate measures. In addition, by taking the user's emotions into consideration, it also provides a safe and secure environment for parents and reduces the burden of childcare.

[1402] The processing flow will be explained below.

[1403] Step 1:

[1404] The server receives real-time data on the baby's movements from the camera device, and analyzes the camera footage to determine whether the baby is moving or quiet.

[1405] Step 2:

[1406] The server acquires the room temperature data and the baby's body temperature data from the temperature sensor device. Specifically, it records the baby's body temperature and the room ambient temperature based on the numerical data obtained from the temperature sensor.

[1407] Step 3:

[1408] The server receives the baby's crying and other sounds in real time from the audio device, analyzes the audio data, and determines whether the baby is crying or quiet.

[1409] Step 4:

[1410] The server uses a camera device and an audio device to analyze the user's (parent's) facial expressions and tone of voice to recognize the user's emotions. Specifically, it obtains data on the user's emotional state, such as whether they are tired, stressed, or relaxed.

[1411] Step 5:

[1412] The server integrates and analyzes the data from steps 1 to 4 to evaluate the baby's current state (e.g., excited, relaxed, normal body temperature, high temperature) and the user's emotional state. For example, if the baby continues to cry and the user is tired, the server determines that the baby is excited and the user is tired.

[1413] Step 6:

[1414] The server uses music generation AI to generate the optimal lullaby based on the analysis results. For example, if the baby is excited and the user is tired, it will generate a lullaby that has a particularly relaxing effect.

[1415] Step 7:

[1416] The device (e.g., a smartphone or tablet) transmits the generated lullaby audio data to a speaker in the baby's room and plays it back. For example, the audio data can be transmitted to the speaker using Wi-Fi, and the relaxing lullaby will be played.

[1417] Step 8:

[1418] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, it sets the air conditioner to cooling or heating mode based on the temperature data and adjusts it to the target temperature.

[1419] Step 9:

[1420] The server then adjusts the lighting in the room appropriately based on the user's emotional data, for example, changing the lighting to softer light to help the user relax.

[1421] Step 10:

[1422] The server checks whether the baby is asleep, and if it is confirmed that the baby has been quiet for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying "The baby has fallen asleep."

[1423] Step 11:

[1424] If something unusual happens to the baby, the server will immediately send an alert to the parent's communication device. For example, if the baby's temperature exceeds a predetermined range or breathing stops, an alert will be sent immediately saying "There is something wrong with the baby."

[1425] This allows the system to monitor the baby's condition from multiple angles and take optimal action while taking into consideration the user's emotions, thereby providing a comfortable sleeping environment for the baby and reducing the burden of childcare on the parents.

[1426] Example 2

[1427] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1428] In modern childcare settings, it is necessary to constantly monitor a baby's health and comfort and respond appropriately. However, parents are not always close to their babies, which increases the burden of childcare. In addition, it is necessary to respond by taking into account not only the baby's condition but also the parent's emotions and stress, but conventional systems do not adequately address this. Therefore, the challenge is to provide a comfortable environment for both babies and parents and reduce the burden of childcare on parents.

[1429] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1430] In this invention, the server includes means for acquiring baby's movements from a camera device, means for acquiring environmental and baby's body temperature data from a temperature sensor device, means for acquiring baby's voice from an audio device, means for recognizing a user's emotion using an emotion engine, means for integrating and analyzing the baby's condition data and the user's emotion data, means for generating a lullaby according to the baby's condition using a music generation AI, means for playing the generated lullaby, means for controlling an air conditioner as needed to adjust the temperature to keep the baby in a comfortable environment, means for adjusting the environment based on the user's emotion data, means for sending a notification to the parent's communication device when the baby falls asleep, and means for immediately sending an alert to the parent's communication device if an abnormality occurs in the baby. This makes it possible to comprehensively monitor and analyze the conditions of both the baby and the parent and automatically take appropriate measures to provide an optimal environment for the baby and the parent and reduce the burden of childcare.

[1431] The "camera device" is a video capture device for monitoring and capturing the baby's movements in real time.

[1432] A "temperature sensor device" is a device for measuring and acquiring the temperature of the environment and the baby's body temperature.

[1433] "Audio Device" refers to a recording device for collecting and capturing a baby's cry or other sounds.

[1434] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to recognize the user's emotional state.

[1435] "Music Generation AI" is an artificial intelligence system that generates optimal lullabies based on the baby's condition and the user's emotions.

[1436] An "air conditioner" is an air conditioning device that adjusts and maintains the temperature in a room.

[1437] "Communication devices" are communication devices such as smartphones and tablets used by parents.

[1438] "Lullaby generation" is the process in which the music generation AI creates a lullaby taking into account the baby's condition and the user's emotions.

[1439] "Playback" is the process of outputting the generated lullaby as sound through a speaker.

[1440] "Notifications" are messages sent to parents' communication devices to inform them of changes in the baby's condition or environment.

[1441] An "alert" is a warning message that sends an emergency notification to parents when there is an abnormality in the baby's health or safety.

[1442] "User emotion data" is data that indicates the user's emotional state analyzed by the emotion engine.

[1443] "Environmental adjustment" is the process of changing settings such as temperature and lighting to provide a comfortable environment for the baby and the user.

[1444] The present invention is a system that appropriately monitors a baby's condition, generates and plays lullabies based on the condition, and adjusts the environment and detects abnormalities as needed, thereby reducing the burden on parents and providing a comfortable sleeping environment for the baby. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the system aims to achieve more effective lullabies and environment adjustments. Specifically, the present invention can be implemented as follows.

[1445] This system consists of a camera device, a temperature sensor device, an audio device, an emotion engine, a server, and a communication terminal (e.g., a smartphone). Each device is connected to the server via a network and transmits data. The server is also connected to the communication terminal and sends notifications and alerts.

[1446] Data collection

[1447] The server receives real-time motion data from the camera device, including the baby's limb movements, rolling over, etc. For example, if the baby is waving its arms, the data is sent to the server.

[1448] The server receives the room temperature data and the baby's temperature data from the temperature sensor device, and the specific values ​​that the room temperature is 25 degrees and the baby's temperature is 37 degrees are sent to the server.

[1449] The server collects the baby's cry and other sounds from the audio device. For example, if the baby is crying loudly, the decibel level of the cry is sent to the server.

[1450] Recognizing user emotions with an emotion engine

[1451] The server uses the camera device to analyze the user's (parent's) facial expressions. For example, when the user's face is facing the camera, the server uses a facial expression recognition algorithm to identify emotions such as smile, anger, sadness, etc.

[1452] The server uses the audio device to analyze the tone of the user's voice. For example, if the user's voice is high-pitched and fast, it determines that the user is stressed. By combining these data, it determines whether the user is tired or not.

[1453] Data analysis and feedback

[1454] The server analyzes the collected data on the baby's movements, temperature, voice, and user's emotions. For example, if a baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, the server will determine that the baby may be crying because of the heat.

[1455] Lullaby generation and playback

[1456] The server uses a music generation AI model to generate the optimal lullaby based on the baby's state and the user's emotions. For example, the prompt text is "My baby is excited. Please generate a lullaby that has a relaxing effect."

[1457] The device receives the generated lullaby audio data and transmits it to a speaker in the baby's room, for example, a smartphone connected to the speaker via Bluetooth, which plays the lullaby.

[1458] environmental adjustment

[1459] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. Specifically, the server sends a command to the air conditioner to set the temperature to 24 degrees.

[1460] The server can also adjust the lighting in the room based on the user's emotional data, for example adjusting the lighting to 30% brightness to help the user relax.

[1461] Notifications and Anomaly Detection

[1462] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's device saying "The baby has fallen asleep." Specifically, it sends a push notification to the parent's smartphone.

[1463] If something abnormal happens to the baby, the server will quickly send an alert to the parent's device. For example, if the baby's temperature exceeds 39 degrees, an alert will be sent saying "There is something abnormal with the baby."

[1464] Example prompts for generative AI models

[1465] "Generate a relaxing lullaby to play when your baby is crying."

[1466] "Right now, your baby is excited and the room temperature is high. Please generate a lullaby that will help your baby relax."

[1467] "When parents are tired, they need a lullaby to calm their baby. Please generate an appropriate lullaby."

[1468] As described above, this system comprehensively monitors the baby's condition and the user's emotions and takes appropriate action to support the baby's comfortable sleep and reduce the burden of childcare.

[1469] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1470] Step 1:

[1471] Data collection

[1472] The server obtains the baby's motion data in real time from the camera device. As input, it receives the video stream from the camera device and processes it with an image analysis algorithm. As output, it obtains motion data such as the baby's limb movements and rolling over. Specifically, the camera captures the baby's movements and sends the data to the server.

[1473] The server receives the room temperature data and the baby's body temperature data from the temperature sensor device. As input, it receives and analyzes the temperature data from the temperature sensor device. As output, it obtains specific numerical data about the room temperature and the baby's body temperature. For example, data that the room temperature is 25 degrees and the baby's body temperature is 37 degrees is sent to the server.

[1474] The server collects the baby's cry and other sounds from the audio device. As input, it receives the audio stream from the audio device and processes it with an audio analysis algorithm. As output, it obtains data about the decibel level and type of sound of the baby's cry. Specifically, if the baby is crying hard, the decibel level of the cry is sent to the server.

[1475] Step 2:

[1476] Recognizing user emotions with an emotion engine

[1477] The server analyzes the user's facial expressions using a camera device. As input, it receives the video stream from the camera device and processes it with a facial expression recognition algorithm. As output, it obtains data about the user's emotional state (e.g., smiling, angry, sad, etc.). Specifically, when the user's face is facing the camera, the server analyzes their facial expressions.

[1478] The server analyzes the tone of the user's voice using the audio device. As input, it receives the audio stream from the audio device and processes it with an audio tone analysis algorithm. As output, it obtains data about the user's emotional state (e.g., stress, fatigue, etc.). Specifically, if the user's voice is high-pitched and fast, it is determined that the user is stressed.

[1479] Step 3:

[1480] Data analysis and feedback

[1481] The server integrates and analyzes the collected baby's movement data, temperature data, voice data, and user's emotional data. As input, it receives the data obtained in each step above and processes it with a data integration algorithm. As output, it obtains judgment data regarding the baby's current state (e.g., crying, warm, etc.) and the user's emotional state. In a specific operation, assuming that the baby is crying, the room temperature is 28 degrees, and the user is feeling stressed, it will make the judgment that "the baby may be crying because of the heat."

[1482] Step 4:

[1483] Lullaby generation and playback

[1484] The server uses a music generation AI model to generate an optimal lullaby based on the baby's state and the user's emotions. The AI ​​model receives the prompt "My baby is excited. Please generate a lullaby that has a relaxing effect." The output is audio data of a lullaby with a relaxing effect. In concrete terms, the AI ​​generates a lullaby based on the prompt, and the server receives the data.

[1485] The device receives the generated lullaby audio data and sends it to the speaker in the baby's room. As input, it receives the audio data sent from the server and sends it to the speaker. As output, the lullaby is played from the speaker. Specifically, the smartphone is connected to the speaker via Bluetooth and the lullaby is played.

[1486] Step 5:

[1487] environmental adjustment

[1488] The server controls the air conditioner as needed to keep the room temperature within an appropriate range. As input, it receives room temperature data and processes it with the air conditioner control algorithm. As output, it obtains commands to send to the air conditioner. In concrete terms, the server sends a command to the air conditioner to "set the temperature to 24 degrees."

[1489] The server can also adjust the lighting in a room based on the user's emotional data. It receives the user's emotional data as input and processes it with a lighting control algorithm. The output is a command to send to the lighting. Specifically, the server sends a command to adjust the lighting to 30% brightness to help the user relax.

[1490] Step 6:

[1491] Notifications and Anomaly Detection

[1492] If the server determines that the baby has been sleeping quietly for a certain period of time (for example, 30 minutes), it sends a notification to the parent's communication device saying, "The baby has fallen asleep." As input, it receives and analyzes the baby's movement data. As output, it generates a notification message and sends it to the parent's device. Specifically, a push notification is sent to the smartphone, displaying the message, "The baby has fallen asleep."

[1493] If something abnormal occurs with the baby, the server quickly sends an alert to the parent's communication device. As input, it receives and analyzes the baby's temperature and breathing data. As output, it generates an alert message informing the parent of the abnormality and sends it to the parent's device. For example, if the body temperature exceeds 39 degrees, an alert saying "There is something abnormal with the baby" is sent to the smartphone.

[1494] (Application example 2)

[1495] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1496] In today's physical stores, customers often feel stressed, which causes a poor customer experience. It is also difficult for store staff to grasp customers' emotional state in real time and respond appropriately, making it a challenge to improve customer satisfaction. Furthermore, some stores do not adequately adjust the environment to maintain customer comfort.

[1497] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1498] In this invention, the server includes means for acquiring motion from a camera device, means for acquiring environmental and body temperature data from a temperature sensor device, means for acquiring audio from an audio device, means for analyzing the acquired data and determining the state, means for generating music according to the state using a music generation AI, means for playing the generated music, means for controlling an air conditioner as needed to adjust the temperature to maintain a comfortable environment, means for sending a notification to a notification terminal when the state changes, means for immediately sending an alert to the notification terminal when an abnormality occurs, means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize emotions, means for adjusting the environment according to the user's emotional state based on data from the audio device, means for the music generation AI to generate music with a high relaxing effect based on the emotional data, and means for controlling lighting to adjust the ambient light, thereby enabling real-time environmental adjustment according to the customer's emotional state.

[1499] A "camera device" is an electronic device for capturing video and image data.

[1500] "Movement" refers to physical movement or action.

[1501] A "temperature sensor device" is an electronic device for measuring the temperature of an environment or object.

[1502] "Environmental data" refers to data that represents the state or condition of the surroundings, such as temperature and light intensity.

[1503] "Body temperature data" refers to data obtained by measuring and recording the temperature of a living body.

[1504] An "audio device" is an electronic device for capturing sound, such as a microphone.

[1505] A "means for determining the status" is a method or system that analyzes and evaluates the current situation or conditions based on acquired data.

[1506] "Music generation AI" is a system that automatically generates music using artificial intelligence technology.

[1507] An "air conditioning device" is a device that adjusts temperature and humidity. Air conditioners are examples of this.

[1508] A "notification terminal" is a device that receives signals and messages, such as a smartphone.

[1509] An "alert" is a warning message or notification of a particular condition.

[1510] The "emotion engine" is a system that analyzes emotions from facial expressions and tone of voice.

[1511] "User" means any person who uses a system or service.

[1512] "Environmental control" is the act of controlling environmental conditions such as temperature and light to maintain or improve comfort.

[1513] "Music with a high relaxing effect" is music that is intended to relax the listener.

[1514] A "means for controlling lighting" is a system or method for adjusting the intensity or color temperature of light.

[1515] This invention is a system that monitors the emotional state of customers in real time in physical stores and adjusts the environment accordingly to improve customer experience. This system consists of a camera device, a temperature sensor device, a voice device, an emotion engine, a music generation AI, an air conditioning system, a lighting system, and a notification terminal.

[1516] System Configuration

[1517] The system includes the following major components:

[1518] 1. Camera Device

[1519] The camera device is used to capture customer behavior in real time, specifically, a common webcam such as the Logitech C920 can be used.

[1520] 2. Temperature sensor device

[1521] Temperature sensor devices are used to obtain environmental and customer body temperature data. For example, DHT22 sensors are available.

[1522] 3. Audio Devices

[1523] The audio device is used to capture the customer's voice. A high-sensitivity microphone such as a Blue Yeti can be used.

[1524] 4. Emotion Engine

[1525] The emotion engine uses data from cameras and audio devices to analyze customers' facial expressions and tone of voice to recognize emotions, using an AI model powered by OpenCV and TensorFlow.

[1526] 5. Music Generation AI

[1527] The music generation AI generates relaxing music based on the customer's emotional state, using AI frameworks such as Google Magenta.

[1528] 6. Air conditioner

[1529] An air conditioning device, such as a Nest Thermostat, to adjust the ambient temperature as needed.

[1530] 7. Lighting System

[1531] A lighting system for adjusting ambient light, compatible with smart lights such as Phillips Hue.

[1532] 8. Notification terminal

[1533] A device that sends notifications when a customer's condition changes or an abnormality occurs. A standard smartphone can be used.

[1534] Program processing overview

[1535] The server acquires real-time customer behavior data from the camera device, environmental and customer body temperature data from the temperature sensor device, and customer voice data from the audio device. Using these data, the emotion engine analyzes the customer's emotional state.

[1536] The acquired data is analyzed by a server to determine the current situation. The music generation AI generates relaxing music according to the customer's emotional state and plays it through the store's speakers. In addition, the air conditioning and lighting systems adjust the environment as needed.

[1537] Specific examples

[1538] For example, if a customer is feeling stressed, that information is collected from the camera and audio device. The emotion engine determines the emotional state as "stressed," and the music generation AI generates relaxing music. This music is played through speakers in the store, and the lighting is adjusted to a softer light. At the same time, the air conditioning is adjusted to an appropriate temperature.

[1539] Prompt Sentence Examples

[1540] "Generate appropriate relaxing music when the customer is feeling stressed. The text data corresponding to the customer's emotional state is as follows: 'Stress'"

[1541] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1542] Step 1:

[1543] The server acquires customer behavior data in real time from the camera device. The acquired video data (input) is used to analyze the customer's location, behavior patterns, etc., and the results (output) become preprocessed data for emotion recognition.

[1544] Step 2:

[1545] The server acquires temperature data of the environment and the customer from the temperature sensor device. The acquired temperature data (input) is analyzed by the server to determine the appropriate range of the environment temperature and the customer's body temperature (output).

[1546] Step 3:

[1547] The server captures the customer's voice from the audio device, and after noise filtering and voice analysis, the captured voice data (input) is used to detect the customer's tone of voice and emotional state (output).

[1548] Step 4:

[1549] The server uses an emotion engine to analyze data acquired from the camera device and audio device to determine the customer's emotional state. Based on the input data (video and audio data), the emotional state (e.g., stress, relaxation) is detected from the customer's facial expression, posture, and tone of voice, and the emotional state (output) is obtained.

[1550] Step 5:

[1551] Based on the output of the emotion engine, the server sends a prompt to the music generation AI. An example of a prompt is, "Please generate appropriate relaxing music for when the customer is feeling stressed." Based on the input prompt, the music generation AI generates (outputs) music with a high relaxing effect, and the music data is obtained.

[1552] Step 6:

[1553] The server then transmits the generated music data to speakers in the store, where it is played, providing an acoustic environment that responds to the customer's emotional state.

[1554] Step 7:

[1555] The server controls the air conditioner based on the environmental data. Based on the acquired temperature data (input), it sets the appropriate room temperature (output) and sends temperature adjustment instructions to the air conditioner. This maintains the optimal ambient temperature.

[1556] Step 8:

[1557] The server controls the lighting system to adjust the ambient light. Based on the output data (input) of the emotion engine, it sets the appropriate lighting settings (output) and sends dimming instructions to the lighting system. This provides a comfortable lighting environment.

[1558] Step 9:

[1559] If the server detects an abnormality, it immediately sends an alert to the notification terminal. Based on the abnormal data (input), it generates an appropriate warning message (output) and sends it to the notification terminal. This enables a quick response.

[1560] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1561] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1562] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1563] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1564] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1565] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1566] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1567] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery...

Claims

1. a means for acquiring baby movements from a camera device; means for acquiring environmental and baby temperature data from a temperature sensor device; a means for obtaining baby audio from an audio device; A means of analyzing the acquired data and determining the baby's condition, A means for generating lullabies according to the baby's condition using music generation AI; a means for playing the generated lullaby; A way to control the air conditioning as needed and adjust the temperature to keep your baby in a comfortable environment; A means of sending a notification to the parent's communication device when the baby falls asleep, and A means to instantly send an alert to the parent's communication device if something abnormal happens to the baby, A system including:

2. The system according to claim 1, wherein the system detects an abnormal condition of the baby based on the acquired data and sends an alert.

3. 10. The system of claim 1, wherein the system determines when the baby is asleep and sends a notification to the parent's communication terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A