System
A monitoring system using unused mobile devices for real-time childcare support addresses the challenge of parental time constraints and financial burden, providing efficient and cost-effective remote childcare monitoring.
Patent Information
- Application Number
- JP2024117323
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
Busy parents in dual-income and single-parent households face challenges in monitoring their infants and young children effectively due to time constraints, and purchasing dedicated monitoring devices is financially burdensome.
A monitoring system utilizing unused mobile communication devices as cameras, enabling real-time video capture, transmission, analysis, and notification, with features like push alerts, music playback, and voice responses to ensure childcare support.
The system allows efficient and cost-effective remote childcare monitoring, reducing the burden on parents while being environmentally friendly by reusing old devices.
Smart Images

Figure 2026016233000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's dual-income households and single-parent households, busy parents find it difficult to devote sufficient time to watching over and raising their children. Monitoring infants and young children in particular requires careful attention, creating a demand for remote monitoring. However, purchasing new dedicated monitoring cameras and devices is a significant financial burden. The objective of this invention is to provide a system that allows parents to efficiently monitor babies and children remotely using unused mobile communication devices, thereby reducing the burden of childcare. [Means for solving the problem]
[0005] The present invention provides a monitoring system that enables parents to remotely monitor their children by using a disused mobile communication device as a camera. Specifically, the system includes a camera means for capturing video using the disused mobile communication device, a communication means for transmitting the video in real time, a user terminal means for receiving and displaying the video via the communication means, an analysis means for analyzing the video and detecting the baby's movements and cries, a notification means for sending push notifications based on the detected information, a music playback means for receiving remote operation instructions from the user terminal and playing music, an automatic call means for making an automatic call when dangerous movements are detected, and a listening means for detecting when the child is speaking and responding gently. This allows for efficient and cost-effective childcare support.
[0006] "Disused mobile communication devices" are portable electronic devices such as old or spare smartphones or tablets that are no longer in use.
[0007] "Camera means" refers to hardware and software components with camera functionality that are used to capture video.
[0008] "Communication means" refers to a function including modules and protocols for sending and receiving data over a network.
[0009] "User terminal means" refers to a device such as a smartphone or tablet used by a parent or supervisor to remotely check video and perform operations.
[0010] "Analysis means" refers to software algorithms or hardware components that analyze video and audio data to detect specific movements or sounds.
[0011] The "notification means" is a system component for notifying designated receiving terminals of event information detected by the analysis means.
[0012] "Music playback means" refers to software and hardware components for playing a specified music file.
[0013] The "automatic call means" is a function for automatically playing a pre-defined voice message when a dangerous situation is detected.
[0014] The "listening means" is a system component that uses speech recognition and synthesis technology to sense what the child is saying in real time and provide an appropriate response. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention relates to a "monitoring and listening app" that allows parents to remotely watch over their babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals").
[0037] Embodiment of camera linkage function
[0038] First, the user installs a dedicated app on their device and configures the camera. The device captures real-time video through the camera and sends the video data to a server. The server then distributes this data as live streaming to the user's smartphone (hereinafter referred to as the "user device"). When the user launches the app on the user device, they can check on their baby's condition in real time from a remote location.
[0039] 1. Embodiment of push notification function
[0040] The device is equipped with a built-in microphone and camera, which are used to detect the baby's movements and cries. The detected data is processed within the device, and if a predetermined threshold is exceeded, an alert is generated. The generated alert information is sent to a server, which then sends a push notification to the user based on the alert information. This allows the user to immediately become aware of any abnormalities in the baby.
[0041] Embodiment of music playback function
[0042] When a user issues a music playback command via the app on their smartphone, the command is sent to the server, which then receives the command and transmits it to the device. The device then plays the specified music file, allowing the baby to listen to the music through the built-in speaker.
[0043] Embodiment of the hazard detection function
[0044] The device is equipped with an AI algorithm that analyzes video and audio data in real time. If it detects that the baby is making dangerous movements, the device will automatically play an audio message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[0045] An embodiment of the listening function
[0046] The device's microphone detects what the child is saying and collects voice data in real time. This voice data is converted into text on the device and analyzed by AI. Based on the analysis results, the AI generates an appropriate response and responds gently to the child through voice synthesis. This allows users to be there for their children even when they are busy.
[0047] Sensor Function Embodiment
[0048] The user installs dedicated sensors in the crib and other dangerous locations. The device collects data from these sensors and processes it as an alert if it detects any dangerous behavior (e.g., a fall or abnormal movement). The alert information is sent to the server, and the device simultaneously plays an appropriate warning sound to protect the baby. The server then sends a push notification to the user to notify them of the abnormality.
[0049] Market deployment implementation example
[0050] The server will offer the app free of charge to new users. The basic features are available for free, but if users want higher resolution video or additional security features, they can purchase premium services.
[0051] As described above, the present invention provides a monitoring system that utilizes unused mobile communication devices, reduces the burden on parents in raising their children, and is also environmentally friendly.
[0052] The processing flow will be explained below.
[0053] Camera linkage function processing steps
[0054] Step 1:
[0055] Users install a dedicated app on their unused devices, open the app, and configure the camera settings, which activates the device's camera.
[0056] Step 2:
[0057] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[0058] Step 3:
[0059] The terminal transmits the captured video data to the server in real time.
[0060] Step 4:
[0061] The server receives the video data and starts streaming it to the user terminal.
[0062] Step 5:
[0063] The user launches the app on the user device and checks the live video being streamed.
[0064] Push notification processing steps
[0065] Step 1:
[0066] The device uses a built-in microphone and camera to detect the baby's movements and cries.
[0067] Step 2:
[0068] The terminal analyzes the sensed data and generates alert information when a certain threshold is exceeded.
[0069] Step 3:
[0070] The terminal transmits the generated alert information to the server.
[0071] Step 4:
[0072] The server receives the alert information and sends a push notification to the user terminal.
[0073] Step 5:
[0074] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[0075] Processing steps for music playback function
[0076] Step 1:
[0077] The user sends a music playback instruction to the server from the user terminal via the app.
[0078] Step 2:
[0079] The server receives the user's instruction to play music and transmits it to the terminal.
[0080] Step 3:
[0081] The terminal plays the specified music file based on the instruction received from the server.
[0082] Step 4:
[0083] The device uses a built-in speaker to play music for your baby.
[0084] Hazard detection function processing steps
[0085] Step 1:
[0086] The device collects data in real time through a camera and microphone.
[0087] Step 2:
[0088] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[0089] Step 3:
[0090] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[0091] Step 4:
[0092] The terminal transmits danger alert information to the server.
[0093] Step 5:
[0094] The server receives the danger alert and sends a push notification to the user terminal.
[0095] Listening function processing steps
[0096] Step 1:
[0097] The device detects what the child is saying through a microphone.
[0098] Step 2:
[0099] The device converts the detected voice data into text and sends it to the AI.
[0100] Step 3:
[0101] The device generates an appropriate response based on the results analyzed by the AI.
[0102] Step 4:
[0103] The device uses voice synthesis to speak gentle words to the child.
[0104] Sensor function processing steps
[0105] Step 1:
[0106] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[0107] Step 2:
[0108] The device continuously collects data from the sensors.
[0109] Step 3:
[0110] The device analyzes the collected data and detects abnormal behavior.
[0111] Step 4:
[0112] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[0113] Step 5:
[0114] The server receives the anomaly alert and sends a push notification to the user terminal, which then plays an appropriate warning sound.
[0115] Go-to-market process steps
[0116] Step 1:
[0117] The server offers the app free of charge to new users.
[0118] Step 2:
[0119] Users can use the basic features free of charge.
[0120] Step 3:
[0121] The server will set up options that offer high-resolution video and additional security features for a fee.
[0122] Step 4:
[0123] If the user selects a paid service, the corresponding function is provided.
[0124] Example 1
[0125] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0126] Today, there are a wide variety of systems available to help parents safely monitor their children. However, many of these systems have drawbacks, such as being expensive or not being able to reuse old mobile communication devices. Furthermore, there are limited means for parents who are in remote locations to check on their children's status in real time and respond appropriately. For these reasons, there is a demand for a monitoring system that is easy to implement, makes effective use of old mobile communication devices, and can ensure the safety of children remotely.
[0127] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0128] In this invention, the server includes a video acquisition means, a video transmission means, a display means, a sensing means, an information notification means, a music playback means, a calling means, a response means, a data collection means, and a means for taking a quick response. This allows parents to use their unused mobile communication devices as surveillance cameras to remotely monitor their children's situations in real time and to respond immediately in the event of an abnormality.
[0129] "Video capture means" refers to technology that uses a camera or other image capture device to capture visual information in real time or at regular intervals.
[0130] "Video transmission means" refers to a technique for transmitting acquired video data to another device via a network.
[0131] "Display means" refers to a display or other display device for visually presenting received video data to a user.
[0132] "Sensing means" refers to technology that analyzes captured video and audio data to detect specific movements and acoustic patterns.
[0133] "Information notification means" refers to notification technology for notifying the user of abnormalities or important information detected by the sensing means.
[0134] "Music playback means" refers to technology including circuitry and software for playing music specified by the user.
[0135] "Hammer" refers to technology that automatically plays a pre-set voice message when a dangerous situation is detected.
[0136] "Response means" refers to technology that detects what the child is saying, generates an appropriate response, and replies in voice.
[0137] "Data collection means" refers to the technology used to collect data from installed sensors and detect abnormal behavior.
[0138] "Means for taking prompt action" refers to technology that uses information notification means to quickly notify the user and prompt them to take the necessary action.
[0139] This invention is a "monitoring and listening system" that allows parents to remotely monitor their children by reusing unused mobile communication devices. The system includes video acquisition means, video transmission means, display means, sensing means, information notification means, music playback means, calling means, response means, data collection means, and means for rapid response.
[0140] First, the user installs a dedicated application on an unused mobile communication device (hereafter referred to as the "device"). This application can be downloaded from the Google Play Store or Apple's App Store. After installing the application, the user configures the camera and enables the device to function as a surveillance camera.
[0141] The device captures real-time video data through the camera and sends it to the server via Wi-Fi or mobile data. The captured video is compressed using an H.264 encoder. The server then delivers the received video data to the user's smartphone as live streaming in an appropriate format (e.g., RTMP). The user can then view the real-time video on their smartphone using a dedicated application.
[0142] The device's built-in microphone and camera detect the baby's movements and cries. Data from these sensors is analyzed in real time using digital signal processing (DSP). Any detected abnormalities that exceed a predetermined threshold are sent to the server as an alert. The server generates a push notification based on the alert information and sends it to the user's smartphone via Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs). This allows the user to quickly learn of any abnormalities in their baby.
[0143] Users can also send instructions to play music from their smartphones through the application. These instructions are sent to the server, which then transmits them to the device. The device receives the instructions, plays the specified music file, and plays the music to the baby through the built-in speaker. The device's built-in media player is used to play the music.
[0144] Furthermore, the device uses AI algorithms to analyze video and audio data in real time, and if it detects dangerous movement, it will automatically play a voice message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[0145] When a child speaks to the device through the microphone, voice data is collected in real time, and the speech recognition engine converts the speech into text. The AI then analyzes the text, generates an appropriate response, and responds gently using the speech synthesis engine.
[0146] Users can also install dedicated sensors in the crib or other dangerous areas. Data from these sensors is collected by the device and processed as an alert when dangerous behavior is detected. The alert information is sent to the server, and the device plays an appropriate warning sound to protect the baby. The server also sends a push notification to the user.
[0147] To illustrate, consider the following prompt:
[0148] 1. "How do I check my baby's real-time video on my smartphone?"
[0149] 2. "How can I get a push notification to my smartphone when my baby cries?"
[0150] 3. "How do I send a music playback command from my smartphone and play music on the device?"
[0151] 4. "How can I detect dangerous baby movements, give voice warnings, and send notifications?"
[0152] 5. "How can I generate appropriate responses and respond to my child's speech using text-to-speech?"
[0153] 6. "How can I have my device sound an alarm and send a notification when the sensor installed in the crib detects dangerous behavior?"
[0154] As described above, this monitoring and listening system provides parents with an effective means of keeping their children safe even from a distance, reducing the burden of childcare. It is also environmentally friendly, as it can be reused from unused mobile communication devices.
[0155] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0156] Camera Linkage Feature Steps
[0157] Step 1: Install and configure the app
[0158] Input: Download a dedicated application to a device that the user no longer uses.
[0159] Specific operation: The user downloads the dedicated application from the Google Play Store or Apple's App Store and installs it on their device.
[0160] Output: The dedicated application is installed on the device and ready to launch.
[0161] Input: The user opens the app's settings screen and configures the camera to launch.
[0162] What it does: When you first open the app, it will show a setup wizard that asks for camera permissions, and you can agree to enable the camera.
[0163] Output: The camera is set to a state where it can be activated on the device.
[0164] Step 2: Capture and transmit real-time video
[0165] Input: Camera powers up and begins collecting real-time footage.
[0166] What it does: The device's camera will start recording real-time footage of your baby or child, and the footage will be compressed using an H.264 encoder.
[0167] Output: Compressed video data is generated.
[0168] Input: Prepare compressed video data for transmission.
[0169] How it works: The device uploads the captured video data to a server in real time via Wi-Fi or mobile data, using HTTP / 2 or WebSocket protocols.
[0170] Output: The compressed video data is sent to the server.
[0171] Step 3: Streaming real-time video
[0172] Input: The server prepares to deliver the received video data to the user's device as live streaming.
[0173] Specific operation: The server converts the video data into an appropriate format and streams it to the user's device using RTMP (Real-time Messaging Protocol).
[0174] Output: The converted video data is delivered to the user's device.
[0175] Input: The user launches the app on their device and checks the real-time video.
[0176] Specific operation: By launching the app on the user's device and selecting the Live Streaming tab, real-time video will be played. The video will be decoded and displayed in a format that is easy for the user to view.
[0177] Output: User can check real-time video.
[0178] Push notification feature steps
[0179] Step 1: Sensing and processing data
[0180] Input: The device's built-in microphone and camera detect baby movements and cries.
[0181] How it works: The device's microphone constantly monitors the surrounding environment, the camera detects movement, and the data is analyzed in real time through digital signal processing (DSP).
[0182] Output: Detected movement and sound data is generated.
[0183] Input: Prepares sensed data to be analyzed.
[0184] How it works: If the baby's crying exceeds a certain decibel level or if any violent movement is detected, the AI module in the device will generate an alert.
[0185] Output: Alert information is generated.
[0186] Step 2: Sending alert information and notifications
[0187] Input: Prepares the generated alert information to be sent to the server.
[0188] Specific operation: The device sends alert information to the server via Wi-Fi or mobile data communication, using the MQTT (Message Queuing Telemetry Transport) protocol.
[0189] Output: Alert information is sent to the server.
[0190] Input: The server prepares to generate a push notification based on the alert information.
[0191] Specific operation: The server analyzes the alert information, generates a push notification, and sends it to the user's device using Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs).
[0192] Output: A push notification is generated and sent to the user device.
[0193] Music playback function steps
[0194] Step 1: Sending a music play command
[0195] Input: The user commands music playback through the app on their smartphone.
[0196] Specific operation: The user selects the music they want to play from the app's music playback menu and taps the play button. Instruction data is generated.
[0197] Output: Music playback instructions are generated.
[0198] Input: Prepares to send music playback instruction data to the server.
[0199] Specific operation: Instruction data generated from the user terminal is sent to the server via the Internet. HTTP / 2 is used as the communication protocol.
[0200] Output: Instruction data is sent to the server.
[0201] Step 2: Transmitting and executing music playback instructions
[0202] Input: The server receives the instruction data and prepares it for transmission to the terminal.
[0203] Specific operation: The server analyzes the received instruction data and sends a playback instruction to the corresponding terminal.
[0204] Output: Playback instructions are sent to the device.
[0205] Input: Prepares the device to play the specified music file.
[0206] Specific operation: The device retrieves and plays the music file from the internal storage or cloud storage based on the received playback command. The device's built-in media player is used to play the audio.
[0207] Output: The specified music file will be played.
[0208] Hazard detection function steps
[0209] Step 1: Analyzing video and audio data
[0210] Input: The device collects video and audio data and analyzes it in real time using AI algorithms.
[0211] Specific operation: The device's AI chip processes video and audio data in real time to detect dangerous movements of the baby.
[0212] Output: Data is generated indicating the dangerous behavior detected.
[0213] Step 2: Play and notify warning messages
[0214] Input: Prepare to play a warning message when dangerous motion is detected.
[0215] Specific operation: Based on the AI detection results, the device plays a pre-set audio message to warn the user.
[0216] Output: A warning audio message is played.
[0217] Input: Sends the danger detection information to the server, which prepares to generate a push notification.
[0218] Specific operation: The danger detection information is sent to the server, and the server generates a push notification and sends it to the user.
[0219] Output: A push notification is generated and sent to the user.
[0220] Steps of active listening function
[0221] Step 1: Detecting speech and collecting voice data
[0222] Input: Detects what your child is saying through the device's microphone.
[0223] How it works: The device's microphone detects what the child is saying and collects audio data, which is then stored in a buffer on the device in real time.
[0224] Output: Speech audio data is collected.
[0225] Step 2: Speech to text conversion and analysis
[0226] Input: Prepare the collected audio data for conversion to text.
[0227] How it works: The device's voice recognition engine converts the voice data into text, and the AI module analyzes the text to generate an appropriate response.
[0228] Output: Parsed text data and response data are generated.
[0229] Step 3: Response generation and speech synthesis
[0230] Input: Converts response data into audio and prepares it for playback.
[0231] Specific operation: The device's speech synthesis engine outputs the generated response as voice and plays it through the speaker.
[0232] Output: An audio response is played to the child.
[0233] Sensor function steps
[0234] Step 1: Sensor installation and data collection
[0235] Input: The user installs special sensors in the crib or other dangerous locations.
[0236] Specific operation: The user places a special sensor in a crib or dangerous location, and the sensor begins to function.
[0237] Output: The sensor is deployed and ready to collect data.
[0238] Input: Prepare to collect data from sensors.
[0239] Specific operation: The device receives data sent from the sensor and detects abnormal movement or falls.
[0240] Output: Data is generated indicating the abnormal behavior detected.
[0241] Step 2: Alerting and Notification
[0242] Input: Prepare to treat anomalous behavior as an alert.
[0243] Specific operation: When the device detects danger, it generates alert information and sends it to the server.
[0244] Output: Alert information is generated and sent to the server.
[0245] Input: The server prepares to generate a push notification based on the alert information.
[0246] What it does: The server sends a push notification to the user, and the device plays a series of warning sounds to protect the baby.
[0247] Output: A push notification is generated and sent to the user. An alert sound is played on the device.
[0248] Market expansion steps
[0249] Step 1: Offer your app and define your market
[0250] Input: The server prepares the app to serve to a new user.
[0251] What happens: The server publishes the app's download link on a website or store, making it free for new users.
[0252] Output: The app download link is published and new users can download the app.
[0253] Step 2: Offering premium services
[0254] Input: The server defines the content of the paid premium service and prepares to provide it to the user.
[0255] Specific operation: Allow users to purchase premium services within the app, and after purchase, the server provides the corresponding services.
[0256] Output: The user purchases a premium service, unlocking additional features.
[0257] (Application example 1)
[0258] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0259] To improve vehicle safety, including that of autonomous vehicles and their drivers, it is necessary to constantly monitor the driver's state and detect abnormalities or dangers early. However, conventional systems lack the ability to detect driver drowsiness or inattention in real time and send warnings. There is also a need for an efficient method of reusing unused mobile communication devices. A system that solves these issues and ensures safety while being economical and environmentally friendly is needed.
[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0261] In this invention, the server is a monitoring system that allows parents to remotely monitor their children by using an unused mobile communication device as a camera, and includes: a camera means for acquiring video; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video; an analysis means for analyzing the video from the mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the detected information; a music playback means for receiving remote operation instructions from the user terminal and playing music; an automatic call means for issuing an automatic call when dangerous movement is detected; a listening means for detecting a child's speech and responding gently; a driver monitoring means for monitoring the driver's state using a camera installed in the vehicle and detecting drowsiness or inattention; and a warning means for sending a warning to the driver and a manager when drowsiness or inattention is detected. This makes it possible to monitor the driver's state in real time and respond immediately if an abnormality is detected.
[0262] A "mobile communications device" is a device used for mobile communications, including smartphones and tablets.
[0263] "Camera means" refers to the camera equipment or functionality used to capture the image.
[0264] "Communication means" refers to communication technology or devices for sending and receiving data or information.
[0265] "User terminal means" refers to a terminal device used by a user, and is a device used for displaying images and remote control.
[0266] "Analysis means" refers to a method or device for analyzing video and audio data and detecting specific movements or situations.
[0267] A "notification means" is a method or device for sending a notification to a user or administrator based on sensed information.
[0268] "Music playback means" refers to a method or device for playing music, including speakers and music playback applications.
[0269] "Automatic calling means" refers to a method or device that automatically generates and plays voice messages under certain circumstances.
[0270] "Listening means" refers to a method or device for detecting and responding to the voice of a subject.
[0271] "Driver monitoring means" refers to a method or device that includes cameras or sensors for monitoring the situation inside the vehicle and, in particular, for checking the driver's condition.
[0272] "Warning means" refers to a method or device for issuing an immediate warning when an abnormality is detected.
[0273] This invention relates to a "monitoring and listening app" that allows parents to remotely monitor their babies and children using unused mobile communication devices. Furthermore, by applying this technology to autonomous vehicles, a system can be realized that monitors the driver and immediately issues a warning in the event of an abnormal or dangerous situation.
[0274] System Configuration
[0275] The server works in conjunction with mobile communication devices, user terminals, sensors, etc. These devices exchange data and analyze it to achieve the following functions:
[0276] Driver status monitoring
[0277] The device is equipped with a camera that captures images of the interior of the vehicle in real time and sends them to a server. This video data is analyzed by the server to detect driver drowsiness and inattention. This detection uses an AI algorithm, specifically TensorFlow, for face detection and movement analysis.
[0278] Warning function
[0279] If an abnormality is detected in the driver, the server will send a warning message to the administrator through the Twilio API, and will use the Google Text-to-Speech (gTTS) library to generate a voice warning and play the warning message aloud, using the pygame library for speech synthesis in this process.
[0280] Music playback function
[0281] The server receives instructions from the user's device and sends a music playback instruction to the device, which then plays the specified music file through its built-in speaker.
[0282] Data analysis
[0283] The video data sent from the device is analyzed by the server. The analysis method is to detect and analyze the driver's face and movements using OpenCV and TensorFlow. If drowsiness or inattention is detected, it is immediately treated as an abnormality and a warning is issued.
[0284] Specific examples
[0285] For example, if a driver falls asleep while driving on a highway, the device's camera will detect this in real time. If the server detects the driver's drowsiness through AI analysis, the system will send an SMS notification to the administrator via the server and simultaneously play an audio warning on the device.
[0286] Prompt Sentence Examples
[0287] "Develop an app that uses mobile communication devices to monitor the driver's status in real time. If the driver becomes inattentive or falls asleep, it should issue a voice warning and send an SMS notification to the manager."
[0288] In this way, the present invention provides a system that can seamlessly monitor and warn drivers, and by reusing unused mobile communication devices, it builds an efficient system that is environmentally friendly.
[0289] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0290] Step 1:
[0291] The device starts the camera and captures the video inside the car. In this step, the camera acquires video data in real time. The input is the video inside the car, and the output is the captured video data. Specifically, the device starts the camera using cv2.VideoCapture(0) and acquires frames using the read() method.
[0292] Step 2:
[0293] The video data captured by the terminal is sent to the server in real time. In this step, data is sent using a communication means. The input is the captured video data, and the output is the video data sent to the server. In concrete terms, the terminal sends the video data to the server using a specified communication protocol (e.g., HTTP POST request).
[0294] Step 3:
[0295] The server analyzes the received video data. In this step, the analysis means is used to determine the driver's state. The input is the transmitted video data, and the output is the analysis result. Specifically, the server processes the video data using OpenCV and TensorFlow to analyze the driver's face and movements.
[0296] Step 4:
[0297] If the server detects an anomaly, it generates a warning message. In this step, the warning means operates based on the analysis results. The input is the analysis result, and the output is the warning message. Specifically, when an anomaly is detected, the server generates a voice warning using the gTTS library and sends an SMS notification using the Twilio API.
[0298] Step 5:
[0299] The terminal plays music according to instructions from the server. In this step, a music playback means is used. The input is a remote control instruction from the user terminal, and the output is the music being played. In concrete terms, the terminal plays music on its built-in speaker based on the instruction received from the server.
[0300] Step 6:
[0301] The device detects what the child is saying and responds appropriately. In this step, the listening means is activated. The input is what the child is saying, and the output is a voice response. Specifically, the device detects the voice through the microphone and sends it to the server. The server converts it into text using a generative AI model and generates an appropriate response.
[0302] Step 7:
[0303] The server comprehensively manages all data and sends notifications to user devices in the event of an abnormality. This step is responsible for overall control. Input is data from various sensors and cameras, and output is notifications to administrators and push notifications to user devices. Specifically, the server receives and analyzes all data, and immediately notifies users if an abnormality is detected.
[0304] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0305] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely watch over babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals"), with an emotion engine that recognizes the user's emotions.
[0306] Embodiment of camera linkage function
[0307] First, the user installs a dedicated app on the device and configures the camera. This activates the device's camera. The device captures real-time video through the camera and transmits the video data to the server. The server then streams the received data to the user's smartphone (hereinafter referred to as the "user device"). The user can then view the live video on their own device.
[0308] 1. Embodiment of push notification function
[0309] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device. The user can then respond promptly via the push notification.
[0310] Embodiment of music playback function
[0311] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server. The device then plays the specified music and plays it to the baby through the built-in speaker.
[0312] Embodiment of the hazard detection function
[0313] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, danger alert information is sent to the server, which then warns the user via push notification.
[0314] An embodiment of the listening function
[0315] The device detects what the child is saying through a microphone and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to gently reply to the child, allowing the user to engage with their child even when they are busy.
[0316] Sensor Function Embodiment
[0317] The user installs dedicated sensors in the crib or other dangerous locations and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The detected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[0318] Embodiment of Emotion Engine
[0319] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The response is converted from text to speech and played through the device's speaker.
[0320] This allows the emotion engine to reduce the user's mental burden and more effectively support the care of babies and children. The combination of the emotion engine improves the functionality and flexibility of the entire system, making it an important tool for supporting childcare.
[0321] The processing flow will be explained below.
[0322] Camera linkage function processing steps
[0323] Step 1:
[0324] Users install a dedicated app on their unused devices and configure the camera settings, which activates the device's camera.
[0325] Step 2:
[0326] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[0327] Step 3:
[0328] The terminal transmits the captured video data to the server in real time.
[0329] Step 4:
[0330] The server receives the video data and starts streaming it to the user terminal.
[0331] Step 5:
[0332] The user launches the app on the user device and checks the live video being streamed.
[0333] Push notification processing steps
[0334] Step 1:
[0335] The device's built-in microphone and camera are used to detect the baby's movements and cries.
[0336] Step 2:
[0337] The terminal analyzes the sensed data in real time and generates alert information when a certain threshold is exceeded.
[0338] Step 3:
[0339] The terminal transmits the generated alert information to the server.
[0340] Step 4:
[0341] The server receives the alert information and sends a push notification to the user terminal.
[0342] Step 5:
[0343] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[0344] Processing steps for music playback function
[0345] Step 1:
[0346] The user sends a music playback instruction to the server from the user terminal via the app.
[0347] Step 2:
[0348] The server receives the user's instruction to play music and transmits it to the terminal.
[0349] Step 3:
[0350] The terminal plays the specified music file based on the instruction received from the server.
[0351] Step 4:
[0352] The device uses a built-in speaker to play music for your baby.
[0353] Hazard detection function processing steps
[0354] Step 1:
[0355] The device collects data in real time through a camera and microphone.
[0356] Step 2:
[0357] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[0358] Step 3:
[0359] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[0360] Step 4:
[0361] The terminal transmits danger alert information to the server.
[0362] Step 5:
[0363] The server receives the danger alert and sends a push notification to the user terminal.
[0364] Listening function processing steps
[0365] Step 1:
[0366] It detects what the child is saying through the device's microphone.
[0367] Step 2:
[0368] The device converts the detected voice data into text and sends it to the AI.
[0369] Step 3:
[0370] The device generates an appropriate response based on the results analyzed by the AI.
[0371] Step 4:
[0372] The device uses voice synthesis to speak soothing words to the child.
[0373] Sensor function processing steps
[0374] Step 1:
[0375] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[0376] Step 2:
[0377] The device continuously collects data from connected sensors.
[0378] Step 3:
[0379] The device analyzes the collected data and detects abnormal behavior.
[0380] Step 4:
[0381] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[0382] Step 5:
[0383] The server receives the anomaly alert and sends a push notification to the user device, which simultaneously plays an appropriate warning sound.
[0384] Emotion Engine Processing Steps
[0385] Step 1:
[0386] The user's voice data is transmitted from the user terminal to the server.
[0387] Step 2:
[0388] The server sends the received voice data to the emotion engine on the device.
[0389] Step 3:
[0390] The device's emotion engine analyzes the voice data and determines the user's emotional state.
[0391] Step 4:
[0392] The terminal's emotion engine generates an appropriate response based on the determined emotional state.
[0393] Step 5:
[0394] The terminal converts the generated response from text to speech and provides feedback to the user.
[0395] Example 2
[0396] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0397] In modern child-rearing, it is difficult for parents to constantly keep an eye on their babies and children, and the burden is particularly heavy when both parents are working or raising children alone. Conventional monitoring systems have limitations in real-time monitoring and voice communication, and lack emotion recognition functionality, making it difficult to reduce the mental burden on parents. Furthermore, they lack the ability to instantly detect dangerous situations and respond appropriately. The purpose of this invention is to solve these problems, more effectively ensure the safety of babies and children, and reduce the burden on parents.
[0398] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a camera means for capturing video from an unused mobile communication device; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video via the communication means; an analysis means for analyzing the video from the unused mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the information detected by the analysis means; a music playback means for receiving remote operation instructions from the user terminal and playing music; an emotion analysis means for analyzing audio data transmitted from the unused mobile communication device and determining the user's emotional state; an automatic call means for issuing an automatic call when dangerous movements are detected; and a listening means for detecting and gently responding to the child's voice. This allows for a multifunctional and flexible monitoring system to be constructed, enabling parents to monitor and respond to the condition of babies and children even remotely. Furthermore, emotion analysis reduces the user's mental burden and enables quick and appropriate responses to dangerous situations.
[0399] "Disused mobile communication devices" are mobile devices such as mobile phones and tablets that have been retired but are now being reused.
[0400] "Camera means" refers to a camera device installed to capture video.
[0401] "Communication means" refers to a means having an internet connection function for sending and receiving images and data.
[0402] "User terminal means" refers to a display device used by a user, such as a smartphone, tablet, or computer.
[0403] The "analysis means" is a means having the function of analyzing the collected data and detecting specific information such as the baby's movements and cries.
[0404] The "notification means" is a means having a function for generating a push notification based on sensed information and transmitting it to the user terminal.
[0405] The "music playback means" is a means having a function of playing music in response to an instruction from a user terminal.
[0406] The "emotion analysis means" is a means having a function of analyzing voice data and determining the emotional state of the user.
[0407] The "automatic calling means" is a means having a function of automatically playing a warning message or the like when a dangerous movement is detected.
[0408] "Listening tools" are tools that have the ability to sense what a child is saying, generate an appropriate response, and respond in a gentle manner.
[0409] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely monitor babies and children by reusing unused mobile communication devices with an emotion engine that recognizes the user's emotions. This system provides a multifunctional and flexible monitoring environment.
[0410] Camera linkage function
[0411] First, the user installs a dedicated app on their device and configures the camera. Once the app is installed and configured, the device's camera is activated. The device captures real-time video through the camera and sends the data to the server. The server receives the video data and streams it to the user's device, allowing the user to view live video on their device.
[0412] Specific examples
[0413] Device used: Old smartphone
[0414] Server used: AWS EC2 instance
[0415] Video streaming software: FFmpeg
[0416] Push notification function
[0417] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device, allowing the user to respond quickly.
[0418] Specific examples
[0419] Hardware used: Android or iOS device
[0420] Server: Google Firebase Cloud Messaging
[0421] App: Real-time analytics with custom mobile app
[0422] Music playback function
[0423] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server, and the device plays the specified music on its built-in speaker.
[0424] Specific examples
[0425] Device used: Old tablet
[0426] Server: AWS Lambda
[0427] Music playback software: VLC
[0428] Hazard detection function
[0429] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, the device sends danger alert information to the server, which then sends a push notification to the user's device.
[0430] Specific examples
[0431] Device used: Raspberry Pi 4 with camera module
[0432] AI algorithm: TensorFlow
[0433] Server: Google Cloud Functions
[0434] Active listening function
[0435] The device detects what the child is saying through a microphone and converts the voice data into text. This text data is analyzed using an AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The generated response is then converted into audio data using speech synthesis technology, which is then played back by the device.
[0436] Specific examples
[0437] Device used: Amazon Echo device
[0438] Speech Recognition: Google Speech-to-Text API
[0439] Response generation: OpenAI GPT-3
[0440] Speech synthesis: Amazon Polly
[0441] Sensor function
[0442] The user installs dedicated sensors in the crib or other dangerous areas and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The collected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[0443] Specific examples
[0444] Sensor used: Xiaomi Smart Home Sensor
[0445] Device used: Old smartphone
[0446] Server: AWS IoT Core
[0447] Emotion Engine
[0448] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The generated response is converted into voice data and played through the device's speaker.
[0449] Specific examples
[0450] Device used: Old smartphone
[0451] Emotion engine: IBM Watson Tone Analyzer
[0452] Speech synthesis: Google Text-to-Speech
[0453] This makes the entire system a versatile and flexible monitoring environment, making it a very useful tool for supporting childcare. Users can monitor the condition of their baby or child in real time, even remotely, and respond quickly if necessary. The emotion engine also reduces mental strain.
[0454] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0455] Camera linkage function processing steps
[0456] Step 1:
[0457] Users install a dedicated app on their device, open the app and configure the camera settings.
[0458] Input: App installation, setting information
[0459] Output: Setup complete, camera ready to start
[0460] Step 2:
[0461] Once the setup is complete, the device will launch the camera.
[0462] Input: Setting completion notification
[0463] Output: Camera starts, video capture begins
[0464] Step 3:
[0465] The device captures real-time video through a camera and transmits the data to a server.
[0466] Input: Real-time video data
[0467] Output: Ready to send, sent to server
[0468] Step 4:
[0469] The server processes the received video data and distributes it as a stream to the user terminal.
[0470] Input: Video data
[0471] Output: Video streaming ready, sent to user device
[0472] Step 5:
[0473] Users can view live footage via the app on their devices.
[0474] Input: Video streaming data
[0475] Output: Video display, real-time monitoring
[0476] Push notification processing steps
[0477] Step 1:
[0478] It analyzes data collected from the device's built-in microphone and camera in real time.
[0479] Input: Audio data, video data
[0480] Output: Analysis results, threshold judgment
[0481] Step 2:
[0482] If the baby's movements or cries exceed the threshold, the device will generate an alert.
[0483] Input: Analysis results
[0484] Output: Alert information
[0485] Step 3:
[0486] The terminal transmits the generated alert information to the server.
[0487] Input: Alert information
[0488] Output: Sent to server
[0489] Step 4:
[0490] The server generates a push notification based on the alert information and sends it to the user's device.
[0491] Input: Alert information
[0492] Output: Push notification
[0493] Step 5:
[0494] Users receive push notifications and can respond quickly.
[0495] Input: Push notification
[0496] Output: Notification confirmation, required action
[0497] Processing steps for music playback function
[0498] Step 1:
[0499] The user operates the app on their smartphone and sends instructions to play music.
[0500] Input: Music playback instructions
[0501] Output: Send instructions
[0502] Step 2:
[0503] The server transfers the playback instruction received from the user to the terminal.
[0504] Input: Playback instructions
[0505] Output: Directed transfer
[0506] Step 3:
[0507] The device will play music through its built-in speaker based on the transferred instructions.
[0508] Input: Playback instructions
[0509] Output: Music playback
[0510] Hazard detection function processing steps
[0511] Step 1:
[0512] The device collects the baby's movements and voice data using a camera and microphone.
[0513] Input: Video data, audio data
[0514] Output: Recorded data, audio data
[0515] Step 2:
[0516] The device analyzes the collected data using AI algorithms to detect dangerous behavior.
[0517] Input: Recorded data, audio data
[0518] Output: Risk assessment result
[0519] Step 3:
[0520] When danger is detected, the device will play an automated voice message such as "Danger! Stop!"
[0521] Input: Risk assessment result
[0522] Output: Automated voice message
[0523] Step 4:
[0524] The terminal transmits danger alert information to the server.
[0525] Input: Risk assessment result
[0526] Output: Alert information
[0527] Step 5:
[0528] The server generates a push notification and sends it to the user device.
[0529] Input: Alert information
[0530] Output: Push notification
[0531] Listening function processing steps
[0532] Step 1:
[0533] The device detects what the child is saying through a microphone.
[0534] Input: Audio data
[0535] Output: Sensor information
[0536] Step 2:
[0537] The device converts the detected voice data into text.
[0538] Input: Audio data
[0539] Output: Text data
[0540] Step 3:
[0541] The converted text data is analyzed by an AI model to generate an appropriate response.
[0542] Input: Text data
[0543] Output: Response data
[0544] Step 4:
[0545] The terminal converts the generated response into voice data using voice synthesis technology.
[0546] Input: Response data
[0547] Output: Audio data
[0548] Step 5:
[0549] The device plays back the audio data and provides an appropriate response to the child.
[0550] Input: Audio data
[0551] Output: Response voice
[0552] Sensor function processing steps
[0553] Step 1:
[0554] Users install special sensors in cribs and dangerous areas and connect them to a terminal.
[0555] Input: Sensor installation
[0556] Output: Connection complete
[0557] Step 2:
[0558] Sensors continuously collect environmental data and detect abnormal behavior.
[0559] Input: Environment data
[0560] Output: Data collection, anomaly detection information
[0561] Step 3:
[0562] The sensor transmits the collected data to the terminal.
[0563] Input: Data collection, anomaly detection information
[0564] Output: Send to terminal
[0565] Step 4:
[0566] The terminal analyzes the received sensor data internally and detects any abnormalities.
[0567] Input: Sensor data
[0568] Output: Abnormality detection result
[0569] Step 5:
[0570] The device generates an alert if an anomaly is detected.
[0571] Input: Abnormality detection result
[0572] Output: Alert information
[0573] Step 6:
[0574] The terminal transmits the alert information to the server.
[0575] Input: Alert information
[0576] Output: Sent to server
[0577] Step 7:
[0578] The server generates a push notification and sends it to the user device.
[0579] Input: Alert information
[0580] Output: Push notification
[0581] Emotion Engine Processing Steps
[0582] Step 1:
[0583] The user transmits voice data to the terminal.
[0584] Input: Audio data
[0585] Output: Send to terminal
[0586] Step 2:
[0587] The device uses an emotion engine to analyze the voice data and determine the user's emotional state.
[0588] Input: Audio data
[0589] Output: Emotion determination result
[0590] Step 3:
[0591] The emotion engine generates an appropriate response based on the judgment result.
[0592] Input: Emotion judgment result
[0593] Output: Response data
[0594] Step 4:
[0595] The terminal converts the generated response into voice data using voice synthesis technology.
[0596] Input: Response data
[0597] Output: Audio data
[0598] Step 5:
[0599] The terminal plays back the audio data and responds appropriately to reduce the user's mental burden.
[0600] Input: Audio data
[0601] Output: Response voice
[0602] (Application example 2)
[0603] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0604] The problem that the present invention aims to solve is to provide a system that effectively utilizes unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children and increase their safety. Furthermore, because conventional monitoring systems do not take emotion recognition into consideration, a means for reducing the mental burden on users was also necessary. Therefore, the present invention aims to compensate for the shortcomings of conventional monitoring systems and improve safety and usability.
[0605] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a communication means for transmitting video in real time, an emotion recognition means for analyzing audio data and recognizing emotions, and a response generation means for generating an appropriate response based on the recognized emotion and outputting it as audio. This makes it possible not only to perform remote real-time monitoring, but also to recognize the emotional state of the user and generate an appropriate response.
[0606] The "camera means" is a means having a function for acquiring an image.
[0607] The "communication means" is a means having a function for transmitting acquired video in real time.
[0608] The "user terminal means" is a means having a function of receiving and displaying video transmitted via a communication means.
[0609] The "analysis means" is a means having the function of analyzing video acquired from a disused mobile communication device and detecting specific movements or sounds.
[0610] The "notification means" is a means having a function of sending a push notification to the user terminal based on the information sensed by the analysis means.
[0611] The "music playback means" is a means having a function of receiving a remote operation instruction from a user terminal and playing music.
[0612] The "automatic call means" is a means having a function of automatically playing back a voice message when a dangerous movement is detected.
[0613] "Listening means" is a means that has the function of detecting what a child is saying and responding gently.
[0614] The "emotion recognition means" is a means having a function of analyzing voice data and recognizing the emotional state of the user.
[0615] The "response generation means" is a means having a function of generating an appropriate response based on the recognized emotion and outputting the response as voice.
[0616] The present invention relates to a system that makes effective use of unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children. This system uses devices and software that include the following elements:
[0617] First, a disused mobile communication device (hereinafter referred to as a terminal) captures video using a camera means. The terminal transmits the video captured by the camera to a server in real time via a communication means. The server can stream the received video to a user terminal means. The user can then view the live video on their own terminal.
[0618] Data collected from the device's built-in microphone and camera is analyzed in real time by the analysis means. This analysis detects the movement and cries of babies and children, and if a threshold is exceeded, an alert is generated by the notification means. These alerts are sent to the user's device via the server as push notifications. The user can respond promptly via the push notifications.
[0619] Furthermore, the user can remotely use the music playback means to send instructions to the terminal to play music. This instruction is transferred to the terminal via the server, and the specified music is played by the baby or child.
[0620] Furthermore, the device is equipped with a danger detection function. Data obtained from the camera and microphone is analyzed using an AI algorithm to detect if the baby is behaving in a dangerous manner. If danger is detected, the device will play an automated voice message such as "Danger! Stop!" At the same time, danger alert information is sent to the server, and a warning is sent to the user.
[0621] The device also has a listening function, which uses a microphone to detect what the child is saying and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to respond gently to the child, allowing the user to interact with their child appropriately even when they are busy.
[0622] The device is also equipped with an emotion recognition means that analyzes the voice data sent from the user device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion recognition means will recognize this and utilize the response generation means to advise an appropriate response. This reduces the user's mental burden and provides support for childcare and monitoring.
[0623] For example, if an employee of Guard AI is wearing smart glasses, the following prompt can trigger the emotion engine and provide appropriate advice:
[0624] Example prompt sentence:
[0625] Employee dictation: "I'm feeling really stressed today..."
[0626] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[0627] In this way, the system of the present invention can achieve both enhanced security and psychological care for the user.
[0628] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0629] Step 1:
[0630] The terminal acquires video using a camera means. The input is real-time video data from the camera, which is captured and stored in a buffer. The output is the captured video data.
[0631] Step 2:
[0632] The device uses a communication means to send the captured video data to the server. The input is the captured video data, which is then encoded into JPEG or MP4 format. The output is the encoded video data.
[0633] Step 3:
[0634] The server streams the received video data to the user terminal means. The input is the video data sent from the terminal and is transferred to the user terminal in real time. The output is the live video displayed on the user terminal.
[0635] Step 4:
[0636] The data collected from the device's built-in microphone and camera is analyzed by an analysis means. The input is audio and video data from the microphone and camera, which is fed into an AI algorithm for analysis. The output is the analysis results, which are the movements and cries of the baby or child.
[0637] Step 5:
[0638] If the analysis result exceeds a certain threshold, an alert is generated by the notification mechanism. The input is the analysis result, which is used to evaluate the alert condition. The output is the generated alert.
[0639] Step 6:
[0640] The server sends the generated alert to the user device as a push notification. The input is the generated alert, which is converted into a push notification format. The output is a push notification that is displayed on the user device.
[0641] Step 7:
[0642] The user remotely uses the music playback means to send a music playback instruction to the terminal. The input is the music playback instruction from the user terminal, which is sent to the terminal via the server. The output is the music played on the terminal.
[0643] Step 8:
[0644] When the terminal detects a dangerous movement, it plays a voice message using the automatic call means. The input is the detection result of the analysis means, and if it is judged to be dangerous, the voice message is played. The output is a warning voice played from the terminal.
[0645] Step 9:
[0646] The terminal uses a listening means to detect what the child is saying and converts the voice data into text. The input is the voice data of the child's speech, which is converted into text data. The output is text data.
[0647] Step 10:
[0648] The text data is analyzed by AI to generate an appropriate response. The input is the converted text data, and the natural language processing algorithm generates a response based on this. The output is the response text.
[0649] Step 11:
[0650] The response generation means outputs the generated response text as voice using speech synthesis technology. The input is the response text, which is converted into voice data by a speech synthesis engine. The output is the response voice played from the terminal.
[0651] Step 12:
[0652] The emotion recognition means analyzes the voice data sent from the user terminal and determines the user's emotional state. The input is the user's voice data, which is supplied to a voice analysis algorithm to recognize the emotional state. The output is the determination result of the emotional state.
[0653] Step 13:
[0654] Based on the recognized emotional state, the response generator generates an appropriate response. The input is the result of the emotional state determination, and the response content is determined based on this. The output is the generated response content.
[0655] Step 14:
[0656] The response is output as voice through the device's speaker. The input is the generated response, which is played back as voice using speech synthesis technology. The output is a voice response played from the device.
[0657] Example prompt sentence:
[0658] Employee dictation: "I'm feeling really stressed today..."
[0659] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[0660] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0661] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0662] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0663] [Second embodiment]
[0664] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0665] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0666] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0667] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0668] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0669] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0670] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0671] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0672] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0673] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0674] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0675] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0676] This invention relates to a "monitoring and listening app" that allows parents to remotely watch over their babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals").
[0677] Embodiment of camera linkage function
[0678] First, the user installs a dedicated app on their device and configures the camera. The device captures real-time video through the camera and sends the video data to a server. The server then distributes this data as live streaming to the user's smartphone (hereinafter referred to as the "user device"). When the user launches the app on the user device, they can check on their baby's condition in real time from a remote location.
[0679] 1. Embodiment of push notification function
[0680] The device is equipped with a built-in microphone and camera, which are used to detect the baby's movements and cries. The detected data is processed within the device, and if a predetermined threshold is exceeded, an alert is generated. The generated alert information is sent to a server, which then sends a push notification to the user based on the alert information. This allows the user to immediately become aware of any abnormalities in the baby.
[0681] Embodiment of music playback function
[0682] When a user issues a music playback command via the app on their smartphone, the command is sent to the server, which then receives the command and transmits it to the device. The device then plays the specified music file, allowing the baby to listen to the music through the built-in speaker.
[0683] Embodiment of the hazard detection function
[0684] The device is equipped with an AI algorithm that analyzes video and audio data in real time. If it detects that the baby is making dangerous movements, the device will automatically play an audio message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[0685] An embodiment of the listening function
[0686] The device's microphone detects what the child is saying and collects voice data in real time. This voice data is converted into text on the device and analyzed by AI. Based on the analysis results, the AI generates an appropriate response and responds gently to the child through voice synthesis. This allows users to be there for their children even when they are busy.
[0687] Sensor Function Embodiment
[0688] The user installs dedicated sensors in the crib and other dangerous locations. The device collects data from these sensors and processes it as an alert if it detects any dangerous behavior (e.g., a fall or abnormal movement). The alert information is sent to the server, and the device simultaneously plays an appropriate warning sound to protect the baby. The server then sends a push notification to the user to notify them of the abnormality.
[0689] Market deployment implementation example
[0690] The server will offer the app free of charge to new users. The basic features are available for free, but if users want higher resolution video or additional security features, they can purchase premium services.
[0691] As described above, the present invention provides a monitoring system that utilizes unused mobile communication devices, reduces the burden on parents in raising their children, and is also environmentally friendly.
[0692] The processing flow will be explained below.
[0693] Camera linkage function processing steps
[0694] Step 1:
[0695] Users install a dedicated app on their unused devices, open the app, and configure the camera settings, which activates the device's camera.
[0696] Step 2:
[0697] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[0698] Step 3:
[0699] The terminal transmits the captured video data to the server in real time.
[0700] Step 4:
[0701] The server receives the video data and starts streaming it to the user terminal.
[0702] Step 5:
[0703] The user launches the app on the user device and checks the live video being streamed.
[0704] Push notification processing steps
[0705] Step 1:
[0706] The device uses a built-in microphone and camera to detect the baby's movements and cries.
[0707] Step 2:
[0708] The terminal analyzes the sensed data and generates alert information when a certain threshold is exceeded.
[0709] Step 3:
[0710] The terminal transmits the generated alert information to the server.
[0711] Step 4:
[0712] The server receives the alert information and sends a push notification to the user terminal.
[0713] Step 5:
[0714] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[0715] Processing steps for music playback function
[0716] Step 1:
[0717] The user sends a music playback instruction to the server from the user terminal via the app.
[0718] Step 2:
[0719] The server receives the user's instruction to play music and transmits it to the terminal.
[0720] Step 3:
[0721] The terminal plays the specified music file based on the instruction received from the server.
[0722] Step 4:
[0723] The device uses a built-in speaker to play music for your baby.
[0724] Hazard detection function processing steps
[0725] Step 1:
[0726] The device collects data in real time through a camera and microphone.
[0727] Step 2:
[0728] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[0729] Step 3:
[0730] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[0731] Step 4:
[0732] The terminal transmits danger alert information to the server.
[0733] Step 5:
[0734] The server receives the danger alert and sends a push notification to the user terminal.
[0735] Listening function processing steps
[0736] Step 1:
[0737] The device detects what the child is saying through a microphone.
[0738] Step 2:
[0739] The device converts the detected voice data into text and sends it to the AI.
[0740] Step 3:
[0741] The device generates an appropriate response based on the results analyzed by the AI.
[0742] Step 4:
[0743] The device uses voice synthesis to speak gentle words to the child.
[0744] Sensor function processing steps
[0745] Step 1:
[0746] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[0747] Step 2:
[0748] The device continuously collects data from the sensors.
[0749] Step 3:
[0750] The device analyzes the collected data and detects abnormal behavior.
[0751] Step 4:
[0752] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[0753] Step 5:
[0754] The server receives the anomaly alert and sends a push notification to the user terminal, which then plays an appropriate warning sound.
[0755] Go-to-market process steps
[0756] Step 1:
[0757] The server offers the app free of charge to new users.
[0758] Step 2:
[0759] Users can use the basic features free of charge.
[0760] Step 3:
[0761] The server will set up options that offer high-resolution video and additional security features for a fee.
[0762] Step 4:
[0763] If the user selects a paid service, the corresponding function is provided.
[0764] Example 1
[0765] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0766] Today, there are a wide variety of systems available to help parents safely monitor their children. However, many of these systems have drawbacks, such as being expensive or not being able to reuse old mobile communication devices. Furthermore, there are limited means for parents who are in remote locations to check on their children's status in real time and respond appropriately. For these reasons, there is a demand for a monitoring system that is easy to implement, makes effective use of old mobile communication devices, and can ensure the safety of children remotely.
[0767] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0768] In this invention, the server includes a video acquisition means, a video transmission means, a display means, a sensing means, an information notification means, a music playback means, a calling means, a response means, a data collection means, and a means for taking a quick response. This allows parents to use their unused mobile communication devices as surveillance cameras to remotely monitor their children's situations in real time and to respond immediately in the event of an abnormality.
[0769] "Video capture means" refers to technology that uses a camera or other image capture device to capture visual information in real time or at regular intervals.
[0770] "Video transmission means" refers to a technique for transmitting acquired video data to another device via a network.
[0771] "Display means" refers to a display or other display device for visually presenting received video data to a user.
[0772] "Sensing means" refers to technology that analyzes captured video and audio data to detect specific movements and acoustic patterns.
[0773] "Information notification means" refers to notification technology for notifying the user of abnormalities or important information detected by the sensing means.
[0774] "Music playback means" refers to technology including circuitry and software for playing music specified by the user.
[0775] "Hammer" refers to technology that automatically plays a pre-set voice message when a dangerous situation is detected.
[0776] "Response means" refers to technology that detects what the child is saying, generates an appropriate response, and replies in voice.
[0777] "Data collection means" refers to the technology used to collect data from installed sensors and detect abnormal behavior.
[0778] "Means for taking prompt action" refers to technology that uses information notification means to quickly notify the user and prompt them to take the necessary action.
[0779] This invention is a "monitoring and listening system" that allows parents to remotely monitor their children by reusing unused mobile communication devices. The system includes video acquisition means, video transmission means, display means, sensing means, information notification means, music playback means, calling means, response means, data collection means, and means for rapid response.
[0780] First, the user installs a dedicated application on an unused mobile communication device (hereafter referred to as the "device"). This application can be downloaded from the Google Play Store or Apple's App Store. After installing the application, the user configures the camera and enables the device to function as a surveillance camera.
[0781] The device captures real-time video data through the camera and sends it to the server via Wi-Fi or mobile data. The captured video is compressed using an H.264 encoder. The server then delivers the received video data to the user's smartphone as live streaming in an appropriate format (e.g., RTMP). The user can then view the real-time video on their smartphone using a dedicated application.
[0782] The device's built-in microphone and camera detect the baby's movements and cries. Data from these sensors is analyzed in real time using digital signal processing (DSP). Any detected abnormalities that exceed a predetermined threshold are sent to the server as an alert. The server generates a push notification based on the alert information and sends it to the user's smartphone via Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs). This allows the user to quickly learn of any abnormalities in their baby.
[0783] Users can also send instructions to play music from their smartphones through the application. These instructions are sent to the server, which then transmits them to the device. The device receives the instructions, plays the specified music file, and plays the music to the baby through the built-in speaker. The device's built-in media player is used to play the music.
[0784] Furthermore, the device uses AI algorithms to analyze video and audio data in real time, and if it detects dangerous movement, it will automatically play a voice message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[0785] When a child speaks to the device through the microphone, voice data is collected in real time, and the speech recognition engine converts the speech into text. The AI then analyzes the text, generates an appropriate response, and responds gently using the speech synthesis engine.
[0786] Users can also install dedicated sensors in the crib or other dangerous areas. Data from these sensors is collected by the device and processed as an alert when dangerous behavior is detected. The alert information is sent to the server, and the device plays an appropriate warning sound to protect the baby. The server also sends a push notification to the user.
[0787] To illustrate, consider the following prompt:
[0788] 1. "How do I check my baby's real-time video on my smartphone?"
[0789] 2. "How can I get a push notification to my smartphone when my baby cries?"
[0790] 3. "How do I send a music playback command from my smartphone and play music on the device?"
[0791] 4. "How can I detect dangerous baby movements, give voice warnings, and send notifications?"
[0792] 5. "How can I generate appropriate responses and respond to my child's speech using text-to-speech?"
[0793] 6. "How can I have my device sound an alarm and send a notification when the sensor installed in the crib detects dangerous behavior?"
[0794] As described above, this monitoring and listening system provides parents with an effective means of keeping their children safe even from a distance, reducing the burden of childcare. It is also environmentally friendly, as it can be reused from unused mobile communication devices.
[0795] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0796] Camera Linkage Feature Steps
[0797] Step 1: Install and configure the app
[0798] Input: Download a dedicated application to a device that the user no longer uses.
[0799] Specific operation: The user downloads the dedicated application from the Google Play Store or Apple's App Store and installs it on their device.
[0800] Output: The dedicated application is installed on the device and ready to launch.
[0801] Input: The user opens the app's settings screen and configures the camera to launch.
[0802] What it does: When you first open the app, it will show a setup wizard that asks for camera permissions, and you can agree to enable the camera.
[0803] Output: The camera is set to a state where it can be activated on the device.
[0804] Step 2: Capture and transmit real-time video
[0805] Input: Camera powers up and begins collecting real-time footage.
[0806] What it does: The device's camera will start recording real-time footage of your baby or child, and the footage will be compressed using an H.264 encoder.
[0807] Output: Compressed video data is generated.
[0808] Input: Prepare compressed video data for transmission.
[0809] How it works: The device uploads the captured video data to a server in real time via Wi-Fi or mobile data, using HTTP / 2 or WebSocket protocols.
[0810] Output: The compressed video data is sent to the server.
[0811] Step 3: Streaming real-time video
[0812] Input: The server prepares to deliver the received video data to the user's device as live streaming.
[0813] Specific operation: The server converts the video data into an appropriate format and streams it to the user's device using RTMP (Real-time Messaging Protocol).
[0814] Output: The converted video data is delivered to the user's device.
[0815] Input: The user launches the app on their device and checks the real-time video.
[0816] Specific operation: By launching the app on the user's device and selecting the Live Streaming tab, real-time video will be played. The video will be decoded and displayed in a format that is easy for the user to view.
[0817] Output: User can check real-time video.
[0818] Push notification feature steps
[0819] Step 1: Sensing and processing data
[0820] Input: The device's built-in microphone and camera detect baby movements and cries.
[0821] How it works: The device's microphone constantly monitors the surrounding environment, the camera detects movement, and the data is analyzed in real time through digital signal processing (DSP).
[0822] Output: Detected movement and sound data is generated.
[0823] Input: Prepares sensed data to be analyzed.
[0824] How it works: If the baby's crying exceeds a certain decibel level or if any violent movement is detected, the AI module in the device will generate an alert.
[0825] Output: Alert information is generated.
[0826] Step 2: Sending alert information and notifications
[0827] Input: Prepares the generated alert information to be sent to the server.
[0828] Specific operation: The device sends alert information to the server via Wi-Fi or mobile data communication, using the MQTT (Message Queuing Telemetry Transport) protocol.
[0829] Output: Alert information is sent to the server.
[0830] Input: The server prepares to generate a push notification based on the alert information.
[0831] Specific operation: The server analyzes the alert information, generates a push notification, and sends it to the user's device using Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs).
[0832] Output: A push notification is generated and sent to the user device.
[0833] Music playback function steps
[0834] Step 1: Sending a music play command
[0835] Input: The user commands music playback through the app on their smartphone.
[0836] Specific operation: The user selects the music they want to play from the app's music playback menu and taps the play button. Instruction data is generated.
[0837] Output: Music playback instructions are generated.
[0838] Input: Prepares to send music playback instruction data to the server.
[0839] Specific operation: Instruction data generated from the user terminal is sent to the server via the Internet. HTTP / 2 is used as the communication protocol.
[0840] Output: Instruction data is sent to the server.
[0841] Step 2: Transmitting and executing music playback instructions
[0842] Input: The server receives the instruction data and prepares it for transmission to the terminal.
[0843] Specific operation: The server analyzes the received instruction data and sends a playback instruction to the corresponding terminal.
[0844] Output: Playback instructions are sent to the device.
[0845] Input: Prepares the device to play the specified music file.
[0846] Specific operation: The device retrieves and plays the music file from the internal storage or cloud storage based on the received playback command. The device's built-in media player is used to play the audio.
[0847] Output: The specified music file will be played.
[0848] Hazard detection function steps
[0849] Step 1: Analyzing video and audio data
[0850] Input: The device collects video and audio data and analyzes it in real time using AI algorithms.
[0851] Specific operation: The device's AI chip processes video and audio data in real time to detect dangerous movements of the baby.
[0852] Output: Data is generated indicating the dangerous behavior detected.
[0853] Step 2: Play and notify warning messages
[0854] Input: Prepare to play a warning message when dangerous motion is detected.
[0855] Specific operation: Based on the AI detection results, the device plays a pre-set audio message to warn the user.
[0856] Output: A warning audio message is played.
[0857] Input: Sends the danger detection information to the server, which prepares to generate a push notification.
[0858] Specific operation: The danger detection information is sent to the server, and the server generates a push notification and sends it to the user.
[0859] Output: A push notification is generated and sent to the user.
[0860] Steps of active listening function
[0861] Step 1: Detecting speech and collecting voice data
[0862] Input: Detects what your child is saying through the device's microphone.
[0863] How it works: The device's microphone detects what the child is saying and collects audio data, which is then stored in a buffer on the device in real time.
[0864] Output: Speech audio data is collected.
[0865] Step 2: Speech to text conversion and analysis
[0866] Input: Prepare the collected audio data for conversion to text.
[0867] How it works: The device's voice recognition engine converts the voice data into text, and the AI module analyzes the text to generate an appropriate response.
[0868] Output: Parsed text data and response data are generated.
[0869] Step 3: Response generation and speech synthesis
[0870] Input: Converts response data into audio and prepares it for playback.
[0871] Specific operation: The device's speech synthesis engine outputs the generated response as voice and plays it through the speaker.
[0872] Output: An audio response is played to the child.
[0873] Sensor function steps
[0874] Step 1: Sensor installation and data collection
[0875] Input: The user installs special sensors in the crib or other dangerous locations.
[0876] Specific operation: The user places a special sensor in a crib or dangerous location, and the sensor begins to function.
[0877] Output: The sensor is deployed and ready to collect data.
[0878] Input: Prepare to collect data from sensors.
[0879] Specific operation: The device receives data sent from the sensor and detects abnormal movement or falls.
[0880] Output: Data is generated indicating the abnormal behavior detected.
[0881] Step 2: Alerting and Notification
[0882] Input: Prepare to treat anomalous behavior as an alert.
[0883] Specific operation: When the device detects danger, it generates alert information and sends it to the server.
[0884] Output: Alert information is generated and sent to the server.
[0885] Input: The server prepares to generate a push notification based on the alert information.
[0886] What it does: The server sends a push notification to the user, and the device plays a series of warning sounds to protect the baby.
[0887] Output: A push notification is generated and sent to the user. An alert sound is played on the device.
[0888] Market expansion steps
[0889] Step 1: Offer your app and define your market
[0890] Input: The server prepares the app to serve to a new user.
[0891] What happens: The server publishes the app's download link on a website or store, making it free for new users.
[0892] Output: The app download link is published and new users can download the app.
[0893] Step 2: Offering premium services
[0894] Input: The server defines the content of the paid premium service and prepares to provide it to the user.
[0895] Specific operation: Allow users to purchase premium services within the app, and after purchase, the server provides the corresponding services.
[0896] Output: The user purchases a premium service, unlocking additional features.
[0897] (Application example 1)
[0898] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0899] To improve vehicle safety, including that of autonomous vehicles and their drivers, it is necessary to constantly monitor the driver's state and detect abnormalities or dangers early. However, conventional systems lack the ability to detect driver drowsiness or inattention in real time and send warnings. There is also a need for an efficient method of reusing unused mobile communication devices. A system that solves these issues and ensures safety while being economical and environmentally friendly is needed.
[0900] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0901] In this invention, the server is a monitoring system that allows parents to remotely monitor their children by using an unused mobile communication device as a camera, and includes: a camera means for acquiring video; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video; an analysis means for analyzing the video from the mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the detected information; a music playback means for receiving remote operation instructions from the user terminal and playing music; an automatic call means for issuing an automatic call when dangerous movement is detected; a listening means for detecting a child's speech and responding gently; a driver monitoring means for monitoring the driver's state using a camera installed in the vehicle and detecting drowsiness or inattention; and a warning means for sending a warning to the driver and a manager when drowsiness or inattention is detected. This makes it possible to monitor the driver's state in real time and respond immediately if an abnormality is detected.
[0902] A "mobile communications device" is a device used for mobile communications, including smartphones and tablets.
[0903] "Camera means" refers to the camera equipment or functionality used to capture the image.
[0904] "Communication means" refers to communication technology or devices for sending and receiving data or information.
[0905] "User terminal means" refers to a terminal device used by a user, and is a device used for displaying images and remote control.
[0906] "Analysis means" refers to a method or device for analyzing video and audio data and detecting specific movements or situations.
[0907] A "notification means" is a method or device for sending a notification to a user or administrator based on sensed information.
[0908] "Music playback means" refers to a method or device for playing music, including speakers and music playback applications.
[0909] "Automatic calling means" refers to a method or device that automatically generates and plays voice messages under certain circumstances.
[0910] "Listening means" refers to a method or device for detecting and responding to the voice of a subject.
[0911] "Driver monitoring means" refers to a method or device that includes cameras or sensors for monitoring the situation inside the vehicle and, in particular, for checking the driver's condition.
[0912] "Warning means" refers to a method or device for issuing an immediate warning when an abnormality is detected.
[0913] This invention relates to a "monitoring and listening app" that allows parents to remotely monitor their babies and children using unused mobile communication devices. Furthermore, by applying this technology to autonomous vehicles, a system can be realized that monitors the driver and immediately issues a warning in the event of an abnormal or dangerous situation.
[0914] System Configuration
[0915] The server works in conjunction with mobile communication devices, user terminals, sensors, etc. These devices exchange data and analyze it to achieve the following functions:
[0916] Driver status monitoring
[0917] The device is equipped with a camera that captures images of the interior of the vehicle in real time and sends them to a server. This video data is analyzed by the server to detect driver drowsiness and inattention. This detection uses an AI algorithm, specifically TensorFlow, for face detection and movement analysis.
[0918] Warning function
[0919] If an abnormality is detected in the driver, the server will send a warning message to the administrator through the Twilio API, and will use the Google Text-to-Speech (gTTS) library to generate a voice warning and play the warning message aloud, using the pygame library for speech synthesis in this process.
[0920] Music playback function
[0921] The server receives instructions from the user's device and sends a music playback instruction to the device, which then plays the specified music file through its built-in speaker.
[0922] Data analysis
[0923] The video data sent from the device is analyzed by the server. The analysis method is to detect and analyze the driver's face and movements using OpenCV and TensorFlow. If drowsiness or inattention is detected, it is immediately treated as an abnormality and a warning is issued.
[0924] Specific examples
[0925] For example, if a driver falls asleep while driving on a highway, the device's camera will detect this in real time. If the server detects the driver's drowsiness through AI analysis, the system will send an SMS notification to the administrator via the server and simultaneously play an audio warning on the device.
[0926] Prompt Sentence Examples
[0927] "Develop an app that uses mobile communication devices to monitor the driver's status in real time. If the driver becomes inattentive or falls asleep, it should issue a voice warning and send an SMS notification to the manager."
[0928] In this way, the present invention provides a system that can seamlessly monitor and warn drivers, and by reusing unused mobile communication devices, it builds an efficient system that is environmentally friendly.
[0929] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0930] Step 1:
[0931] The device starts the camera and captures the video inside the car. In this step, the camera acquires video data in real time. The input is the video inside the car, and the output is the captured video data. Specifically, the device starts the camera using cv2.VideoCapture(0) and acquires frames using the read() method.
[0932] Step 2:
[0933] The video data captured by the terminal is sent to the server in real time. In this step, data is sent using a communication means. The input is the captured video data, and the output is the video data sent to the server. In concrete terms, the terminal sends the video data to the server using a specified communication protocol (e.g., HTTP POST request).
[0934] Step 3:
[0935] The server analyzes the received video data. In this step, the analysis means is used to determine the driver's state. The input is the transmitted video data, and the output is the analysis result. Specifically, the server processes the video data using OpenCV and TensorFlow to analyze the driver's face and movements.
[0936] Step 4:
[0937] If the server detects an anomaly, it generates a warning message. In this step, the warning means operates based on the analysis results. The input is the analysis result, and the output is the warning message. Specifically, when an anomaly is detected, the server generates a voice warning using the gTTS library and sends an SMS notification using the Twilio API.
[0938] Step 5:
[0939] The terminal plays music according to instructions from the server. In this step, a music playback means is used. The input is a remote control instruction from the user terminal, and the output is the music being played. In concrete terms, the terminal plays music on its built-in speaker based on the instruction received from the server.
[0940] Step 6:
[0941] The device detects what the child is saying and responds appropriately. In this step, the listening means is activated. The input is what the child is saying, and the output is a voice response. Specifically, the device detects the voice through the microphone and sends it to the server. The server converts it into text using a generative AI model and generates an appropriate response.
[0942] Step 7:
[0943] The server comprehensively manages all data and sends notifications to user devices in the event of an abnormality. This step is responsible for overall control. Input is data from various sensors and cameras, and output is notifications to administrators and push notifications to user devices. Specifically, the server receives and analyzes all data, and immediately notifies users if an abnormality is detected.
[0944] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0945] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely watch over babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals"), with an emotion engine that recognizes the user's emotions.
[0946] Embodiment of camera linkage function
[0947] First, the user installs a dedicated app on the device and configures the camera. This activates the device's camera. The device captures real-time video through the camera and transmits the video data to the server. The server then streams the received data to the user's smartphone (hereinafter referred to as the "user device"). The user can then view the live video on their own device.
[0948] 1. Embodiment of push notification function
[0949] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device. The user can then respond promptly via the push notification.
[0950] Embodiment of music playback function
[0951] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server. The device then plays the specified music and plays it to the baby through the built-in speaker.
[0952] Embodiment of the hazard detection function
[0953] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, danger alert information is sent to the server, which then warns the user via push notification.
[0954] An embodiment of the listening function
[0955] The device detects what the child is saying through a microphone and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to gently reply to the child, allowing the user to engage with their child even when they are busy.
[0956] Sensor Function Embodiment
[0957] The user installs dedicated sensors in the crib or other dangerous locations and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The detected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[0958] Embodiment of Emotion Engine
[0959] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The response is converted from text to speech and played through the device's speaker.
[0960] This allows the emotion engine to reduce the user's mental burden and more effectively support the care of babies and children. The combination of the emotion engine improves the functionality and flexibility of the entire system, making it an important tool for supporting childcare.
[0961] The processing flow will be explained below.
[0962] Camera linkage function processing steps
[0963] Step 1:
[0964] Users install a dedicated app on their unused devices and configure the camera settings, which activates the device's camera.
[0965] Step 2:
[0966] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[0967] Step 3:
[0968] The terminal transmits the captured video data to the server in real time.
[0969] Step 4:
[0970] The server receives the video data and starts streaming it to the user terminal.
[0971] Step 5:
[0972] The user launches the app on the user device and checks the live video being streamed.
[0973] Push notification processing steps
[0974] Step 1:
[0975] The device's built-in microphone and camera are used to detect the baby's movements and cries.
[0976] Step 2:
[0977] The terminal analyzes the sensed data in real time and generates alert information when a certain threshold is exceeded.
[0978] Step 3:
[0979] The terminal transmits the generated alert information to the server.
[0980] Step 4:
[0981] The server receives the alert information and sends a push notification to the user terminal.
[0982] Step 5:
[0983] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[0984] Processing steps for music playback function
[0985] Step 1:
[0986] The user sends a music playback instruction to the server from the user terminal via the app.
[0987] Step 2:
[0988] The server receives the user's instruction to play music and transmits it to the terminal.
[0989] Step 3:
[0990] The terminal plays the specified music file based on the instruction received from the server.
[0991] Step 4:
[0992] The device uses a built-in speaker to play music for your baby.
[0993] Hazard detection function processing steps
[0994] Step 1:
[0995] The device collects data in real time through a camera and microphone.
[0996] Step 2:
[0997] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[0998] Step 3:
[0999] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[1000] Step 4:
[1001] The terminal transmits danger alert information to the server.
[1002] Step 5:
[1003] The server receives the danger alert and sends a push notification to the user terminal.
[1004] Listening function processing steps
[1005] Step 1:
[1006] It detects what the child is saying through the device's microphone.
[1007] Step 2:
[1008] The device converts the detected voice data into text and sends it to the AI.
[1009] Step 3:
[1010] The device generates an appropriate response based on the results analyzed by the AI.
[1011] Step 4:
[1012] The device uses voice synthesis to speak soothing words to the child.
[1013] Sensor function processing steps
[1014] Step 1:
[1015] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[1016] Step 2:
[1017] The device continuously collects data from connected sensors.
[1018] Step 3:
[1019] The device analyzes the collected data and detects abnormal behavior.
[1020] Step 4:
[1021] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[1022] Step 5:
[1023] The server receives the anomaly alert and sends a push notification to the user device, which simultaneously plays an appropriate warning sound.
[1024] Emotion Engine Processing Steps
[1025] Step 1:
[1026] The user's voice data is transmitted from the user terminal to the server.
[1027] Step 2:
[1028] The server sends the received voice data to the emotion engine on the device.
[1029] Step 3:
[1030] The device's emotion engine analyzes the voice data and determines the user's emotional state.
[1031] Step 4:
[1032] The terminal's emotion engine generates an appropriate response based on the determined emotional state.
[1033] Step 5:
[1034] The terminal converts the generated response from text to speech and provides feedback to the user.
[1035] Example 2
[1036] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1037] In modern child-rearing, it is difficult for parents to constantly keep an eye on their babies and children, and the burden is particularly heavy when both parents are working or raising children alone. Conventional monitoring systems have limitations in real-time monitoring and voice communication, and lack emotion recognition functionality, making it difficult to reduce the mental burden on parents. Furthermore, they lack the ability to instantly detect dangerous situations and respond appropriately. The purpose of this invention is to solve these problems, more effectively ensure the safety of babies and children, and reduce the burden on parents.
[1038] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a camera means for capturing video from an unused mobile communication device; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video via the communication means; an analysis means for analyzing the video from the unused mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the information detected by the analysis means; a music playback means for receiving remote operation instructions from the user terminal and playing music; an emotion analysis means for analyzing audio data transmitted from the unused mobile communication device and determining the user's emotional state; an automatic call means for issuing an automatic call when dangerous movements are detected; and a listening means for detecting and gently responding to the child's voice. This allows for a multifunctional and flexible monitoring system to be constructed, enabling parents to monitor and respond to the condition of babies and children even remotely. Furthermore, emotion analysis reduces the user's mental burden and enables quick and appropriate responses to dangerous situations.
[1039] "Disused mobile communication devices" are mobile devices such as mobile phones and tablets that have been retired but are now being reused.
[1040] "Camera means" refers to a camera device installed to capture video.
[1041] "Communication means" refers to a means having an internet connection function for sending and receiving images and data.
[1042] "User terminal means" refers to a display device used by a user, such as a smartphone, tablet, or computer.
[1043] The "analysis means" is a means having the function of analyzing the collected data and detecting specific information such as the baby's movements and cries.
[1044] The "notification means" is a means having a function for generating a push notification based on sensed information and transmitting it to the user terminal.
[1045] The "music playback means" is a means having a function of playing music in response to an instruction from a user terminal.
[1046] The "emotion analysis means" is a means having a function of analyzing voice data and determining the emotional state of the user.
[1047] The "automatic calling means" is a means having a function of automatically playing a warning message or the like when a dangerous movement is detected.
[1048] "Listening tools" are tools that have the ability to sense what a child is saying, generate an appropriate response, and respond in a gentle manner.
[1049] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely monitor babies and children by reusing unused mobile communication devices with an emotion engine that recognizes the user's emotions. This system provides a multifunctional and flexible monitoring environment.
[1050] Camera linkage function
[1051] First, the user installs a dedicated app on their device and configures the camera. Once the app is installed and configured, the device's camera is activated. The device captures real-time video through the camera and sends the data to the server. The server receives the video data and streams it to the user's device, allowing the user to view live video on their device.
[1052] Specific examples
[1053] Device used: Old smartphone
[1054] Server used: AWS EC2 instance
[1055] Video streaming software: FFmpeg
[1056] Push notification function
[1057] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device, allowing the user to respond quickly.
[1058] Specific examples
[1059] Hardware used: Android or iOS device
[1060] Server: Google Firebase Cloud Messaging
[1061] App: Real-time analytics with custom mobile app
[1062] Music playback function
[1063] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server, and the device plays the specified music on its built-in speaker.
[1064] Specific examples
[1065] Device used: Old tablet
[1066] Server: AWS Lambda
[1067] Music playback software: VLC
[1068] Hazard detection function
[1069] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, the device sends danger alert information to the server, which then sends a push notification to the user's device.
[1070] Specific examples
[1071] Device used: Raspberry Pi 4 with camera module
[1072] AI algorithm: TensorFlow
[1073] Server: Google Cloud Functions
[1074] Active listening function
[1075] The device detects what the child is saying through a microphone and converts the voice data into text. This text data is analyzed using an AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The generated response is then converted into audio data using speech synthesis technology, which is then played back by the device.
[1076] Specific examples
[1077] Device used: Amazon Echo device
[1078] Speech Recognition: Google Speech-to-Text API
[1079] Response generation: OpenAI GPT-3
[1080] Speech synthesis: Amazon Polly
[1081] Sensor function
[1082] The user installs dedicated sensors in the crib or other dangerous areas and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The collected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[1083] Specific examples
[1084] Sensor used: Xiaomi Smart Home Sensor
[1085] Device used: Old smartphone
[1086] Server: AWS IoT Core
[1087] Emotion Engine
[1088] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The generated response is converted into voice data and played through the device's speaker.
[1089] Specific examples
[1090] Device used: Old smartphone
[1091] Emotion engine: IBM Watson Tone Analyzer
[1092] Speech synthesis: Google Text-to-Speech
[1093] This makes the entire system a versatile and flexible monitoring environment, making it a very useful tool for supporting childcare. Users can monitor the condition of their baby or child in real time, even remotely, and respond quickly if necessary. The emotion engine also reduces mental strain.
[1094] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1095] Camera linkage function processing steps
[1096] Step 1:
[1097] Users install a dedicated app on their device, open the app and configure the camera settings.
[1098] Input: App installation, setting information
[1099] Output: Setup complete, camera ready to start
[1100] Step 2:
[1101] Once the setup is complete, the device will launch the camera.
[1102] Input: Setting completion notification
[1103] Output: Camera starts, video capture begins
[1104] Step 3:
[1105] The device captures real-time video through a camera and transmits the data to a server.
[1106] Input: Real-time video data
[1107] Output: Ready to send, sent to server
[1108] Step 4:
[1109] The server processes the received video data and distributes it as a stream to the user terminal.
[1110] Input: Video data
[1111] Output: Video streaming ready, sent to user device
[1112] Step 5:
[1113] Users can view live footage via the app on their devices.
[1114] Input: Video streaming data
[1115] Output: Video display, real-time monitoring
[1116] Push notification processing steps
[1117] Step 1:
[1118] It analyzes data collected from the device's built-in microphone and camera in real time.
[1119] Input: Audio data, video data
[1120] Output: Analysis results, threshold judgment
[1121] Step 2:
[1122] If the baby's movements or cries exceed the threshold, the device will generate an alert.
[1123] Input: Analysis results
[1124] Output: Alert information
[1125] Step 3:
[1126] The terminal transmits the generated alert information to the server.
[1127] Input: Alert information
[1128] Output: Sent to server
[1129] Step 4:
[1130] The server generates a push notification based on the alert information and sends it to the user's device.
[1131] Input: Alert information
[1132] Output: Push notification
[1133] Step 5:
[1134] Users receive push notifications and can respond quickly.
[1135] Input: Push notification
[1136] Output: Notification confirmation, required action
[1137] Processing steps for music playback function
[1138] Step 1:
[1139] The user operates the app on their smartphone and sends instructions to play music.
[1140] Input: Music playback instructions
[1141] Output: Send instructions
[1142] Step 2:
[1143] The server transfers the playback instruction received from the user to the terminal.
[1144] Input: Playback instructions
[1145] Output: Directed transfer
[1146] Step 3:
[1147] The device will play music through its built-in speaker based on the transferred instructions.
[1148] Input: Playback instructions
[1149] Output: Music playback
[1150] Hazard detection function processing steps
[1151] Step 1:
[1152] The device collects the baby's movements and voice data using a camera and microphone.
[1153] Input: Video data, audio data
[1154] Output: Recorded data, audio data
[1155] Step 2:
[1156] The device analyzes the collected data using AI algorithms to detect dangerous behavior.
[1157] Input: Recorded data, audio data
[1158] Output: Risk assessment result
[1159] Step 3:
[1160] When danger is detected, the device will play an automated voice message such as "Danger! Stop!"
[1161] Input: Risk assessment result
[1162] Output: Automated voice message
[1163] Step 4:
[1164] The terminal transmits danger alert information to the server.
[1165] Input: Risk assessment result
[1166] Output: Alert information
[1167] Step 5:
[1168] The server generates a push notification and sends it to the user device.
[1169] Input: Alert information
[1170] Output: Push notification
[1171] Listening function processing steps
[1172] Step 1:
[1173] The device detects what the child is saying through a microphone.
[1174] Input: Audio data
[1175] Output: Sensor information
[1176] Step 2:
[1177] The device converts the detected voice data into text.
[1178] Input: Audio data
[1179] Output: Text data
[1180] Step 3:
[1181] The converted text data is analyzed by an AI model to generate an appropriate response.
[1182] Input: Text data
[1183] Output: Response data
[1184] Step 4:
[1185] The terminal converts the generated response into voice data using voice synthesis technology.
[1186] Input: Response data
[1187] Output: Audio data
[1188] Step 5:
[1189] The device plays back the audio data and provides an appropriate response to the child.
[1190] Input: Audio data
[1191] Output: Response voice
[1192] Sensor function processing steps
[1193] Step 1:
[1194] Users install special sensors in cribs and dangerous areas and connect them to a terminal.
[1195] Input: Sensor installation
[1196] Output: Connection complete
[1197] Step 2:
[1198] Sensors continuously collect environmental data and detect abnormal behavior.
[1199] Input: Environment data
[1200] Output: Data collection, anomaly detection information
[1201] Step 3:
[1202] The sensor transmits the collected data to the terminal.
[1203] Input: Data collection, anomaly detection information
[1204] Output: Send to terminal
[1205] Step 4:
[1206] The terminal analyzes the received sensor data internally and detects any abnormalities.
[1207] Input: Sensor data
[1208] Output: Abnormality detection result
[1209] Step 5:
[1210] The device generates an alert if an anomaly is detected.
[1211] Input: Abnormality detection result
[1212] Output: Alert information
[1213] Step 6:
[1214] The terminal transmits the alert information to the server.
[1215] Input: Alert information
[1216] Output: Sent to server
[1217] Step 7:
[1218] The server generates a push notification and sends it to the user device.
[1219] Input: Alert information
[1220] Output: Push notification
[1221] Emotion Engine Processing Steps
[1222] Step 1:
[1223] The user transmits voice data to the terminal.
[1224] Input: Audio data
[1225] Output: Send to terminal
[1226] Step 2:
[1227] The device uses an emotion engine to analyze the voice data and determine the user's emotional state.
[1228] Input: Audio data
[1229] Output: Emotion determination result
[1230] Step 3:
[1231] The emotion engine generates an appropriate response based on the judgment result.
[1232] Input: Emotion judgment result
[1233] Output: Response data
[1234] Step 4:
[1235] The terminal converts the generated response into voice data using voice synthesis technology.
[1236] Input: Response data
[1237] Output: Audio data
[1238] Step 5:
[1239] The terminal plays back the audio data and responds appropriately to reduce the user's mental burden.
[1240] Input: Audio data
[1241] Output: Response voice
[1242] (Application example 2)
[1243] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1244] The problem that the present invention aims to solve is to provide a system that effectively utilizes unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children and increase their safety. Furthermore, because conventional monitoring systems do not take emotion recognition into consideration, a means for reducing the mental burden on users was also necessary. Therefore, the present invention aims to compensate for the shortcomings of conventional monitoring systems and improve safety and usability.
[1245] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a communication means for transmitting video in real time, an emotion recognition means for analyzing audio data and recognizing emotions, and a response generation means for generating an appropriate response based on the recognized emotion and outputting it as audio. This makes it possible not only to perform remote real-time monitoring, but also to recognize the emotional state of the user and generate an appropriate response.
[1246] The "camera means" is a means having a function for acquiring an image.
[1247] The "communication means" is a means having a function for transmitting acquired video in real time.
[1248] The "user terminal means" is a means having a function of receiving and displaying video transmitted via a communication means.
[1249] The "analysis means" is a means having the function of analyzing video acquired from a disused mobile communication device and detecting specific movements or sounds.
[1250] The "notification means" is a means having a function of sending a push notification to the user terminal based on the information sensed by the analysis means.
[1251] The "music playback means" is a means having a function of receiving a remote operation instruction from a user terminal and playing music.
[1252] The "automatic call means" is a means having a function of automatically playing back a voice message when a dangerous movement is detected.
[1253] "Listening means" is a means that has the function of detecting what a child is saying and responding gently.
[1254] The "emotion recognition means" is a means having a function of analyzing voice data and recognizing the emotional state of the user.
[1255] The "response generation means" is a means having a function of generating an appropriate response based on the recognized emotion and outputting the response as voice.
[1256] The present invention relates to a system that makes effective use of unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children. This system uses devices and software that include the following elements:
[1257] First, a disused mobile communication device (hereinafter referred to as a terminal) captures video using a camera means. The terminal transmits the video captured by the camera to a server in real time via a communication means. The server can stream the received video to a user terminal means. The user can then view the live video on their own terminal.
[1258] Data collected from the device's built-in microphone and camera is analyzed in real time by the analysis means. This analysis detects the movement and cries of babies and children, and if a threshold is exceeded, an alert is generated by the notification means. These alerts are sent to the user's device via the server as push notifications. The user can respond promptly via the push notifications.
[1259] Furthermore, the user can remotely use the music playback means to send instructions to the terminal to play music. This instruction is transferred to the terminal via the server, and the specified music is played by the baby or child.
[1260] Furthermore, the device is equipped with a danger detection function. Data obtained from the camera and microphone is analyzed using an AI algorithm to detect if the baby is behaving in a dangerous manner. If danger is detected, the device will play an automated voice message such as "Danger! Stop!" At the same time, danger alert information is sent to the server, and a warning is sent to the user.
[1261] The device also has a listening function, which uses a microphone to detect what the child is saying and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to respond gently to the child, allowing the user to interact with their child appropriately even when they are busy.
[1262] The device is also equipped with an emotion recognition means that analyzes the voice data sent from the user device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion recognition means will recognize this and utilize the response generation means to advise an appropriate response. This reduces the user's mental burden and provides support for childcare and monitoring.
[1263] For example, if an employee of Guard AI is wearing smart glasses, the following prompt can trigger the emotion engine and provide appropriate advice:
[1264] Example prompt sentence:
[1265] Employee dictation: "I'm feeling really stressed today..."
[1266] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[1267] In this way, the system of the present invention can achieve both enhanced security and psychological care for the user.
[1268] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1269] Step 1:
[1270] The terminal acquires video using a camera means. The input is real-time video data from the camera, which is captured and stored in a buffer. The output is the captured video data.
[1271] Step 2:
[1272] The device uses a communication means to send the captured video data to the server. The input is the captured video data, which is then encoded into JPEG or MP4 format. The output is the encoded video data.
[1273] Step 3:
[1274] The server streams the received video data to the user terminal means. The input is the video data sent from the terminal and is transferred to the user terminal in real time. The output is the live video displayed on the user terminal.
[1275] Step 4:
[1276] The data collected from the device's built-in microphone and camera is analyzed by an analysis means. The input is audio and video data from the microphone and camera, which is fed into an AI algorithm for analysis. The output is the analysis results, which are the movements and cries of the baby or child.
[1277] Step 5:
[1278] If the analysis result exceeds a certain threshold, an alert is generated by the notification mechanism. The input is the analysis result, which is used to evaluate the alert condition. The output is the generated alert.
[1279] Step 6:
[1280] The server sends the generated alert to the user device as a push notification. The input is the generated alert, which is converted into a push notification format. The output is a push notification that is displayed on the user device.
[1281] Step 7:
[1282] The user remotely uses the music playback means to send a music playback instruction to the terminal. The input is the music playback instruction from the user terminal, which is sent to the terminal via the server. The output is the music played on the terminal.
[1283] Step 8:
[1284] When the terminal detects a dangerous movement, it plays a voice message using the automatic call means. The input is the detection result of the analysis means, and if it is judged to be dangerous, the voice message is played. The output is a warning voice played from the terminal.
[1285] Step 9:
[1286] The terminal uses a listening means to detect what the child is saying and converts the voice data into text. The input is the voice data of the child's speech, which is converted into text data. The output is text data.
[1287] Step 10:
[1288] The text data is analyzed by AI to generate an appropriate response. The input is the converted text data, and the natural language processing algorithm generates a response based on this. The output is the response text.
[1289] Step 11:
[1290] The response generation means outputs the generated response text as voice using speech synthesis technology. The input is the response text, which is converted into voice data by a speech synthesis engine. The output is the response voice played from the terminal.
[1291] Step 12:
[1292] The emotion recognition means analyzes the voice data sent from the user terminal and determines the user's emotional state. The input is the user's voice data, which is supplied to a voice analysis algorithm to recognize the emotional state. The output is the determination result of the emotional state.
[1293] Step 13:
[1294] Based on the recognized emotional state, the response generator generates an appropriate response. The input is the result of the emotional state determination, and the response content is determined based on this. The output is the generated response content.
[1295] Step 14:
[1296] The response is output as voice through the device's speaker. The input is the generated response, which is played back as voice using speech synthesis technology. The output is a voice response played from the device.
[1297] Example prompt sentence:
[1298] Employee dictation: "I'm feeling really stressed today..."
[1299] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[1300] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1301] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1302] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1303] [Third embodiment]
[1304] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1305] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1306] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1307] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1308] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1309] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1310] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1311] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1312] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1313] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1314] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1315] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1316] This invention relates to a "monitoring and listening app" that allows parents to remotely watch over their babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals").
[1317] Embodiment of camera linkage function
[1318] First, the user installs a dedicated app on their device and configures the camera. The device captures real-time video through the camera and sends the video data to a server. The server then distributes this data as live streaming to the user's smartphone (hereinafter referred to as the "user device"). When the user launches the app on the user device, they can check on their baby's condition in real time from a remote location.
[1319] 1. Embodiment of push notification function
[1320] The device is equipped with a built-in microphone and camera, which are used to detect the baby's movements and cries. The detected data is processed within the device, and if a predetermined threshold is exceeded, an alert is generated. The generated alert information is sent to a server, which then sends a push notification to the user based on the alert information. This allows the user to immediately become aware of any abnormalities in the baby.
[1321] Embodiment of music playback function
[1322] When a user issues a music playback command via the app on their smartphone, the command is sent to the server, which then receives the command and transmits it to the device. The device then plays the specified music file, allowing the baby to listen to the music through the built-in speaker.
[1323] Embodiment of the hazard detection function
[1324] The device is equipped with an AI algorithm that analyzes video and audio data in real time. If it detects that the baby is making dangerous movements, the device will automatically play an audio message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[1325] An embodiment of the listening function
[1326] The device's microphone detects what the child is saying and collects voice data in real time. This voice data is converted into text on the device and analyzed by AI. Based on the analysis results, the AI generates an appropriate response and responds gently to the child through voice synthesis. This allows users to be there for their children even when they are busy.
[1327] Sensor Function Embodiment
[1328] The user installs dedicated sensors in the crib and other dangerous locations. The device collects data from these sensors and processes it as an alert if it detects any dangerous behavior (e.g., a fall or abnormal movement). The alert information is sent to the server, and the device simultaneously plays an appropriate warning sound to protect the baby. The server then sends a push notification to the user to notify them of the abnormality.
[1329] Market deployment implementation example
[1330] The server will offer the app free of charge to new users. The basic features are available for free, but if users want higher resolution video or additional security features, they can purchase premium services.
[1331] As described above, the present invention provides a monitoring system that utilizes unused mobile communication devices, reduces the burden on parents in raising their children, and is also environmentally friendly.
[1332] The processing flow will be explained below.
[1333] Camera linkage function processing steps
[1334] Step 1:
[1335] Users install a dedicated app on their unused devices, open the app, and configure the camera settings, which activates the device's camera.
[1336] Step 2:
[1337] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[1338] Step 3:
[1339] The terminal transmits the captured video data to the server in real time.
[1340] Step 4:
[1341] The server receives the video data and starts streaming it to the user terminal.
[1342] Step 5:
[1343] The user launches the app on the user device and checks the live video being streamed.
[1344] Push notification processing steps
[1345] Step 1:
[1346] The device uses a built-in microphone and camera to detect the baby's movements and cries.
[1347] Step 2:
[1348] The terminal analyzes the sensed data and generates alert information when a certain threshold is exceeded.
[1349] Step 3:
[1350] The terminal transmits the generated alert information to the server.
[1351] Step 4:
[1352] The server receives the alert information and sends a push notification to the user terminal.
[1353] Step 5:
[1354] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[1355] Processing steps for music playback function
[1356] Step 1:
[1357] The user sends a music playback instruction to the server from the user terminal via the app.
[1358] Step 2:
[1359] The server receives the user's instruction to play music and transmits it to the terminal.
[1360] Step 3:
[1361] The terminal plays the specified music file based on the instruction received from the server.
[1362] Step 4:
[1363] The device uses a built-in speaker to play music for your baby.
[1364] Hazard detection function processing steps
[1365] Step 1:
[1366] The device collects data in real time through a camera and microphone.
[1367] Step 2:
[1368] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[1369] Step 3:
[1370] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[1371] Step 4:
[1372] The terminal transmits danger alert information to the server.
[1373] Step 5:
[1374] The server receives the danger alert and sends a push notification to the user terminal.
[1375] Listening function processing steps
[1376] Step 1:
[1377] The device detects what the child is saying through a microphone.
[1378] Step 2:
[1379] The device converts the detected voice data into text and sends it to the AI.
[1380] Step 3:
[1381] The device generates an appropriate response based on the results analyzed by the AI.
[1382] Step 4:
[1383] The device uses voice synthesis to speak gentle words to the child.
[1384] Sensor function processing steps
[1385] Step 1:
[1386] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[1387] Step 2:
[1388] The device continuously collects data from the sensors.
[1389] Step 3:
[1390] The device analyzes the collected data and detects abnormal behavior.
[1391] Step 4:
[1392] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[1393] Step 5:
[1394] The server receives the anomaly alert and sends a push notification to the user terminal, which then plays an appropriate warning sound.
[1395] Go-to-market process steps
[1396] Step 1:
[1397] The server offers the app free of charge to new users.
[1398] Step 2:
[1399] Users can use the basic features free of charge.
[1400] Step 3:
[1401] The server will set up options that offer high-resolution video and additional security features for a fee.
[1402] Step 4:
[1403] If the user selects a paid service, the corresponding function is provided.
[1404] Example 1
[1405] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1406] Today, there are a wide variety of systems available to help parents safely monitor their children. However, many of these systems have drawbacks, such as being expensive or not being able to reuse old mobile communication devices. Furthermore, there are limited means for parents who are in remote locations to check on their children's status in real time and respond appropriately. For these reasons, there is a demand for a monitoring system that is easy to implement, makes effective use of old mobile communication devices, and can ensure the safety of children remotely.
[1407] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1408] In this invention, the server includes a video acquisition means, a video transmission means, a display means, a sensing means, an information notification means, a music playback means, a calling means, a response means, a data collection means, and a means for taking a quick response. This allows parents to use their unused mobile communication devices as surveillance cameras to remotely monitor their children's situations in real time and to respond immediately in the event of an abnormality.
[1409] "Video capture means" refers to technology that uses a camera or other image capture device to capture visual information in real time or at regular intervals.
[1410] "Video transmission means" refers to a technique for transmitting acquired video data to another device via a network.
[1411] "Display means" refers to a display or other display device for visually presenting received video data to a user.
[1412] "Sensing means" refers to technology that analyzes captured video and audio data to detect specific movements and acoustic patterns.
[1413] "Information notification means" refers to notification technology for notifying the user of abnormalities or important information detected by the sensing means.
[1414] "Music playback means" refers to technology including circuitry and software for playing music specified by the user.
[1415] "Hammer" refers to technology that automatically plays a pre-set voice message when a dangerous situation is detected.
[1416] "Response means" refers to technology that detects what the child is saying, generates an appropriate response, and replies in voice.
[1417] "Data collection means" refers to the technology used to collect data from installed sensors and detect abnormal behavior.
[1418] "Means for taking prompt action" refers to technology that uses information notification means to quickly notify the user and prompt them to take the necessary action.
[1419] This invention is a "monitoring and listening system" that allows parents to remotely monitor their children by reusing unused mobile communication devices. The system includes video acquisition means, video transmission means, display means, sensing means, information notification means, music playback means, calling means, response means, data collection means, and means for rapid response.
[1420] First, the user installs a dedicated application on an unused mobile communication device (hereafter referred to as the "device"). This application can be downloaded from the Google Play Store or Apple's App Store. After installing the application, the user configures the camera and enables the device to function as a surveillance camera.
[1421] The device captures real-time video data through the camera and sends it to the server via Wi-Fi or mobile data. The captured video is compressed using an H.264 encoder. The server then delivers the received video data to the user's smartphone as live streaming in an appropriate format (e.g., RTMP). The user can then view the real-time video on their smartphone using a dedicated application.
[1422] The device's built-in microphone and camera detect the baby's movements and cries. Data from these sensors is analyzed in real time using digital signal processing (DSP). Any detected abnormalities that exceed a predetermined threshold are sent to the server as an alert. The server generates a push notification based on the alert information and sends it to the user's smartphone via Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs). This allows the user to quickly learn of any abnormalities in their baby.
[1423] Users can also send instructions to play music from their smartphones through the application. These instructions are sent to the server, which then transmits them to the device. The device receives the instructions, plays the specified music file, and plays the music to the baby through the built-in speaker. The device's built-in media player is used to play the music.
[1424] Furthermore, the device uses AI algorithms to analyze video and audio data in real time, and if it detects dangerous movement, it will automatically play a voice message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[1425] When a child speaks to the device through the microphone, voice data is collected in real time, and the speech recognition engine converts the speech into text. The AI then analyzes the text, generates an appropriate response, and responds gently using the speech synthesis engine.
[1426] Users can also install dedicated sensors in the crib or other dangerous areas. Data from these sensors is collected by the device and processed as an alert when dangerous behavior is detected. The alert information is sent to the server, and the device plays an appropriate warning sound to protect the baby. The server also sends a push notification to the user.
[1427] To illustrate, consider the following prompt:
[1428] 1. "How do I check my baby's real-time video on my smartphone?"
[1429] 2. "How can I get a push notification to my smartphone when my baby cries?"
[1430] 3. "How do I send a music playback command from my smartphone and play music on the device?"
[1431] 4. "How can I detect dangerous baby movements, give voice warnings, and send notifications?"
[1432] 5. "How can I generate appropriate responses and respond to my child's speech using text-to-speech?"
[1433] 6. "How can I have my device sound an alarm and send a notification when the sensor installed in the crib detects dangerous behavior?"
[1434] As described above, this monitoring and listening system provides parents with an effective means of keeping their children safe even from a distance, reducing the burden of childcare. It is also environmentally friendly, as it can be reused from unused mobile communication devices.
[1435] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1436] Camera Linkage Feature Steps
[1437] Step 1: Install and configure the app
[1438] Input: Download a dedicated application to a device that the user no longer uses.
[1439] Specific operation: The user downloads the dedicated application from the Google Play Store or Apple's App Store and installs it on their device.
[1440] Output: The dedicated application is installed on the device and ready to launch.
[1441] Input: The user opens the app's settings screen and configures the camera to launch.
[1442] What it does: When you first open the app, it will show a setup wizard that asks for camera permissions, and you can agree to enable the camera.
[1443] Output: The camera is set to a state where it can be activated on the device.
[1444] Step 2: Capture and transmit real-time video
[1445] Input: Camera powers up and begins collecting real-time footage.
[1446] What it does: The device's camera will start recording real-time footage of your baby or child, and the footage will be compressed using an H.264 encoder.
[1447] Output: Compressed video data is generated.
[1448] Input: Prepare compressed video data for transmission.
[1449] How it works: The device uploads the captured video data to a server in real time via Wi-Fi or mobile data, using HTTP / 2 or WebSocket protocols.
[1450] Output: The compressed video data is sent to the server.
[1451] Step 3: Streaming real-time video
[1452] Input: The server prepares to deliver the received video data to the user's device as live streaming.
[1453] Specific operation: The server converts the video data into an appropriate format and streams it to the user's device using RTMP (Real-time Messaging Protocol).
[1454] Output: The converted video data is delivered to the user's device.
[1455] Input: The user launches the app on their device and checks the real-time video.
[1456] Specific operation: By launching the app on the user's device and selecting the Live Streaming tab, real-time video will be played. The video will be decoded and displayed in a format that is easy for the user to view.
[1457] Output: User can check real-time video.
[1458] Push notification feature steps
[1459] Step 1: Sensing and processing data
[1460] Input: The device's built-in microphone and camera detect baby movements and cries.
[1461] How it works: The device's microphone constantly monitors the surrounding environment, the camera detects movement, and the data is analyzed in real time through digital signal processing (DSP).
[1462] Output: Detected movement and sound data is generated.
[1463] Input: Prepares sensed data to be analyzed.
[1464] How it works: If the baby's crying exceeds a certain decibel level or if any violent movement is detected, the AI module in the device will generate an alert.
[1465] Output: Alert information is generated.
[1466] Step 2: Sending alert information and notifications
[1467] Input: Prepares the generated alert information to be sent to the server.
[1468] Specific operation: The device sends alert information to the server via Wi-Fi or mobile data communication, using the MQTT (Message Queuing Telemetry Transport) protocol.
[1469] Output: Alert information is sent to the server.
[1470] Input: The server prepares to generate a push notification based on the alert information.
[1471] Specific operation: The server analyzes the alert information, generates a push notification, and sends it to the user's device using Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs).
[1472] Output: A push notification is generated and sent to the user device.
[1473] Music playback function steps
[1474] Step 1: Sending a music play command
[1475] Input: The user commands music playback through the app on their smartphone.
[1476] Specific operation: The user selects the music they want to play from the app's music playback menu and taps the play button. Instruction data is generated.
[1477] Output: Music playback instructions are generated.
[1478] Input: Prepares to send music playback instruction data to the server.
[1479] Specific operation: Instruction data generated from the user terminal is sent to the server via the Internet. HTTP / 2 is used as the communication protocol.
[1480] Output: Instruction data is sent to the server.
[1481] Step 2: Transmitting and executing music playback instructions
[1482] Input: The server receives the instruction data and prepares it for transmission to the terminal.
[1483] Specific operation: The server analyzes the received instruction data and sends a playback instruction to the corresponding terminal.
[1484] Output: Playback instructions are sent to the device.
[1485] Input: Prepares the device to play the specified music file.
[1486] Specific operation: The device retrieves and plays the music file from the internal storage or cloud storage based on the received playback command. The device's built-in media player is used to play the audio.
[1487] Output: The specified music file will be played.
[1488] Hazard detection function steps
[1489] Step 1: Analyzing video and audio data
[1490] Input: The device collects video and audio data and analyzes it in real time using AI algorithms.
[1491] Specific operation: The device's AI chip processes video and audio data in real time to detect dangerous movements of the baby.
[1492] Output: Data is generated indicating the dangerous behavior detected.
[1493] Step 2: Play and notify warning messages
[1494] Input: Prepare to play a warning message when dangerous motion is detected.
[1495] Specific operation: Based on the AI detection results, the device plays a pre-set audio message to warn the user.
[1496] Output: A warning audio message is played.
[1497] Input: Sends the danger detection information to the server, which prepares to generate a push notification.
[1498] Specific operation: The danger detection information is sent to the server, and the server generates a push notification and sends it to the user.
[1499] Output: A push notification is generated and sent to the user.
[1500] Steps of active listening function
[1501] Step 1: Detecting speech and collecting voice data
[1502] Input: Detects what your child is saying through the device's microphone.
[1503] How it works: The device's microphone detects what the child is saying and collects audio data, which is then stored in a buffer on the device in real time.
[1504] Output: Speech audio data is collected.
[1505] Step 2: Speech to text conversion and analysis
[1506] Input: Prepare the collected audio data for conversion to text.
[1507] How it works: The device's voice recognition engine converts the voice data into text, and the AI module analyzes the text to generate an appropriate response.
[1508] Output: Parsed text data and response data are generated.
[1509] Step 3: Response generation and speech synthesis
[1510] Input: Converts response data into audio and prepares it for playback.
[1511] Specific operation: The device's speech synthesis engine outputs the generated response as voice and plays it through the speaker.
[1512] Output: An audio response is played to the child.
[1513] Sensor function steps
[1514] Step 1: Sensor installation and data collection
[1515] Input: The user installs special sensors in the crib or other dangerous locations.
[1516] Specific operation: The user places a special sensor in a crib or dangerous location, and the sensor begins to function.
[1517] Output: The sensor is deployed and ready to collect data.
[1518] Input: Prepare to collect data from sensors.
[1519] Specific operation: The device receives data sent from the sensor and detects abnormal movement or falls.
[1520] Output: Data is generated indicating the abnormal behavior detected.
[1521] Step 2: Alerting and Notification
[1522] Input: Prepare to treat anomalous behavior as an alert.
[1523] Specific operation: When the device detects danger, it generates alert information and sends it to the server.
[1524] Output: Alert information is generated and sent to the server.
[1525] Input: The server prepares to generate a push notification based on the alert information.
[1526] What it does: The server sends a push notification to the user, and the device plays a series of warning sounds to protect the baby.
[1527] Output: A push notification is generated and sent to the user. An alert sound is played on the device.
[1528] Market expansion steps
[1529] Step 1: Offer your app and define your market
[1530] Input: The server prepares the app to serve to a new user.
[1531] What happens: The server publishes the app's download link on a website or store, making it free for new users.
[1532] Output: The app download link is published and new users can download the app.
[1533] Step 2: Offering premium services
[1534] Input: The server defines the content of the paid premium service and prepares to provide it to the user.
[1535] Specific operation: Allow users to purchase premium services within the app, and after purchase, the server provides the corresponding services.
[1536] Output: The user purchases a premium service, unlocking additional features.
[1537] (Application example 1)
[1538] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1539] To improve vehicle safety, including that of autonomous vehicles and their drivers, it is necessary to constantly monitor the driver's state and detect abnormalities or dangers early. However, conventional systems lack the ability to detect driver drowsiness or inattention in real time and send warnings. There is also a need for an efficient method of reusing unused mobile communication devices. A system that solves these issues and ensures safety while being economical and environmentally friendly is needed.
[1540] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1541] In this invention, the server is a monitoring system that allows parents to remotely monitor their children by using an unused mobile communication device as a camera, and includes: a camera means for acquiring video; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video; an analysis means for analyzing the video from the mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the detected information; a music playback means for receiving remote operation instructions from the user terminal and playing music; an automatic call means for issuing an automatic call when dangerous movement is detected; a listening means for detecting a child's speech and responding gently; a driver monitoring means for monitoring the driver's state using a camera installed in the vehicle and detecting drowsiness or inattention; and a warning means for sending a warning to the driver and a manager when drowsiness or inattention is detected. This makes it possible to monitor the driver's state in real time and respond immediately if an abnormality is detected.
[1542] A "mobile communications device" is a device used for mobile communications, including smartphones and tablets.
[1543] "Camera means" refers to the camera equipment or functionality used to capture the image.
[1544] "Communication means" refers to communication technology or devices for sending and receiving data or information.
[1545] "User terminal means" refers to a terminal device used by a user, and is a device used for displaying images and remote control.
[1546] "Analysis means" refers to a method or device for analyzing video and audio data and detecting specific movements or situations.
[1547] A "notification means" is a method or device for sending a notification to a user or administrator based on sensed information.
[1548] "Music playback means" refers to a method or device for playing music, including speakers and music playback applications.
[1549] "Automatic calling means" refers to a method or device that automatically generates and plays voice messages under certain circumstances.
[1550] "Listening means" refers to a method or device for detecting and responding to the voice of a subject.
[1551] "Driver monitoring means" refers to a method or device that includes cameras or sensors for monitoring the situation inside the vehicle and, in particular, for checking the driver's condition.
[1552] "Warning means" refers to a method or device for issuing an immediate warning when an abnormality is detected.
[1553] This invention relates to a "monitoring and listening app" that allows parents to remotely monitor their babies and children using unused mobile communication devices. Furthermore, by applying this technology to autonomous vehicles, a system can be realized that monitors the driver and immediately issues a warning in the event of an abnormal or dangerous situation.
[1554] System Configuration
[1555] The server works in conjunction with mobile communication devices, user terminals, sensors, etc. These devices exchange data and analyze it to achieve the following functions:
[1556] Driver status monitoring
[1557] The device is equipped with a camera that captures images of the interior of the vehicle in real time and sends them to a server. This video data is analyzed by the server to detect driver drowsiness and inattention. This detection uses an AI algorithm, specifically TensorFlow, for face detection and movement analysis.
[1558] Warning function
[1559] If an abnormality is detected in the driver, the server will send a warning message to the administrator through the Twilio API, and will use the Google Text-to-Speech (gTTS) library to generate a voice warning and play the warning message aloud, using the pygame library for speech synthesis in this process.
[1560] Music playback function
[1561] The server receives instructions from the user's device and sends a music playback instruction to the device, which then plays the specified music file through its built-in speaker.
[1562] Data analysis
[1563] The video data sent from the device is analyzed by the server. The analysis method is to detect and analyze the driver's face and movements using OpenCV and TensorFlow. If drowsiness or inattention is detected, it is immediately treated as an abnormality and a warning is issued.
[1564] Specific examples
[1565] For example, if a driver falls asleep while driving on a highway, the device's camera will detect this in real time. If the server detects the driver's drowsiness through AI analysis, the system will send an SMS notification to the administrator via the server and simultaneously play an audio warning on the device.
[1566] Prompt Sentence Examples
[1567] "Develop an app that uses mobile communication devices to monitor the driver's status in real time. If the driver becomes inattentive or falls asleep, it should issue a voice warning and send an SMS notification to the manager."
[1568] In this way, the present invention provides a system that can seamlessly monitor and warn drivers, and by reusing unused mobile communication devices, it builds an efficient system that is environmentally friendly.
[1569] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1570] Step 1:
[1571] The device starts the camera and captures the video inside the car. In this step, the camera acquires video data in real time. The input is the video inside the car, and the output is the captured video data. Specifically, the device starts the camera using cv2.VideoCapture(0) and acquires frames using the read() method.
[1572] Step 2:
[1573] The video data captured by the terminal is sent to the server in real time. In this step, data is sent using a communication means. The input is the captured video data, and the output is the video data sent to the server. In concrete terms, the terminal sends the video data to the server using a specified communication protocol (e.g., HTTP POST request).
[1574] Step 3:
[1575] The server analyzes the received video data. In this step, the analysis means is used to determine the driver's state. The input is the transmitted video data, and the output is the analysis result. Specifically, the server processes the video data using OpenCV and TensorFlow to analyze the driver's face and movements.
[1576] Step 4:
[1577] If the server detects an anomaly, it generates a warning message. In this step, the warning means operates based on the analysis results. The input is the analysis result, and the output is the warning message. Specifically, when an anomaly is detected, the server generates a voice warning using the gTTS library and sends an SMS notification using the Twilio API.
[1578] Step 5:
[1579] The terminal plays music according to instructions from the server. In this step, a music playback means is used. The input is a remote control instruction from the user terminal, and the output is the music being played. In concrete terms, the terminal plays music on its built-in speaker based on the instruction received from the server.
[1580] Step 6:
[1581] The device detects what the child is saying and responds appropriately. In this step, the listening means is activated. The input is what the child is saying, and the output is a voice response. Specifically, the device detects the voice through the microphone and sends it to the server. The server converts it into text using a generative AI model and generates an appropriate response.
[1582] Step 7:
[1583] The server comprehensively manages all data and sends notifications to user devices in the event of an abnormality. This step is responsible for overall control. Input is data from various sensors and cameras, and output is notifications to administrators and push notifications to user devices. Specifically, the server receives and analyzes all data, and immediately notifies users if an abnormality is detected.
[1584] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1585] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely watch over babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals"), with an emotion engine that recognizes the user's emotions.
[1586] Embodiment of camera linkage function
[1587] First, the user installs a dedicated app on the device and configures the camera. This activates the device's camera. The device captures real-time video through the camera and transmits the video data to the server. The server then streams the received data to the user's smartphone (hereinafter referred to as the "user device"). The user can then view the live video on their own device.
[1588] 1. Embodiment of push notification function
[1589] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device. The user can then respond promptly via the push notification.
[1590] Embodiment of music playback function
[1591] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server. The device then plays the specified music and plays it to the baby through the built-in speaker.
[1592] Embodiment of the hazard detection function
[1593] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, danger alert information is sent to the server, which then warns the user via push notification.
[1594] An embodiment of the listening function
[1595] The device detects what the child is saying through a microphone and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to gently reply to the child, allowing the user to engage with their child even when they are busy.
[1596] Sensor Function Embodiment
[1597] The user installs dedicated sensors in the crib or other dangerous locations and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The detected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[1598] Embodiment of Emotion Engine
[1599] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The response is converted from text to speech and played through the device's speaker.
[1600] This allows the emotion engine to reduce the user's mental burden and more effectively support the care of babies and children. The combination of the emotion engine improves the functionality and flexibility of the entire system, making it an important tool for supporting childcare.
[1601] The processing flow will be explained below.
[1602] Camera linkage function processing steps
[1603] Step 1:
[1604] Users install a dedicated app on their unused devices and configure the camera settings, which activates the device's camera.
[1605] Step 2:
[1606] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[1607] Step 3:
[1608] The terminal transmits the captured video data to the server in real time.
[1609] Step 4:
[1610] The server receives the video data and starts streaming it to the user terminal.
[1611] Step 5:
[1612] The user launches the app on the user device and checks the live video being streamed.
[1613] Push notification processing steps
[1614] Step 1:
[1615] The device's built-in microphone and camera are used to detect the baby's movements and cries.
[1616] Step 2:
[1617] The terminal analyzes the sensed data in real time and generates alert information when a certain threshold is exceeded.
[1618] Step 3:
[1619] The terminal transmits the generated alert information to the server.
[1620] Step 4:
[1621] The server receives the alert information and sends a push notification to the user terminal.
[1622] Step 5:
[1623] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[1624] Processing steps for music playback function
[1625] Step 1:
[1626] The user sends a music playback instruction to the server from the user terminal via the app.
[1627] Step 2:
[1628] The server receives the user's instruction to play music and transmits it to the terminal.
[1629] Step 3:
[1630] The terminal plays the specified music file based on the instruction received from the server.
[1631] Step 4:
[1632] The device uses a built-in speaker to play music for your baby.
[1633] Hazard detection function processing steps
[1634] Step 1:
[1635] The device collects data in real time through a camera and microphone.
[1636] Step 2:
[1637] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[1638] Step 3:
[1639] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[1640] Step 4:
[1641] The terminal transmits danger alert information to the server.
[1642] Step 5:
[1643] The server receives the danger alert and sends a push notification to the user terminal.
[1644] Listening function processing steps
[1645] Step 1:
[1646] It detects what the child is saying through the device's microphone.
[1647] Step 2:
[1648] The device converts the detected voice data into text and sends it to the AI.
[1649] Step 3:
[1650] The device generates an appropriate response based on the results analyzed by the AI.
[1651] Step 4:
[1652] The device uses voice synthesis to speak soothing words to the child.
[1653] Sensor function processing steps
[1654] Step 1:
[1655] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[1656] Step 2:
[1657] The device continuously collects data from connected sensors.
[1658] Step 3:
[1659] The device analyzes the collected data and detects abnormal behavior.
[1660] Step 4:
[1661] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[1662] Step 5:
[1663] The server receives the anomaly alert and sends a push notification to the user device, which simultaneously plays an appropriate warning sound.
[1664] Emotion Engine Processing Steps
[1665] Step 1:
[1666] The user's voice data is transmitted from the user terminal to the server.
[1667] Step 2:
[1668] The server sends the received voice data to the emotion engine on the device.
[1669] Step 3:
[1670] The device's emotion engine analyzes the voice data and determines the user's emotional state.
[1671] Step 4:
[1672] The terminal's emotion engine generates an appropriate response based on the determined emotional state.
[1673] Step 5:
[1674] The terminal converts the generated response from text to speech and provides feedback to the user.
[1675] Example 2
[1676] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1677] In modern child-rearing, it is difficult for parents to constantly keep an eye on their babies and children, and the burden is particularly heavy when both parents are working or raising children alone. Conventional monitoring systems have limitations in real-time monitoring and voice communication, and lack emotion recognition functionality, making it difficult to reduce the mental burden on parents. Furthermore, they lack the ability to instantly detect dangerous situations and respond appropriately. The purpose of this invention is to solve these problems, more effectively ensure the safety of babies and children, and reduce the burden on parents.
[1678] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a camera means for capturing video from an unused mobile communication device; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video via the communication means; an analysis means for analyzing the video from the unused mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the information detected by the analysis means; a music playback means for receiving remote operation instructions from the user terminal and playing music; an emotion analysis means for analyzing audio data transmitted from the unused mobile communication device and determining the user's emotional state; an automatic call means for issuing an automatic call when dangerous movements are detected; and a listening means for detecting and gently responding to the child's voice. This allows for a multifunctional and flexible monitoring system to be constructed, enabling parents to monitor and respond to the condition of babies and children even remotely. Furthermore, emotion analysis reduces the user's mental burden and enables quick and appropriate responses to dangerous situations.
[1679] "Disused mobile communication devices" are mobile devices such as mobile phones and tablets that have been retired but are now being reused.
[1680] "Camera means" refers to a camera device installed to capture video.
[1681] "Communication means" refers to a means having an internet connection function for sending and receiving images and data.
[1682] "User terminal means" refers to a display device used by a user, such as a smartphone, tablet, or computer.
[1683] The "analysis means" is a means having the function of analyzing the collected data and detecting specific information such as the baby's movements and cries.
[1684] The "notification means" is a means having a function for generating a push notification based on sensed information and transmitting it to the user terminal.
[1685] The "music playback means" is a means having a function of playing music in response to an instruction from a user terminal.
[1686] The "emotion analysis means" is a means having a function of analyzing voice data and determining the emotional state of the user.
[1687] The "automatic calling means" is a means having a function of automatically playing a warning message or the like when a dangerous movement is detected.
[1688] "Listening tools" are tools that have the ability to sense what a child is saying, generate an appropriate response, and respond in a gentle manner.
[1689] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely monitor babies and children by reusing unused mobile communication devices with an emotion engine that recognizes the user's emotions. This system provides a multifunctional and flexible monitoring environment.
[1690] Camera linkage function
[1691] First, the user installs a dedicated app on their device and configures the camera. Once the app is installed and configured, the device's camera is activated. The device captures real-time video through the camera and sends the data to the server. The server receives the video data and streams it to the user's device, allowing the user to view live video on their device.
[1692] Specific examples
[1693] Device used: Old smartphone
[1694] Server used: AWS EC2 instance
[1695] Video streaming software: FFmpeg
[1696] Push notification function
[1697] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device, allowing the user to respond quickly.
[1698] Specific examples
[1699] Hardware used: Android or iOS device
[1700] Server: Google Firebase Cloud Messaging
[1701] App: Real-time analytics with custom mobile app
[1702] Music playback function
[1703] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server, and the device plays the specified music on its built-in speaker.
[1704] Specific examples
[1705] Device used: Old tablet
[1706] Server: AWS Lambda
[1707] Music playback software: VLC
[1708] Hazard detection function
[1709] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, the device sends danger alert information to the server, which then sends a push notification to the user's device.
[1710] Specific examples
[1711] Device used: Raspberry Pi 4 with camera module
[1712] AI algorithm: TensorFlow
[1713] Server: Google Cloud Functions
[1714] Active listening function
[1715] The device detects what the child is saying through a microphone and converts the voice data into text. This text data is analyzed using an AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The generated response is then converted into audio data using speech synthesis technology, which is then played back by the device.
[1716] Specific examples
[1717] Device used: Amazon Echo device
[1718] Speech Recognition: Google Speech-to-Text API
[1719] Response generation: OpenAI GPT-3
[1720] Speech synthesis: Amazon Polly
[1721] Sensor function
[1722] The user installs dedicated sensors in the crib or other dangerous areas and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The collected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[1723] Specific examples
[1724] Sensor used: Xiaomi Smart Home Sensor
[1725] Device used: Old smartphone
[1726] Server: AWS IoT Core
[1727] Emotion Engine
[1728] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The generated response is converted into voice data and played through the device's speaker.
[1729] Specific examples
[1730] Device used: Old smartphone
[1731] Emotion engine: IBM Watson Tone Analyzer
[1732] Speech synthesis: Google Text-to-Speech
[1733] This makes the entire system a versatile and flexible monitoring environment, making it a very useful tool for supporting childcare. Users can monitor the condition of their baby or child in real time, even remotely, and respond quickly if necessary. The emotion engine also reduces mental strain.
[1734] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1735] Camera linkage function processing steps
[1736] Step 1:
[1737] Users install a dedicated app on their device, open the app and configure the camera settings.
[1738] Input: App installation, setting information
[1739] Output: Setup complete, camera ready to start
[1740] Step 2:
[1741] Once the setup is complete, the device will launch the camera.
[1742] Input: Setting completion notification
[1743] Output: Camera starts, video capture begins
[1744] Step 3:
[1745] The device captures real-time video through a camera and transmits the data to a server.
[1746] Input: Real-time video data
[1747] Output: Ready to send, sent to server
[1748] Step 4:
[1749] The server processes the received video data and distributes it as a stream to the user terminal.
[1750] Input: Video data
[1751] Output: Video streaming ready, sent to user device
[1752] Step 5:
[1753] Users can view live footage via the app on their devices.
[1754] Input: Video streaming data
[1755] Output: Video display, real-time monitoring
[1756] Push notification processing steps
[1757] Step 1:
[1758] It analyzes data collected from the device's built-in microphone and camera in real time.
[1759] Input: Audio data, video data
[1760] Output: Analysis results, threshold judgment
[1761] Step 2:
[1762] If the baby's movements or cries exceed the threshold, the device will generate an alert.
[1763] Input: Analysis results
[1764] Output: Alert information
[1765] Step 3:
[1766] The terminal transmits the generated alert information to the server.
[1767] Input: Alert information
[1768] Output: Sent to server
[1769] Step 4:
[1770] The server generates a push notification based on the alert information and sends it to the user's device.
[1771] Input: Alert information
[1772] Output: Push notification
[1773] Step 5:
[1774] Users receive push notifications and can respond quickly.
[1775] Input: Push notification
[1776] Output: Notification confirmation, required action
[1777] Processing steps for music playback function
[1778] Step 1:
[1779] The user operates the app on their smartphone and sends instructions to play music.
[1780] Input: Music playback instructions
[1781] Output: Send instructions
[1782] Step 2:
[1783] The server transfers the playback instruction received from the user to the terminal.
[1784] Input: Playback instructions
[1785] Output: Directed transfer
[1786] Step 3:
[1787] The device will play music through its built-in speaker based on the transferred instructions.
[1788] Input: Playback instructions
[1789] Output: Music playback
[1790] Hazard detection function processing steps
[1791] Step 1:
[1792] The device collects the baby's movements and voice data using a camera and microphone.
[1793] Input: Video data, audio data
[1794] Output: Recorded data, audio data
[1795] Step 2:
[1796] The device analyzes the collected data using AI algorithms to detect dangerous behavior.
[1797] Input: Recorded data, audio data
[1798] Output: Risk assessment result
[1799] Step 3:
[1800] When danger is detected, the device will play an automated voice message such as "Danger! Stop!"
[1801] Input: Risk assessment result
[1802] Output: Automated voice message
[1803] Step 4:
[1804] The terminal transmits danger alert information to the server.
[1805] Input: Risk assessment result
[1806] Output: Alert information
[1807] Step 5:
[1808] The server generates a push notification and sends it to the user device.
[1809] Input: Alert information
[1810] Output: Push notification
[1811] Listening function processing steps
[1812] Step 1:
[1813] The device detects what the child is saying through a microphone.
[1814] Input: Audio data
[1815] Output: Sensor information
[1816] Step 2:
[1817] The device converts the detected voice data into text.
[1818] Input: Audio data
[1819] Output: Text data
[1820] Step 3:
[1821] The converted text data is analyzed by an AI model to generate an appropriate response.
[1822] Input: Text data
[1823] Output: Response data
[1824] Step 4:
[1825] The terminal converts the generated response into voice data using voice synthesis technology.
[1826] Input: Response data
[1827] Output: Audio data
[1828] Step 5:
[1829] The device plays back the audio data and provides an appropriate response to the child.
[1830] Input: Audio data
[1831] Output: Response voice
[1832] Sensor function processing steps
[1833] Step 1:
[1834] Users install special sensors in cribs and dangerous areas and connect them to a terminal.
[1835] Input: Sensor installation
[1836] Output: Connection complete
[1837] Step 2:
[1838] Sensors continuously collect environmental data and detect abnormal behavior.
[1839] Input: Environment data
[1840] Output: Data collection, anomaly detection information
[1841] Step 3:
[1842] The sensor transmits the collected data to the terminal.
[1843] Input: Data collection, anomaly detection information
[1844] Output: Send to terminal
[1845] Step 4:
[1846] The terminal analyzes the received sensor data internally and detects any abnormalities.
[1847] Input: Sensor data
[1848] Output: Abnormality detection result
[1849] Step 5:
[1850] The device generates an alert if an anomaly is detected.
[1851] Input: Abnormality detection result
[1852] Output: Alert information
[1853] Step 6:
[1854] The terminal transmits the alert information to the server.
[1855] Input: Alert information
[1856] Output: Sent to server
[1857] Step 7:
[1858] The server generates a push notification and sends it to the user device.
[1859] Input: Alert information
[1860] Output: Push notification
[1861] Emotion Engine Processing Steps
[1862] Step 1:
[1863] The user transmits voice data to the terminal.
[1864] Input: Audio data
[1865] Output: Send to terminal
[1866] Step 2:
[1867] The device uses an emotion engine to analyze the voice data and determine the user's emotional state.
[1868] Input: Audio data
[1869] Output: Emotion determination result
[1870] Step 3:
[1871] The emotion engine generates an appropriate response based on the judgment result.
[1872] Input: Emotion judgment result
[1873] Output: Response data
[1874] Step 4:
[1875] The terminal converts the generated response into voice data using voice synthesis technology.
[1876] Input: Response data
[1877] Output: Audio data
[1878] Step 5:
[1879] The terminal plays back the audio data and responds appropriately to reduce the user's mental burden.
[1880] Input: Audio data
[1881] Output: Response voice
[1882] (Application example 2)
[1883] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1884] The problem that the present invention aims to solve is to provide a system that effectively utilizes unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children and increase their safety. Furthermore, because conventional monitoring systems do not take emotion recognition into consideration, a means for reducing the mental burden on users was also necessary. Therefore, the present invention aims to compensate for the shortcomings of conventional monitoring systems and improve safety and usability.
[1885] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a communication means for transmitting video in real time, an emotion recognition means for analyzing audio data and recognizing emotions, and a response generation means for generating an appropriate response based on the recognized emotion and outputting it as audio. This makes it possible not only to perform remote real-time monitoring, but also to recognize the emotional state of the user and generate an appropriate response.
[1886] The "camera means" is a means having a function for acquiring an image.
[1887] The "communication means" is a means having a function for transmitting acquired video in real time.
[1888] The "user terminal means" is a means having a function of receiving and displaying video transmitted via a communication means.
[1889] The "analysis means" is a means having the function of analyzing video acquired from a disused mobile communication device and detecting specific movements or sounds.
[1890] The "notification means" is a means having a function of sending a push notification to the user terminal based on the information sensed by the analysis means.
[1891] The "music playback means" is a means having a function of receiving a remote operation instruction from a user terminal and playing music.
[1892] The "automatic call means" is a means having a function of automatically playing back a voice message when a dangerous movement is detected.
[1893] "Listening means" is a means that has the function of detecting what a child is saying and responding gently.
[1894] The "emotion recognition means" is a means having a function of analyzing voice data and recognizing the emotional state of the user.
[1895] The "response generation means" is a means having a function of generating an appropriate response based on the recognized emotion and outputting the response as voice.
[1896] The present invention relates to a system that makes effective use of unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children. This system uses devices and software that include the following elements:
[1897] First, a disused mobile communication device (hereinafter referred to as a terminal) captures video using a camera means. The terminal transmits the video captured by the camera to a server in real time via a communication means. The server can stream the received video to a user terminal means. The user can then view the live video on their own terminal.
[1898] Data collected from the device's built-in microphone and camera is analyzed in real time by the analysis means. This analysis detects the movement and cries of babies and children, and if a threshold is exceeded, an alert is generated by the notification means. These alerts are sent to the user's device via the server as push notifications. The user can respond promptly via the push notifications.
[1899] Furthermore, the user can remotely use the music playback means to send instructions to the terminal to play music. This instruction is transferred to the terminal via the server, and the specified music is played by the baby or child.
[1900] Furthermore, the device is equipped with a danger detection function. Data obtained from the camera and microphone is analyzed using an AI algorithm to detect if the baby is behaving in a dangerous manner. If danger is detected, the device will play an automated voice message such as "Danger! Stop!" At the same time, danger alert information is sent to the server, and a warning is sent to the user.
[1901] The device also has a listening function, which uses a microphone to detect what the child is saying and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to respond gently to the child, allowing the user to interact with their child appropriately even when they are busy.
[1902] The device is also equipped with an emotion recognition means that analyzes the voice data sent from the user device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion recognition means will recognize this and utilize the response generation means to advise an appropriate response. This reduces the user's mental burden and provides support for childcare and monitoring.
[1903] For example, if an employee of Guard AI is wearing smart glasses, the following prompt can trigger the emotion engine and provide appropriate advice:
[1904] Example prompt sentence:
[1905] Employee dictation: "I'm feeling really stressed today..."
[1906] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[1907] In this way, the system of the present invention can achieve both enhanced security and psychological care for the user.
[1908] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1909] Step 1:
[1910] The terminal acquires video using a camera means. The input is real-time video data from the camera, which is captured and stored in a buffer. The output is the captured video data.
[1911] Step 2:
[1912] The device uses a communication means to send the captured video data to the server. The input is the captured video data, which is then encoded into JPEG or MP4 format. The output is the encoded video data.
[1913] Step 3:
[1914] The server streams the received video data to the user terminal means. The input is the video data sent from the terminal and is transferred to the user terminal in real time. The output is the live video displayed on the user terminal.
[1915] Step 4:
[1916] The data collected from the device's built-in microphone and camera is analyzed by an analysis means. The input is audio and video data from the microphone and camera, which is fed into an AI algorithm for analysis. The output is the analysis results, which are the movements and cries of the baby or child.
[1917] Step 5:
[1918] If the analysis result exceeds a certain threshold, an alert is generated by the notification mechanism. The input is the analysis result, which is used to evaluate the alert condition. The output is the generated alert.
[1919] Step 6:
[1920] The server sends the generated alert to the user device as a push notification. The input is the generated alert, which is converted into a push notification format. The output is a push notification that is displayed on the user device.
[1921] Step 7:
[1922] The user remotely uses the music playback means to send a music playback instruction to the terminal. The input is the music playback instruction from the user terminal, which is sent to the terminal via the server. The output is the music played on the terminal.
[1923] Step 8:
[1924] When the terminal detects a dangerous movement, it plays a voice message using the automatic call means. The input is the detection result of the analysis means, and if it is judged to be dangerous, the voice message is played. The output is a warning voice played from the terminal.
[1925] Step 9:
[1926] The terminal uses a listening means to detect what the child is saying and converts the voice data into text. The input is the voice data of the child's speech, which is converted into text data. The output is text data.
[1927] Step 10:
[1928] The text data is analyzed by AI to generate an appropriate response. The input is the converted text data, and the natural language processing algorithm generates a response based on this. The output is the response text.
[1929] Step 11:
[1930] The response generation means outputs the generated response text as voice using speech synthesis technology. The input is the response text, which is converted into voice data by a speech synthesis engine. The output is the response voice played from the terminal.
[1931] Step 12:
[1932] The emotion recognition means analyzes the voice data sent from the user terminal and determines the user's emotional state. The input is the user's voice data, which is supplied to a voice analysis algorithm to recognize the emotional state. The output is the determination result of the emotional state.
[1933] Step 13:
[1934] Based on the recognized emotional state, the response generator generates an appropriate response. The input is the result of the emotional state determination, and the response content is determined based on this. The output is the generated response content.
[1935] Step 14:
[1936] The response is output as voice through the device's speaker. The input is the generated response, which is played back as voice using speech synthesis technology. The output is a voice response played from the device.
[1937] Example prompt sentence:
[1938] Employee dictation: "I'm feeling really stressed today..."
[1939] AI response: "You seem tired lately. Please take a break to refresh yourself and ask for help if you need it."
[1940] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1941] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1942] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1943] [Fourth embodiment]
[1944] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1945] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1946] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1947] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1948] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1949] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1950] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1951] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1952] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1953] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1954] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1955] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1956] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1957] This invention relates to a "monitoring and listening app" that allows parents to remotely watch over their babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals").
[1958] Embodiment of camera linkage function
[1959] First, the user installs a dedicated app on their device and configures the camera. The device captures real-time video through the camera and sends the video data to a server. The server then distributes this data as live streaming to the user's smartphone (hereinafter referred to as the "user device"). When the user launches the app on the user device, they can check on their baby's condition in real time from a remote location.
[1960] 1. Embodiment of push notification function
[1961] The device is equipped with a built-in microphone and camera, which are used to detect the baby's movements and cries. The detected data is processed within the device, and if a predetermined threshold is exceeded, an alert is generated. The generated alert information is sent to a server, which then sends a push notification to the user based on the alert information. This allows the user to immediately become aware of any abnormalities in the baby.
[1962] Embodiment of music playback function
[1963] When a user issues a music playback command via the app on their smartphone, the command is sent to the server, which then receives the command and transmits it to the device. The device then plays the specified music file, allowing the baby to listen to the music through the built-in speaker.
[1964] Embodiment of the hazard detection function
[1965] The device is equipped with an AI algorithm that analyzes video and audio data in real time. If it detects that the baby is making dangerous movements, the device will automatically play an audio message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[1966] An embodiment of the listening function
[1967] The device's microphone detects what the child is saying and collects voice data in real time. This voice data is converted into text on the device and analyzed by AI. Based on the analysis results, the AI generates an appropriate response and responds gently to the child through voice synthesis. This allows users to be there for their children even when they are busy.
[1968] Sensor Function Embodiment
[1969] The user installs dedicated sensors in the crib and other dangerous locations. The device collects data from these sensors and processes it as an alert if it detects any dangerous behavior (e.g., a fall or abnormal movement). The alert information is sent to the server, and the device simultaneously plays an appropriate warning sound to protect the baby. The server then sends a push notification to the user to notify them of the abnormality.
[1970] Market deployment implementation example
[1971] The server will offer the app free of charge to new users. The basic features are available for free, but if users want higher resolution video or additional security features, they can purchase premium services.
[1972] As described above, the present invention provides a monitoring system that utilizes unused mobile communication devices, reduces the burden on parents in raising their children, and is also environmentally friendly.
[1973] The processing flow will be explained below.
[1974] Camera linkage function processing steps
[1975] Step 1:
[1976] Users install a dedicated app on their unused devices, open the app, and configure the camera settings, which activates the device's camera.
[1977] Step 2:
[1978] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[1979] Step 3:
[1980] The terminal transmits the captured video data to the server in real time.
[1981] Step 4:
[1982] The server receives the video data and starts streaming it to the user terminal.
[1983] Step 5:
[1984] The user launches the app on the user device and checks the live video being streamed.
[1985] Push notification processing steps
[1986] Step 1:
[1987] The device uses a built-in microphone and camera to detect the baby's movements and cries.
[1988] Step 2:
[1989] The terminal analyzes the sensed data and generates alert information when a certain threshold is exceeded.
[1990] Step 3:
[1991] The terminal transmits the generated alert information to the server.
[1992] Step 4:
[1993] The server receives the alert information and sends a push notification to the user terminal.
[1994] Step 5:
[1995] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[1996] Processing steps for music playback function
[1997] Step 1:
[1998] The user sends a music playback instruction to the server from the user terminal via the app.
[1999] Step 2:
[2000] The server receives the user's instruction to play music and transmits it to the terminal.
[2001] Step 3:
[2002] The terminal plays the specified music file based on the instruction received from the server.
[2003] Step 4:
[2004] The device uses a built-in speaker to play music for your baby.
[2005] Hazard detection function processing steps
[2006] Step 1:
[2007] The device collects data in real time through a camera and microphone.
[2008] Step 2:
[2009] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[2010] Step 3:
[2011] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[2012] Step 4:
[2013] The terminal transmits danger alert information to the server.
[2014] Step 5:
[2015] The server receives the danger alert and sends a push notification to the user terminal.
[2016] Listening function processing steps
[2017] Step 1:
[2018] The device detects what the child is saying through a microphone.
[2019] Step 2:
[2020] The device converts the detected voice data into text and sends it to the AI.
[2021] Step 3:
[2022] The device generates an appropriate response based on the results analyzed by the AI.
[2023] Step 4:
[2024] The device uses voice synthesis to speak gentle words to the child.
[2025] Sensor function processing steps
[2026] Step 1:
[2027] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[2028] Step 2:
[2029] The device continuously collects data from the sensors.
[2030] Step 3:
[2031] The device analyzes the collected data and detects abnormal behavior.
[2032] Step 4:
[2033] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[2034] Step 5:
[2035] The server receives the anomaly alert and sends a push notification to the user terminal, which then plays an appropriate warning sound.
[2036] Go-to-market process steps
[2037] Step 1:
[2038] The server offers the app free of charge to new users.
[2039] Step 2:
[2040] Users can use the basic features free of charge.
[2041] Step 3:
[2042] The server will set up options that offer high-resolution video and additional security features for a fee.
[2043] Step 4:
[2044] If the user selects a paid service, the corresponding function is provided.
[2045] Example 1
[2046] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2047] Today, there are a wide variety of systems available to help parents safely monitor their children. However, many of these systems have drawbacks, such as being expensive or not being able to reuse old mobile communication devices. Furthermore, there are limited means for parents who are in remote locations to check on their children's status in real time and respond appropriately. For these reasons, there is a demand for a monitoring system that is easy to implement, makes effective use of old mobile communication devices, and can ensure the safety of children remotely.
[2048] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2049] In this invention, the server includes a video acquisition means, a video transmission means, a display means, a sensing means, an information notification means, a music playback means, a calling means, a response means, a data collection means, and a means for taking a quick response. This allows parents to use their unused mobile communication devices as surveillance cameras to remotely monitor their children's situations in real time and to respond immediately in the event of an abnormality.
[2050] "Video capture means" refers to technology that uses a camera or other image capture device to capture visual information in real time or at regular intervals.
[2051] "Video transmission means" refers to a technique for transmitting acquired video data to another device via a network.
[2052] "Display means" refers to a display or other display device for visually presenting received video data to a user.
[2053] "Sensing means" refers to technology that analyzes captured video and audio data to detect specific movements and acoustic patterns.
[2054] "Information notification means" refers to notification technology for notifying the user of abnormalities or important information detected by the sensing means.
[2055] "Music playback means" refers to technology including circuitry and software for playing music specified by the user.
[2056] "Hammer" refers to technology that automatically plays a pre-set voice message when a dangerous situation is detected.
[2057] "Response means" refers to technology that detects what the child is saying, generates an appropriate response, and replies in voice.
[2058] "Data collection means" refers to the technology used to collect data from installed sensors and detect abnormal behavior.
[2059] "Means for taking prompt action" refers to technology that uses information notification means to quickly notify the user and prompt them to take the necessary action.
[2060] This invention is a "monitoring and listening system" that allows parents to remotely monitor their children by reusing unused mobile communication devices. The system includes video acquisition means, video transmission means, display means, sensing means, information notification means, music playback means, calling means, response means, data collection means, and means for rapid response.
[2061] First, the user installs a dedicated application on an unused mobile communication device (hereafter referred to as the "device"). This application can be downloaded from the Google Play Store or Apple's App Store. After installing the application, the user configures the camera and enables the device to function as a surveillance camera.
[2062] The device captures real-time video data through the camera and sends it to the server via Wi-Fi or mobile data. The captured video is compressed using an H.264 encoder. The server then delivers the received video data to the user's smartphone as live streaming in an appropriate format (e.g., RTMP). The user can then view the real-time video on their smartphone using a dedicated application.
[2063] The device's built-in microphone and camera detect the baby's movements and cries. Data from these sensors is analyzed in real time using digital signal processing (DSP). Any detected abnormalities that exceed a predetermined threshold are sent to the server as an alert. The server generates a push notification based on the alert information and sends it to the user's smartphone via Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs). This allows the user to quickly learn of any abnormalities in their baby.
[2064] Users can also send instructions to play music from their smartphones through the application. These instructions are sent to the server, which then transmits them to the device. The device receives the instructions, plays the specified music file, and plays the music to the baby through the built-in speaker. The device's built-in media player is used to play the music.
[2065] Furthermore, the device uses AI algorithms to analyze video and audio data in real time, and if it detects dangerous movement, it will automatically play a voice message saying, "Danger! Stop!" This information is also sent to the server, which then sends a push notification to the user to warn them.
[2066] When a child speaks to the device through the microphone, voice data is collected in real time, and the speech recognition engine converts the speech into text. The AI then analyzes the text, generates an appropriate response, and responds gently using the speech synthesis engine.
[2067] Users can also install dedicated sensors in the crib or other dangerous areas. Data from these sensors is collected by the device and processed as an alert when dangerous behavior is detected. The alert information is sent to the server, and the device plays an appropriate warning sound to protect the baby. The server also sends a push notification to the user.
[2068] To illustrate, consider the following prompt:
[2069] 1. "How do I check my baby's real-time video on my smartphone?"
[2070] 2. "How can I get a push notification to my smartphone when my baby cries?"
[2071] 3. "How do I send a music playback command from my smartphone and play music on the device?"
[2072] 4. "How can I detect dangerous baby movements, give voice warnings, and send notifications?"
[2073] 5. "How can I generate appropriate responses and respond to my child's speech using text-to-speech?"
[2074] 6. "How can I have my device sound an alarm and send a notification when the sensor installed in the crib detects dangerous behavior?"
[2075] As described above, this monitoring and listening system provides parents with an effective means of keeping their children safe even from a distance, reducing the burden of childcare. It is also environmentally friendly, as it can be reused from unused mobile communication devices.
[2076] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2077] Camera Linkage Feature Steps
[2078] Step 1: Install and configure the app
[2079] Input: Download a dedicated application to a device that the user no longer uses.
[2080] Specific operation: The user downloads the dedicated application from the Google Play Store or Apple's App Store and installs it on their device.
[2081] Output: The dedicated application is installed on the device and ready to launch.
[2082] Input: The user opens the app's settings screen and configures the camera to launch.
[2083] What it does: When you first open the app, it will show a setup wizard that asks for camera permissions, and you can agree to enable the camera.
[2084] Output: The camera is set to a state where it can be activated on the device.
[2085] Step 2: Capture and transmit real-time video
[2086] Input: Camera powers up and begins collecting real-time footage.
[2087] What it does: The device's camera will start recording real-time footage of your baby or child, and the footage will be compressed using an H.264 encoder.
[2088] Output: Compressed video data is generated.
[2089] Input: Prepare compressed video data for transmission.
[2090] How it works: The device uploads the captured video data to a server in real time via Wi-Fi or mobile data, using HTTP / 2 or WebSocket protocols.
[2091] Output: The compressed video data is sent to the server.
[2092] Step 3: Streaming real-time video
[2093] Input: The server prepares to deliver the received video data to the user's device as live streaming.
[2094] Specific operation: The server converts the video data into an appropriate format and streams it to the user's device using RTMP (Real-time Messaging Protocol).
[2095] Output: The converted video data is delivered to the user's device.
[2096] Input: The user launches the app on their device and checks the real-time video.
[2097] Specific operation: By launching the app on the user's device and selecting the Live Streaming tab, real-time video will be played. The video will be decoded and displayed in a format that is easy for the user to view.
[2098] Output: User can check real-time video.
[2099] Push notification feature steps
[2100] Step 1: Sensing and processing data
[2101] Input: The device's built-in microphone and camera detect baby movements and cries.
[2102] How it works: The device's microphone constantly monitors the surrounding environment, the camera detects movement, and the data is analyzed in real time through digital signal processing (DSP).
[2103] Output: Detected movement and sound data is generated.
[2104] Input: Prepares sensed data to be analyzed.
[2105] How it works: If the baby's crying exceeds a certain decibel level or if any violent movement is detected, the AI module in the device will generate an alert.
[2106] Output: Alert information is generated.
[2107] Step 2: Sending alert information and notifications
[2108] Input: Prepares the generated alert information to be sent to the server.
[2109] Specific operation: The device sends alert information to the server via Wi-Fi or mobile data communication, using the MQTT (Message Queuing Telemetry Transport) protocol.
[2110] Output: Alert information is sent to the server.
[2111] Input: The server prepares to generate a push notification based on the alert information.
[2112] Specific operation: The server analyzes the alert information, generates a push notification, and sends it to the user's device using Firebase Cloud Messaging (FCM) or Apple Push Notification Service (APNs).
[2113] Output: A push notification is generated and sent to the user device.
[2114] Music playback function steps
[2115] Step 1: Sending a music play command
[2116] Input: The user commands music playback through the app on their smartphone.
[2117] Specific operation: The user selects the music they want to play from the app's music playback menu and taps the play button. Instruction data is generated.
[2118] Output: Music playback instructions are generated.
[2119] Input: Prepares to send music playback instruction data to the server.
[2120] Specific operation: Instruction data generated from the user terminal is sent to the server via the Internet. HTTP / 2 is used as the communication protocol.
[2121] Output: Instruction data is sent to the server.
[2122] Step 2: Transmitting and executing music playback instructions
[2123] Input: The server receives the instruction data and prepares it for transmission to the terminal.
[2124] Specific operation: The server analyzes the received instruction data and sends a playback instruction to the corresponding terminal.
[2125] Output: Playback instructions are sent to the device.
[2126] Input: Prepares the device to play the specified music file.
[2127] Specific operation: The device retrieves and plays the music file from the internal storage or cloud storage based on the received playback command. The device's built-in media player is used to play the audio.
[2128] Output: The specified music file will be played.
[2129] Hazard detection function steps
[2130] Step 1: Analyzing video and audio data
[2131] Input: The device collects video and audio data and analyzes it in real time using AI algorithms.
[2132] Specific operation: The device's AI chip processes video and audio data in real time to detect dangerous movements of the baby.
[2133] Output: Data is generated indicating the dangerous behavior detected.
[2134] Step 2: Play and notify warning messages
[2135] Input: Prepare to play a warning message when dangerous motion is detected.
[2136] Specific operation: Based on the AI detection results, the device plays a pre-set audio message to warn the user.
[2137] Output: A warning audio message is played.
[2138] Input: Sends the danger detection information to the server, which prepares to generate a push notification.
[2139] Specific operation: The danger detection information is sent to the server, and the server generates a push notification and sends it to the user.
[2140] Output: A push notification is generated and sent to the user.
[2141] Steps of active listening function
[2142] Step 1: Detecting speech and collecting voice data
[2143] Input: Detects what your child is saying through the device's microphone.
[2144] How it works: The device's microphone detects what the child is saying and collects audio data, which is then stored in a buffer on the device in real time.
[2145] Output: Speech audio data is collected.
[2146] Step 2: Speech to text conversion and analysis
[2147] Input: Prepare the collected audio data for conversion to text.
[2148] How it works: The device's voice recognition engine converts the voice data into text, and the AI module analyzes the text to generate an appropriate response.
[2149] Output: Parsed text data and response data are generated.
[2150] Step 3: Response generation and speech synthesis
[2151] Input: Converts response data into audio and prepares it for playback.
[2152] Specific operation: The device's speech synthesis engine outputs the generated response as voice and plays it through the speaker.
[2153] Output: An audio response is played to the child.
[2154] Sensor function steps
[2155] Step 1: Sensor installation and data collection
[2156] Input: The user installs special sensors in the crib or other dangerous locations.
[2157] Specific operation: The user places a special sensor in a crib or dangerous location, and the sensor begins to function.
[2158] Output: The sensor is deployed and ready to collect data.
[2159] Input: Prepare to collect data from sensors.
[2160] Specific operation: The device receives data sent from the sensor and detects abnormal movement or falls.
[2161] Output: Data is generated indicating the abnormal behavior detected.
[2162] Step 2: Alerting and Notification
[2163] Input: Prepare to treat anomalous behavior as an alert.
[2164] Specific operation: When the device detects danger, it generates alert information and sends it to the server.
[2165] Output: Alert information is generated and sent to the server.
[2166] Input: The server prepares to generate a push notification based on the alert information.
[2167] What it does: The server sends a push notification to the user, and the device plays a series of warning sounds to protect the baby.
[2168] Output: A push notification is generated and sent to the user. An alert sound is played on the device.
[2169] Market expansion steps
[2170] Step 1: Offer your app and define your market
[2171] Input: The server prepares the app to serve to a new user.
[2172] What happens: The server publishes the app's download link on a website or store, making it free for new users.
[2173] Output: The app download link is published and new users can download the app.
[2174] Step 2: Offering premium services
[2175] Input: The server defines the content of the paid premium service and prepares to provide it to the user.
[2176] Specific operation: Allow users to purchase premium services within the app, and after purchase, the server provides the corresponding services.
[2177] Output: The user purchases a premium service, unlocking additional features.
[2178] (Application example 1)
[2179] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2180] To improve vehicle safety, including that of autonomous vehicles and their drivers, it is necessary to constantly monitor the driver's state and detect abnormalities or dangers early. However, conventional systems lack the ability to detect driver drowsiness or inattention in real time and send warnings. There is also a need for an efficient method of reusing unused mobile communication devices. A system that solves these issues and ensures safety while being economical and environmentally friendly is needed.
[2181] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2182] In this invention, the server is a monitoring system that allows parents to remotely monitor their children by using an unused mobile communication device as a camera, and includes: a camera means for acquiring video; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video; an analysis means for analyzing the video from the mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the detected information; a music playback means for receiving remote operation instructions from the user terminal and playing music; an automatic call means for issuing an automatic call when dangerous movement is detected; a listening means for detecting a child's speech and responding gently; a driver monitoring means for monitoring the driver's state using a camera installed in the vehicle and detecting drowsiness or inattention; and a warning means for sending a warning to the driver and a manager when drowsiness or inattention is detected. This makes it possible to monitor the driver's state in real time and respond immediately if an abnormality is detected.
[2183] A "mobile communications device" is a device used for mobile communications, including smartphones and tablets.
[2184] "Camera means" refers to the camera equipment or functionality used to capture the image.
[2185] "Communication means" refers to communication technology or devices for sending and receiving data or information.
[2186] "User terminal means" refers to a terminal device used by a user, and is a device used for displaying images and remote control.
[2187] "Analysis means" refers to a method or device for analyzing video and audio data and detecting specific movements or situations.
[2188] A "notification means" is a method or device for sending a notification to a user or administrator based on sensed information.
[2189] "Music playback means" refers to a method or device for playing music, including speakers and music playback applications.
[2190] "Automatic calling means" refers to a method or device that automatically generates and plays voice messages under certain circumstances.
[2191] "Listening means" refers to a method or device for detecting and responding to the voice of a subject.
[2192] "Driver monitoring means" refers to a method or device that includes cameras or sensors for monitoring the situation inside the vehicle and, in particular, for checking the driver's condition.
[2193] "Warning means" refers to a method or device for issuing an immediate warning when an abnormality is detected.
[2194] This invention relates to a "monitoring and listening app" that allows parents to remotely monitor their babies and children using unused mobile communication devices. Furthermore, by applying this technology to autonomous vehicles, a system can be realized that monitors the driver and immediately issues a warning in the event of an abnormal or dangerous situation.
[2195] System Configuration
[2196] The server works in conjunction with mobile communication devices, user terminals, sensors, etc. These devices exchange data and analyze it to achieve the following functions:
[2197] Driver status monitoring
[2198] The device is equipped with a camera that captures images of the interior of the vehicle in real time and sends them to a server. This video data is analyzed by the server to detect driver drowsiness and inattention. This detection uses an AI algorithm, specifically TensorFlow, for face detection and movement analysis.
[2199] Warning function
[2200] If an abnormality is detected in the driver, the server will send a warning message to the administrator through the Twilio API, and will use the Google Text-to-Speech (gTTS) library to generate a voice warning and play the warning message aloud, using the pygame library for speech synthesis in this process.
[2201] Music playback function
[2202] The server receives instructions from the user's device and sends a music playback instruction to the device, which then plays the specified music file through its built-in speaker.
[2203] Data analysis
[2204] The video data sent from the device is analyzed by the server. The analysis method is to detect and analyze the driver's face and movements using OpenCV and TensorFlow. If drowsiness or inattention is detected, it is immediately treated as an abnormality and a warning is issued.
[2205] Specific examples
[2206] For example, if a driver falls asleep while driving on a highway, the device's camera will detect this in real time. If the server detects the driver's drowsiness through AI analysis, the system will send an SMS notification to the administrator via the server and simultaneously play an audio warning on the device.
[2207] Prompt Sentence Examples
[2208] "Develop an app that uses mobile communication devices to monitor the driver's status in real time. If the driver becomes inattentive or falls asleep, it should issue a voice warning and send an SMS notification to the manager."
[2209] In this way, the present invention provides a system that can seamlessly monitor and warn drivers, and by reusing unused mobile communication devices, it builds an efficient system that is environmentally friendly.
[2210] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2211] Step 1:
[2212] The device starts the camera and captures the video inside the car. In this step, the camera acquires video data in real time. The input is the video inside the car, and the output is the captured video data. Specifically, the device starts the camera using cv2.VideoCapture(0) and acquires frames using the read() method.
[2213] Step 2:
[2214] The video data captured by the terminal is sent to the server in real time. In this step, data is sent using a communication means. The input is the captured video data, and the output is the video data sent to the server. In concrete terms, the terminal sends the video data to the server using a specified communication protocol (e.g., HTTP POST request).
[2215] Step 3:
[2216] The server analyzes the received video data. In this step, the analysis means is used to determine the driver's state. The input is the transmitted video data, and the output is the analysis result. Specifically, the server processes the video data using OpenCV and TensorFlow to analyze the driver's face and movements.
[2217] Step 4:
[2218] If the server detects an anomaly, it generates a warning message. In this step, the warning means operates based on the analysis results. The input is the analysis result, and the output is the warning message. Specifically, when an anomaly is detected, the server generates a voice warning using the gTTS library and sends an SMS notification using the Twilio API.
[2219] Step 5:
[2220] The terminal plays music according to instructions from the server. In this step, a music playback means is used. The input is a remote control instruction from the user terminal, and the output is the music being played. In concrete terms, the terminal plays music on its built-in speaker based on the instruction received from the server.
[2221] Step 6:
[2222] The device detects what the child is saying and responds appropriately. In this step, the listening means is activated. The input is what the child is saying, and the output is a voice response. Specifically, the device detects the voice through the microphone and sends it to the server. The server converts it into text using a generative AI model and generates an appropriate response.
[2223] Step 7:
[2224] The server comprehensively manages all data and sends notifications to user devices in the event of an abnormality. This step is responsible for overall control. Input is data from various sensors and cameras, and output is notifications to administrators and push notifications to user devices. Specifically, the server receives and analyzes all data, and immediately notifies users if an abnormality is detected.
[2225] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2226] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely watch over babies and children by utilizing unused mobile communication devices (hereinafter referred to as "terminals"), with an emotion engine that recognizes the user's emotions.
[2227] Embodiment of camera linkage function
[2228] First, the user installs a dedicated app on the device and configures the camera. This activates the device's camera. The device captures real-time video through the camera and transmits the video data to the server. The server then streams the received data to the user's smartphone (hereinafter referred to as the "user device"). The user can then view the live video on their own device.
[2229] 1. Embodiment of push notification function
[2230] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device. The user can then respond promptly via the push notification.
[2231] Embodiment of music playback function
[2232] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server. The device then plays the specified music and plays it to the baby through the built-in speaker.
[2233] Embodiment of the hazard detection function
[2234] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, danger alert information is sent to the server, which then warns the user via push notification.
[2235] An embodiment of the listening function
[2236] The device detects what the child is saying through a microphone and converts the voice data into text. The converted text data is then analyzed by AI to generate an appropriate response. This response uses voice synthesis technology to gently reply to the child, allowing the user to engage with their child even when they are busy.
[2237] Sensor Function Embodiment
[2238] The user installs dedicated sensors in the crib or other dangerous locations and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The detected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[2239] Embodiment of Emotion Engine
[2240] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The response is converted from text to speech and played through the device's speaker.
[2241] This allows the emotion engine to reduce the user's mental burden and more effectively support the care of babies and children. The combination of the emotion engine improves the functionality and flexibility of the entire system, making it an important tool for supporting childcare.
[2242] The processing flow will be explained below.
[2243] Camera linkage function processing steps
[2244] Step 1:
[2245] Users install a dedicated app on their unused devices and configure the camera settings, which activates the device's camera.
[2246] Step 2:
[2247] The device captures real-time video through the camera, and the captured video is encoded into an appropriate format.
[2248] Step 3:
[2249] The terminal transmits the captured video data to the server in real time.
[2250] Step 4:
[2251] The server receives the video data and starts streaming it to the user terminal.
[2252] Step 5:
[2253] The user launches the app on the user device and checks the live video being streamed.
[2254] Push notification processing steps
[2255] Step 1:
[2256] The device's built-in microphone and camera are used to detect the baby's movements and cries.
[2257] Step 2:
[2258] The terminal analyzes the sensed data in real time and generates alert information when a certain threshold is exceeded.
[2259] Step 3:
[2260] The terminal transmits the generated alert information to the server.
[2261] Step 4:
[2262] The server receives the alert information and sends a push notification to the user terminal.
[2263] Step 5:
[2264] The user checks the push notification on the user device and takes appropriate action depending on the situation.
[2265] Processing steps for music playback function
[2266] Step 1:
[2267] The user sends a music playback instruction to the server from the user terminal via the app.
[2268] Step 2:
[2269] The server receives the user's instruction to play music and transmits it to the terminal.
[2270] Step 3:
[2271] The terminal plays the specified music file based on the instruction received from the server.
[2272] Step 4:
[2273] The device uses a built-in speaker to play music for your baby.
[2274] Hazard detection function processing steps
[2275] Step 1:
[2276] The device collects data in real time through a camera and microphone.
[2277] Step 2:
[2278] The device analyzes the collected data using AI algorithms to detect dangerous movements.
[2279] Step 3:
[2280] If danger is detected, the device will automatically play a voice message saying, "Danger! Stop!"
[2281] Step 4:
[2282] The terminal transmits danger alert information to the server.
[2283] Step 5:
[2284] The server receives the danger alert and sends a push notification to the user terminal.
[2285] Listening function processing steps
[2286] Step 1:
[2287] It detects what the child is saying through the device's microphone.
[2288] Step 2:
[2289] The device converts the detected voice data into text and sends it to the AI.
[2290] Step 3:
[2291] The device generates an appropriate response based on the results analyzed by the AI.
[2292] Step 4:
[2293] The device uses voice synthesis to speak soothing words to the child.
[2294] Sensor function processing steps
[2295] Step 1:
[2296] The user installs sensors in the crib or other dangerous locations and connects them to a terminal.
[2297] Step 2:
[2298] The device continuously collects data from connected sensors.
[2299] Step 3:
[2300] The device analyzes the collected data and detects abnormal behavior.
[2301] Step 4:
[2302] If an abnormality is detected, the terminal generates an alert and sends it to the server.
[2303] Step 5:
[2304] The server receives the anomaly alert and sends a push notification to the user device, which simultaneously plays an appropriate warning sound.
[2305] Emotion Engine Processing Steps
[2306] Step 1:
[2307] The user's voice data is transmitted from the user terminal to the server.
[2308] Step 2:
[2309] The server sends the received voice data to the emotion engine on the device.
[2310] Step 3:
[2311] The device's emotion engine analyzes the voice data and determines the user's emotional state.
[2312] Step 4:
[2313] The terminal's emotion engine generates an appropriate response based on the determined emotional state.
[2314] Step 5:
[2315] The terminal converts the generated response from text to speech and provides feedback to the user.
[2316] Example 2
[2317] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2318] In modern child-rearing, it is difficult for parents to constantly keep an eye on their babies and children, and the burden is particularly heavy when both parents are working or raising children alone. Conventional monitoring systems have limitations in real-time monitoring and voice communication, and lack emotion recognition functionality, making it difficult to reduce the mental burden on parents. Furthermore, they lack the ability to instantly detect dangerous situations and respond appropriately. The purpose of this invention is to solve these problems, more effectively ensure the safety of babies and children, and reduce the burden on parents.
[2319] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a camera means for capturing video from an unused mobile communication device; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video via the communication means; an analysis means for analyzing the video from the unused mobile communication device and detecting baby movements and cries; a notification means for sending a push notification based on the information detected by the analysis means; a music playback means for receiving remote operation instructions from the user terminal and playing music; an emotion analysis means for analyzing audio data transmitted from the unused mobile communication device and determining the user's emotional state; an automatic call means for issuing an automatic call when dangerous movements are detected; and a listening means for detecting and gently responding to the child's voice. This allows for a multifunctional and flexible monitoring system to be constructed, enabling parents to monitor and respond to the condition of babies and children even remotely. Furthermore, emotion analysis reduces the user's mental burden and enables quick and appropriate responses to dangerous situations.
[2320] "Disused mobile communication devices" are mobile devices such as mobile phones and tablets that have been retired but are now being reused.
[2321] "Camera means" refers to a camera device installed to capture video.
[2322] "Communication means" refers to a means having an internet connection function for sending and receiving images and data.
[2323] "User terminal means" refers to a display device used by a user, such as a smartphone, tablet, or computer.
[2324] The "analysis means" is a means having the function of analyzing the collected data and detecting specific information such as the baby's movements and cries.
[2325] The "notification means" is a means having a function for generating a push notification based on sensed information and transmitting it to the user terminal.
[2326] The "music playback means" is a means having a function of playing music in response to an instruction from a user terminal.
[2327] The "emotion analysis means" is a means having a function of analyzing voice data and determining the emotional state of the user.
[2328] The "automatic calling means" is a means having a function of automatically playing a warning message or the like when a dangerous movement is detected.
[2329] "Listening tools" are tools that have the ability to sense what a child is saying, generate an appropriate response, and respond in a gentle manner.
[2330] This invention relates to a system that combines a "monitoring and listening app" that allows parents to remotely monitor babies and children by reusing unused mobile communication devices with an emotion engine that recognizes the user's emotions. This system provides a multifunctional and flexible monitoring environment.
[2331] Camera linkage function
[2332] First, the user installs a dedicated app on their device and configures the camera. Once the app is installed and configured, the device's camera is activated. The device captures real-time video through the camera and sends the data to the server. The server receives the video data and streams it to the user's device, allowing the user to view live video on their device.
[2333] Specific examples
[2334] Device used: Old smartphone
[2335] Server used: AWS EC2 instance
[2336] Video streaming software: FFmpeg
[2337] Push notification function
[2338] Data collected from the device's built-in microphone and camera is analyzed in real time to detect baby movements and cries. If a threshold is exceeded, the device generates an alert and sends it to the server. The server then generates a push notification based on the alert information and sends it to the user's device, allowing the user to respond quickly.
[2339] Specific examples
[2340] Hardware used: Android or iOS device
[2341] Server: Google Firebase Cloud Messaging
[2342] App: Real-time analytics with custom mobile app
[2343] Music playback function
[2344] The user operates the app on their smartphone and sends instructions to play music. The instructions are then transferred to the device via the server, and the device plays the specified music on its built-in speaker.
[2345] Specific examples
[2346] Device used: Old tablet
[2347] Server: AWS Lambda
[2348] Music playback software: VLC
[2349] Hazard detection function
[2350] The device uses AI algorithms to analyze data from the camera and microphone to detect when the baby is behaving dangerously. When danger is detected, the device plays an automated voice message such as "Danger! Stop!". At the same time, the device sends danger alert information to the server, which then sends a push notification to the user's device.
[2351] Specific examples
[2352] Device used: Raspberry Pi 4 with camera module
[2353] AI algorithm: TensorFlow
[2354] Server: Google Cloud Functions
[2355] Active listening function
[2356] The device detects what the child is saying through a microphone and converts the voice data into text. This text data is analyzed using an AI model (e.g., OpenAI GPT-3) to generate an appropriate response. The generated response is then converted into audio data using speech synthesis technology, which is then played back by the device.
[2357] Specific examples
[2358] Device used: Amazon Echo device
[2359] Speech Recognition: Google Speech-to-Text API
[2360] Response generation: OpenAI GPT-3
[2361] Speech synthesis: Amazon Polly
[2362] Sensor function
[2363] The user installs dedicated sensors in the crib or other dangerous areas and connects them to the device. The sensors continuously collect environmental data and detect abnormal behavior. The collected data is sent to the device, where it is analyzed internally. If an abnormality is detected, the device generates an alert and sends it to the server. The server then sends the alert information to the user's device as a push notification, and the device simultaneously plays an appropriate warning sound.
[2364] Specific examples
[2365] Sensor used: Xiaomi Smart Home Sensor
[2366] Device used: Old smartphone
[2367] Server: AWS IoT Core
[2368] Emotion Engine
[2369] The device is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the voice data sent from the user's device and determines the user's emotional state. For example, if the user is feeling stressed, the emotion engine will recognize this and use AI to generate an appropriate response based on the situation. The generated response is converted into voice data and played through the device's speaker.
[2370] Specific examples
[2371] Device used: Old smartphone
[2372] Emotion engine: IBM Watson Tone Analyzer
[2373] Speech synthesis: Google Text-to-Speech
[2374] This makes the entire system a versatile and flexible monitoring environment, making it a very useful tool for supporting childcare. Users can monitor the condition of their baby or child in real time, even remotely, and respond quickly if necessary. The emotion engine also reduces mental strain.
[2375] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2376] Camera linkage function processing steps
[2377] Step 1:
[2378] Users install a dedicated app on their device, open the app and configure the camera settings.
[2379] Input: App installation, setting information
[2380] Output: Setup complete, camera ready to start
[2381] Step 2:
[2382] Once the setup is complete, the device will launch the camera.
[2383] Input: Setting completion notification
[2384] Output: Camera starts, video capture begins
[2385] Step 3:
[2386] The device captures real-time video through a camera and transmits the data to a server.
[2387] Input: Real-time video data
[2388] Output: Ready to send, sent to server
[2389] Step 4:
[2390] The server processes the received video data and distributes it as a stream to the user terminal.
[2391] Input: Video data
[2392] Output: Video streaming ready, sent to user device
[2393] Step 5:
[2394] Users can view live footage via the app on their devices.
[2395] Input: Video streaming data
[2396] Output: Video display, real-time monitoring
[2397] Push notification processing steps
[2398] Step 1:
[2399] It analyzes data collected from the device's built-in microphone and camera in real time.
[2400] Input: Audio data, video data
[2401] Output: Analysis results, threshold judgment
[2402] Step 2:
[2403] If the baby's movements or cries exceed the threshold, the device will generate an alert.
[2404] Input: Analysis results
[2405] Output: Alert information
[2406] Step 3:
[2407] The terminal transmits the generated alert information to the server.
[2408] Input: Alert information
[2409] Output: Sent to server
[2410] Step 4:
[2411] The server generates a push notification based on the alert information and sends it to the user's device.
[2412] Input: Alert information
[2413] Output: Push notification
[2414] Step 5:
[2415] Users receive push notifications and can respond quickly.
[2416] Input: Push notification
[2417] Output: Notification confirmation, required action
[2418] Processing steps for music playback function
[2419] Step 1:
[2420] The user operates the app on their smartphone and sends instructions to play music.
[2421] Input: Music playback instructions
[2422] Output: Send instructions
[2423] Step 2:
[2424] The server transfers the playback instruction received from the user to the terminal.
[2425] Input: Playback instructions
[2426] Output: Directed transfer
[2427] Step 3:
[2428] The device will play music through its built-in speaker based on the transferred instructions.
[2429] Input: Playback instructions
[2430] Output: Music playback
[2431] Hazard detection function processing steps
[2432] Step 1:
[2433] The device collects the baby's movements and voice data using a camera and microphone.
[2434] Input: Video data, audio data
[2435] Output: Recorded data, audio data
[2436] Step 2:
[2437] The device analyzes the collected data using AI algorithms to detect dangerous behavior.
[2438] Input: Recorded data, audio data
[2439] Output: Risk assessment result
[2440] Step 3:
[2441] When danger is detected, the device will play an automated voice message such as "Danger! Stop!"
[2442] Input: Risk assessment result
[2443] Output: Automated voice message
[2444] Step 4:
[2445] The terminal transmits danger alert information to the server.
[2446] Input: Risk assessment result
[2447] Output: Alert information
[2448] Step 5:
[2449] The server generates a push notification and sends it to the user device.
[2450] Input: Alert information
[2451] Output: Push notification
[2452] Listening function processing steps
[2453] Step 1:
[2454] The device detects what the child is saying through a microphone.
[2455] Input: Audio data
[2456] Output: Sensor information
[2457] Step 2:
[2458] The device converts the detected voice data into text.
[2459] Input: Audio data
[2460] Output: Text data
[2461] Step 3:
[2462] The converted text data is analyzed by an AI model to generate an appropriate response.
[2463] Input: Text data
[2464] Output: Response data
[2465] Step 4:
[2466] The terminal converts the generated response into voice data using voice synthesis technology.
[2467] Input: Response data
[2468] Output: Audio data
[2469] Step 5:
[2470] The device plays back the audio data and provides an appropriate response to the child.
[2471] Input: Audio data
[2472] Output: Response voice
[2473] Sensor function processing steps
[2474] Step 1:
[2475] Users install special sensors in cribs and dangerous areas and connect them to a terminal.
[2476] Input: Sensor installation
[2477] Output: Connection complete
[2478] Step 2:
[2479] Sensors continuously collect environmental data and detect abnormal behavior.
[2480] Input: Environment data
[2481] Output: Data collection, anomaly detection information
[2482] Step 3:
[2483] The sensor transmits the collected data to the terminal.
[2484] Input: Data collection, anomaly detection information
[2485] Output: Send to terminal
[2486] Step 4:
[2487] The terminal analyzes the received sensor data internally and detects any abnormalities.
[2488] Input: Sensor data
[2489] Output: Abnormality detection result
[2490] Step 5:
[2491] The device generates an alert if an anomaly is detected.
[2492] Input: Abnormality detection result
[2493] Output: Alert information
[2494] Step 6:
[2495] The terminal transmits the alert information to the server.
[2496] Input: Alert information
[2497] Output: Sent to server
[2498] Step 7:
[2499] The server generates a push notification and sends it to the user device.
[2500] Input: Alert information
[2501] Output: Push notification
[2502] Emotion Engine Processing Steps
[2503] Step 1:
[2504] The user transmits voice data to the terminal.
[2505] Input: Audio data
[2506] Output: Send to terminal
[2507] Step 2:
[2508] The device uses an emotion engine to analyze the voice data and determine the user's emotional state.
[2509] Input: Audio data
[2510] Output: Emotion determination result
[2511] Step 3:
[2512] The emotion engine generates an appropriate response based on the judgment result.
[2513] Input: Emotion judgment result
[2514] Output: Response data
[2515] Step 4:
[2516] The terminal converts the generated response into voice data using voice synthesis technology.
[2517] Input: Response data
[2518] Output: Audio data
[2519] Step 5:
[2520] The terminal plays back the audio data and responds appropriately to reduce the user's mental burden.
[2521] Input: Audio data
[2522] Output: Response voice
[2523] (Application example 2)
[2524] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2525] The problem that the present invention aims to solve is to provide a system that effectively utilizes unused mobile communication devices, allowing parents or guardians to remotely monitor babies and children and increase their safety. Furthermore, because conventional monitoring systems do not take emotion recognition into consideration, a means for reducing the mental burden on users was also necessary. Therefore, the present invention aims to compensate for the shortcomings of conventional monitoring systems and improve safety and usability.
[2526] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a communication means for transmitting video in real time, an emotion recognition means for analyzing audio data and recognizing emotions, and a response generation means for generating an appropriate response based on the recognized emotion and outputting it as audio. This makes it possible not only to perform remote real-time monitoring, but also to recognize the emotional state of the user and generate an appropriate response. 【2527...
Claims
1. A monitoring system that allows parents to monitor their children remotely by using unused mobile communication devices as cameras. The unused mobile communication device has a camera means for acquiring an image; a communication means for transmitting the video in real time; a user terminal means for receiving and displaying the video via the communication means; an analysis means for analyzing the video from the unused mobile communication device and detecting the baby's movements and cries; a notification means for transmitting a push notification based on the information sensed by the analysis means; a music playback means for receiving a remote operation instruction from the user terminal and playing music; an automatic call means for making an automatic call when a dangerous movement is detected; A listening method that detects what the child is saying and responds gently, A system including:
2. 2. The system according to claim 1, wherein the analyzing means detects abnormal behavior based on data from a sensor installed in the baby bed.
3. The system according to claim 1 , wherein the notification means sends a push notification to the user terminal based on the sensed information, and takes a prompt action based on the information displayed on the user terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A