system
The system addresses the challenge of real-time danger detection for children by using AI to analyze surroundings, generate voice messages, and notify guardians, ensuring effective safety support during commutes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing systems fail to detect dangers around children in real-time and provide direct support.
A system comprising an analysis unit, generation unit, and notification unit that utilizes multimodal generative AI to analyze the surroundings of a child, generate voice messages, and deliver them to the child while notifying guardians, ensuring real-time safety support.
The system effectively detects and responds to potential dangers around children, providing real-time safety support by generating and delivering voice messages and notifying guardians, enhancing child safety during commutes.
Smart Images

Figure 2026073616000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that it is difficult to detect the danger around a child in real time and provide direct support.
[0005] The system according to the embodiment aims to detect the danger around a child in real time and provide direct support.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an analysis unit, a generation unit, a transmission unit, and a notification unit. The analysis unit analyzes the surrounding environment of the child. The generation unit detects danger based on the data analyzed by the analysis unit and generates an audio message. The transmission unit conveys the audio message generated by the generation unit to the child. The notification unit notifies the guardian of the danger detected by the analysis unit. [Effects of the Invention]
[0007] The system according to this embodiment can detect dangers around a child in real time and provide direct support. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server. [[ID=????]]
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] It seems there is a missing ID in the original text for the line "
[0018] ". I've left it as "????" in the translation for now. If you can provide the correct ID for that line, I can update the translation.The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The monitoring system according to an embodiment of the present invention is a system that provides real-time support for the safety of children when they are going to and from school. This monitoring system utilizes multimodal generative AI to analyze the situation around the child, detect danger, generate voice messages, and deliver them to the child and notify the guardian. For example, if the monitoring system sees a child about to cross a red light, it analyzes the situation through a camera and microphone and generates a gentle voice message in the voice of a parent or teacher. For example, it might generate a voice message such as, "○○-chan, it's a red light. Please stop," and deliver it to the child. Next, if the monitoring system sees a suspicious person following the child, it detects the situation and contacts the guardian in real time. Guardians can check the child's current location and situation through a dedicated app and, if necessary, report it to the police. Furthermore, if an unknown adult speaks to the child, the monitoring system analyzes the conversation and responds on behalf of the child. For example, it might generate a voice message such as, "I'm sorry, I can't talk right now," and respond on behalf of the child. In this way, the monitoring system provides advanced monitoring functions as if an adult were accompanying the child, ensuring the child's safety. Furthermore, by running the generation AI on edge devices, the monitoring system completes processing locally without sending data to the cloud, achieving both privacy protection and real-time capabilities. As a result, the monitoring system can support the safety of children during their commutes to and from school in real time, ensuring their safety.
[0029] The monitoring system according to this embodiment comprises an analysis unit, a generation unit, a transmission unit, and a notification unit. The analysis unit analyzes the situation around the child. The analysis unit collects data, for example, through a camera and a microphone, and analyzes the situation in detail. For example, the analysis unit uses a camera to acquire and analyze video of the child's surroundings in real time. The analysis unit can also collect and analyze audio data from the surroundings using a microphone. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to grasp the situation in more detail. The generation unit detects danger based on the data analyzed by the analysis unit and generates an audio message. For example, the generation unit uses a generation AI to generate an appropriate audio message based on the analysis results. For example, if the child is about to cross against a red light, the generation unit generates an audio message such as, "○○-chan, it's a red light. Please stop." The generation unit can also generate an audio message such as, "A suspicious person is following you. Please be careful." if a suspicious person is following the child. Furthermore, the generation unit can also generate an audio message such as, "I'm sorry, I can't talk to you right now." if an unknown adult approaches the child. The transmission unit transmits the voice message generated by the generation unit to the child. The transmission unit transmits the voice message to the child using, for example, a speaker. The transmission unit can also transmit the voice message to the child using earphones. Furthermore, the transmission unit can also transmit the voice message to the child using a vibration function. The notification unit notifies the guardian of the danger detected by the analysis unit. The notification unit notifies the guardian, for example, through a dedicated app. The notification unit sends a notification to the guardian's smartphone, for example, displaying the child's current location and situation in real time. The notification unit can also notify the guardian using email. Furthermore, the notification unit can also notify the guardian using SMS. Thus, the monitoring system according to the embodiment can ensure the child's safety by analyzing the situation around the child, detecting danger, generating a voice message, transmitting it to the child, and notifying the guardian.
[0030] The analysis unit analyzes the child's surroundings. For example, the analysis unit collects data through cameras and microphones and analyzes the surroundings in detail. Specifically, the camera acquires high-resolution video in real time, capturing the child's movements and surrounding environment in detail. The camera is equipped with facial recognition and object detection functions, allowing it to identify the child's facial expressions and movements, as well as surrounding objects and people. For example, the camera captures the situation in a park where a child is playing, and the analysis unit analyzes the video in real time. The microphone collects ambient audio data, and the analysis unit analyzes this audio data to identify the child's voice and surrounding sounds. For example, if a child cries out for help or if suspicious sounds are heard in the surroundings, the analysis unit can analyze the audio data and detect danger. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to understand the situation in more detail. For example, if the camera footage detects that a child is approaching a road, and the microphone simultaneously detects the sound of a car engine, the analysis unit can understand that the child is attempting to cross the road, creating a dangerous situation. In this way, the analysis unit can analyze the child's surroundings from multiple angles and detect danger quickly and accurately.
[0031] The generation unit detects danger based on data analyzed by the analysis unit and generates a voice message. For example, the generation unit uses a generation AI to generate an appropriate voice message based on the analysis results. The generation AI uses natural language processing technology to generate specific messages corresponding to the analysis results. For example, if a child is about to cross against a red light, the generation AI will generate a voice message such as, "○○-chan, it's a red light. Please stop." The generation AI can select and generate an appropriate message based on pre-set scenarios. Furthermore, if a suspicious person is following a child, the generation unit can generate a voice message such as, "A suspicious person is following you. Please be careful." Because the generation AI generates the optimal message for the situation based on the data provided by the analysis unit, it can respond quickly and accurately. Additionally, if an unknown adult approaches a child, the generation unit can generate a voice message such as, "I'm sorry, I can't talk right now." The generation AI can generate messages for various scenarios to ensure the child's safety. This allows the generation unit to quickly generate an appropriate voice message based on the data provided by the analysis unit and deliver it to the child.
[0032] The transmission unit delivers the voice message generated by the generation unit to the child. The transmission unit can deliver the voice message to the child using, for example, a speaker. The speaker can play the message at a volume and sound quality that is easy for the child to hear. For example, it can reliably deliver the voice message in various environments, such as a park where the child is playing or on their way to school. The transmission unit can also deliver the voice message to the child using earphones. The earphones are designed to be easy for children to wear, block out external noise, and deliver the message clearly. Furthermore, the transmission unit can deliver the voice message to the child using a vibration function. The vibration function works in conjunction with the voice message to attract the child's attention. For example, even if the child is not wearing earphones or the surroundings are noisy, the message can be reliably delivered by vibration. This allows the transmission unit to quickly and reliably deliver the voice message generated by the generation unit to the child using various means. Furthermore, the transmission unit can monitor the child's reaction and, if necessary, play the message or give additional instructions. This allows the transmission unit to deliver voice messages flexibly and effectively to ensure the child's safety.
[0033] The notification unit alerts parents to dangers detected by the analysis unit. For example, the notification unit can notify parents through a dedicated app. This app is installed on the parent's smartphone and can display the child's current location and situation in real time. For instance, if a child is in a dangerous situation, the app immediately sends a notification to the parent, issuing a warning. The notification unit can also notify parents via email. Emails can include detailed descriptions, images, and audio data, helping parents accurately understand the situation. Furthermore, the notification unit can notify parents via SMS. SMS allows for quick information transmission via short messages, making it suitable for immediate responses in emergencies. This enables the notification unit to quickly and reliably notify parents of dangers detected by the analysis unit using various means. Additionally, the notification unit can receive feedback from parents and adjust the system's operation accordingly. For example, if a parent rushes to their child's location, the notification unit transmits this information to the analysis and generation units, allowing for appropriate action. This enables the notification unit to communicate information flexibly and effectively in cooperation with parents to ensure the child's safety.
[0034] The analysis unit can collect data through cameras and microphones. For example, the analysis unit can use a camera to acquire and analyze video footage of the child's surroundings in real time. The analysis unit can also use a microphone to collect and analyze audio data from the surroundings. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to gain a more detailed understanding of the situation. This allows for a detailed analysis of the surrounding situation by collecting data through cameras and microphones. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform the data analysis.
[0035] The generation unit can generate voice messages based on the analysis results. For example, the generation unit uses a generation AI to generate appropriate voice messages based on the analysis results. For example, if a child is about to cross the street against a red light, the generation unit can generate a voice message such as, "○○-chan, the light is red. Please stop." The generation unit can also generate a voice message such as, "A suspicious person is following you. Please be careful," if a suspicious person is following a child. Furthermore, if an unknown adult approaches a child, the generation unit can generate a voice message such as, "I'm sorry, I can't talk to you right now." In this way, by generating voice messages based on the analysis results, appropriate messages can be conveyed to children. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input the data analyzed by the analysis unit into the generation AI and have the generation AI perform the generation of voice messages.
[0036] The transmission unit can deliver the generated voice message to the child. The transmission unit can deliver the voice message to the child using, for example, a speaker. It can also deliver the voice message to the child using earphones. Furthermore, the transmission unit can deliver the voice message to the child using a vibration function. This allows the child to receive appropriate attention by receiving the generated voice message. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the voice message generated by the generation unit into a generation AI and have the generation AI perform the delivery of the voice message.
[0037] The notification unit can notify parents in real time. For example, the notification unit can notify parents through a dedicated app. For example, the notification unit can send a notification to the parent's smartphone, displaying the child's current location and situation in real time. The notification unit can also notify parents via email. Furthermore, the notification unit can notify parents via SMS. This allows for a quick response by notifying parents in real time. Some or all of the above-described processes in the notification unit may be performed using AI or not. For example, the notification unit can input the danger detected by the analysis unit into a generation AI and have the generation AI generate the notification content.
[0038] The analysis unit can optimize its analysis algorithm by referring to past data when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations at a specific location based on past data and adjust the analysis algorithm accordingly. It can also predict dangerous situations during a specific time period by referring to past data and optimize the analysis algorithm accordingly. Furthermore, the analysis unit can adjust the analysis algorithm based on past data to make it easier to detect specific behavioral patterns. This improves the accuracy of the analysis by optimizing the analysis algorithm by referring to past data. Some or all of the above processes in the analysis unit may be performed using AI or not. For example, the analysis unit can input past data into a generating AI and have the generating AI perform the optimization of the analysis algorithm.
[0039] The analysis unit can analyze camera and microphone data in real time and detect abnormal behavioral patterns. For example, the analysis unit can analyze camera footage in real time to detect when a child is being chased by a suspicious person. The analysis unit can also analyze microphone audio data in real time to detect when a child is calling for help. Furthermore, the analysis unit can combine and analyze camera and microphone data to detect when a child is in a dangerous situation. This allows for the rapid detection of abnormal behavioral patterns by analyzing camera and microphone data in real time. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform the detection of abnormal behavioral patterns.
[0040] The analysis unit can improve the accuracy of its analysis by considering the child's location information when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations in specific locations based on the child's location information and improve the accuracy of its analysis. It can also predict dangerous situations during specific time periods by referring to the child's location information and improve the accuracy of its analysis. Furthermore, the analysis unit can improve the accuracy of its analysis by making it easier to detect specific behavioral patterns based on the child's location information. In this way, improving the accuracy of the analysis by considering the child's location information enables more accurate analysis. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input the child's location information into a generating AI and have the generating AI perform the analysis accuracy improvement.
[0041] The analysis unit can improve the accuracy of its analysis by filtering ambient noise when analyzing data from cameras and microphones. For example, the analysis unit can filter ambient noise to clearly detect children's voices or the voices of suspicious individuals. The analysis unit can also filter ambient noise to make it easier to analyze specific voice patterns. Furthermore, the analysis unit can filter ambient noise to make it easier to detect abnormal behavioral patterns. In this way, filtering ambient noise improves the accuracy of the analysis. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform ambient noise filtering.
[0042] The generation unit can optimize the content of the voice messages it generates based on past analysis results. For example, the generation unit can generate voice messages suitable for a specific situation based on past analysis results. It can also generate voice messages suitable for a specific time period by referring to past analysis results. Furthermore, it can generate voice messages suitable for a specific behavioral pattern based on past analysis results. In this way, by optimizing the content of voice messages based on past analysis results, more effective messages can be generated. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input past analysis results into the generation AI and have the generation AI perform the optimization of voice messages.
[0043] The voice generation unit can generate more natural-sounding speech by meticulously analyzing the characteristics of voices, such as those of parents or teachers, when imitating their voices. For example, the voice generation unit can meticulously analyze the tone and rhythm of parents' and teachers' voices to generate natural-sounding speech. Furthermore, the voice generation unit can generate individual voice messages based on the characteristics of parents' and teachers' voices. In addition, the voice generation unit can analyze the intonation of parents' and teachers' voices to generate even more natural-sounding speech. This allows for the generation of more natural-sounding speech by meticulously analyzing voice characteristics. Some or all of the above-described processes in the voice generation unit are performed using a generation AI. For example, the voice generation unit can input data of parents' and teachers' voices into the generation AI, allowing the AI to generate natural-sounding speech.
[0044] The generation unit can customize the content of the generated voice messages according to the child's age and gender. For example, the generation unit can generate voice messages using easy-to-understand language according to the child's age. It can also generate voice messages with appropriate tone and content according to the child's gender. Furthermore, the generation unit can generate individual voice messages based on the child's age and gender. This allows for more appropriate messages to be conveyed by customizing voice messages according to the child's age and gender. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input data on the child's age and gender into the generation AI and have the generation AI perform the voice message customization.
[0045] The voice generation unit can handle diverse scenarios by generating different voice variations when imitating the voices of parents and teachers. For example, the voice generation unit can generate different variations of parents' and teachers' voices and provide voice messages suitable for specific scenarios. The voice generation unit can also generate individual voice messages based on the variations of parents' and teachers' voices. Furthermore, the voice generation unit can analyze the different variations of parents' and teachers' voices and generate more natural-sounding voices. This allows it to handle diverse scenarios by generating different voice variations. Some or all of the above processing in the voice generation unit is performed using a generation AI. For example, the voice generation unit can input data of parents' and teachers' voices into the generation AI and have the generation AI perform the generation of different voice variations.
[0046] The transmission unit can automatically adjust the volume of voice messages by considering the ambient noise level. For example, if the surroundings are noisy, the transmission unit can automatically increase the volume to transmit the voice message. Conversely, if the surroundings are quiet, the transmission unit can automatically decrease the volume to transmit the voice message. Furthermore, the transmission unit can analyze the ambient noise level in real time and transmit the voice message at the optimal volume. This ensures that voice messages are properly conveyed by automatically adjusting the volume in consideration of the ambient noise level. Some or all of the above processing in the transmission unit may be performed using AI, or it may be performed without AI. For example, the transmission unit can input ambient noise level data into a generating AI and have the generating AI perform automatic volume adjustment.
[0047] The transmission unit can monitor the child's response in real time when transmitting a voice message and regenerate the message as needed. For example, if the child does not respond to the voice message, the transmission unit will regenerate and transmit the message again. The transmission unit can also modify the content and regenerate the message if the child gives an inappropriate response to the voice message. Furthermore, the transmission unit can monitor the child's response in real time and regenerate the message at the optimal timing. This allows for the regeneration of an appropriate message by monitoring the child's response in real time. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the child's response data into a generation AI and have the generation AI perform message regeneration.
[0048] The transmission unit can select the optimal transmission method when transmitting a voice message, taking into account the child's location information. For example, if the child is in a specific location, the transmission unit will transmit a voice message appropriate for that location. The transmission unit can also select the optimal transmission method based on the child's location information. Furthermore, if the child is moving, the transmission unit can update the location information in real time and select the optimal transmission method. This allows for more effective message delivery by selecting the optimal transmission method while considering the child's location information. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the child's location information into a generating AI and have the generating AI select the optimal transmission method.
[0049] The transmission unit can select the optimal transmission method when transmitting voice messages, taking into account the child's device information. For example, if the child is using a smartphone, the transmission unit will transmit a voice message optimized for the smartphone. It can also transmit a voice message optimized for a tablet if the child is using a tablet. Furthermore, if the child is using a smartwatch, the transmission unit can transmit a voice message optimized for the smartwatch. This allows for more effective message delivery by selecting the optimal transmission method based on the child's device information. Some or all of the above processing in the transmission unit may be performed using AI, or without AI. For example, the transmission unit can input the child's device information into a generating AI and have the generating AI select the optimal transmission method.
[0050] The notification unit can select the most suitable notification method by referring to past notification history when notifying parents. For example, the notification unit can select the notification method that parents are most likely to respond to based on past notification history. The notification unit can also select a notification method suitable for a specific time period by referring to past notification history. Furthermore, the notification unit can select a notification method that allows parents to respond most quickly based on past notification history. This allows parents to respond quickly by selecting the most suitable notification method by referring to past notification history. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input past notification history into a generating AI and have the generating AI select the most suitable notification method.
[0051] The notification unit can provide additional information to parents when notifying them, providing a detailed report of the child's current situation. For example, the notification unit may send a notification that includes the child's current location information. It may also send a notification that provides a detailed explanation of the child's surroundings. Furthermore, the notification unit may analyze the child's behavior patterns and send a notification that includes a detailed explanation of the situation. This allows parents to respond appropriately by providing a detailed report of the child's current situation. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit may input the child's current situation data into a generating AI and have the generating AI generate the additional information.
[0052] The notification unit can select the most appropriate notification method when notifying a parent, taking into account the parent's location. For example, if the parent is at home, the notification unit may prioritize phone notification. If the parent is at work, the notification unit may prioritize email notification. Furthermore, if the parent is on the move, the notification unit may prioritize SMS notification. This allows parents to respond quickly by selecting the most appropriate notification method based on their location. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input the parent's location information into a generating AI and have the generating AI select the most appropriate notification method.
[0053] The notification unit can provide a function that allows parents to monitor the situation in real time through a dedicated app when a notification is sent. For example, the notification unit can display the child's current location in real time through the dedicated app. The notification unit can also monitor the child's surroundings in real time through the dedicated app. Furthermore, the notification unit can analyze the child's behavior patterns in real time through the dedicated app and notify the parents. This allows parents to respond quickly by monitoring the situation in real time through the dedicated app. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input the dedicated app into a generating AI and have the generating AI perform the provision of the real-time monitoring function.
[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0055] The analysis unit can be equipped with the ability to learn and predict children's behavior patterns. For example, it can predict how a child will behave at a specific time based on past behavioral data. It can also learn children's behavior patterns and detect abnormal behavior. Furthermore, it can predict dangerous situations based on children's behavior patterns and issue warnings in advance. This allows for more effective child safety by learning and predicting children's behavior patterns.
[0056] The analysis unit can improve the accuracy of its analysis by taking weather information into account when analyzing the surrounding conditions. For example, the analysis unit adjusts the accuracy of its camera image analysis when visibility is poor in rainy weather. It can also adjust the accuracy of its audio data analysis when there is strong wind. Furthermore, the analysis unit can predict specific dangerous situations based on weather information and improve the accuracy of its analysis. By improving the accuracy of the analysis by taking weather information into account, more accurate analysis becomes possible.
[0057] The notification unit can analyze a child's behavior patterns and adjust the notification content based on predicted behavior when notifying parents. For example, if a child is often in a specific location at a specific time, the notification unit will send a notification containing information about that location. The notification unit can also send notifications containing warnings about specific behaviors based on the child's behavior patterns. Furthermore, the notification unit can analyze a child's behavior patterns and adjust the timing of notifications based on predicted behavior. In this way, by analyzing a child's behavior patterns and adjusting the notification content based on predicted behavior, the notification unit can provide parents with appropriate information.
[0058] The generation unit can customize the content of voice messages according to the child's learning progress. For example, if a child is at a specific learning stage, the generation unit will generate a voice message with content appropriate for that stage. The generation unit can also generate voice messages using easy-to-understand language based on the child's learning progress. Furthermore, the generation unit can generate individual voice messages according to the child's learning progress. This allows for the delivery of more appropriate messages by customizing voice messages according to the child's learning progress.
[0059] The analysis unit can improve the accuracy of its analysis by considering the child's location information when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations in specific locations based on the child's location information, thereby improving the accuracy of the analysis. Furthermore, the analysis unit can also predict dangerous situations during specific time periods by referring to the child's location information, thereby improving the accuracy of the analysis. In addition, the analysis unit can improve the accuracy of the analysis by making it easier to detect specific behavioral patterns based on the child's location information. By improving the accuracy of the analysis by considering the child's location information, a more accurate analysis becomes possible.
[0060] The following briefly describes the processing flow for example form 1.
[0061] Step 1: The analysis unit analyzes the child's surroundings. The analysis unit collects data through cameras and microphones and analyzes the surroundings in detail. For example, it uses cameras to acquire and analyze video footage of the child's surroundings in real time. It can also collect and analyze audio data from the surroundings using microphones. Furthermore, by combining and analyzing the data from cameras and microphones, a more detailed understanding of the situation can be obtained. Step 2: The generation unit detects danger based on the data analyzed by the analysis unit and generates a voice message. The generation unit uses a generation AI to generate an appropriate voice message based on the analysis results. For example, if a child is about to cross against a red light, it will generate a voice message such as, "○○-chan, it's a red light. Please stop." It can also generate a voice message such as, "A suspicious person is following you. Please be careful." Furthermore, if an unknown adult approaches a child, it can generate a voice message such as, "I'm sorry, I can't talk to you right now." Step 3: The transmission unit transmits the voice message generated by the generation unit to the child. The transmission unit transmits the voice message to the child using a speaker. It can also transmit the voice message to the child using earphones. Furthermore, it can transmit the voice message to the child using a vibration function. Step 4: The notification unit notifies parents of the dangers detected by the analysis unit. The notification unit notifies parents through a dedicated app. For example, it sends a notification to the parent's smartphone, displaying the child's current location and situation in real time. It can also notify parents via email or SMS.
[0062] (Example of form 2) The monitoring system according to an embodiment of the present invention is a system that provides real-time support for the safety of children when they are going to and from school. This monitoring system utilizes multimodal generative AI to analyze the situation around the child, detect danger, generate voice messages, and deliver them to the child and notify the guardian. For example, if the monitoring system sees a child about to cross a red light, it analyzes the situation through a camera and microphone and generates a gentle voice message in the voice of a parent or teacher. For example, it might generate a voice message such as, "○○-chan, it's a red light. Please stop," and deliver it to the child. Next, if the monitoring system sees a suspicious person following the child, it detects the situation and contacts the guardian in real time. Guardians can check the child's current location and situation through a dedicated app and, if necessary, report it to the police. Furthermore, if an unknown adult speaks to the child, the monitoring system analyzes the conversation and responds on behalf of the child. For example, it might generate a voice message such as, "I'm sorry, I can't talk right now," and respond on behalf of the child. In this way, the monitoring system provides advanced monitoring functions as if an adult were accompanying the child, ensuring the child's safety. Furthermore, by running the generation AI on edge devices, the monitoring system completes processing locally without sending data to the cloud, achieving both privacy protection and real-time capabilities. As a result, the monitoring system can support the safety of children during their commutes to and from school in real time, ensuring their safety.
[0063] The monitoring system according to this embodiment comprises an analysis unit, a generation unit, a transmission unit, and a notification unit. The analysis unit analyzes the situation around the child. The analysis unit collects data, for example, through a camera and a microphone, and analyzes the situation in detail. For example, the analysis unit uses a camera to acquire and analyze video of the child's surroundings in real time. The analysis unit can also collect and analyze audio data from the surroundings using a microphone. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to grasp the situation in more detail. The generation unit detects danger based on the data analyzed by the analysis unit and generates an audio message. For example, the generation unit uses a generation AI to generate an appropriate audio message based on the analysis results. For example, if the child is about to cross against a red light, the generation unit generates an audio message such as, "○○-chan, it's a red light. Please stop." The generation unit can also generate an audio message such as, "A suspicious person is following you. Please be careful." if a suspicious person is following the child. Furthermore, the generation unit can also generate an audio message such as, "I'm sorry, I can't talk to you right now." if an unknown adult approaches the child. The transmission unit transmits the voice message generated by the generation unit to the child. The transmission unit transmits the voice message to the child using, for example, a speaker. The transmission unit can also transmit the voice message to the child using earphones. Furthermore, the transmission unit can also transmit the voice message to the child using a vibration function. The notification unit notifies the guardian of the danger detected by the analysis unit. The notification unit notifies the guardian, for example, through a dedicated app. The notification unit sends a notification to the guardian's smartphone, for example, displaying the child's current location and situation in real time. The notification unit can also notify the guardian using email. Furthermore, the notification unit can also notify the guardian using SMS. Thus, the monitoring system according to the embodiment can ensure the child's safety by analyzing the situation around the child, detecting danger, generating a voice message, transmitting it to the child, and notifying the guardian.
[0064] The analysis unit analyzes the child's surroundings. For example, the analysis unit collects data through cameras and microphones and analyzes the surroundings in detail. Specifically, the camera acquires high-resolution video in real time, capturing the child's movements and surrounding environment in detail. The camera is equipped with facial recognition and object detection functions, allowing it to identify the child's facial expressions and movements, as well as surrounding objects and people. For example, the camera captures the situation in a park where a child is playing, and the analysis unit analyzes the video in real time. The microphone collects ambient audio data, and the analysis unit analyzes this audio data to identify the child's voice and surrounding sounds. For example, if a child cries out for help or if suspicious sounds are heard in the surroundings, the analysis unit can analyze the audio data and detect danger. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to understand the situation in more detail. For example, if the camera footage detects that a child is approaching a road, and the microphone simultaneously detects the sound of a car engine, the analysis unit can understand that the child is attempting to cross the road, creating a dangerous situation. In this way, the analysis unit can analyze the child's surroundings from multiple angles and detect danger quickly and accurately.
[0065] The generation unit detects danger based on data analyzed by the analysis unit and generates a voice message. For example, the generation unit uses a generation AI to generate an appropriate voice message based on the analysis results. The generation AI uses natural language processing technology to generate specific messages corresponding to the analysis results. For example, if a child is about to cross against a red light, the generation AI will generate a voice message such as, "○○-chan, it's a red light. Please stop." The generation AI can select and generate an appropriate message based on pre-set scenarios. Furthermore, if a suspicious person is following a child, the generation unit can generate a voice message such as, "A suspicious person is following you. Please be careful." Because the generation AI generates the optimal message for the situation based on the data provided by the analysis unit, it can respond quickly and accurately. Additionally, if an unknown adult approaches a child, the generation unit can generate a voice message such as, "I'm sorry, I can't talk right now." The generation AI can generate messages for various scenarios to ensure the child's safety. This allows the generation unit to quickly generate an appropriate voice message based on the data provided by the analysis unit and deliver it to the child.
[0066] The transmission unit delivers the voice message generated by the generation unit to the child. The transmission unit can deliver the voice message to the child using, for example, a speaker. The speaker can play the message at a volume and sound quality that is easy for the child to hear. For example, it can reliably deliver the voice message in various environments, such as a park where the child is playing or on their way to school. The transmission unit can also deliver the voice message to the child using earphones. The earphones are designed to be easy for children to wear, block out external noise, and deliver the message clearly. Furthermore, the transmission unit can deliver the voice message to the child using a vibration function. The vibration function works in conjunction with the voice message to attract the child's attention. For example, even if the child is not wearing earphones or the surroundings are noisy, the message can be reliably delivered by vibration. This allows the transmission unit to quickly and reliably deliver the voice message generated by the generation unit to the child using various means. Furthermore, the transmission unit can monitor the child's reaction and, if necessary, play the message or give additional instructions. This allows the transmission unit to deliver voice messages flexibly and effectively to ensure the child's safety.
[0067] The notification unit alerts parents to dangers detected by the analysis unit. For example, the notification unit can notify parents through a dedicated app. This app is installed on the parent's smartphone and can display the child's current location and situation in real time. For instance, if a child is in a dangerous situation, the app immediately sends a notification to the parent, issuing a warning. The notification unit can also notify parents via email. Emails can include detailed descriptions, images, and audio data, helping parents accurately understand the situation. Furthermore, the notification unit can notify parents via SMS. SMS allows for quick information transmission via short messages, making it suitable for immediate responses in emergencies. This enables the notification unit to quickly and reliably notify parents of dangers detected by the analysis unit using various means. Additionally, the notification unit can receive feedback from parents and adjust the system's operation accordingly. For example, if a parent rushes to their child's location, the notification unit transmits this information to the analysis and generation units, allowing for appropriate action. This enables the notification unit to communicate information flexibly and effectively in cooperation with parents to ensure the child's safety.
[0068] The analysis unit can collect data through cameras and microphones. For example, the analysis unit can use a camera to acquire and analyze video footage of the child's surroundings in real time. The analysis unit can also use a microphone to collect and analyze audio data from the surroundings. Furthermore, the analysis unit can combine and analyze the data from the camera and microphone to gain a more detailed understanding of the situation. This allows for a detailed analysis of the surrounding situation by collecting data through cameras and microphones. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform the data analysis.
[0069] The generation unit can generate voice messages based on the analysis results. For example, the generation unit uses a generation AI to generate appropriate voice messages based on the analysis results. For example, if a child is about to cross the street against a red light, the generation unit can generate a voice message such as, "○○-chan, the light is red. Please stop." The generation unit can also generate a voice message such as, "A suspicious person is following you. Please be careful," if a suspicious person is following a child. Furthermore, if an unknown adult approaches a child, the generation unit can generate a voice message such as, "I'm sorry, I can't talk to you right now." In this way, by generating voice messages based on the analysis results, appropriate messages can be conveyed to children. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input the data analyzed by the analysis unit into the generation AI and have the generation AI perform the generation of voice messages.
[0070] The transmission unit can deliver the generated voice message to the child. The transmission unit can deliver the voice message to the child using, for example, a speaker. It can also deliver the voice message to the child using earphones. Furthermore, the transmission unit can deliver the voice message to the child using a vibration function. This allows the child to receive appropriate attention by receiving the generated voice message. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the voice message generated by the generation unit into a generation AI and have the generation AI perform the delivery of the voice message.
[0071] The notification unit can notify parents in real time. For example, the notification unit can notify parents through a dedicated app. For example, the notification unit can send a notification to the parent's smartphone, displaying the child's current location and situation in real time. The notification unit can also notify parents via email. Furthermore, the notification unit can notify parents via SMS. This allows for a quick response by notifying parents in real time. Some or all of the above-described processes in the notification unit may be performed using AI or not. For example, the notification unit can input the danger detected by the analysis unit into a generation AI and have the generation AI generate the notification content.
[0072] The analysis unit can estimate the child's emotions and adjust the accuracy of the analysis based on the estimated emotions. For example, if the child is anxious, the analysis unit can collect more detailed data to improve the accuracy of the analysis. Alternatively, if the child is relaxed, the analysis unit can perform normal data collection and standard analysis. Furthermore, if the child is excited, the analysis unit can focus on detecting abnormal behavior and adjust the accuracy of the analysis accordingly. This allows for more accurate analysis by adjusting the accuracy of the analysis based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI or not. For example, the analysis unit can input the child's emotion data into a generative AI and have the generative AI perform emotion estimation.
[0073] The analysis unit can optimize its analysis algorithm by referring to past data when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations at a specific location based on past data and adjust the analysis algorithm accordingly. It can also predict dangerous situations during a specific time period by referring to past data and optimize the analysis algorithm accordingly. Furthermore, the analysis unit can adjust the analysis algorithm based on past data to make it easier to detect specific behavioral patterns. This improves the accuracy of the analysis by optimizing the analysis algorithm by referring to past data. Some or all of the above processes in the analysis unit may be performed using AI or not. For example, the analysis unit can input past data into a generating AI and have the generating AI perform the optimization of the analysis algorithm.
[0074] The analysis unit can analyze camera and microphone data in real time and detect abnormal behavioral patterns. For example, the analysis unit can analyze camera footage in real time to detect when a child is being chased by a suspicious person. The analysis unit can also analyze microphone audio data in real time to detect when a child is calling for help. Furthermore, the analysis unit can combine and analyze camera and microphone data to detect when a child is in a dangerous situation. This allows for the rapid detection of abnormal behavioral patterns by analyzing camera and microphone data in real time. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform the detection of abnormal behavioral patterns.
[0075] The analysis unit can estimate a child's emotions and determine the priority of analysis based on the estimated emotions. For example, if a child is feeling fear, the analysis unit will prioritize analyzing that situation. Alternatively, if the child is feeling safe, the analysis unit can perform a normal analysis and prioritize other important situations. Furthermore, if the child is agitated, the analysis unit can prioritize detecting abnormal behavior and adjust the order of analysis accordingly. This allows for prioritizing analysis of important situations by determining the priority of analysis based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI or not. For example, the analysis unit can input child emotion data into a generative AI and have the generative AI perform emotion estimation.
[0076] The analysis unit can improve the accuracy of its analysis by considering the child's location information when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations in specific locations based on the child's location information and improve the accuracy of its analysis. It can also predict dangerous situations during specific time periods by referring to the child's location information and improve the accuracy of its analysis. Furthermore, the analysis unit can improve the accuracy of its analysis by making it easier to detect specific behavioral patterns based on the child's location information. In this way, improving the accuracy of the analysis by considering the child's location information enables more accurate analysis. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input the child's location information into a generating AI and have the generating AI perform the analysis accuracy improvement.
[0077] The analysis unit can improve the accuracy of its analysis by filtering ambient noise when analyzing data from cameras and microphones. For example, the analysis unit can filter ambient noise to clearly detect children's voices or the voices of suspicious individuals. The analysis unit can also filter ambient noise to make it easier to analyze specific voice patterns. Furthermore, the analysis unit can filter ambient noise to make it easier to detect abnormal behavioral patterns. In this way, filtering ambient noise improves the accuracy of the analysis. Some or all of the above processing in the analysis unit may be performed using AI or not. For example, the analysis unit can input data acquired from cameras and microphones into a generating AI and have the generating AI perform ambient noise filtering.
[0078] The generation unit can estimate a child's emotions and adjust the tone of the voice message based on the estimated emotions. For example, if a child is frightened, the generation unit can generate a voice message in a calm tone. It can also generate a voice message in a bright tone if the child is relaxed. Furthermore, if the child is excited, it can generate a voice message in a gentle tone. This allows for more appropriate messaging by adjusting the tone of the voice message based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit is performed using the generative AI. For example, the generation unit can input child emotion data into the generative AI and have the generative AI perform tone adjustment of the voice message.
[0079] The generation unit can optimize the content of the voice messages it generates based on past analysis results. For example, the generation unit can generate voice messages suitable for a specific situation based on past analysis results. It can also generate voice messages suitable for a specific time period by referring to past analysis results. Furthermore, it can generate voice messages suitable for a specific behavioral pattern based on past analysis results. In this way, by optimizing the content of voice messages based on past analysis results, more effective messages can be generated. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input past analysis results into the generation AI and have the generation AI perform the optimization of voice messages.
[0080] The voice generation unit can generate more natural-sounding speech by meticulously analyzing the characteristics of voices, such as those of parents or teachers, when imitating their voices. For example, the voice generation unit can meticulously analyze the tone and rhythm of parents' and teachers' voices to generate natural-sounding speech. Furthermore, the voice generation unit can generate individual voice messages based on the characteristics of parents' and teachers' voices. In addition, the voice generation unit can analyze the intonation of parents' and teachers' voices to generate even more natural-sounding speech. This allows for the generation of more natural-sounding speech by meticulously analyzing voice characteristics. Some or all of the above-described processes in the voice generation unit are performed using a generation AI. For example, the voice generation unit can input data of parents' and teachers' voices into the generation AI, allowing the AI to generate natural-sounding speech.
[0081] The generation unit can estimate a child's emotions and adjust the length of the voice message based on the estimated emotions. For example, if the child is frightened, the generation unit can generate a short, concise voice message. If the child is relaxed, the generation unit can also generate a longer voice message with more detailed explanations. Furthermore, if the child is excited, the generation unit can generate a voice message with visually stimulating effects. By adjusting the length of the voice message based on the child's emotions, a more appropriate message can be conveyed. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit is performed using the generative AI. For example, the generation unit can input child emotion data into the generative AI and have the generative AI perform the length adjustment of the voice message.
[0082] The generation unit can customize the content of the generated voice messages according to the child's age and gender. For example, the generation unit can generate voice messages using easy-to-understand language according to the child's age. It can also generate voice messages with appropriate tone and content according to the child's gender. Furthermore, the generation unit can generate individual voice messages based on the child's age and gender. This allows for more appropriate messages to be conveyed by customizing voice messages according to the child's age and gender. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can input data on the child's age and gender into the generation AI and have the generation AI perform the voice message customization.
[0083] The voice generation unit can handle diverse scenarios by generating different voice variations when imitating the voices of parents and teachers. For example, the voice generation unit can generate different variations of parents' and teachers' voices and provide voice messages suitable for specific scenarios. The voice generation unit can also generate individual voice messages based on the variations of parents' and teachers' voices. Furthermore, the voice generation unit can analyze the different variations of parents' and teachers' voices and generate more natural-sounding voices. This allows it to handle diverse scenarios by generating different voice variations. Some or all of the above processing in the voice generation unit is performed using a generation AI. For example, the voice generation unit can input data of parents' and teachers' voices into the generation AI and have the generation AI perform the generation of different voice variations.
[0084] The transmission unit can estimate the child's emotions and adjust the delivery method of the voice message based on the estimated emotions. For example, if the child is frightened, the transmission unit will deliver the voice message in a calm tone. It can also deliver the voice message in a bright tone if the child is relaxed. Furthermore, if the child is excited, it can deliver the voice message in a gentle tone. This allows for the delivery of a more appropriate message by adjusting the delivery method based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the child's emotion data into the generative AI and have the generative AI adjust the delivery method of the voice message.
[0085] The transmission unit can automatically adjust the volume of voice messages by considering the ambient noise level. For example, if the surroundings are noisy, the transmission unit can automatically increase the volume to transmit the voice message. Conversely, if the surroundings are quiet, the transmission unit can automatically decrease the volume to transmit the voice message. Furthermore, the transmission unit can analyze the ambient noise level in real time and transmit the voice message at the optimal volume. This ensures that voice messages are properly conveyed by automatically adjusting the volume in consideration of the ambient noise level. Some or all of the above processing in the transmission unit may be performed using AI, or it may be performed without AI. For example, the transmission unit can input ambient noise level data into a generating AI and have the generating AI perform automatic volume adjustment.
[0086] The transmission unit can monitor the child's response in real time when transmitting a voice message and regenerate the message as needed. For example, if the child does not respond to the voice message, the transmission unit will regenerate and transmit the message again. The transmission unit can also modify the content and regenerate the message if the child gives an inappropriate response to the voice message. Furthermore, the transmission unit can monitor the child's response in real time and regenerate the message at the optimal timing. This allows for the regeneration of an appropriate message by monitoring the child's response in real time. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the child's response data into a generation AI and have the generation AI perform message regeneration.
[0087] The communication unit can estimate the child's emotions and adjust the timing of voice message delivery based on the estimated emotions. For example, if the child is feeling frightened, the communication unit will deliver the voice message immediately. It can also deliver the voice message at an appropriate time if the child is relaxed. Furthermore, if the child is excited, the communication unit can deliver the voice message at a calmer time. This allows for more appropriate timing of message delivery by adjusting the timing based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the communication unit may be performed using AI or not. For example, the communication unit can input the child's emotion data into the generative AI and have the generative AI adjust the timing of voice message delivery.
[0088] The transmission unit can select the optimal transmission method when transmitting a voice message, taking into account the child's location information. For example, if the child is in a specific location, the transmission unit will transmit a voice message appropriate for that location. The transmission unit can also select the optimal transmission method based on the child's location information. Furthermore, if the child is moving, the transmission unit can update the location information in real time and select the optimal transmission method. This allows for more effective message delivery by selecting the optimal transmission method while considering the child's location information. Some or all of the above processing in the transmission unit may be performed using AI or not. For example, the transmission unit can input the child's location information into a generating AI and have the generating AI select the optimal transmission method.
[0089] The transmission unit can select the optimal transmission method when transmitting voice messages, taking into account the child's device information. For example, if the child is using a smartphone, the transmission unit will transmit a voice message optimized for the smartphone. It can also transmit a voice message optimized for a tablet if the child is using a tablet. Furthermore, if the child is using a smartwatch, the transmission unit can transmit a voice message optimized for the smartwatch. This allows for more effective message delivery by selecting the optimal transmission method based on the child's device information. Some or all of the above processing in the transmission unit may be performed using AI, or without AI. For example, the transmission unit can input the child's device information into a generating AI and have the generating AI select the optimal transmission method.
[0090] The notification unit can estimate a child's emotions and adjust the notification content based on the estimated emotions. For example, if a child is feeling frightened, the notification unit can send an emergency notification to the parent. It can also send a normal notification if the child is relaxed. Furthermore, if the child is agitated, the notification unit can send a notification that includes a detailed description of the situation. This allows the notification unit to provide parents with appropriate information by adjusting the notification content based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input child emotion data into a generative AI and have the generative AI adjust the notification content.
[0091] The notification unit can select the most suitable notification method by referring to past notification history when notifying parents. For example, the notification unit can select the notification method that parents are most likely to respond to based on past notification history. The notification unit can also select a notification method suitable for a specific time period by referring to past notification history. Furthermore, the notification unit can select a notification method that allows parents to respond most quickly based on past notification history. This allows parents to respond quickly by selecting the most suitable notification method by referring to past notification history. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input past notification history into a generating AI and have the generating AI select the most suitable notification method.
[0092] The notification unit can provide additional information to parents when notifying them, providing a detailed report of the child's current situation. For example, the notification unit may send a notification that includes the child's current location information. It may also send a notification that provides a detailed explanation of the child's surroundings. Furthermore, the notification unit may analyze the child's behavior patterns and send a notification that includes a detailed explanation of the situation. This allows parents to respond appropriately by providing a detailed report of the child's current situation. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit may input the child's current situation data into a generating AI and have the generating AI generate the additional information.
[0093] The notification unit can estimate a child's emotions and determine the priority of notifications based on the estimated emotions. For example, if a child is feeling frightened, the notification unit can set the notification priority to the highest level. It can also send notifications with normal priority if the child is relaxed. Furthermore, if a child is agitated, the notification unit can prioritize sending notifications that include detailed situational descriptions. This ensures that important notifications are sent preferentially by prioritizing them based on the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input child emotion data into a generative AI and have the generative AI determine the notification priority.
[0094] The notification unit can select the most appropriate notification method when notifying a parent, taking into account the parent's location. For example, if the parent is at home, the notification unit may prioritize phone notification. If the parent is at work, the notification unit may prioritize email notification. Furthermore, if the parent is on the move, the notification unit may prioritize SMS notification. This allows parents to respond quickly by selecting the most appropriate notification method based on their location. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input the parent's location information into a generating AI and have the generating AI select the most appropriate notification method.
[0095] The notification unit can provide a function that allows parents to monitor the situation in real time through a dedicated app when a notification is sent. For example, the notification unit can display the child's current location in real time through the dedicated app. The notification unit can also monitor the child's surroundings in real time through the dedicated app. Furthermore, the notification unit can analyze the child's behavior patterns in real time through the dedicated app and notify the parents. This allows parents to respond quickly by monitoring the situation in real time through the dedicated app. Some or all of the above processing in the notification unit may be performed using AI or not. For example, the notification unit can input the dedicated app into a generating AI and have the generating AI perform the provision of the real-time monitoring function.
[0096] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0097] The analysis unit can be equipped with the ability to learn and predict children's behavior patterns. For example, it can predict how a child will behave at a specific time based on past behavioral data. It can also learn children's behavior patterns and detect abnormal behavior. Furthermore, it can predict dangerous situations based on children's behavior patterns and issue warnings in advance. This allows for more effective child safety by learning and predicting children's behavior patterns.
[0098] The generation unit can estimate a child's emotions and adjust the content of the voice message based on those emotions. For example, if a child is feeling frightened, the generation unit will generate a reassuring voice message. It can also generate an encouraging voice message if the child is relaxed. Furthermore, if the child is agitated, it can generate a calming voice message. This allows for more appropriate messages to be conveyed by adjusting the content of the voice message based on the child's emotions.
[0099] The notification unit can estimate the child's emotions when notifying parents and adjust the urgency of the notification based on those emotions. For example, if the child is feeling frightened, the notification unit will send a high-urgency notification. It can also send a normal notification if the child is relaxed. Furthermore, if the child is agitated, the notification unit can send a notification that includes a detailed description of the situation. This allows the system to provide parents with appropriate information by adjusting the urgency of notifications based on the child's emotions.
[0100] The analysis unit can improve the accuracy of its analysis by taking weather information into account when analyzing the surrounding conditions. For example, the analysis unit adjusts the accuracy of its camera image analysis when visibility is poor in rainy weather. It can also adjust the accuracy of its audio data analysis when there is strong wind. Furthermore, the analysis unit can predict specific dangerous situations based on weather information and improve the accuracy of its analysis. By improving the accuracy of the analysis by taking weather information into account, more accurate analysis becomes possible.
[0101] The generation unit can estimate a child's emotions and select the language of the voice message based on those emotions. For example, if a child is feeling frightened, the generation unit can generate a reassuring voice message in the child's native language. It can also generate an encouraging voice message in the child's second language if the child is relaxed. Furthermore, if the child is agitated, the generation unit can generate a calming voice message in the child's native language. This allows for the delivery of more appropriate messages by selecting the language of the voice message based on the child's emotions.
[0102] The notification unit can analyze a child's behavior patterns and adjust the notification content based on predicted behavior when notifying parents. For example, if a child is often in a specific location at a specific time, the notification unit will send a notification containing information about that location. The notification unit can also send notifications containing warnings about specific behaviors based on the child's behavior patterns. Furthermore, the notification unit can analyze a child's behavior patterns and adjust the timing of notifications based on predicted behavior. In this way, by analyzing a child's behavior patterns and adjusting the notification content based on predicted behavior, the notification unit can provide parents with appropriate information.
[0103] The analysis unit can estimate the child's emotions and determine the priority of the analysis based on the estimated emotions. For example, if the child is feeling fear, the analysis unit will prioritize analyzing that situation. If the child is feeling safe, the analysis unit can perform a normal analysis and prioritize other important situations. Furthermore, if the child is agitated, the analysis unit can prioritize the detection of abnormal behavior and adjust the order of analysis accordingly. In this way, by determining the priority of analysis based on the child's emotions, important situations can be analyzed preferentially.
[0104] The generation unit can customize the content of voice messages according to the child's learning progress. For example, if a child is at a specific learning stage, the generation unit will generate a voice message with content appropriate for that stage. The generation unit can also generate voice messages using easy-to-understand language based on the child's learning progress. Furthermore, the generation unit can generate individual voice messages according to the child's learning progress. This allows for the delivery of more appropriate messages by customizing voice messages according to the child's learning progress.
[0105] The notification unit can estimate the child's emotions when notifying parents and prioritize notifications based on those emotions. For example, if the child is feeling frightened, the notification unit will set the notification priority to the highest level. It can also send notifications with normal priority if the child is relaxed. Furthermore, if the child is agitated, the notification unit can prioritize sending notifications that include detailed situational descriptions. This allows important notifications to be sent first by prioritizing them based on the child's emotions.
[0106] The analysis unit can improve the accuracy of its analysis by considering the child's location information when analyzing the surrounding environment. For example, the analysis unit can predict dangerous situations in specific locations based on the child's location information, thereby improving the accuracy of the analysis. Furthermore, the analysis unit can also predict dangerous situations during specific time periods by referring to the child's location information, thereby improving the accuracy of the analysis. In addition, the analysis unit can improve the accuracy of the analysis by making it easier to detect specific behavioral patterns based on the child's location information. By improving the accuracy of the analysis by considering the child's location information, a more accurate analysis becomes possible.
[0107] The following briefly describes the processing flow for example form 2.
[0108] Step 1: The analysis unit analyzes the child's surroundings. The analysis unit collects data through cameras and microphones and analyzes the surroundings in detail. For example, it uses cameras to acquire and analyze video footage of the child's surroundings in real time. It can also collect and analyze audio data from the surroundings using microphones. Furthermore, by combining and analyzing the data from cameras and microphones, a more detailed understanding of the situation can be obtained. Step 2: The generation unit detects danger based on the data analyzed by the analysis unit and generates a voice message. The generation unit uses a generation AI to generate an appropriate voice message based on the analysis results. For example, if a child is about to cross against a red light, it will generate a voice message such as, "○○-chan, it's a red light. Please stop." It can also generate a voice message such as, "A suspicious person is following you. Please be careful." Furthermore, if an unknown adult approaches a child, it can generate a voice message such as, "I'm sorry, I can't talk to you right now." Step 3: The transmission unit transmits the voice message generated by the generation unit to the child. The transmission unit transmits the voice message to the child using a speaker. It can also transmit the voice message to the child using earphones. Furthermore, it can transmit the voice message to the child using a vibration function. Step 4: The notification unit notifies parents of the dangers detected by the analysis unit. The notification unit notifies parents through a dedicated app. For example, it sends a notification to the parent's smartphone, displaying the child's current location and situation in real time. It can also notify parents via email or SMS.
[0109] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0110] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0111] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0112] Each of the multiple elements described above, including the analysis unit, generation unit, transmission unit, and notification unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the analysis unit analyzes the child's surroundings using the camera 42 and microphone 38B of the smart device 14, and the control unit 46A processes the analysis results. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, and generates an appropriate voice message based on the analysis results. The transmission unit transmits the voice message to the child using the speaker 40B or earphones of the smart device 14, for example. The notification unit is implemented in the specific processing unit 290 of the data processing unit 12, and notifies the guardian through a dedicated app. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0113] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0114] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0116] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0120] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0123] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0125] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0127] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0128] Each of the multiple elements described above, including the analysis unit, generation unit, transmission unit, and notification unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the analysis unit analyzes the child's surroundings using the camera 42 and microphone 238 of the smart glasses 214, and the control unit 46A processes the analysis results. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and generates an appropriate voice message based on the analysis results. The transmission unit transmits the voice message to the child using the speaker 240 or earphones of the smart glasses 214, for example. The notification unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and notifies the guardian through a dedicated app. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0129] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0130] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0132] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0136] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0139] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0141] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0143] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0144] Each of the multiple elements described above, including the analysis unit, generation unit, transmission unit, and notification unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the analysis unit analyzes the child's surroundings using the camera 42 and microphone 238 of the headset terminal 314, and the control unit 46A processes the analysis results. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and generates an appropriate voice message based on the analysis results. The transmission unit transmits the voice message to the child using the speaker 240 or earphones of the headset terminal 314, for example. The notification unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and notifies the guardian through a dedicated application. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0145] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0146] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0152] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0153] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0154] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0155] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0156] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0157] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0158] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0159] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0160] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0161] Each of the multiple elements described above, including the analysis unit, generation unit, transmission unit, and notification unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the analysis unit analyzes the child's surroundings using the robot 414's camera 42 and microphone 238, and the control unit 46A processes the analysis results. The generation unit is implemented in the data processing unit 12's specific processing unit 290 and generates an appropriate voice message based on the analysis results. The transmission unit transmits the voice message to the child using the robot 414's speaker 240 or earphones. The notification unit is implemented in the data processing unit 12's specific processing unit 290 and notifies the guardian through a dedicated app. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.
[0162] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0163] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0164] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0165] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0166] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0167] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0169] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0170] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0172] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0173] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0174] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0175] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0176] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0177] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0178] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0179] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0180] (Note 1) An analysis unit that analyzes the surrounding environment of the child, A generation unit detects danger based on the data analyzed by the aforementioned analysis unit and generates an audio message, A transmission unit that conveys the voice message generated by the generation unit to the child, The system includes a notification unit that notifies the guardian of the danger detected by the analysis unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, Collect data through cameras and microphones. The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is Generate a voice message based on the analysis results. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned transmission unit is The generated voice message is conveyed to the child. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned notification unit, Notify parents in real time. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned analysis unit, The system estimates the child's emotions and adjusts the accuracy of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned analysis unit, When analyzing the surrounding environment, we optimize the analysis algorithm by referring to past data. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned analysis unit, It analyzes camera and microphone data in real time to detect abnormal behavioral patterns. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned analysis unit, The system estimates the child's emotions and determines the priority of analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, When analyzing the surrounding environment, the accuracy of the analysis is improved by taking into account the child's location information. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, When analyzing camera and microphone data, filtering out ambient noise improves the accuracy of the analysis. The system described in Appendix 1, characterized by the features described herein. (Note 12) The generating unit is It estimates the child's emotions and adjusts the tone of voice messages based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The generating unit is The content of the generated voice messages is optimized based on past analysis results. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is When imitating the voices of parents or teachers, the system analyzes the characteristics of the voice in detail to generate more natural-sounding speech. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is It estimates the child's emotions and adjusts the length of the voice message based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is Customize the content of the generated voice messages according to the child's age and gender. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is When imitating the voices of parents or teachers, it generates different voice variations to accommodate a variety of scenarios. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned transmission unit is It estimates the child's emotions and adjusts how voice messages are delivered based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned transmission unit is When transmitting voice messages, the volume is automatically adjusted to take into account the surrounding noise level. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned transmission unit is When delivering voice messages, the system monitors the child's response in real time and regenerates the message as needed. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned transmission unit is The system estimates the child's emotions and adjusts the timing of voice messages based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned transmission unit is When sending voice messages, the optimal method of transmission is selected considering the child's location. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned transmission unit is When delivering voice messages, the optimal delivery method is selected considering the child's device information. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned notification unit, The system estimates the child's emotions and adjusts the notification content based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned notification unit, When notifying parents, refer to past notification history to select the most appropriate notification method. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned notification unit, When notifying parents, provide additional information to report in detail on the child's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned notification unit, The system estimates the child's emotions and prioritizes notifications based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned notification unit, When notifying parents, the most appropriate notification method will be selected, taking into account the parents' location information. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned notification unit, When notifying parents, the system provides a feature that allows them to monitor the situation in real time through a dedicated app. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0181] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. An analysis unit that analyzes the surrounding environment of the child, A generation unit detects danger based on the data analyzed by the aforementioned analysis unit and generates an audio message, A transmission unit that conveys the voice message generated by the generation unit to the child, The system includes a notification unit that notifies the guardian of the danger detected by the analysis unit. A system characterized by the following features.
2. The aforementioned analysis unit, Collect data through cameras and microphones. The system according to feature 1.
3. The generating unit is Generate a voice message based on the analysis results. The system according to feature 1.
4. The aforementioned transmission unit is The generated voice message is conveyed to the child. The system according to feature 1.
5. The aforementioned notification unit, Notify parents in real time. The system according to feature 1.
6. The aforementioned analysis unit, The system estimates the child's emotions and adjusts the accuracy of the analysis based on the estimated emotions. The system according to feature 1.
7. The aforementioned analysis unit, When analyzing the surrounding environment, we optimize the analysis algorithm by referring to past data. The system according to feature 1.
8. The aforementioned analysis unit, It analyzes camera and microphone data in real time to detect abnormal behavioral patterns. The system according to feature 1.
9. The aforementioned analysis unit, The system estimates the child's emotions and determines the priority of analysis based on the estimated emotions. The system according to feature 1.
10. The aforementioned analysis unit, When analyzing the surrounding environment, the accuracy of the analysis is improved by taking into account the child's location information. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A