system

A system that monitors and analyzes parental interactions to generate and play back appropriate video and audio responses addresses the challenge of caring for babies when parents are unavailable, effectively soothing them.

JP2026041383APending Publication Date: 2026-03-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Parents often struggle to find time to care for their babies, leading to feelings of loneliness and anxiety in babies, and there is a lack of effective systems to quickly identify and respond to their crying or irritability.

Method used

A system that monitors a baby's condition and records parental words, actions, voice, and facial expressions, analyzes this data to generate appropriate video and audio responses using generative AI, and plays them back to the baby in real-time to provide reassurance.

Benefits of technology

The system effectively soothes babies by providing timely and appropriate responses even when parents are away or busy, alleviating feelings of loneliness and anxiety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026041383000001_ABST
    Figure 2026041383000001_ABST
Patent Text Reader

Abstract

Provide a system. A means for monitoring the condition of a baby; A means of recording the words, actions, voices, and facial expressions of parents; means for analyzing the monitored baby condition data and the recorded parent data; means for generating parental reaction video and audio based on the analysis; A means of playing the generated video and audio to the baby; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, it is becoming increasingly common for parents to find it difficult to find time to care for their babies. As a result, babies are more likely to feel lonely and anxious, but there is a lack of ways to reassure babies when their parents are away or busy. It is also difficult to quickly identify the cause of a baby's crying or irritability and provide the most appropriate response. A system to solve this problem is needed. [Means for solving the problem]

[0005] The present invention provides a system for monitoring a baby's condition and recording the parent's words, actions, voice, and facial expressions. Specifically, the system includes a means for monitoring the baby's crying and movements, a means for recording the parent's facial expressions and voice, a means for analyzing the monitoring data and recorded data, a means for generating a video and audio of the parent's reaction based on the analysis results, and a means for playing the generated video and audio for the baby. The system also includes a means for analyzing the tone of the baby's crying and behavior, identifying the parent's reaction pattern, and a means for selecting the most appropriate reaction video from the generated video and playing it in a timely manner, enabling the system to provide the baby with the necessary care quickly and appropriately.

[0006] "Means for monitoring the baby's condition" refers to devices such as sensors, cameras, and microphones that detect and record the baby's crying and behavior in real time.

[0007] "Means for recording parental behavior, voice, and facial expressions" refers to devices that collect parents' facial expressions, tone of voice, content of statements, gestures, etc. through cameras and microphones and store them as data.

[0008] "Means for analyzing monitored baby status data and recorded parent data" refers to software or algorithms that analyze collected baby crying and behavior data and parent verbal and behavior data to infer the baby's emotions and needs.

[0009] The "means for generating video and audio of parental reactions based on analysis" refers to a generative AI technology that generates video and audio that mimics the parent's facial expressions and voice based on the analysis results, providing an appropriate response to the baby.

[0010] The "means for playing the generated video and audio to the baby" refers to a display or speaker that plays the generated video and audio of the parent's reaction on a device such as a tablet or smartphone and shows it to the baby.

[0011] "Means for analyzing the tone and behavior of a baby's cries" are algorithms that analyze the sound waveforms of a baby's cries and behavioral patterns to identify emotions and needs.

[0012] The "means for identifying common parental response patterns" is a machine learning algorithm that identifies specific parental responses to babies (e.g., smiling, gentle voice, etc.) based on data analysis and learns response patterns.

[0013] "Means for selecting the most appropriate video from among the multiple generated parent reaction videos and playing it for the baby in a timely manner" is a system for selecting the video that best suits the baby's state and emotions from among the multiple generated parent reaction videos and playing it in real time. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] This invention is a system that soothes a baby in place of the parent, mainly monitoring the baby's condition, collecting and analyzing the parent's words, actions, voice, and facial expressions, and based on the analysis results, generating and playing appropriate video and audio responses of the parent. This system is designed to reassure the baby when the parent is not around, even when the baby feels lonely or anxious.

[0036] System configuration and operation

[0037] The system consists of the following major components:

[0038] 1. Data Collection Module

[0039] 2. Sentiment Analysis Module

[0040] 3. Generative AI Module

[0041] 4. Video Playback Module

[0042] 1. Data Collection Module

[0043] To monitor the baby's condition, the device (e.g., a smartphone or tablet) is equipped with a camera and microphone. It also contains sensors and a recording device to record the parent's words, actions, voice, and facial expressions. As the parent interacts with the baby, the device collects this data in real time and sends it to a server.

[0044] Examples:

[0045] The device records the parent smiling and speaking to the child in a gentle voice, and then uploads the video and audio data to a server.

[0046] 2. Sentiment Analysis Module

[0047] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs from the tone of the crying and body language.

[0048] Examples:

[0049] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[0050] 3. Generative AI Module

[0051] Based on the results of the emotion analysis, the server uses generative AI to generate video and audio of the parent's reaction. At this stage, technology is used to realistically reproduce the parent's facial expressions and voice based on past data.

[0052] Examples:

[0053] If the baby "requests a diaper change," the generative AI will create a video of the parent taking the baby to the changing table.

[0054] 4. Video Playback Module

[0055] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[0056] Examples:

[0057] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0058] Natural language explanation of the system's program

[0059] 1. Data Collection

[0060] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0061] 2. Sentiment analysis

[0062] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[0063] 3. Generation AI

[0064] Server: Based on sentiment analysis data, it identifies the most appropriate parental response and uses generative AI to generate video and audio that replicates that facial expression and voice.

[0065] 4. Video playback

[0066] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, giving the baby a sense of security.

[0067] Through the above processing, the present invention provides a system that can effectively soothe a baby even when the parent is away or busy.

[0068] The processing flow will be explained below.

[0069] Step 1: Data collection

[0070] Device: Records interactions between the baby and parent using a camera and microphone. Data such as the parent's facial expressions, tone of voice, what they say, and gestures are collected and sent to the server in real time.

[0071] Example: A scene in which a parent speaks to a baby in a gentle voice is recorded using a camera and microphone, and the data is sent to a server.

[0072] Step 2: Monitor your baby

[0073] Device: Monitors the baby's crying and behavior in real time, collecting data using a camera and microphone. When the baby starts crying, the data is sent to the server.

[0074] Example: When a baby starts crying, the crying is recorded and a video is also taken and sent to a server.

[0075] Step 3: Parental sentiment analysis

[0076] Server: Analyzes the parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." Text mining is performed on the tone of voice and the content of the speech, and emotion labels are assigned.

[0077] Example: The server analyzes a parent's "smile" and "gentle voice" and assigns them the emotion label "gentle."

[0078] Step 4: Analyze the baby's condition

[0079] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[0080] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[0081] Step 5: Parent Reaction Generation

[0082] Server: Combines the analyzed parent's emotional patterns with the baby's state to predict the most appropriate parental response. Using generative AI, the predicted parental response is generated as video and audio.

[0083] Example: If a baby is judged to be "lonely," the AI ​​will generate a video of the parent speaking to the baby in a gentle voice.

[0084] Step 6: Select video

[0085] Server: Select the most appropriate parental reaction video from the generated videos.

[0086] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice.

[0087] Step 7: Send your video

[0088] Server: Sends the selected parent's video to the device.

[0089] Example: Send a video of someone speaking to you in a gentle voice to your device.

[0090] Step 8: Play the video to your baby

[0091] Device: The parent's video and audio are played in real time on the tablet screen and through the speaker.

[0092] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[0093] This series of steps allows parents to provide a quick and appropriate response to their baby's emotions and needs, making it possible for them to feel safe even when parents are away or busy.

[0094] Example 1

[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0096] In modern homes, there is a need to reassure babies even when parents are away or busy. However, there is no established system that can analyze parents' words, actions, and emotions, as well as babies' cries and behavior, in real time and provide appropriate responses based on that analysis. Conventional methods have difficulty reproducing the parent's presence and reactions in real time, making it difficult to respond immediately when a baby feels lonely or anxious. New technology is needed to solve these problems.

[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0098] In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating a video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio to the baby, means for analyzing the tone and behavior of the baby's crying, means for identifying the parent's emotional pattern, means for generating a parent's reaction using a generative AI model, and means for selecting the parent's reaction video most appropriate for the baby's condition from among the multiple generated videos and playing it to the baby in a timely manner. This makes it possible to instantly provide an appropriate response according to the baby's condition even when the parent is not present, effectively reassuring the baby.

[0099] The "means for monitoring the baby's condition" is a device that records the baby's crying, behavior, and facial expressions in real time and collects them as data.

[0100] "Means for recording parental behavior, voice, and facial expressions" refers to a device that uses a camera or microphone to record a parent's speech, facial expressions, gestures, etc. as digital data.

[0101] "Monitored baby condition data" refers to data collected in real time, including information on the tone of a baby's crying, behavioral patterns, and facial expressions.

[0102] "Recorded parent data" refers to digital data that records the parent's speech, facial expressions, behavior, etc.

[0103] The "means for analyzing" is a software and hardware system for analyzing the collected data and determining the emotions and states of the baby and parent.

[0104] The "means for generating parental reaction video and audio" is a system including a generative AI model for generating parental video and audio based on the results of emotion analysis.

[0105] The "means for playing the generated video and audio to the baby" is a terminal for showing and playing the generated video and audio of the parent's reaction to the baby.

[0106] The "means for analyzing the tone and behavior of a baby's crying" is a software and hardware system for analyzing the audio data of a baby's crying and behavioral patterns and estimating its condition.

[0107] The "means for identifying parental emotional patterns" is a system for identifying a parent's emotional state from their facial expressions, tone of voice, etc., and identifying it as a pattern.

[0108] A "generative AI model" is an algorithm that uses artificial intelligence to learn from past data and generate parental reaction videos and audio.

[0109] The "means for playing back to the baby in a timely manner" refers to a terminal and control system for playing back the parent's reaction video and audio, which are generated at the optimal timing depending on the baby's condition.

[0110] MODE FOR CARRYING OUT THE INVENTION

[0111] This system monitors the baby's condition, collects and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate parental reaction video and audio based on the analysis results. This system is designed to reassure the baby when the parent is not around, even when the baby is feeling lonely or anxious.

[0112] The system consists of the following main components:

[0113] 1. Data Collection Module

[0114] 2. Sentiment Analysis Module

[0115] 3. Generative AI Module

[0116] 4. Video Playback Module

[0117] Data Collection Module

[0118] Device: A mobile device such as a smartphone or tablet is used. Using a camera and microphone, the device collects the baby's crying and behavior, as well as the parent's facial expressions and voice, in real time and sends the data to a server. For example, the device can record a parent smiling or speaking to the baby in a gentle voice, and upload the video and audio data to the server.

[0119] Sentiment Analysis Module

[0120] Server: Receives and analyzes the baby and parent data sent from the device. It identifies emotional patterns from the parent's facial expressions and tone of voice, and infers the baby's state by analyzing the tone of the baby's cry and behavioral patterns. For example, if it analyzes a baby's cry and determines that the baby is feeling lonely, it will identify a pattern from previously collected parent data that says, "Speak to the baby in a gentle voice when the baby feels lonely."

[0121] Generative AI Module

[0122] Server: Based on the results of the sentiment analysis, a generative AI model is used to generate a video and audio of the parent's reaction. The generative AI model uses past data to realistically reproduce the parent's facial expressions and voice. For example, if it determines that the baby is "requesting a diaper change," the generative AI model generates a video of the parent taking the baby to the changing table.

[0123] Video Playback Module

[0124] Device: Receives the parent's video and audio responses sent from the server and plays them in a timely manner to reassure the baby. For example, if a baby is "crying in fear," the device will play a video of the parent reassuring the baby with a gentle smile.

[0125] Examples of prompt statements

[0126] Prompt statements are used to instruct the system to perform a specific operation. For example:

[0127] "Analyze a baby's cry and generate a video of the parent's reaction if they are feeling lonely."

[0128] "Generate a video of a parent's reaction when it is determined that the child is requesting a diaper change."

[0129] "Generate a parent's reaction video to comfort a scared, crying baby."

[0130] In this way, the present invention provides a system that can effectively soothe babies even when parents are away or busy. The data collection, emotion analysis, generative AI, and video playback modules work together to quickly provide appropriate responses according to the baby's condition.

[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0132] Step 1: Data collection

[0133] Device: Uses a camera and microphone to record the baby's cries, movements, and the parent's words, voice, and facial expressions in real time.

[0134] Input: Baby's crying, behavior, parent's words, voice, facial expressions

[0135] Output: Recorded video data, audio data

[0136] Specific behavior:

[0137] The device activates the camera and captures video of the baby's behavior and the parent's facial expressions.

[0138] The device activates a microphone and collects the baby's cries and the parent's voice as audio data.

[0139] The collected data is compressed and sent to the server in real time.

[0140] Step 2: Sentiment analysis

[0141] Server: Receives and analyzes the baby and parent data sent from the device.

[0142] Input: Video and audio data sent from the device

[0143] Output: Baby's emotional state data, parent's emotional pattern data

[0144] Specific behavior:

[0145] The server analyzes the video data and identifies the baby's facial expressions and movements.

[0146] The server analyzes the audio data and identifies the tone of the baby's cry and the tone of the parent's voice.

[0147] Identify the baby's emotional state (e.g., lonely, scared) and the parent's emotional patterns (e.g., smiling, soft voice).

[0148] Step 3: Prompt generation

[0149] Server: Based on the results of sentiment analysis, generate prompt sentences to be input to the generative AI.

[0150] Input: Baby's emotional state data, parent's emotional pattern data

[0151] Output: Prompt text to be input to the generation AI

[0152] Specific behavior:

[0153] The server generates a corresponding prompt sentence based on the analysis results (e.g., "Generate a parent's response to a lonely baby").

[0154] Step 4: Parent Reaction Generation

[0155] Server: Uses a generative AI model to generate video and audio parental responses based on the prompt.

[0156] Input: Prompt sentence, past parent response data

[0157] Output: Parents' reaction video data, audio data

[0158] Specific behavior:

[0159] A prompt sentence is input into the generative AI model, and a video and audio of the parent's reaction is generated.

[0160] The generated video and audio are encoded and prepared for transmission to the device.

[0161] Step 5: Play the video

[0162] Device: Receives the parent's reaction video and audio sent from the server and plays them according to the baby's condition.

[0163] Input: Parent reaction video data, audio data

[0164] Output: Video and audio played to the baby

[0165] Specific behavior:

[0166] The terminal receives the data from the server.

[0167] Video and audio are played at the optimal time for your baby's current condition.

[0168] The baby's reactions during and after playback are recorded again and sent as feedback to the server.

[0169] In this way, the system can effectively guide the baby through the entire process from data collection to video playback.

[0170] (Application example 1)

[0171] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0172] In food delivery services, when delivery personnel are busy or unable to communicate directly with customers during deliveries, they tend to inadequately explain the delivery status or problems to customers. This results in customer anxiety and dissatisfaction, leading to a decline in the quality of service. It is also necessary to improve the situation where delivery personnel are unable to respond appropriately and in a timely manner when they encounter unexpected problems.

[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0174] In this invention, the server includes a means for monitoring the status of the delivery person, a means for recording the speech, behavior, voice, and facial expression of the delivery person, a means for analyzing the monitored status data of the delivery person and the recorded data of the delivery person, a means for generating a video and audio response of the delivery person, and a means for playing the generated video and audio to the customer. This makes it possible to provide an appropriate response to the customer even when the delivery person is not present, thereby reducing the customer's anxiety and dissatisfaction.

[0175] "Delivery Person" means an individual or entity responsible for delivering goods to a customer.

[0176] "Means for monitoring the status" refers to the function of checking the target's behavior and voice in real time using devices such as cameras and microphones.

[0177] "Means for recording speech, voice, and facial expressions" refers to devices and software that store the delivery person's speech, voice, and facial expressions as digital data.

[0178] "Means of analysis" refers to the algorithms and programs used to analyze collected data and convert it into meaningful information.

[0179] "Means for generating responsive video and audio" refers to technology that creates visual and audio data to appropriately reproduce the delivery person's condition and emotions based on the analysis results.

[0180] "Means for displaying video and audio to customers" means any device or software used to display or play the generated visual and audio data to customers.

[0181] "Means for identifying response patterns" refers to algorithms or methods for detecting and identifying patterns of delivery personnel's behavior and speech in multiple situations.

[0182] "Means for timely playback" refers to technology or equipment that allows appropriate reactions and responses to be played back in real time at the required timing.

[0183] The system for implementing this invention aims to alleviate customer anxiety and dissatisfaction by monitoring the status of delivery personnel and providing appropriate responses to customers. A specific implementation method for this system is described below.

[0184] 1. System Configuration

[0185] This system consists of the following main modules:

[0186] a. Data Collection Module

[0187] The device (e.g., smartphone or tablet) is equipped with a camera and microphone, which records the delivery person's words, actions, voice, and facial expressions. This data is collected in real time and sent to a server. For example, the device records a scene in which the delivery person has a "confused expression" or "explains something to a customer over the phone," and uploads the video and audio data to the server.

[0188] b. Sentiment Analysis Module

[0189] The server analyzes the facial expressions, voice, and behavioral data of the delivery person sent from the device to identify the delivery person's emotional patterns. At the same time, it also analyzes the behavioral data of the delivery person sent from the device to infer the reason for the delivery person's actions. For example, it can determine that the delivery person is delayed due to "traffic congestion" based on a confused expression on their face.

[0190] c. Generative AI module

[0191] The server uses generative AI to generate a video and audio response from the delivery person based on the results of the emotion analysis. At this stage, technology is used to realistically reproduce the delivery person's facial expressions and voice based on past data. For example, a video can be generated in which the delivery person explains to the customer that the delivery will be five minutes late.

[0192] d. Video playback module

[0193] The generated video and audio data of the delivery person is sent from the server to the terminal. The terminal selects and plays the delivery person's response video that is most appropriate for the customer. This allows the customer to understand the delivery status and feel reassured. For example, if the customer is confused, a video of the delivery person explaining in a calm tone will be played.

[0194] 2. Hardware and software used

[0195] Terminal (smartphone, tablet): Records the delivery person's status using a camera and microphone and sends it to the server.

[0196] Server: Analyzes the collected data using deep learning models (e.g., face recognition models using Keras). For generative AI, frameworks such as TensorFlow are used.

[0197] Generative AI: Technology that uses facial recognition and voice data to generate realistic facial expressions and voices of delivery personnel.

[0198] Playback device: Software installed on the device that plays the generated video and audio to the customer.

[0199] 3. Examples of prompt sentences

[0200] For example, the prompt for a scene in which a delivery person is stuck in traffic and looks confused is as follows:

[0201] "The delivery person is stuck in traffic and has a confused expression. Based on that, generate a sentence that will convey an appropriate message to the customer."

[0202] With the above components and processes, the present invention makes it possible to provide appropriate responses to customers even when a delivery person is not present, thereby reducing customer anxiety and dissatisfaction.

[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0204] Step 1:

[0205] The terminal collects the delivery person's status (video and audio).

[0206] Specifically, the device's camera and microphone are used to record the delivery person's movements, voice, and facial expressions in real time, and this data is recorded as video and audio files that include the delivery person's speech and facial expressions.

[0207] Input: Real-time video and audio of the delivery person

[0208] Output: Video file (e.g. mp4 format), audio file (e.g. wav format)

[0209] Step 2:

[0210] The data collected by the device is sent to the server in real time.

[0211] The collected video and audio files are uploaded to a server using an appropriate communication protocol (e.g., HTTP).

[0212] Input: Video files, audio files

[0213] Output: Data stored on the server

[0214] Step 3:

[0215] The server analyzes the received data and identifies the emotional patterns of the delivery person.

[0216] Specifically, an emotion analysis model using Keras is used to analyze the facial expressions and tone of the delivery person's voice from the received video and audio data, and this analysis identifies the delivery person's emotions, such as "confused" or "happy."

[0217] Input: Video and audio files on the server

[0218] Output: Emotional pattern (e.g., confusion, joy)

[0219] Step 4:

[0220] Generative AI generates response video and audio based on emotional patterns.

[0221] Based on the results of the sentiment analysis, a generative AI model (e.g., a text-to-video generation model) is used to create a realistic video or audio response from the delivery person. For example, if the delivery person is confused, a video containing the message "Delivery will be 5 minutes late" is generated.

[0222] Input: Emotion pattern, prompt sentence

[0223] Output: Response video file, response audio file

[0224] Step 5:

[0225] The server generates a response video and audio and sends it to the terminal.

[0226] The generated response video and audio are then transmitted to the terminal again using an appropriate communication protocol.

[0227] Input: Response video file, response audio file

[0228] Output: Data stored on the device

[0229] Step 6:

[0230] The terminal plays the most appropriate response video and audio to the customer.

[0231] The device will then play the video and audio responses it receives to help customers understand the situation and feel at ease. For example, if a video explaining a delivery delay is played, customers will be able to understand the delivery person's situation and their anxiety about waiting will be alleviated.

[0232] Input: Response video file, response audio file

[0233] Output: Video and audio response played

[0234] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0235] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby in place of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[0236] System configuration and operation

[0237] The system consists of the following major components:

[0238] 1. Data Collection Module

[0239] 2. Sentiment Analysis Module

[0240] 3. Generative AI Module

[0241] 4. Video Playback Module

[0242] 5. Emotion Engine

[0243] 1. Data Collection Module

[0244] The devices (smartphones and tablets) are equipped with cameras and microphones to monitor the baby's condition and record the parent's behavior, voice, and facial expressions. Using these devices, the devices transmit the baby-parent interaction data to a server in real time.

[0245] Examples:

[0246] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[0247] 2. Sentiment Analysis Module

[0248] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs.

[0249] Examples:

[0250] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[0251] 3. Generative AI Module

[0252] The server uses AI to generate the parent's reaction video and audio based on the emotion analysis results. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[0253] Examples:

[0254] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0255] 4. Video Playback Module

[0256] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[0257] Examples:

[0258] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0259] 5. Emotion Engine

[0260] The emotion engine recognizes the emotions of the parent user in real time and reflects that emotional data in the operation of the entire system. The emotion engine analyzes video and audio data to identify emotional states such as "the parent is feeling stressed." Based on these results, the generative AI then generates more appropriate and detailed videos of the parent's reactions.

[0261] Examples:

[0262] If the parent is feeling stressed, the generated video will adjust the parent's facial expression to appear slightly calmer.

[0263] Natural language explanation of the system's program

[0264] 1. Data Collection

[0265] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0266] 2. Sentiment analysis

[0267] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[0268] 3. Emotion Engine

[0269] Server: The emotion engine recognizes the parent's real-time emotions and reflects the results in the analysis data, thereby understanding the parent's current emotional state.

[0270] 4. Generation AI

[0271] Server: Based on the sentiment analysis data and the results of the emotion engine, it identifies the most appropriate parental response and uses generative AI to generate video and audio that reproduces that facial expression and voice.

[0272] 5. Video playback

[0273] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0274] Through the above processing, the present invention can effectively soothe a baby even when the parent is away or busy, and by reflecting the emotions of the user (parent), it is possible to provide a more realistic and appropriate response.

[0275] The processing flow will be explained below.

[0276] Step 1: Data collection

[0277] Device: The camera and microphone record the parent's facial expressions, tone of voice, speech, gestures, etc. This data is sent to the server in real time.

[0278] Example: A parent speaking to their baby in a gentle voice is recorded using a camera and microphone, and the video and audio data is sent to a server.

[0279] Step 2: Monitor your baby

[0280] Device: Monitors the baby's crying and behavior in real time. Cameras and microphones capture the crying and video and send it to the server.

[0281] Example: When a baby starts crying, the crying is recorded and the video is sent to a server.

[0282] Step 3: Parental sentiment analysis

[0283] Server: Analyzes the received parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." It also analyzes the tone of voice and the content of the speech to assign an emotional label.

[0284] Example: The server analyzes data on a parent's "smile" and "gentle voice" and assigns the emotion label "gentle" to them.

[0285] Step 4: Analyze the baby's condition

[0286] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[0287] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[0288] Step 5: Parental emotion recognition by emotion engine

[0289] Server: Recognizes the parent's real-time emotions using an emotion engine. Analyzes the parent's video and audio data to detect their current emotional state (e.g., stress, relief).

[0290] Example: The emotion engine analyzes the video and audio data of a parent and recognizes that the parent is feeling stressed.

[0291] Step 6: Parent Reaction Generation

[0292] Server: Based on the results of the parent's emotion analysis and the emotion engine, the server predicts the most appropriate parental response. Using generative AI, it generates video and audio that realistically reproduces the parent's facial expressions and voice.

[0293] Example: If a baby is judged to be "lonely" and the parent is recognized as "stressed," the AI ​​will generate a video of the parent speaking to the baby in a calm, gentle voice.

[0294] Step 7: Select and send your video

[0295] Server: Selects the most appropriate parental reaction video from the generated videos and sends it to the device.

[0296] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice and send it to the device.

[0297] Step 8: Play the video

[0298] Device: The parent's video and audio are sent from the server and played in real time on the tablet or smartphone screen, with the audio played through the speaker.

[0299] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[0300] This series of steps enables the system to respond quickly and appropriately to the baby's emotions and requests, and by reflecting the emotions of the user (parent), it can provide a more realistic and effective way to soothe a baby.

[0301] Example 2

[0302] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0303] In today's busy home environment, it is often difficult for parents to spend all their time with their babies. In particular, if a baby starts crying or feels anxious, parents' inability to respond quickly can have an impact on the baby's emotions. Another issue is that parents themselves may find it difficult to respond optimally if they are stressed.

[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio for the baby, and means for recognizing the parent's real-time emotions and reflecting the data in the operation of the entire system. This allows the baby to be soothed effectively even when the parent is away or busy, and by reflecting the parent's emotions, a more realistic and appropriate response is possible.

[0305] A "means for monitoring the baby's condition" is a device or system that uses a camera or microphone to record the baby's movements and cries and analyzes the data.

[0306] "Means for recording parental behavior, voice, and facial expressions" refers to a device or system that uses a camera or microphone to record the words spoken by a parent, their tone of voice, and their facial expressions.

[0307] "Monitored baby status data" refers to data on the baby's movements and crying collected by monitoring means such as cameras and microphones.

[0308] "Recorded parent data" refers to data such as parental words, tone of voice, and facial expressions collected through means that record parental behavior, voice, and facial expressions.

[0309] "Means of analysis" refers to software and algorithms that analyze the collected data and identify the emotions and states of the baby and parent.

[0310] The "means for generating parental reaction video and audio" is a device or system that generates video clips and audio to reproduce the parent's actions and voice based on the analysis results.

[0311] The "means for playing the generated video and audio to the baby" is a device or system for playing the generated parent's reaction video and audio to the baby in a timely manner.

[0312] "Means for recognizing the parent's real-time emotions and reflecting that data in the operation of the entire system" refers to a device or system that analyzes the parent's current emotional state using an emotion engine and takes the results into account and reflects them in the operation of the system.

[0313] An "emotion engine" is software or an algorithm that analyzes video and audio data to identify the emotional state of parents and babies and reflect that information in the system's operation.

[0314] "Generative AI" refers to the artificial intelligence models and techniques used to generate parental reaction videos and audio.

[0315] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby on behalf of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[0316] The system consists of the following major components:

[0317] 1. Data Collection Module

[0318] 2. Data transmission module

[0319] 3. Sentiment Analysis Module

[0320] 4. Emotion Engine Module

[0321] 5. Generative AI Module

[0322] 6. Video Playback Module

[0323] 1. Data Collection Module

[0324] The device uses a camera and microphone to record the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and the parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0325] Examples:

[0326] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[0327] 2. Data transmission module

[0328] The device transmits the collected data to the server in real time, which reflects the real-time status of the baby and parent and is analyzed in the next processing step.

[0329] Examples:

[0330] The device records the parent's facial expressions and voice data and sends it to the server.

[0331] 3. Sentiment Analysis Module

[0332] The server analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," and "surprised." It also analyzes the baby's data and estimates the baby's state (e.g., loneliness, fear, or request for a diaper change) from the tone of the crying and behavior.

[0333] Examples:

[0334] The server analyzes the tone of the baby's cry to determine whether the baby is feeling lonely, and then, based on past data collected from parents, identifies a pattern of parents speaking to their babies in a gentle voice when they feel lonely.

[0335] 4. Emotion Engine Module

[0336] The server uses an emotion engine to recognize the parent's real-time emotions and reflects the results in the analysis data, thereby making it possible to understand the parent's current emotional state.

[0337] Examples:

[0338] The emotion engine analyzes the parent's facial expressions and tone of voice and determines that the parent is stressed. This information is reflected in subsequent processing.

[0339] 5. Generative AI Module

[0340] The server uses generative AI to generate the parent's reaction video and audio based on the emotion analysis data and the results of the emotion engine. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[0341] Examples:

[0342] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0343] 6. Video Playback Module

[0344] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0345] Examples:

[0346] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0347] Example prompts to be input to the generative AI model

[0348] "If your baby is feeling lonely, generate a video of the parent talking to them in a gentle voice."

[0349] "If a baby is scared and crying, generate a video of the parent reassuring the baby with a soft smile."

[0350] "If a baby requests a diaper change, generate a video of the parent taking the baby to the changing table."

[0351] This system allows babies to be soothed effectively even when parents are away or busy, and by reflecting the parents' emotions, it allows for more realistic and appropriate responses.

[0352] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0353] Step 1:

[0354] Data collection

[0355] The device uses a camera and microphone to record the baby's state and the parent's words, actions, voice, and facial expressions, thereby collecting scenes of interaction between the baby and the parent.

[0356] Input: Baby's cries and movements, parent's facial expressions and words.

[0357] Output: Collected video and audio data.

[0358] Specific operation: The device captures real-time footage of the baby crying and the parent's gentle response, and records this data.

[0359] Step 2:

[0360] Data transmission

[0361] The device transmits the collected data in real time to a server, which prepares the data for the next analysis step.

[0362] Input: Collected video and audio data.

[0363] Output: The data sent to the server.

[0364] Specific operation: The device records the baby's crying and the parent's reaction, and sends the video and audio data to a server via the network.

[0365] Step 3:

[0366] sentiment analysis

[0367] The server analyzes the received parent data to identify the parent's emotional patterns, and analyzes the baby data to estimate the baby's state from the tone of the crying and behavior.

[0368] Input: Video and audio data sent to the server.

[0369] Output: Parent and baby emotion pattern data.

[0370] Specific operation: The server analyzes the baby's cry and determines that the baby is "lonely," and at the same time identifies patterns from the parent's facial expressions and tone of voice that indicate "the parent has the intention to comfort the baby."

[0371] Step 4:

[0372] Emotion Engine Recognition

[0373] The server recognizes the parent's real-time emotions using an emotion engine and reflects that data in the overall system behavior, which takes the parent's current emotional state into account when proceeding to the next step.

[0374] Input: Parent video and audio data.

[0375] Output: Parent emotion data from the emotion engine.

[0376] How it works: The emotion engine analyzes the parent's facial expressions and tone of voice to determine that the parent is stressed. This information is used as input for the generative AI.

[0377] Step 5:

[0378] Generative AI video creation

[0379] The server uses generative AI to generate parental reaction videos and audio based on the emotion analysis results and emotion engine data, reflecting the user's emotions and recreating realistic facial expressions and voices.

[0380] Input: Sentiment analysis data and sentiment engine data.

[0381] Output: Generated reaction video and audio data.

[0382] Specific behavior: If it is determined that the baby "requires a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0383] Step 6:

[0384] Video playback

[0385] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0386] Input: Generated reaction video and audio data.

[0387] Output: Video and audio played to the baby.

[0388] What it does: The device calms a crying baby by playing a video of a parent comforting the baby with a soft smile.

[0389] (Application example 2)

[0390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0391] In the workplace, reducing worker stress and fatigue and creating a safe and efficient work environment are important challenges. However, conventional methods make it difficult to grasp workers' emotions and physical conditions in real time and take appropriate measures. As a result, workers' health risks increase and work efficiency may decline. For this reason, there is a demand for a system that can monitor workers' emotions and physical conditions in real time and take appropriate measures.

[0392] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0393] In this invention, the server includes a means for monitoring the state of workers, a means for recording the words, actions, voices, and facial expressions of workers, a means for analyzing the monitored state data of workers and the recorded data of workers, a means for generating appropriate reaction video and audio based on the analysis, and a means for playing the generated video and audio to the workers. This makes it possible to monitor the emotions and physical condition of workers in real time and take appropriate measures.

[0394] "Means for monitoring worker status" refers to devices or systems for monitoring the physical status of workers, such as their complexion, posture, and movements, in real time.

[0395] "Means for recording the words, actions, voice, and facial expressions of workers" refers to devices such as cameras and microphones that record the content of workers' words, tone of voice, facial expressions, etc.

[0396] "Means for analyzing monitored worker condition data and recorded worker data" refers to software and hardware for analyzing collected worker physical and emotional data.

[0397] "Means for generating appropriate response videos and audio based on analysis" refers to artificial intelligence (AI) technology for automatically generating appropriate response messages and videos according to the emotional state of workers.

[0398] "Means for playing the generated video and audio to workers" refers to devices such as displays and speakers used to communicate the generated response messages and videos to workers.

[0399] The system that realizes this application example, a worker emotion recognition support robot, consists of the following main components:

[0400] System configuration and operation

[0401] 1. Means of monitoring the condition of workers

[0402] Cameras suitable for the work environment (e.g., general surveillance cameras or high-resolution cameras) are used to monitor the status of workers in real time, and highly sensitive microphones are installed to collect the voices of workers.

[0403] 2. Means of recording the words, actions, voices, and facial expressions of workers

[0404] The robots will be fitted with devices that record the workers' facial expressions and tone of voice, allowing for detailed recording of their behavior and emotions.

[0405] 3. Means of analyzing monitored worker status data and recorded worker data

[0406] The server is installed with software that uses GOOGLE TENSOR® FLOW® to analyze worker status data, analyzing facial expressions and voice to identify the worker's level of stress and fatigue.

[0407] 4. A method for generating appropriate reaction video and audio based on the analysis

[0408] The server uses OpenAI's GPT-4 to generate appropriate video and audio responses tailored to the worker's state. This generative AI model uses past data to create more natural and effective responses.

[0409] 5. A means of playing the generated video and audio to workers

[0410] The generated video and audio are then played back to the worker through the robot's on-board display and speakers, allowing the worker to receive encouragement and break suggestions at appropriate times.

[0411] Examples of the system and how to use it

[0412] How to use this system will be explained with concrete examples.

[0413] 1. Data Collection

[0414] The device is equipped with a camera and microphone that records the worker's facial expressions and voice in real time and sends the data to a server. For example, if a worker says "I'm tired" while working, the microphone will collect the voice.

[0415] 2. Sentiment analysis

[0416] The data is then analyzed on the server, and TensorFlow is used to identify the worker's fatigue and stress levels for the day based on subtle changes in facial expressions and tone of voice.

[0417] 3. Emotion Engine

[0418] The server uses Amazon Comprehend to recognize the worker's real-time emotions and incorporates that information into the next steps, allowing for more accurate responses.

[0419] 4. Generation AI

[0420] Using GPT-4, the system generates voice messages and videos based on the results of sentiment analysis, such as encouraging workers or suggesting them to take a break. For example, it generates a message saying, "It's a good idea to take a short break."

[0421] 5. Video playback

[0422] The generated messages and videos are played through the robot's displays and speakers, and if a worker shows signs of fatigue, a video message such as "Take a deep breath to relax" will be displayed at that point.

[0423] Examples of prompt statements

[0424] An example of an input prompt for the generative AI model is as follows:

[0425] If a worker looks tired, generate a message saying, "Good work. Maybe you should take a break."

[0426] If a worker is feeling stressed, generate a voice message saying, "Take a deep breath to relax. It's okay."

[0427] This system allows workers' emotions and physical condition to be monitored in real time, and appropriate measures to be taken to reduce health risks and improve work efficiency.

[0428] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0429] Step 1:

[0430] The camera and microphone installed on the terminal collect the facial expressions and voices of the workers. The system records what the workers say and their facial expressions while working and sends this data to a server in real time. The input is the workers' audio and video data, and the output is the transmission of this data to the server.

[0431] Step 2:

[0432] The server analyzes the received facial and voice data. Using Google® TensorFlow, it analyzes the worker's emotional state (e.g., fatigue, stress, concentration, etc.) from their facial expressions. It also analyzes the tone and content of their voice from the audio data to determine their emotional state. The input is the worker's video and audio data, and the output is the worker's emotional state data as an analysis result.

[0433] Step 3:

[0434] The server uses an emotion engine to perform a more detailed analysis of the worker's real-time emotions and incorporates the results into the overall data. It uses Amazon Comprehend to identify the emotional state and output intermediate data for application to generative AI. The input is facial expression and voice analysis results, and the output is detailed emotional data for application to the generative AI model.

[0435] Step 4:

[0436] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate response messages and videos based on the worker's emotional state. The generative AI model creates messages and videos that are optimal for the worker's situation based on the prompt text. The input is detailed emotional data and the prompt text, and the output is the generated response message and video data.

[0437] Step 5:

[0438] The generated messages and videos are sent from the server to the terminal and played to the worker through the terminal's display and speaker. The user watches the presented message or video and takes appropriate action (e.g., take a break or take a deep breath). The input is the generated response message and video data, and the output is to provide them to the worker.

[0439] This series of steps will create a system that supports worker health and work efficiency in real time.

[0440] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0442] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0443] [Second embodiment]

[0444] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0445] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0446] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0447] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0448] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0450] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0451] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0452] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0453] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0454] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0455] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0456] This invention is a system that soothes a baby in place of the parent, mainly monitoring the baby's condition, collecting and analyzing the parent's words, actions, voice, and facial expressions, and based on the analysis results, generating and playing appropriate video and audio responses of the parent. This system is designed to reassure the baby when the parent is not around, even when the baby feels lonely or anxious.

[0457] System configuration and operation

[0458] The system consists of the following major components:

[0459] 1. Data Collection Module

[0460] 2. Sentiment Analysis Module

[0461] 3. Generative AI Module

[0462] 4. Video Playback Module

[0463] 1. Data Collection Module

[0464] To monitor the baby's condition, the device (e.g., a smartphone or tablet) is equipped with a camera and microphone. It also contains sensors and a recording device to record the parent's words, actions, voice, and facial expressions. As the parent interacts with the baby, the device collects this data in real time and sends it to a server.

[0465] Examples:

[0466] The device records the parent smiling and speaking to the child in a gentle voice, and then uploads the video and audio data to a server.

[0467] 2. Sentiment Analysis Module

[0468] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs from the tone of the crying and body language.

[0469] Examples:

[0470] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[0471] 3. Generative AI Module

[0472] Based on the results of the emotion analysis, the server uses generative AI to generate video and audio of the parent's reaction. At this stage, technology is used to realistically reproduce the parent's facial expressions and voice based on past data.

[0473] Examples:

[0474] If the baby "requests a diaper change," the generative AI will create a video of the parent taking the baby to the changing table.

[0475] 4. Video Playback Module

[0476] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[0477] Examples:

[0478] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0479] Natural language explanation of the system's program

[0480] 1. Data Collection

[0481] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0482] 2. Sentiment analysis

[0483] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[0484] 3. Generation AI

[0485] Server: Based on sentiment analysis data, it identifies the most appropriate parental response and uses generative AI to generate video and audio that replicates that facial expression and voice.

[0486] 4. Video playback

[0487] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, giving the baby a sense of security.

[0488] Through the above processing, the present invention provides a system that can effectively soothe a baby even when the parent is away or busy.

[0489] The processing flow will be explained below.

[0490] Step 1: Data collection

[0491] Device: Records interactions between the baby and parent using a camera and microphone. Data such as the parent's facial expressions, tone of voice, what they say, and gestures are collected and sent to the server in real time.

[0492] Example: A scene in which a parent speaks to a baby in a gentle voice is recorded using a camera and microphone, and the data is sent to a server.

[0493] Step 2: Monitor your baby

[0494] Device: Monitors the baby's crying and behavior in real time, collecting data using a camera and microphone. When the baby starts crying, the data is sent to the server.

[0495] Example: When a baby starts crying, the crying is recorded and a video is also taken and sent to a server.

[0496] Step 3: Parental sentiment analysis

[0497] Server: Analyzes the parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." Text mining is performed on the tone of voice and the content of the speech, and emotion labels are assigned.

[0498] Example: The server analyzes a parent's "smile" and "gentle voice" and assigns them the emotion label "gentle."

[0499] Step 4: Analyze the baby's condition

[0500] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[0501] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[0502] Step 5: Parent Reaction Generation

[0503] Server: Combines the analyzed parent's emotional patterns with the baby's state to predict the most appropriate parental response. Using generative AI, the predicted parental response is generated as video and audio.

[0504] Example: If a baby is judged to be "lonely," the AI ​​will generate a video of the parent speaking to the baby in a gentle voice.

[0505] Step 6: Select video

[0506] Server: Select the most appropriate parental reaction video from the generated videos.

[0507] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice.

[0508] Step 7: Send your video

[0509] Server: Sends the selected parent's video to the device.

[0510] Example: Send a video of someone speaking to you in a gentle voice to your device.

[0511] Step 8: Play the video to your baby

[0512] Device: The parent's video and audio are played in real time on the tablet screen and through the speaker.

[0513] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[0514] This series of steps allows parents to provide a quick and appropriate response to their baby's emotions and needs, making it possible for them to feel safe even when parents are away or busy.

[0515] Example 1

[0516] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0517] In modern homes, there is a need to reassure babies even when parents are away or busy. However, there is no established system that can analyze parents' words, actions, and emotions, as well as babies' cries and behavior, in real time and provide appropriate responses based on that analysis. Conventional methods have difficulty reproducing the parent's presence and reactions in real time, making it difficult to respond immediately when a baby feels lonely or anxious. New technology is needed to solve these problems.

[0518] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0519] In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating a video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio to the baby, means for analyzing the tone and behavior of the baby's crying, means for identifying the parent's emotional pattern, means for generating a parent's reaction using a generative AI model, and means for selecting the parent's reaction video most appropriate for the baby's condition from among the multiple generated videos and playing it to the baby in a timely manner. This makes it possible to instantly provide an appropriate response according to the baby's condition even when the parent is not present, effectively reassuring the baby.

[0520] The "means for monitoring the baby's condition" is a device that records the baby's crying, behavior, and facial expressions in real time and collects them as data.

[0521] "Means for recording parental behavior, voice, and facial expressions" refers to a device that uses a camera or microphone to record a parent's speech, facial expressions, gestures, etc. as digital data.

[0522] "Monitored baby condition data" refers to data collected in real time, including information on the tone of a baby's crying, behavioral patterns, and facial expressions.

[0523] "Recorded parent data" refers to digital data that records the parent's speech, facial expressions, behavior, etc.

[0524] The "means for analyzing" is a software and hardware system for analyzing the collected data and determining the emotions and states of the baby and parent.

[0525] The "means for generating parental reaction video and audio" is a system including a generative AI model for generating parental video and audio based on the results of emotion analysis.

[0526] The "means for playing the generated video and audio to the baby" is a terminal for showing and playing the generated video and audio of the parent's reaction to the baby.

[0527] The "means for analyzing the tone and behavior of a baby's crying" is a software and hardware system for analyzing the audio data of a baby's crying and behavioral patterns and estimating its condition.

[0528] The "means for identifying parental emotional patterns" is a system for identifying a parent's emotional state from their facial expressions, tone of voice, etc., and identifying it as a pattern.

[0529] A "generative AI model" is an algorithm that uses artificial intelligence to learn from past data and generate parental reaction videos and audio.

[0530] The "means for playing back to the baby in a timely manner" refers to a terminal and control system for playing back the parent's reaction video and audio, which are generated at the optimal timing depending on the baby's condition.

[0531] MODE FOR CARRYING OUT THE INVENTION

[0532] This system monitors the baby's condition, collects and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate parental reaction video and audio based on the analysis results. This system is designed to reassure the baby when the parent is not around, even when the baby is feeling lonely or anxious.

[0533] The system consists of the following main components:

[0534] 1. Data Collection Module

[0535] 2. Sentiment Analysis Module

[0536] 3. Generative AI Module

[0537] 4. Video Playback Module

[0538] Data Collection Module

[0539] Device: A mobile device such as a smartphone or tablet is used. Using a camera and microphone, the device collects the baby's crying and behavior, as well as the parent's facial expressions and voice, in real time and sends the data to a server. For example, the device can record a parent smiling or speaking to the baby in a gentle voice, and upload the video and audio data to the server.

[0540] Sentiment Analysis Module

[0541] Server: Receives and analyzes the baby and parent data sent from the device. It identifies emotional patterns from the parent's facial expressions and tone of voice, and infers the baby's state by analyzing the tone of the baby's cry and behavioral patterns. For example, if it analyzes a baby's cry and determines that the baby is feeling lonely, it will identify a pattern from previously collected parent data that says, "Speak to the baby in a gentle voice when the baby feels lonely."

[0542] Generative AI Module

[0543] Server: Based on the results of the sentiment analysis, a generative AI model is used to generate a video and audio of the parent's reaction. The generative AI model uses past data to realistically reproduce the parent's facial expressions and voice. For example, if it determines that the baby is "requesting a diaper change," the generative AI model generates a video of the parent taking the baby to the changing table.

[0544] Video Playback Module

[0545] Device: Receives the parent's video and audio responses sent from the server and plays them in a timely manner to reassure the baby. For example, if a baby is "crying in fear," the device will play a video of the parent reassuring the baby with a gentle smile.

[0546] Examples of prompt statements

[0547] Prompt statements are used to instruct the system to perform a specific operation. For example:

[0548] "Analyze a baby's cry and generate a video of the parent's reaction if they are feeling lonely."

[0549] "Generate a video of a parent's reaction when it is determined that the child is requesting a diaper change."

[0550] "Generate a parent's reaction video to comfort a scared, crying baby."

[0551] In this way, the present invention provides a system that can effectively soothe babies even when parents are away or busy. The data collection, emotion analysis, generative AI, and video playback modules work together to quickly provide appropriate responses according to the baby's condition.

[0552] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0553] Step 1: Data collection

[0554] Device: Uses a camera and microphone to record the baby's cries, movements, and the parent's words, voice, and facial expressions in real time.

[0555] Input: Baby's crying, behavior, parent's words, voice, facial expressions

[0556] Output: Recorded video data, audio data

[0557] Specific behavior:

[0558] The device activates the camera and captures video of the baby's behavior and the parent's facial expressions.

[0559] The device activates a microphone and collects the baby's cries and the parent's voice as audio data.

[0560] The collected data is compressed and sent to the server in real time.

[0561] Step 2: Sentiment analysis

[0562] Server: Receives and analyzes the baby and parent data sent from the device.

[0563] Input: Video and audio data sent from the device

[0564] Output: Baby's emotional state data, parent's emotional pattern data

[0565] Specific behavior:

[0566] The server analyzes the video data and identifies the baby's facial expressions and movements.

[0567] The server analyzes the audio data and identifies the tone of the baby's cry and the tone of the parent's voice.

[0568] Identify the baby's emotional state (e.g., lonely, scared) and the parent's emotional patterns (e.g., smiling, soft voice).

[0569] Step 3: Prompt generation

[0570] Server: Based on the results of sentiment analysis, generate prompt sentences to be input to the generative AI.

[0571] Input: Baby's emotional state data, parent's emotional pattern data

[0572] Output: Prompt text to be input to the generation AI

[0573] Specific behavior:

[0574] The server generates a corresponding prompt sentence based on the analysis results (e.g., "Generate a parent's response to a lonely baby").

[0575] Step 4: Parent Reaction Generation

[0576] Server: Uses a generative AI model to generate video and audio parental responses based on the prompt.

[0577] Input: Prompt sentence, past parent response data

[0578] Output: Parents' reaction video data, audio data

[0579] Specific behavior:

[0580] A prompt sentence is input into the generative AI model, and a video and audio of the parent's reaction is generated.

[0581] The generated video and audio are encoded and prepared for transmission to the device.

[0582] Step 5: Play the video

[0583] Device: Receives the parent's reaction video and audio sent from the server and plays them according to the baby's condition.

[0584] Input: Parent reaction video data, audio data

[0585] Output: Video and audio played to the baby

[0586] Specific behavior:

[0587] The terminal receives the data from the server.

[0588] Video and audio are played at the optimal time for your baby's current condition.

[0589] The baby's reactions during and after playback are recorded again and sent as feedback to the server.

[0590] In this way, the system can effectively guide the baby through the entire process from data collection to video playback.

[0591] (Application example 1)

[0592] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0593] In food delivery services, when delivery personnel are busy or unable to communicate directly with customers during deliveries, they tend to inadequately explain the delivery status or problems to customers. This results in customer anxiety and dissatisfaction, leading to a decline in the quality of service. It is also necessary to improve the situation where delivery personnel are unable to respond appropriately and in a timely manner when they encounter unexpected problems.

[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0595] In this invention, the server includes a means for monitoring the status of the delivery person, a means for recording the speech, behavior, voice, and facial expression of the delivery person, a means for analyzing the monitored status data of the delivery person and the recorded data of the delivery person, a means for generating a video and audio response of the delivery person, and a means for playing the generated video and audio to the customer. This makes it possible to provide an appropriate response to the customer even when the delivery person is not present, thereby reducing the customer's anxiety and dissatisfaction.

[0596] "Delivery Person" means an individual or entity responsible for delivering goods to a customer.

[0597] "Means for monitoring the status" refers to the function of checking the target's behavior and voice in real time using devices such as cameras and microphones.

[0598] "Means for recording speech, voice, and facial expressions" refers to devices and software that store the delivery person's speech, voice, and facial expressions as digital data.

[0599] "Means of analysis" refers to the algorithms and programs used to analyze collected data and convert it into meaningful information.

[0600] "Means for generating responsive video and audio" refers to technology that creates visual and audio data to appropriately reproduce the delivery person's condition and emotions based on the analysis results.

[0601] "Means for displaying video and audio to customers" means any device or software used to display or play the generated visual and audio data to customers.

[0602] "Means for identifying response patterns" refers to algorithms or methods for detecting and identifying patterns of delivery personnel's behavior and speech in multiple situations.

[0603] "Means for timely playback" refers to technology or equipment that allows appropriate reactions and responses to be played back in real time at the required timing.

[0604] The system for implementing this invention aims to alleviate customer anxiety and dissatisfaction by monitoring the status of delivery personnel and providing appropriate responses to customers. A specific implementation method for this system is described below.

[0605] 1. System Configuration

[0606] This system consists of the following main modules:

[0607] a. Data Collection Module

[0608] The device (e.g., smartphone or tablet) is equipped with a camera and microphone, which records the delivery person's words, actions, voice, and facial expressions. This data is collected in real time and sent to a server. For example, the device records a scene in which the delivery person has a "confused expression" or "explains something to a customer over the phone," and uploads the video and audio data to the server.

[0609] b. Sentiment Analysis Module

[0610] The server analyzes the facial expressions, voice, and behavioral data of the delivery person sent from the device to identify the delivery person's emotional patterns. At the same time, it also analyzes the behavioral data of the delivery person sent from the device to infer the reason for the delivery person's actions. For example, it can determine that the delivery person is delayed due to "traffic congestion" based on a confused expression on their face.

[0611] c. Generative AI module

[0612] The server uses generative AI to generate a video and audio response from the delivery person based on the results of the emotion analysis. At this stage, technology is used to realistically reproduce the delivery person's facial expressions and voice based on past data. For example, a video can be generated in which the delivery person explains to the customer that the delivery will be five minutes late.

[0613] d. Video playback module

[0614] The generated video and audio data of the delivery person is sent from the server to the terminal. The terminal selects and plays the delivery person's response video that is most appropriate for the customer. This allows the customer to understand the delivery status and feel reassured. For example, if the customer is confused, a video of the delivery person explaining in a calm tone will be played.

[0615] 2. Hardware and software used

[0616] Terminal (smartphone, tablet): Records the delivery person's status using a camera and microphone and sends it to the server.

[0617] Server: Analyzes the collected data using deep learning models (e.g., face recognition models using Keras). For generative AI, frameworks such as TensorFlow are used.

[0618] Generative AI: Technology that uses facial recognition and voice data to generate realistic facial expressions and voices of delivery personnel.

[0619] Playback device: Software installed on the device that plays the generated video and audio to the customer.

[0620] 3. Examples of prompt sentences

[0621] For example, the prompt for a scene in which a delivery person is stuck in traffic and looks confused is as follows:

[0622] "The delivery person is stuck in traffic and has a confused expression. Based on that, generate a sentence that will convey an appropriate message to the customer."

[0623] With the above components and processes, the present invention makes it possible to provide appropriate responses to customers even when a delivery person is not present, thereby reducing customer anxiety and dissatisfaction.

[0624] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0625] Step 1:

[0626] The terminal collects the delivery person's status (video and audio).

[0627] Specifically, the device's camera and microphone are used to record the delivery person's movements, voice, and facial expressions in real time, and this data is recorded as video and audio files that include the delivery person's speech and facial expressions.

[0628] Input: Real-time video and audio of the delivery person

[0629] Output: Video file (e.g. mp4 format), audio file (e.g. wav format)

[0630] Step 2:

[0631] The data collected by the device is sent to the server in real time.

[0632] The collected video and audio files are uploaded to a server using an appropriate communication protocol (e.g., HTTP).

[0633] Input: Video files, audio files

[0634] Output: Data stored on the server

[0635] Step 3:

[0636] The server analyzes the received data and identifies the emotional patterns of the delivery person.

[0637] Specifically, an emotion analysis model using Keras is used to analyze the facial expressions and tone of the delivery person's voice from the received video and audio data, and this analysis identifies the delivery person's emotions, such as "confused" or "happy."

[0638] Input: Video and audio files on the server

[0639] Output: Emotional pattern (e.g., confusion, joy)

[0640] Step 4:

[0641] Generative AI generates response video and audio based on emotional patterns.

[0642] Based on the results of the sentiment analysis, a generative AI model (e.g., a text-to-video generation model) is used to create a realistic video or audio response from the delivery person. For example, if the delivery person is confused, a video containing the message "Delivery will be 5 minutes late" is generated.

[0643] Input: Emotion pattern, prompt sentence

[0644] Output: Response video file, response audio file

[0645] Step 5:

[0646] The server generates a response video and audio and sends it to the terminal.

[0647] The generated response video and audio are then transmitted to the terminal again using an appropriate communication protocol.

[0648] Input: Response video file, response audio file

[0649] Output: Data stored on the device

[0650] Step 6:

[0651] The terminal plays the most appropriate response video and audio to the customer.

[0652] The device will then play the video and audio responses it receives to help customers understand the situation and feel at ease. For example, if a video explaining a delivery delay is played, customers will be able to understand the delivery person's situation and their anxiety about waiting will be alleviated.

[0653] Input: Response video file, response audio file

[0654] Output: Video and audio response played

[0655] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0656] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby in place of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[0657] System configuration and operation

[0658] The system consists of the following major components:

[0659] 1. Data Collection Module

[0660] 2. Sentiment Analysis Module

[0661] 3. Generative AI Module

[0662] 4. Video Playback Module

[0663] 5. Emotion Engine

[0664] 1. Data Collection Module

[0665] The devices (smartphones and tablets) are equipped with cameras and microphones to monitor the baby's condition and record the parent's behavior, voice, and facial expressions. Using these devices, the devices transmit the baby-parent interaction data to a server in real time.

[0666] Examples:

[0667] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[0668] 2. Sentiment Analysis Module

[0669] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs.

[0670] Examples:

[0671] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[0672] 3. Generative AI Module

[0673] The server uses AI to generate the parent's reaction video and audio based on the emotion analysis results. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[0674] Examples:

[0675] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0676] 4. Video Playback Module

[0677] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[0678] Examples:

[0679] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0680] 5. Emotion Engine

[0681] The emotion engine recognizes the emotions of the parent user in real time and reflects that emotional data in the operation of the entire system. The emotion engine analyzes video and audio data to identify emotional states such as "the parent is feeling stressed." Based on these results, the generative AI then generates more appropriate and detailed videos of the parent's reactions.

[0682] Examples:

[0683] If the parent is feeling stressed, the generated video will adjust the parent's facial expression to appear slightly calmer.

[0684] Natural language explanation of the system's program

[0685] 1. Data Collection

[0686] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0687] 2. Sentiment analysis

[0688] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[0689] 3. Emotion Engine

[0690] Server: The emotion engine recognizes the parent's real-time emotions and reflects the results in the analysis data, thereby understanding the parent's current emotional state.

[0691] 4. Generation AI

[0692] Server: Based on the sentiment analysis data and the results of the emotion engine, it identifies the most appropriate parental response and uses generative AI to generate video and audio that reproduces that facial expression and voice.

[0693] 5. Video playback

[0694] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0695] Through the above processing, the present invention can effectively soothe a baby even when the parent is away or busy, and by reflecting the emotions of the user (parent), it is possible to provide a more realistic and appropriate response.

[0696] The processing flow will be explained below.

[0697] Step 1: Data collection

[0698] Device: The camera and microphone record the parent's facial expressions, tone of voice, speech, gestures, etc. This data is sent to the server in real time.

[0699] Example: A parent speaking to their baby in a gentle voice is recorded using a camera and microphone, and the video and audio data is sent to a server.

[0700] Step 2: Monitor your baby

[0701] Device: Monitors the baby's crying and behavior in real time. Cameras and microphones capture the crying and video and send it to the server.

[0702] Example: When a baby starts crying, the crying is recorded and the video is sent to a server.

[0703] Step 3: Parental sentiment analysis

[0704] Server: Analyzes the received parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." It also analyzes the tone of voice and the content of the speech to assign an emotional label.

[0705] Example: The server analyzes data on a parent's "smile" and "gentle voice" and assigns the emotion label "gentle" to them.

[0706] Step 4: Analyze the baby's condition

[0707] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[0708] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[0709] Step 5: Parental emotion recognition by emotion engine

[0710] Server: Recognizes the parent's real-time emotions using an emotion engine. Analyzes the parent's video and audio data to detect their current emotional state (e.g., stress, relief).

[0711] Example: The emotion engine analyzes the video and audio data of a parent and recognizes that the parent is feeling stressed.

[0712] Step 6: Parent Reaction Generation

[0713] Server: Based on the results of the parent's emotion analysis and the emotion engine, the server predicts the most appropriate parental response. Using generative AI, it generates video and audio that realistically reproduces the parent's facial expressions and voice.

[0714] Example: If a baby is judged to be "lonely" and the parent is recognized as "stressed," the AI ​​will generate a video of the parent speaking to the baby in a calm, gentle voice.

[0715] Step 7: Select and send your video

[0716] Server: Selects the most appropriate parental reaction video from the generated videos and sends it to the device.

[0717] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice and send it to the device.

[0718] Step 8: Play the video

[0719] Device: The parent's video and audio are sent from the server and played in real time on the tablet or smartphone screen, with the audio played through the speaker.

[0720] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[0721] This series of steps enables the system to respond quickly and appropriately to the baby's emotions and requests, and by reflecting the emotions of the user (parent), it can provide a more realistic and effective way to soothe a baby.

[0722] Example 2

[0723] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0724] In today's busy home environment, it is often difficult for parents to spend all their time with their babies. In particular, if a baby starts crying or feels anxious, parents' inability to respond quickly can have an impact on the baby's emotions. Another issue is that parents themselves may find it difficult to respond optimally if they are stressed.

[0725] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio for the baby, and means for recognizing the parent's real-time emotions and reflecting the data in the operation of the entire system. This allows the baby to be soothed effectively even when the parent is away or busy, and by reflecting the parent's emotions, a more realistic and appropriate response is possible.

[0726] A "means for monitoring the baby's condition" is a device or system that uses a camera or microphone to record the baby's movements and cries and analyzes the data.

[0727] "Means for recording parental behavior, voice, and facial expressions" refers to a device or system that uses a camera or microphone to record the words spoken by a parent, their tone of voice, and their facial expressions.

[0728] "Monitored baby status data" refers to data on the baby's movements and crying collected by monitoring means such as cameras and microphones.

[0729] "Recorded parent data" refers to data such as parental words, tone of voice, and facial expressions collected through means that record parental behavior, voice, and facial expressions.

[0730] "Means of analysis" refers to software and algorithms that analyze the collected data and identify the emotions and states of the baby and parent.

[0731] The "means for generating parental reaction video and audio" is a device or system that generates video clips and audio to reproduce the parent's actions and voice based on the analysis results.

[0732] The "means for playing the generated video and audio to the baby" is a device or system for playing the generated parent's reaction video and audio to the baby in a timely manner.

[0733] "Means for recognizing the parent's real-time emotions and reflecting that data in the operation of the entire system" refers to a device or system that analyzes the parent's current emotional state using an emotion engine and takes the results into account and reflects them in the operation of the system.

[0734] An "emotion engine" is software or an algorithm that analyzes video and audio data to identify the emotional state of parents and babies and reflect that information in the system's operation.

[0735] "Generative AI" refers to the artificial intelligence models and techniques used to generate parental reaction videos and audio.

[0736] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby on behalf of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[0737] The system consists of the following major components:

[0738] 1. Data Collection Module

[0739] 2. Data transmission module

[0740] 3. Sentiment Analysis Module

[0741] 4. Emotion Engine Module

[0742] 5. Generative AI Module

[0743] 6. Video Playback Module

[0744] 1. Data Collection Module

[0745] The device uses a camera and microphone to record the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and the parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0746] Examples:

[0747] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[0748] 2. Data transmission module

[0749] The device transmits the collected data to the server in real time, which reflects the real-time status of the baby and parent and is analyzed in the next processing step.

[0750] Examples:

[0751] The device records the parent's facial expressions and voice data and sends it to the server.

[0752] 3. Sentiment Analysis Module

[0753] The server analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," and "surprised." It also analyzes the baby's data and estimates the baby's state (e.g., loneliness, fear, or request for a diaper change) from the tone of the crying and behavior.

[0754] Examples:

[0755] The server analyzes the tone of the baby's cry to determine whether the baby is feeling lonely, and then, based on past data collected from parents, identifies a pattern of parents speaking to their babies in a gentle voice when they feel lonely.

[0756] 4. Emotion Engine Module

[0757] The server uses an emotion engine to recognize the parent's real-time emotions and reflects the results in the analysis data, thereby making it possible to understand the parent's current emotional state.

[0758] Examples:

[0759] The emotion engine analyzes the parent's facial expressions and tone of voice and determines that the parent is stressed. This information is reflected in subsequent processing.

[0760] 5. Generative AI Module

[0761] The server uses generative AI to generate the parent's reaction video and audio based on the emotion analysis data and the results of the emotion engine. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[0762] Examples:

[0763] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0764] 6. Video Playback Module

[0765] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0766] Examples:

[0767] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0768] Example prompts to be input to the generative AI model

[0769] "If your baby is feeling lonely, generate a video of the parent talking to them in a gentle voice."

[0770] "If a baby is scared and crying, generate a video of the parent reassuring the baby with a soft smile."

[0771] "If a baby requests a diaper change, generate a video of the parent taking the baby to the changing table."

[0772] This system allows babies to be soothed effectively even when parents are away or busy, and by reflecting the parents' emotions, it allows for more realistic and appropriate responses.

[0773] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0774] Step 1:

[0775] Data collection

[0776] The device uses a camera and microphone to record the baby's state and the parent's words, actions, voice, and facial expressions, thereby collecting scenes of interaction between the baby and the parent.

[0777] Input: Baby's cries and movements, parent's facial expressions and words.

[0778] Output: Collected video and audio data.

[0779] Specific operation: The device captures real-time footage of the baby crying and the parent's gentle response, and records this data.

[0780] Step 2:

[0781] Data transmission

[0782] The device transmits the collected data in real time to a server, which prepares the data for the next analysis step.

[0783] Input: Collected video and audio data.

[0784] Output: The data sent to the server.

[0785] Specific operation: The device records the baby's crying and the parent's reaction, and sends the video and audio data to a server via the network.

[0786] Step 3:

[0787] sentiment analysis

[0788] The server analyzes the received parent data to identify the parent's emotional patterns, and analyzes the baby data to estimate the baby's state from the tone of the crying and behavior.

[0789] Input: Video and audio data sent to the server.

[0790] Output: Parent and baby emotion pattern data.

[0791] Specific operation: The server analyzes the baby's cry and determines that the baby is "lonely," and at the same time identifies patterns from the parent's facial expressions and tone of voice that indicate "the parent has the intention to comfort the baby."

[0792] Step 4:

[0793] Emotion Engine Recognition

[0794] The server recognizes the parent's real-time emotions using an emotion engine and reflects that data in the overall system behavior, which takes the parent's current emotional state into account when proceeding to the next step.

[0795] Input: Parent video and audio data.

[0796] Output: Parent emotion data from the emotion engine.

[0797] How it works: The emotion engine analyzes the parent's facial expressions and tone of voice to determine that the parent is stressed. This information is used as input for the generative AI.

[0798] Step 5:

[0799] Generative AI video creation

[0800] The server uses generative AI to generate parental reaction videos and audio based on the emotion analysis results and emotion engine data, reflecting the user's emotions and recreating realistic facial expressions and voices.

[0801] Input: Sentiment analysis data and sentiment engine data.

[0802] Output: Generated reaction video and audio data.

[0803] Specific behavior: If it is determined that the baby "requires a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[0804] Step 6:

[0805] Video playback

[0806] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[0807] Input: Generated reaction video and audio data.

[0808] Output: Video and audio played to the baby.

[0809] What it does: The device calms a crying baby by playing a video of a parent comforting the baby with a soft smile.

[0810] (Application example 2)

[0811] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0812] In the workplace, reducing worker stress and fatigue and creating a safe and efficient work environment are important challenges. However, conventional methods make it difficult to grasp workers' emotions and physical conditions in real time and take appropriate measures. As a result, workers' health risks increase and work efficiency may decline. For this reason, there is a demand for a system that can monitor workers' emotions and physical conditions in real time and take appropriate measures.

[0813] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0814] In this invention, the server includes a means for monitoring the state of workers, a means for recording the words, actions, voices, and facial expressions of workers, a means for analyzing the monitored state data of workers and the recorded data of workers, a means for generating appropriate reaction video and audio based on the analysis, and a means for playing the generated video and audio to the workers. This makes it possible to monitor the emotions and physical condition of workers in real time and take appropriate measures.

[0815] "Means for monitoring worker status" refers to devices or systems for monitoring the physical status of workers, such as their complexion, posture, and movements, in real time.

[0816] "Means for recording the words, actions, voice, and facial expressions of workers" refers to devices such as cameras and microphones that record the content of workers' words, tone of voice, facial expressions, etc.

[0817] "Means for analyzing monitored worker condition data and recorded worker data" refers to software and hardware for analyzing collected worker physical and emotional data.

[0818] "Means for generating appropriate response videos and audio based on analysis" refers to artificial intelligence (AI) technology for automatically generating appropriate response messages and videos according to the emotional state of workers.

[0819] "Means for playing the generated video and audio to workers" refers to devices such as displays and speakers used to communicate the generated response messages and videos to workers.

[0820] The system that realizes this application example, a worker emotion recognition support robot, consists of the following main components:

[0821] System configuration and operation

[0822] 1. Means of monitoring the condition of workers

[0823] Cameras suitable for the work environment (e.g., general surveillance cameras or high-resolution cameras) are used to monitor the status of workers in real time, and highly sensitive microphones are installed to collect the voices of workers.

[0824] 2. Means of recording the words, actions, voices, and facial expressions of workers

[0825] The robots will be fitted with devices that record the workers' facial expressions and tone of voice, allowing for detailed recording of their behavior and emotions.

[0826] 3. Means of analyzing monitored worker status data and recorded worker data

[0827] The server will be installed with software that uses Google TensorFlow to analyze worker status data, analyzing facial expressions and voice to identify the worker's level of stress and fatigue.

[0828] 4. A method for generating appropriate reaction video and audio based on the analysis

[0829] The server uses OpenAI's GPT-4 to generate appropriate reaction videos and audio messages tailored to the worker's state. This generative AI model uses past data to create more natural and effective responses.

[0830] 5. A means of playing the generated video and audio to workers

[0831] The generated video and audio are then played back to the worker through the robot's on-board display and speakers, allowing the worker to receive encouragement and break suggestions at appropriate times.

[0832] Examples of the system and how to use it

[0833] How to use this system will be explained with concrete examples.

[0834] 1. Data Collection

[0835] The device is equipped with a camera and microphone that records the worker's facial expressions and voice in real time and sends the data to a server. For example, if a worker says "I'm tired" while working, the microphone will collect the voice.

[0836] 2. Sentiment analysis

[0837] The data is then analyzed on the server, and TensorFlow is used to identify the worker's fatigue and stress levels for the day based on subtle changes in facial expressions and tone of voice.

[0838] 3. Emotion Engine

[0839] The server uses Amazon Comprehend to recognize the worker's real-time emotions and incorporates that information into the next steps, allowing for more accurate responses.

[0840] 4. Generation AI

[0841] Using GPT-4, the system generates voice messages and videos based on the results of sentiment analysis, such as encouraging workers or suggesting them to take a break. For example, it generates a message saying, "It's a good idea to take a short break."

[0842] 5. Video playback

[0843] The generated messages and videos are played through the robot's displays and speakers, and if a worker shows signs of fatigue, a video message such as "Take a deep breath to relax" will be displayed at that point.

[0844] Examples of prompt statements

[0845] An example of an input prompt for the generative AI model is as follows:

[0846] If a worker looks tired, generate a message saying, "Good work. Maybe you should take a break."

[0847] If a worker is feeling stressed, generate a voice message saying, "Take a deep breath to relax. It's okay."

[0848] This system allows workers' emotions and physical condition to be monitored in real time, and appropriate measures to be taken to reduce health risks and improve work efficiency.

[0849] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0850] Step 1:

[0851] The camera and microphone installed on the terminal collect the facial expressions and voices of the workers. The system records what the workers say and their facial expressions while working and sends this data to a server in real time. The input is the workers' audio and video data, and the output is the transmission of this data to the server.

[0852] Step 2:

[0853] The server analyzes the received facial and voice data. Using Google TensorFlow, it analyzes the worker's emotional state (e.g., fatigue, stress, concentration, etc.) from their facial expressions. It also analyzes the tone and content of their voice from the audio data to determine their emotional state. The input is the worker's video and audio data, and the output is the worker's emotional state data as an analysis result.

[0854] Step 3:

[0855] The server uses an emotion engine to perform a more detailed analysis of the worker's real-time emotions and incorporates the results into the overall data. It uses Amazon Comprehend to identify the emotional state and output intermediate data for application to generative AI. The input is facial expression and voice analysis results, and the output is detailed emotional data for application to the generative AI model.

[0856] Step 4:

[0857] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate response messages and videos based on the worker's emotional state. The generative AI model creates messages and videos that are optimal for the worker's situation based on the prompt text. The input is detailed emotional data and the prompt text, and the output is the generated response message and video data.

[0858] Step 5:

[0859] The generated messages and videos are sent from the server to the terminal and played to the worker through the terminal's display and speaker. The user watches the presented message or video and takes appropriate action (e.g., take a break or take a deep breath). The input is the generated response message and video data, and the output is to provide them to the worker.

[0860] This series of steps will create a system that supports worker health and work efficiency in real time.

[0861] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0862] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0863] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0864] [Third embodiment]

[0865] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0866] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0867] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0868] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0869] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0870] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0871] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0872] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0873] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0874] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0875] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0876] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0877] This invention is a system that soothes a baby in place of the parent, mainly monitoring the baby's condition, collecting and analyzing the parent's words, actions, voice, and facial expressions, and based on the analysis results, generating and playing appropriate video and audio responses of the parent. This system is designed to reassure the baby when the parent is not around, even when the baby feels lonely or anxious.

[0878] System configuration and operation

[0879] The system consists of the following major components:

[0880] 1. Data Collection Module

[0881] 2. Sentiment Analysis Module

[0882] 3. Generative AI Module

[0883] 4. Video Playback Module

[0884] 1. Data Collection Module

[0885] To monitor the baby's condition, the device (e.g., a smartphone or tablet) is equipped with a camera and microphone. It also contains sensors and a recording device to record the parent's words, actions, voice, and facial expressions. As the parent interacts with the baby, the device collects this data in real time and sends it to a server.

[0886] Examples:

[0887] The device records the parent smiling and speaking to the child in a gentle voice, and then uploads the video and audio data to a server.

[0888] 2. Sentiment Analysis Module

[0889] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs from the tone of the crying and body language.

[0890] Examples:

[0891] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[0892] 3. Generative AI Module

[0893] Based on the results of the emotion analysis, the server uses generative AI to generate video and audio of the parent's reaction. At this stage, technology is used to realistically reproduce the parent's facial expressions and voice based on past data.

[0894] Examples:

[0895] If the baby "requests a diaper change," the generative AI will create a video of the parent taking the baby to the changing table.

[0896] 4. Video Playback Module

[0897] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[0898] Examples:

[0899] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[0900] Natural language explanation of the system's program

[0901] 1. Data Collection

[0902] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[0903] 2. Sentiment analysis

[0904] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[0905] 3. Generation AI

[0906] Server: Based on sentiment analysis data, it identifies the most appropriate parental response and uses generative AI to generate video and audio that replicates that facial expression and voice.

[0907] 4. Video playback

[0908] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, giving the baby a sense of security.

[0909] Through the above processing, the present invention provides a system that can effectively soothe a baby even when the parent is away or busy.

[0910] The processing flow will be explained below.

[0911] Step 1: Data collection

[0912] Device: Records interactions between the baby and parent using a camera and microphone. Data such as the parent's facial expressions, tone of voice, what they say, and gestures are collected and sent to the server in real time.

[0913] Example: A scene in which a parent speaks to a baby in a gentle voice is recorded using a camera and microphone, and the data is sent to a server.

[0914] Step 2: Monitor your baby

[0915] Device: Monitors the baby's crying and behavior in real time, collecting data using a camera and microphone. When the baby starts crying, the data is sent to the server.

[0916] Example: When a baby starts crying, the crying is recorded and a video is also taken and sent to a server.

[0917] Step 3: Parental sentiment analysis

[0918] Server: Analyzes the parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." Text mining is performed on the tone of voice and the content of the speech, and emotion labels are assigned.

[0919] Example: The server analyzes a parent's "smile" and "gentle voice" and assigns them the emotion label "gentle."

[0920] Step 4: Analyze the baby's condition

[0921] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[0922] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[0923] Step 5: Parent Reaction Generation

[0924] Server: Combines the analyzed parent's emotional patterns with the baby's state to predict the most appropriate parental response. Using generative AI, the predicted parental response is generated as video and audio.

[0925] Example: If a baby is judged to be "lonely," the AI ​​will generate a video of the parent speaking to the baby in a gentle voice.

[0926] Step 6: Select video

[0927] Server: Select the most appropriate parental reaction video from the generated videos.

[0928] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice.

[0929] Step 7: Send your video

[0930] Server: Sends the selected parent's video to the device.

[0931] Example: Send a video of someone speaking to you in a gentle voice to your device.

[0932] Step 8: Play the video to your baby

[0933] Device: The parent's video and audio are played in real time on the tablet screen and through the speaker.

[0934] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[0935] This series of steps allows parents to provide a quick and appropriate response to their baby's emotions and needs, making it possible for them to feel safe even when parents are away or busy.

[0936] Example 1

[0937] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0938] In modern homes, there is a need to reassure babies even when parents are away or busy. However, there is no established system that can analyze parents' words, actions, and emotions, as well as babies' cries and behavior, in real time and provide appropriate responses based on that analysis. Conventional methods have difficulty reproducing the parent's presence and reactions in real time, making it difficult to respond immediately when a baby feels lonely or anxious. New technology is needed to solve these problems.

[0939] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0940] In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating a video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio to the baby, means for analyzing the tone and behavior of the baby's crying, means for identifying the parent's emotional pattern, means for generating a parent's reaction using a generative AI model, and means for selecting the parent's reaction video most appropriate for the baby's condition from among the multiple generated videos and playing it to the baby in a timely manner. This makes it possible to instantly provide an appropriate response according to the baby's condition even when the parent is not present, effectively reassuring the baby.

[0941] The "means for monitoring the baby's condition" is a device that records the baby's crying, behavior, and facial expressions in real time and collects them as data.

[0942] "Means for recording parental behavior, voice, and facial expressions" refers to a device that uses a camera or microphone to record a parent's speech, facial expressions, gestures, etc. as digital data.

[0943] "Monitored baby condition data" refers to data collected in real time, including information on the tone of a baby's crying, behavioral patterns, and facial expressions.

[0944] "Recorded parent data" refers to digital data that records the parent's speech, facial expressions, behavior, etc.

[0945] The "means for analyzing" is a software and hardware system for analyzing the collected data and determining the emotions and states of the baby and parent.

[0946] The "means for generating parental reaction video and audio" is a system including a generative AI model for generating parental video and audio based on the results of emotion analysis.

[0947] The "means for playing the generated video and audio to the baby" is a terminal for showing and playing the generated video and audio of the parent's reaction to the baby.

[0948] The "means for analyzing the tone and behavior of a baby's crying" is a software and hardware system for analyzing the audio data of a baby's crying and behavioral patterns and estimating its condition.

[0949] The "means for identifying parental emotional patterns" is a system for identifying a parent's emotional state from their facial expressions, tone of voice, etc., and identifying it as a pattern.

[0950] A "generative AI model" is an algorithm that uses artificial intelligence to learn from past data and generate parental reaction videos and audio.

[0951] The "means for playing back to the baby in a timely manner" refers to a terminal and control system for playing back the parent's reaction video and audio, which are generated at the optimal timing depending on the baby's condition.

[0952] MODE FOR CARRYING OUT THE INVENTION

[0953] This system monitors the baby's condition, collects and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate parental reaction video and audio based on the analysis results. This system is designed to reassure the baby when the parent is not around, even when the baby is feeling lonely or anxious.

[0954] The system consists of the following main components:

[0955] 1. Data Collection Module

[0956] 2. Sentiment Analysis Module

[0957] 3. Generative AI Module

[0958] 4. Video Playback Module

[0959] Data Collection Module

[0960] Device: A mobile device such as a smartphone or tablet is used. Using a camera and microphone, the device collects the baby's crying and behavior, as well as the parent's facial expressions and voice, in real time and sends the data to a server. For example, the device can record a parent smiling or speaking to the baby in a gentle voice, and upload the video and audio data to the server.

[0961] Sentiment Analysis Module

[0962] Server: Receives and analyzes the baby and parent data sent from the device. It identifies emotional patterns from the parent's facial expressions and tone of voice, and infers the baby's state by analyzing the tone of the baby's cry and behavioral patterns. For example, if it analyzes a baby's cry and determines that the baby is feeling lonely, it will identify a pattern from previously collected parent data that says, "Speak to the baby in a gentle voice when the baby feels lonely."

[0963] Generative AI Module

[0964] Server: Based on the results of the sentiment analysis, a generative AI model is used to generate a video and audio of the parent's reaction. The generative AI model uses past data to realistically reproduce the parent's facial expressions and voice. For example, if it determines that the baby is "requesting a diaper change," the generative AI model generates a video of the parent taking the baby to the changing table.

[0965] Video Playback Module

[0966] Device: Receives the parent's video and audio responses sent from the server and plays them in a timely manner to reassure the baby. For example, if a baby is "crying in fear," the device will play a video of the parent reassuring the baby with a gentle smile.

[0967] Examples of prompt statements

[0968] Prompt statements are used to instruct the system to perform a specific operation. For example:

[0969] "Analyze a baby's cry and generate a video of the parent's reaction if they are feeling lonely."

[0970] "Generate a video of a parent's reaction when it is determined that the child is requesting a diaper change."

[0971] "Generate a parent's reaction video to comfort a scared, crying baby."

[0972] In this way, the present invention provides a system that can effectively soothe babies even when parents are away or busy. The data collection, emotion analysis, generative AI, and video playback modules work together to quickly provide appropriate responses according to the baby's condition.

[0973] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0974] Step 1: Data collection

[0975] Device: Uses a camera and microphone to record the baby's cries, movements, and the parent's words, voice, and facial expressions in real time.

[0976] Input: Baby's crying, behavior, parent's words, voice, facial expressions

[0977] Output: Recorded video data, audio data

[0978] Specific behavior:

[0979] The device activates the camera and captures video of the baby's behavior and the parent's facial expressions.

[0980] The device activates a microphone and collects the baby's cries and the parent's voice as audio data.

[0981] The collected data is compressed and sent to the server in real time.

[0982] Step 2: Sentiment analysis

[0983] Server: Receives and analyzes the baby and parent data sent from the device.

[0984] Input: Video and audio data sent from the device

[0985] Output: Baby's emotional state data, parent's emotional pattern data

[0986] Specific behavior:

[0987] The server analyzes the video data and identifies the baby's facial expressions and movements.

[0988] The server analyzes the audio data and identifies the tone of the baby's cry and the tone of the parent's voice.

[0989] Identify the baby's emotional state (e.g., lonely, scared) and the parent's emotional patterns (e.g., smiling, soft voice).

[0990] Step 3: Prompt generation

[0991] Server: Based on the results of sentiment analysis, generate prompt sentences to be input to the generative AI.

[0992] Input: Baby's emotional state data, parent's emotional pattern data

[0993] Output: Prompt text to be input to the generation AI

[0994] Specific behavior:

[0995] The server generates a corresponding prompt sentence based on the analysis results (e.g., "Generate a parent's response to a lonely baby").

[0996] Step 4: Parent Reaction Generation

[0997] Server: Uses a generative AI model to generate video and audio parental responses based on the prompt.

[0998] Input: Prompt sentence, past parent response data

[0999] Output: Parents' reaction video data, audio data

[1000] Specific behavior:

[1001] A prompt sentence is input into the generative AI model, and a video and audio of the parent's reaction is generated.

[1002] The generated video and audio are encoded and prepared for transmission to the device.

[1003] Step 5: Play the video

[1004] Device: Receives the parent's reaction video and audio sent from the server and plays them according to the baby's condition.

[1005] Input: Parent reaction video data, audio data

[1006] Output: Video and audio played to the baby

[1007] Specific behavior:

[1008] The terminal receives the data from the server.

[1009] Video and audio are played at the optimal time for your baby's current condition.

[1010] The baby's reactions during and after playback are recorded again and sent as feedback to the server.

[1011] In this way, the system can effectively guide the baby through the entire process from data collection to video playback.

[1012] (Application example 1)

[1013] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1014] In food delivery services, when delivery personnel are busy or unable to communicate directly with customers during deliveries, they tend to inadequately explain the delivery status or problems to customers. This results in customer anxiety and dissatisfaction, leading to a decline in the quality of service. It is also necessary to improve the situation where delivery personnel are unable to respond appropriately and in a timely manner when they encounter unexpected problems.

[1015] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1016] In this invention, the server includes a means for monitoring the status of the delivery person, a means for recording the speech, behavior, voice, and facial expression of the delivery person, a means for analyzing the monitored status data of the delivery person and the recorded data of the delivery person, a means for generating a video and audio response of the delivery person, and a means for playing the generated video and audio to the customer. This makes it possible to provide an appropriate response to the customer even when the delivery person is not present, thereby reducing the customer's anxiety and dissatisfaction.

[1017] "Delivery Person" means an individual or entity responsible for delivering goods to a customer.

[1018] "Means for monitoring the status" refers to the function of checking the target's behavior and voice in real time using devices such as cameras and microphones.

[1019] "Means for recording speech, voice, and facial expressions" refers to devices and software that store the delivery person's speech, voice, and facial expressions as digital data.

[1020] "Means of analysis" refers to the algorithms and programs used to analyze collected data and convert it into meaningful information.

[1021] "Means for generating responsive video and audio" refers to technology that creates visual and audio data to appropriately reproduce the delivery person's condition and emotions based on the analysis results.

[1022] "Means for displaying video and audio to customers" means any device or software used to display or play the generated visual and audio data to customers.

[1023] "Means for identifying response patterns" refers to algorithms or methods for detecting and identifying patterns of delivery personnel's behavior and speech in multiple situations.

[1024] "Means for timely playback" refers to technology or equipment that allows appropriate reactions and responses to be played back in real time at the required timing.

[1025] The system for implementing this invention aims to alleviate customer anxiety and dissatisfaction by monitoring the status of delivery personnel and providing appropriate responses to customers. A specific implementation method for this system is described below.

[1026] 1. System Configuration

[1027] This system consists of the following main modules:

[1028] a. Data Collection Module

[1029] The device (e.g., smartphone or tablet) is equipped with a camera and microphone, which records the delivery person's words, actions, voice, and facial expressions. This data is collected in real time and sent to a server. For example, the device records a scene in which the delivery person has a "confused expression" or "explains something to a customer over the phone," and uploads the video and audio data to the server.

[1030] b. Sentiment Analysis Module

[1031] The server analyzes the facial expressions, voice, and behavioral data of the delivery person sent from the device to identify the delivery person's emotional patterns. At the same time, it also analyzes the behavioral data of the delivery person sent from the device to infer the reason for the delivery person's actions. For example, it can determine that the delivery person is delayed due to "traffic congestion" based on a confused expression on their face.

[1032] c. Generative AI module

[1033] The server uses generative AI to generate a video and audio response from the delivery person based on the results of the emotion analysis. At this stage, technology is used to realistically reproduce the delivery person's facial expressions and voice based on past data. For example, a video can be generated in which the delivery person explains to the customer that the delivery will be five minutes late.

[1034] d. Video playback module

[1035] The generated video and audio data of the delivery person is sent from the server to the terminal. The terminal selects and plays the delivery person's response video that is most appropriate for the customer. This allows the customer to understand the delivery status and feel reassured. For example, if the customer is confused, a video of the delivery person explaining in a calm tone will be played.

[1036] 2. Hardware and software used

[1037] Terminal (smartphone, tablet): Records the delivery person's status using a camera and microphone and sends it to the server.

[1038] Server: Analyzes the collected data using deep learning models (e.g., face recognition models using Keras). For generative AI, frameworks such as TensorFlow are used.

[1039] Generative AI: Technology that uses facial recognition and voice data to generate realistic facial expressions and voices of delivery personnel.

[1040] Playback device: Software installed on the device that plays the generated video and audio to the customer.

[1041] 3. Examples of prompt sentences

[1042] For example, the prompt for a scene in which a delivery person is stuck in traffic and looks confused is as follows:

[1043] "The delivery person is stuck in traffic and has a confused expression. Based on that, generate a sentence that will convey an appropriate message to the customer."

[1044] With the above components and processes, the present invention makes it possible to provide appropriate responses to customers even when a delivery person is not present, thereby reducing customer anxiety and dissatisfaction.

[1045] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1046] Step 1:

[1047] The terminal collects the delivery person's status (video and audio).

[1048] Specifically, the device's camera and microphone are used to record the delivery person's movements, voice, and facial expressions in real time, and this data is recorded as video and audio files that include the delivery person's speech and facial expressions.

[1049] Input: Real-time video and audio of the delivery person

[1050] Output: Video file (e.g. mp4 format), audio file (e.g. wav format)

[1051] Step 2:

[1052] The data collected by the device is sent to the server in real time.

[1053] The collected video and audio files are uploaded to a server using an appropriate communication protocol (e.g., HTTP).

[1054] Input: Video files, audio files

[1055] Output: Data stored on the server

[1056] Step 3:

[1057] The server analyzes the received data and identifies the emotional patterns of the delivery person.

[1058] Specifically, an emotion analysis model using Keras is used to analyze the facial expressions and tone of the delivery person's voice from the received video and audio data, and this analysis identifies the delivery person's emotions, such as "confused" or "happy."

[1059] Input: Video and audio files on the server

[1060] Output: Emotional pattern (e.g., confusion, joy)

[1061] Step 4:

[1062] Generative AI generates response video and audio based on emotional patterns.

[1063] Based on the results of the sentiment analysis, a generative AI model (e.g., a text-to-video generation model) is used to create a realistic video or audio response from the delivery person. For example, if the delivery person is confused, a video containing the message "Delivery will be 5 minutes late" is generated.

[1064] Input: Emotion pattern, prompt sentence

[1065] Output: Response video file, response audio file

[1066] Step 5:

[1067] The server generates a response video and audio and sends it to the terminal.

[1068] The generated response video and audio are then transmitted to the terminal again using an appropriate communication protocol.

[1069] Input: Response video file, response audio file

[1070] Output: Data stored on the device

[1071] Step 6:

[1072] The terminal plays the most appropriate response video and audio to the customer.

[1073] The device will then play the video and audio responses it receives to help customers understand the situation and feel at ease. For example, if a video explaining a delivery delay is played, customers will be able to understand the delivery person's situation and their anxiety about waiting will be alleviated.

[1074] Input: Response video file, response audio file

[1075] Output: Video and audio response played

[1076] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1077] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby in place of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[1078] System configuration and operation

[1079] The system consists of the following major components:

[1080] 1. Data Collection Module

[1081] 2. Sentiment Analysis Module

[1082] 3. Generative AI Module

[1083] 4. Video Playback Module

[1084] 5. Emotion Engine

[1085] 1. Data Collection Module

[1086] The devices (smartphones and tablets) are equipped with cameras and microphones to monitor the baby's condition and record the parent's behavior, voice, and facial expressions. Using these devices, the devices transmit the baby-parent interaction data to a server in real time.

[1087] Examples:

[1088] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[1089] 2. Sentiment Analysis Module

[1090] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs.

[1091] Examples:

[1092] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[1093] 3. Generative AI Module

[1094] The server uses AI to generate the parent's reaction video and audio based on the emotion analysis results. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[1095] Examples:

[1096] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1097] 4. Video Playback Module

[1098] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[1099] Examples:

[1100] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[1101] 5. Emotion Engine

[1102] The emotion engine recognizes the emotions of the parent user in real time and reflects that emotional data in the operation of the entire system. The emotion engine analyzes video and audio data to identify emotional states such as "the parent is feeling stressed." Based on these results, the generative AI then generates more appropriate and detailed videos of the parent's reactions.

[1103] Examples:

[1104] If the parent is feeling stressed, the generated video will adjust the parent's facial expression to appear slightly calmer.

[1105] Natural language explanation of the system's program

[1106] 1. Data Collection

[1107] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[1108] 2. Sentiment analysis

[1109] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[1110] 3. Emotion Engine

[1111] Server: The emotion engine recognizes the parent's real-time emotions and reflects the results in the analysis data, thereby understanding the parent's current emotional state.

[1112] 4. Generation AI

[1113] Server: Based on the sentiment analysis data and the results of the emotion engine, it identifies the most appropriate parental response and uses generative AI to generate video and audio that reproduces that facial expression and voice.

[1114] 5. Video playback

[1115] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1116] Through the above processing, the present invention can effectively soothe a baby even when the parent is away or busy, and by reflecting the emotions of the user (parent), it is possible to provide a more realistic and appropriate response.

[1117] The processing flow will be explained below.

[1118] Step 1: Data collection

[1119] Device: The camera and microphone record the parent's facial expressions, tone of voice, speech, gestures, etc. This data is sent to the server in real time.

[1120] Example: A parent speaking to their baby in a gentle voice is recorded using a camera and microphone, and the video and audio data is sent to a server.

[1121] Step 2: Monitor your baby

[1122] Device: Monitors the baby's crying and behavior in real time. Cameras and microphones capture the crying and video and send it to the server.

[1123] Example: When a baby starts crying, the crying is recorded and the video is sent to a server.

[1124] Step 3: Parental sentiment analysis

[1125] Server: Analyzes the received parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." It also analyzes the tone of voice and the content of the speech to assign an emotional label.

[1126] Example: The server analyzes data on a parent's "smile" and "gentle voice" and assigns the emotion label "gentle" to them.

[1127] Step 4: Analyze the baby's condition

[1128] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[1129] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[1130] Step 5: Parental emotion recognition by emotion engine

[1131] Server: Recognizes the parent's real-time emotions using an emotion engine. Analyzes the parent's video and audio data to detect their current emotional state (e.g., stress, relief).

[1132] Example: The emotion engine analyzes the video and audio data of a parent and recognizes that the parent is feeling stressed.

[1133] Step 6: Parent Reaction Generation

[1134] Server: Based on the results of the parent's emotion analysis and the emotion engine, the server predicts the most appropriate parental response. Using generative AI, it generates video and audio that realistically reproduces the parent's facial expressions and voice.

[1135] Example: If a baby is judged to be "lonely" and the parent is recognized as "stressed," the AI ​​will generate a video of the parent speaking to the baby in a calm, gentle voice.

[1136] Step 7: Select and send your video

[1137] Server: Selects the most appropriate parental reaction video from the generated videos and sends it to the device.

[1138] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice and send it to the device.

[1139] Step 8: Play the video

[1140] Device: The parent's video and audio are sent from the server and played in real time on the tablet or smartphone screen, with the audio played through the speaker.

[1141] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[1142] This series of steps enables the system to respond quickly and appropriately to the baby's emotions and requests, and by reflecting the emotions of the user (parent), it can provide a more realistic and effective way to soothe a baby.

[1143] Example 2

[1144] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1145] In today's busy home environment, it is often difficult for parents to spend all their time with their babies. In particular, if a baby starts crying or feels anxious, parents' inability to respond quickly can have an impact on the baby's emotions. Another issue is that parents themselves may find it difficult to respond optimally if they are stressed.

[1146] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio for the baby, and means for recognizing the parent's real-time emotions and reflecting the data in the operation of the entire system. This allows the baby to be soothed effectively even when the parent is away or busy, and by reflecting the parent's emotions, a more realistic and appropriate response is possible.

[1147] A "means for monitoring the baby's condition" is a device or system that uses a camera or microphone to record the baby's movements and cries and analyzes the data.

[1148] "Means for recording parental behavior, voice, and facial expressions" refers to a device or system that uses a camera or microphone to record the words spoken by a parent, their tone of voice, and their facial expressions.

[1149] "Monitored baby status data" refers to data on the baby's movements and crying collected by monitoring means such as cameras and microphones.

[1150] "Recorded parent data" refers to data such as parental words, tone of voice, and facial expressions collected through means that record parental behavior, voice, and facial expressions.

[1151] "Means of analysis" refers to software and algorithms that analyze the collected data and identify the emotions and states of the baby and parent.

[1152] The "means for generating parental reaction video and audio" is a device or system that generates video clips and audio to reproduce the parent's actions and voice based on the analysis results.

[1153] The "means for playing the generated video and audio to the baby" is a device or system for playing the generated parent's reaction video and audio to the baby in a timely manner.

[1154] "Means for recognizing the parent's real-time emotions and reflecting that data in the operation of the entire system" refers to a device or system that analyzes the parent's current emotional state using an emotion engine and takes the results into account and reflects them in the operation of the system.

[1155] An "emotion engine" is software or an algorithm that analyzes video and audio data to identify the emotional state of parents and babies and reflect that information in the system's operation.

[1156] "Generative AI" refers to the artificial intelligence models and techniques used to generate parental reaction videos and audio.

[1157] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby on behalf of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[1158] The system consists of the following major components:

[1159] 1. Data Collection Module

[1160] 2. Data transmission module

[1161] 3. Sentiment Analysis Module

[1162] 4. Emotion Engine Module

[1163] 5. Generative AI Module

[1164] 6. Video Playback Module

[1165] 1. Data Collection Module

[1166] The device uses a camera and microphone to record the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and the parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[1167] Examples:

[1168] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[1169] 2. Data transmission module

[1170] The device transmits the collected data to the server in real time, which reflects the real-time status of the baby and parent and is analyzed in the next processing step.

[1171] Examples:

[1172] The device records the parent's facial expressions and voice data and sends it to the server.

[1173] 3. Sentiment Analysis Module

[1174] The server analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," and "surprised." It also analyzes the baby's data and estimates the baby's state (e.g., loneliness, fear, or request for a diaper change) from the tone of the crying and behavior.

[1175] Examples:

[1176] The server analyzes the tone of the baby's cry to determine whether the baby is feeling lonely, and then, based on past data collected from parents, identifies a pattern of parents speaking to their babies in a gentle voice when they feel lonely.

[1177] 4. Emotion Engine Module

[1178] The server uses an emotion engine to recognize the parent's real-time emotions and reflects the results in the analysis data, thereby making it possible to understand the parent's current emotional state.

[1179] Examples:

[1180] The emotion engine analyzes the parent's facial expressions and tone of voice and determines that the parent is stressed. This information is reflected in subsequent processing.

[1181] 5. Generative AI Module

[1182] The server uses generative AI to generate the parent's reaction video and audio based on the emotion analysis data and the results of the emotion engine. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[1183] Examples:

[1184] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1185] 6. Video Playback Module

[1186] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1187] Examples:

[1188] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[1189] Example prompts to be input to the generative AI model

[1190] "If your baby is feeling lonely, generate a video of the parent talking to them in a gentle voice."

[1191] "If a baby is scared and crying, generate a video of the parent reassuring the baby with a soft smile."

[1192] "If a baby requests a diaper change, generate a video of the parent taking the baby to the changing table."

[1193] This system allows babies to be soothed effectively even when parents are away or busy, and by reflecting the parents' emotions, it allows for more realistic and appropriate responses.

[1194] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1195] Step 1:

[1196] Data collection

[1197] The device uses a camera and microphone to record the baby's state and the parent's words, actions, voice, and facial expressions, thereby collecting scenes of interaction between the baby and the parent.

[1198] Input: Baby's cries and movements, parent's facial expressions and words.

[1199] Output: Collected video and audio data.

[1200] Specific operation: The device captures real-time footage of the baby crying and the parent's gentle response, and records this data.

[1201] Step 2:

[1202] Data transmission

[1203] The device transmits the collected data in real time to a server, which prepares the data for the next analysis step.

[1204] Input: Collected video and audio data.

[1205] Output: The data sent to the server.

[1206] Specific operation: The device records the baby's crying and the parent's reaction, and sends the video and audio data to a server via the network.

[1207] Step 3:

[1208] sentiment analysis

[1209] The server analyzes the received parent data to identify the parent's emotional patterns, and analyzes the baby data to estimate the baby's state from the tone of the crying and behavior.

[1210] Input: Video and audio data sent to the server.

[1211] Output: Parent and baby emotion pattern data.

[1212] Specific operation: The server analyzes the baby's cry and determines that the baby is "lonely," and at the same time identifies patterns from the parent's facial expressions and tone of voice that indicate "the parent has the intention to comfort the baby."

[1213] Step 4:

[1214] Emotion Engine Recognition

[1215] The server recognizes the parent's real-time emotions using an emotion engine and reflects that data in the overall system behavior, which takes the parent's current emotional state into account when proceeding to the next step.

[1216] Input: Parent video and audio data.

[1217] Output: Parent emotion data from the emotion engine.

[1218] How it works: The emotion engine analyzes the parent's facial expressions and tone of voice to determine that the parent is stressed. This information is used as input for the generative AI.

[1219] Step 5:

[1220] Generative AI video creation

[1221] The server uses generative AI to generate parental reaction videos and audio based on the emotion analysis results and emotion engine data, reflecting the user's emotions and recreating realistic facial expressions and voices.

[1222] Input: Sentiment analysis data and sentiment engine data.

[1223] Output: Generated reaction video and audio data.

[1224] Specific behavior: If it is determined that the baby "requires a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1225] Step 6:

[1226] Video playback

[1227] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1228] Input: Generated reaction video and audio data.

[1229] Output: Video and audio played to the baby.

[1230] What it does: The device calms a crying baby by playing a video of a parent comforting the baby with a soft smile.

[1231] (Application example 2)

[1232] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1233] In the workplace, reducing worker stress and fatigue and creating a safe and efficient work environment are important challenges. However, conventional methods make it difficult to grasp workers' emotions and physical conditions in real time and take appropriate measures. As a result, workers' health risks increase and work efficiency may decline. For this reason, there is a demand for a system that can monitor workers' emotions and physical conditions in real time and take appropriate measures.

[1234] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1235] In this invention, the server includes a means for monitoring the state of workers, a means for recording the words, actions, voices, and facial expressions of workers, a means for analyzing the monitored state data of workers and the recorded data of workers, a means for generating appropriate reaction video and audio based on the analysis, and a means for playing the generated video and audio to the workers. This makes it possible to monitor the emotions and physical condition of workers in real time and take appropriate measures.

[1236] "Means for monitoring worker status" refers to devices or systems for monitoring the physical status of workers, such as their complexion, posture, and movements, in real time.

[1237] "Means for recording the words, actions, voice, and facial expressions of workers" refers to devices such as cameras and microphones that record the content of workers' words, tone of voice, facial expressions, etc.

[1238] "Means for analyzing monitored worker condition data and recorded worker data" refers to software and hardware for analyzing collected worker physical and emotional data.

[1239] "Means for generating appropriate response videos and audio based on analysis" refers to artificial intelligence (AI) technology for automatically generating appropriate response messages and videos according to the emotional state of workers.

[1240] "Means for playing the generated video and audio to workers" refers to devices such as displays and speakers used to communicate the generated response messages and videos to workers.

[1241] The system that realizes this application example, a worker emotion recognition support robot, consists of the following main components:

[1242] System configuration and operation

[1243] 1. Means of monitoring the condition of workers

[1244] Cameras suitable for the work environment (e.g., general surveillance cameras or high-resolution cameras) are used to monitor the status of workers in real time, and highly sensitive microphones are installed to collect the voices of workers.

[1245] 2. Means of recording the words, actions, voices, and facial expressions of workers

[1246] The robots will be fitted with devices that record the workers' facial expressions and tone of voice, allowing for detailed recording of their behavior and emotions.

[1247] 3. Means of analyzing monitored worker status data and recorded worker data

[1248] The server will be installed with software that uses Google TensorFlow to analyze worker status data, analyzing facial expressions and voice to identify the worker's level of stress and fatigue.

[1249] 4. A method for generating appropriate reaction video and audio based on the analysis

[1250] The server uses OpenAI's GPT-4 to generate appropriate reaction videos and audio messages tailored to the worker's state. This generative AI model uses past data to create more natural and effective responses.

[1251] 5. A means of playing the generated video and audio to workers

[1252] The generated video and audio are then played back to the worker through the robot's on-board display and speakers, allowing the worker to receive encouragement and break suggestions at appropriate times.

[1253] Examples of the system and how to use it

[1254] How to use this system will be explained with concrete examples.

[1255] 1. Data Collection

[1256] The device is equipped with a camera and microphone that records the worker's facial expressions and voice in real time and sends the data to a server. For example, if a worker says "I'm tired" while working, the microphone will collect the voice.

[1257] 2. Sentiment analysis

[1258] The data is then analyzed on the server, and TensorFlow is used to identify the worker's fatigue and stress levels for the day based on subtle changes in facial expressions and tone of voice.

[1259] 3. Emotion Engine

[1260] The server uses Amazon Comprehend to recognize the worker's real-time emotions and incorporates that information into the next steps, allowing for more accurate responses.

[1261] 4. Generation AI

[1262] Using GPT-4, the system generates voice messages and videos based on the results of sentiment analysis, such as encouraging workers or suggesting them to take a break. For example, it generates a message saying, "It's a good idea to take a short break."

[1263] 5. Video playback

[1264] The generated messages and videos are played through the robot's displays and speakers, and if a worker shows signs of fatigue, a video message such as "Take a deep breath to relax" will be displayed at that point.

[1265] Examples of prompt statements

[1266] An example of an input prompt for the generative AI model is as follows:

[1267] If a worker looks tired, generate a message saying, "Good work. Maybe you should take a break."

[1268] If a worker is feeling stressed, generate a voice message saying, "Take a deep breath to relax. It's okay."

[1269] This system allows workers' emotions and physical condition to be monitored in real time, and appropriate measures to be taken to reduce health risks and improve work efficiency.

[1270] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1271] Step 1:

[1272] The camera and microphone installed on the terminal collect the facial expressions and voices of the workers. The system records what the workers say and their facial expressions while working and sends this data to a server in real time. The input is the workers' audio and video data, and the output is the transmission of this data to the server.

[1273] Step 2:

[1274] The server analyzes the received facial and voice data. Using Google TensorFlow, it analyzes the worker's emotional state (e.g., fatigue, stress, concentration, etc.) from their facial expressions. It also analyzes the tone and content of their voice from the audio data to determine their emotional state. The input is the worker's video and audio data, and the output is the worker's emotional state data as an analysis result.

[1275] Step 3:

[1276] The server uses an emotion engine to perform a more detailed analysis of the worker's real-time emotions and incorporates the results into the overall data. It uses Amazon Comprehend to identify the emotional state and output intermediate data for application to generative AI. The input is facial expression and voice analysis results, and the output is detailed emotional data for application to the generative AI model.

[1277] Step 4:

[1278] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate response messages and videos based on the worker's emotional state. The generative AI model creates messages and videos that are optimal for the worker's situation based on the prompt text. The input is detailed emotional data and the prompt text, and the output is the generated response message and video data.

[1279] Step 5:

[1280] The generated messages and videos are sent from the server to the terminal and played to the worker through the terminal's display and speaker. The user watches the presented message or video and takes appropriate action (e.g., take a break or take a deep breath). The input is the generated response message and video data, and the output is to provide them to the worker.

[1281] This series of steps will create a system that supports worker health and work efficiency in real time.

[1282] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1283] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1284] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1285] [Fourth embodiment]

[1286] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1287] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1288] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1289] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1290] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1291] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1292] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1293] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1294] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1295] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1296] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1297] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1298] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1299] This invention is a system that soothes a baby in place of the parent, mainly monitoring the baby's condition, collecting and analyzing the parent's words, actions, voice, and facial expressions, and based on the analysis results, generating and playing appropriate video and audio responses of the parent. This system is designed to reassure the baby when the parent is not around, even when the baby feels lonely or anxious.

[1300] System configuration and operation

[1301] The system consists of the following major components:

[1302] 1. Data Collection Module

[1303] 2. Sentiment Analysis Module

[1304] 3. Generative AI Module

[1305] 4. Video Playback Module

[1306] 1. Data Collection Module

[1307] To monitor the baby's condition, the device (e.g., a smartphone or tablet) is equipped with a camera and microphone. It also contains sensors and a recording device to record the parent's words, actions, voice, and facial expressions. As the parent interacts with the baby, the device collects this data in real time and sends it to a server.

[1308] Examples:

[1309] The device records the parent smiling and speaking to the child in a gentle voice, and then uploads the video and audio data to a server.

[1310] 2. Sentiment Analysis Module

[1311] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs from the tone of the crying and body language.

[1312] Examples:

[1313] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[1314] 3. Generative AI Module

[1315] Based on the results of the emotion analysis, the server uses generative AI to generate video and audio of the parent's reaction. At this stage, technology is used to realistically reproduce the parent's facial expressions and voice based on past data.

[1316] Examples:

[1317] If the baby "requests a diaper change," the generative AI will create a video of the parent taking the baby to the changing table.

[1318] 4. Video Playback Module

[1319] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[1320] Examples:

[1321] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[1322] Natural language explanation of the system's program

[1323] 1. Data Collection

[1324] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[1325] 2. Sentiment analysis

[1326] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[1327] 3. Generation AI

[1328] Server: Based on sentiment analysis data, it identifies the most appropriate parental response and uses generative AI to generate video and audio that replicates that facial expression and voice.

[1329] 4. Video playback

[1330] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, giving the baby a sense of security.

[1331] Through the above processing, the present invention provides a system that can effectively soothe a baby even when the parent is away or busy.

[1332] The processing flow will be explained below.

[1333] Step 1: Data collection

[1334] Device: Records interactions between the baby and parent using a camera and microphone. Data such as the parent's facial expressions, tone of voice, what they say, and gestures are collected and sent to the server in real time.

[1335] Example: A scene in which a parent speaks to a baby in a gentle voice is recorded using a camera and microphone, and the data is sent to a server.

[1336] Step 2: Monitor your baby

[1337] Device: Monitors the baby's crying and behavior in real time, collecting data using a camera and microphone. When the baby starts crying, the data is sent to the server.

[1338] Example: When a baby starts crying, the crying is recorded and a video is also taken and sent to a server.

[1339] Step 3: Parental sentiment analysis

[1340] Server: Analyzes the parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." Text mining is performed on the tone of voice and the content of the speech, and emotion labels are assigned.

[1341] Example: The server analyzes a parent's "smile" and "gentle voice" and assigns them the emotion label "gentle."

[1342] Step 4: Analyze the baby's condition

[1343] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[1344] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[1345] Step 5: Parent Reaction Generation

[1346] Server: Combines the analyzed parent's emotional patterns with the baby's state to predict the most appropriate parental response. Using generative AI, the predicted parental response is generated as video and audio.

[1347] Example: If a baby is judged to be "lonely," the AI ​​will generate a video of the parent speaking to the baby in a gentle voice.

[1348] Step 6: Select video

[1349] Server: Select the most appropriate parental reaction video from the generated videos.

[1350] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice.

[1351] Step 7: Send your video

[1352] Server: Sends the selected parent's video to the device.

[1353] Example: Send a video of someone speaking to you in a gentle voice to your device.

[1354] Step 8: Play the video to your baby

[1355] Device: The parent's video and audio are played in real time on the tablet screen and through the speaker.

[1356] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[1357] This series of steps allows parents to provide a quick and appropriate response to their baby's emotions and needs, making it possible for them to feel safe even when parents are away or busy.

[1358] Example 1

[1359] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1360] In modern homes, there is a need to reassure babies even when parents are away or busy. However, there is no established system that can analyze parents' words, actions, and emotions, as well as babies' cries and behavior, in real time and provide appropriate responses based on that analysis. Conventional methods have difficulty reproducing the parent's presence and reactions in real time, making it difficult to respond immediately when a baby feels lonely or anxious. New technology is needed to solve these problems.

[1361] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1362] In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating a video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio to the baby, means for analyzing the tone and behavior of the baby's crying, means for identifying the parent's emotional pattern, means for generating a parent's reaction using a generative AI model, and means for selecting the parent's reaction video most appropriate for the baby's condition from among the multiple generated videos and playing it to the baby in a timely manner. This makes it possible to instantly provide an appropriate response according to the baby's condition even when the parent is not present, effectively reassuring the baby.

[1363] The "means for monitoring the baby's condition" is a device that records the baby's crying, behavior, and facial expressions in real time and collects them as data.

[1364] "Means for recording parental behavior, voice, and facial expressions" refers to a device that uses a camera or microphone to record a parent's speech, facial expressions, gestures, etc. as digital data.

[1365] "Monitored baby condition data" refers to data collected in real time, including information on the tone of a baby's crying, behavioral patterns, and facial expressions.

[1366] "Recorded parent data" refers to digital data that records the parent's speech, facial expressions, behavior, etc.

[1367] The "means for analyzing" is a software and hardware system for analyzing the collected data and determining the emotions and states of the baby and parent.

[1368] The "means for generating parental reaction video and audio" is a system including a generative AI model for generating parental video and audio based on the results of emotion analysis.

[1369] The "means for playing the generated video and audio to the baby" is a terminal for showing and playing the generated video and audio of the parent's reaction to the baby.

[1370] The "means for analyzing the tone and behavior of a baby's crying" is a software and hardware system for analyzing the audio data of a baby's crying and behavioral patterns and estimating its condition.

[1371] The "means for identifying parental emotional patterns" is a system for identifying a parent's emotional state from their facial expressions, tone of voice, etc., and identifying it as a pattern.

[1372] A "generative AI model" is an algorithm that uses artificial intelligence to learn from past data and generate parental reaction videos and audio.

[1373] The "means for playing back to the baby in a timely manner" refers to a terminal and control system for playing back the parent's reaction video and audio, which are generated at the optimal timing depending on the baby's condition.

[1374] MODE FOR CARRYING OUT THE INVENTION

[1375] This system monitors the baby's condition, collects and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate parental reaction video and audio based on the analysis results. This system is designed to reassure the baby when the parent is not around, even when the baby is feeling lonely or anxious.

[1376] The system consists of the following main components:

[1377] 1. Data Collection Module

[1378] 2. Sentiment Analysis Module

[1379] 3. Generative AI Module

[1380] 4. Video Playback Module

[1381] Data Collection Module

[1382] Device: A mobile device such as a smartphone or tablet is used. Using a camera and microphone, the device collects the baby's crying and behavior, as well as the parent's facial expressions and voice, in real time and sends the data to a server. For example, the device can record a parent smiling or speaking to the baby in a gentle voice, and upload the video and audio data to the server.

[1383] Sentiment Analysis Module

[1384] Server: Receives and analyzes the baby and parent data sent from the device. It identifies emotional patterns from the parent's facial expressions and tone of voice, and infers the baby's state by analyzing the tone of the baby's cry and behavioral patterns. For example, if it analyzes a baby's cry and determines that the baby is feeling lonely, it will identify a pattern from previously collected parent data that says, "Speak to the baby in a gentle voice when the baby feels lonely."

[1385] Generative AI Module

[1386] Server: Based on the results of the sentiment analysis, a generative AI model is used to generate a video and audio of the parent's reaction. The generative AI model uses past data to realistically reproduce the parent's facial expressions and voice. For example, if it determines that the baby is "requesting a diaper change," the generative AI model generates a video of the parent taking the baby to the changing table.

[1387] Video Playback Module

[1388] Device: Receives the parent's video and audio responses sent from the server and plays them in a timely manner to reassure the baby. For example, if a baby is "crying in fear," the device will play a video of the parent reassuring the baby with a gentle smile.

[1389] Examples of prompt statements

[1390] Prompt statements are used to instruct the system to perform a specific operation. For example:

[1391] "Analyze a baby's cry and generate a video of the parent's reaction if they are feeling lonely."

[1392] "Generate a video of a parent's reaction when it is determined that the child is requesting a diaper change."

[1393] "Generate a parent's reaction video to comfort a scared, crying baby."

[1394] In this way, the present invention provides a system that can effectively soothe babies even when parents are away or busy. The data collection, emotion analysis, generative AI, and video playback modules work together to quickly provide appropriate responses according to the baby's condition.

[1395] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1396] Step 1: Data collection

[1397] Device: Uses a camera and microphone to record the baby's cries, movements, and the parent's words, voice, and facial expressions in real time.

[1398] Input: Baby's crying, behavior, parent's words, voice, facial expressions

[1399] Output: Recorded video data, audio data

[1400] Specific behavior:

[1401] The device activates the camera and captures video of the baby's behavior and the parent's facial expressions.

[1402] The device activates a microphone and collects the baby's cries and the parent's voice as audio data.

[1403] The collected data is compressed and sent to the server in real time.

[1404] Step 2: Sentiment analysis

[1405] Server: Receives and analyzes the baby and parent data sent from the device.

[1406] Input: Video and audio data sent from the device

[1407] Output: Baby's emotional state data, parent's emotional pattern data

[1408] Specific behavior:

[1409] The server analyzes the video data and identifies the baby's facial expressions and movements.

[1410] The server analyzes the audio data and identifies the tone of the baby's cry and the tone of the parent's voice.

[1411] Identify the baby's emotional state (e.g., lonely, scared) and the parent's emotional patterns (e.g., smiling, soft voice).

[1412] Step 3: Prompt generation

[1413] Server: Based on the results of sentiment analysis, generate prompt sentences to be input to the generative AI.

[1414] Input: Baby's emotional state data, parent's emotional pattern data

[1415] Output: Prompt text to be input to the generation AI

[1416] Specific behavior:

[1417] The server generates a corresponding prompt sentence based on the analysis results (e.g., "Generate a parent's response to a lonely baby").

[1418] Step 4: Parent Reaction Generation

[1419] Server: Uses a generative AI model to generate video and audio parental responses based on the prompt.

[1420] Input: Prompt sentence, past parent response data

[1421] Output: Parents' reaction video data, audio data

[1422] Specific behavior:

[1423] A prompt sentence is input into the generative AI model, and a video and audio of the parent's reaction is generated.

[1424] The generated video and audio are encoded and prepared for transmission to the device.

[1425] Step 5: Play the video

[1426] Device: Receives the parent's reaction video and audio sent from the server and plays them according to the baby's condition.

[1427] Input: Parent reaction video data, audio data

[1428] Output: Video and audio played to the baby

[1429] Specific behavior:

[1430] The terminal receives the data from the server.

[1431] Video and audio are played at the optimal time for your baby's current condition.

[1432] The baby's reactions during and after playback are recorded again and sent as feedback to the server.

[1433] In this way, the system can effectively guide the baby through the entire process from data collection to video playback.

[1434] (Application example 1)

[1435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1436] In food delivery services, when delivery personnel are busy or unable to communicate directly with customers during deliveries, they tend to inadequately explain the delivery status or problems to customers. This results in customer anxiety and dissatisfaction, leading to a decline in the quality of service. It is also necessary to improve the situation where delivery personnel are unable to respond appropriately and in a timely manner when they encounter unexpected problems.

[1437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1438] In this invention, the server includes a means for monitoring the status of the delivery person, a means for recording the speech, behavior, voice, and facial expression of the delivery person, a means for analyzing the monitored status data of the delivery person and the recorded data of the delivery person, a means for generating a video and audio response of the delivery person, and a means for playing the generated video and audio to the customer. This makes it possible to provide an appropriate response to the customer even when the delivery person is not present, thereby reducing the customer's anxiety and dissatisfaction.

[1439] "Delivery Person" means an individual or entity responsible for delivering goods to a customer.

[1440] "Means for monitoring the status" refers to the function of checking the target's behavior and voice in real time using devices such as cameras and microphones.

[1441] "Means for recording speech, voice, and facial expressions" refers to devices and software that store the delivery person's speech, voice, and facial expressions as digital data.

[1442] "Means of analysis" refers to the algorithms and programs used to analyze collected data and convert it into meaningful information.

[1443] "Means for generating responsive video and audio" refers to technology that creates visual and audio data to appropriately reproduce the delivery person's condition and emotions based on the analysis results.

[1444] "Means for displaying video and audio to customers" means any device or software used to display or play the generated visual and audio data to customers.

[1445] "Means for identifying response patterns" refers to algorithms or methods for detecting and identifying patterns of delivery personnel's behavior and speech in multiple situations.

[1446] "Means for timely playback" refers to technology or equipment that allows appropriate reactions and responses to be played back in real time at the required timing.

[1447] The system for implementing this invention aims to alleviate customer anxiety and dissatisfaction by monitoring the status of delivery personnel and providing appropriate responses to customers. A specific implementation method for this system is described below.

[1448] 1. System Configuration

[1449] This system consists of the following main modules:

[1450] a. Data Collection Module

[1451] The device (e.g., smartphone or tablet) is equipped with a camera and microphone, which records the delivery person's words, actions, voice, and facial expressions. This data is collected in real time and sent to a server. For example, the device records a scene in which the delivery person has a "confused expression" or "explains something to a customer over the phone," and uploads the video and audio data to the server.

[1452] b. Sentiment Analysis Module

[1453] The server analyzes the facial expressions, voice, and behavioral data of the delivery person sent from the device to identify the delivery person's emotional patterns. At the same time, it also analyzes the behavioral data of the delivery person sent from the device to infer the reason for the delivery person's actions. For example, it can determine that the delivery person is delayed due to "traffic congestion" based on a confused expression on their face.

[1454] c. Generative AI module

[1455] The server uses generative AI to generate a video and audio response from the delivery person based on the results of the emotion analysis. At this stage, technology is used to realistically reproduce the delivery person's facial expressions and voice based on past data. For example, a video can be generated in which the delivery person explains to the customer that the delivery will be five minutes late.

[1456] d. Video playback module

[1457] The generated video and audio data of the delivery person is sent from the server to the terminal. The terminal selects and plays the delivery person's response video that is most appropriate for the customer. This allows the customer to understand the delivery status and feel reassured. For example, if the customer is confused, a video of the delivery person explaining in a calm tone will be played.

[1458] 2. Hardware and software used

[1459] Terminal (smartphone, tablet): Records the delivery person's status using a camera and microphone and sends it to the server.

[1460] Server: Analyzes the collected data using deep learning models (e.g., face recognition models using Keras). For generative AI, frameworks such as TensorFlow are used.

[1461] Generative AI: Technology that uses facial recognition and voice data to generate realistic facial expressions and voices of delivery personnel.

[1462] Playback device: Software installed on the device that plays the generated video and audio to the customer.

[1463] 3. Examples of prompt sentences

[1464] For example, the prompt for a scene in which a delivery person is stuck in traffic and looks confused is as follows:

[1465] "The delivery person is stuck in traffic and has a confused expression. Based on that, generate a sentence that will convey an appropriate message to the customer."

[1466] With the above components and processes, the present invention makes it possible to provide appropriate responses to customers even when a delivery person is not present, thereby reducing customer anxiety and dissatisfaction.

[1467] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1468] Step 1:

[1469] The terminal collects the delivery person's status (video and audio).

[1470] Specifically, the device's camera and microphone are used to record the delivery person's movements, voice, and facial expressions in real time, and this data is recorded as video and audio files that include the delivery person's speech and facial expressions.

[1471] Input: Real-time video and audio of the delivery person

[1472] Output: Video file (e.g. mp4 format), audio file (e.g. wav format)

[1473] Step 2:

[1474] The data collected by the device is sent to the server in real time.

[1475] The collected video and audio files are uploaded to a server using an appropriate communication protocol (e.g., HTTP).

[1476] Input: Video files, audio files

[1477] Output: Data stored on the server

[1478] Step 3:

[1479] The server analyzes the received data and identifies the emotional patterns of the delivery person.

[1480] Specifically, an emotion analysis model using Keras is used to analyze the facial expressions and tone of the delivery person's voice from the received video and audio data, and this analysis identifies the delivery person's emotions, such as "confused" or "happy."

[1481] Input: Video and audio files on the server

[1482] Output: Emotional pattern (e.g., confusion, joy)

[1483] Step 4:

[1484] Generative AI generates response video and audio based on emotional patterns.

[1485] Based on the results of the sentiment analysis, a generative AI model (e.g., a text-to-video generation model) is used to create a realistic video or audio response from the delivery person. For example, if the delivery person is confused, a video containing the message "Delivery will be 5 minutes late" is generated.

[1486] Input: Emotion pattern, prompt sentence

[1487] Output: Response video file, response audio file

[1488] Step 5:

[1489] The server generates a response video and audio and sends it to the terminal.

[1490] The generated response video and audio are then transmitted to the terminal again using an appropriate communication protocol.

[1491] Input: Response video file, response audio file

[1492] Output: Data stored on the device

[1493] Step 6:

[1494] The terminal plays the most appropriate response video and audio to the customer.

[1495] The device will then play the video and audio responses it receives to help customers understand the situation and feel at ease. For example, if a video explaining a delivery delay is played, customers will be able to understand the delivery person's situation and their anxiety about waiting will be alleviated.

[1496] Input: Response video file, response audio file

[1497] Output: Video and audio response played

[1498] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1499] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby in place of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[1500] System configuration and operation

[1501] The system consists of the following major components:

[1502] 1. Data Collection Module

[1503] 2. Sentiment Analysis Module

[1504] 3. Generative AI Module

[1505] 4. Video Playback Module

[1506] 5. Emotion Engine

[1507] 1. Data Collection Module

[1508] The devices (smartphones and tablets) are equipped with cameras and microphones to monitor the baby's condition and record the parent's behavior, voice, and facial expressions. Using these devices, the devices transmit the baby-parent interaction data to a server in real time.

[1509] Examples:

[1510] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[1511] 2. Sentiment Analysis Module

[1512] The server analyzes the parent's facial expressions, voice, and behavioral data sent from the device to identify the parent's emotional patterns, while also analyzing the baby's crying and behavioral data sent from the device to infer the baby's emotions and needs.

[1513] Examples:

[1514] The server analyzes the baby's cries and, if it determines that the baby is lonely, it uses previously collected data on parents to identify a pattern of "speaking to a baby who feels lonely in a gentle voice."

[1515] 3. Generative AI Module

[1516] The server uses AI to generate the parent's reaction video and audio based on the emotion analysis results. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[1517] Examples:

[1518] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1519] 4. Video Playback Module

[1520] The generated video and audio data of the parent are sent from the server to the device, which then selects and plays the video of the parent's reaction that best suits the baby's state. This allows the baby to feel the parent's presence and feel reassured.

[1521] Examples:

[1522] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[1523] 5. Emotion Engine

[1524] The emotion engine recognizes the emotions of the parent user in real time and reflects that emotional data in the operation of the entire system. The emotion engine analyzes video and audio data to identify emotional states such as "the parent is feeling stressed." Based on these results, the generative AI then generates more appropriate and detailed videos of the parent's reactions.

[1525] Examples:

[1526] If the parent is feeling stressed, the generated video will adjust the parent's facial expression to appear slightly calmer.

[1527] Natural language explanation of the system's program

[1528] 1. Data Collection

[1529] Device: Using a camera and microphone, the device records the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[1530] 2. Sentiment analysis

[1531] Server: Analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," "surprised," etc. Additionally, analyzes the baby data and infers the baby's state (e.g., loneliness, fear, request for diaper change) from the tone of the crying and behavior.

[1532] 3. Emotion Engine

[1533] Server: The emotion engine recognizes the parent's real-time emotions and reflects the results in the analysis data, thereby understanding the parent's current emotional state.

[1534] 4. Generation AI

[1535] Server: Based on the sentiment analysis data and the results of the emotion engine, it identifies the most appropriate parental response and uses generative AI to generate video and audio that reproduces that facial expression and voice.

[1536] 5. Video playback

[1537] Device: Receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1538] Through the above processing, the present invention can effectively soothe a baby even when the parent is away or busy, and by reflecting the emotions of the user (parent), it is possible to provide a more realistic and appropriate response.

[1539] The processing flow will be explained below.

[1540] Step 1: Data collection

[1541] Device: The camera and microphone record the parent's facial expressions, tone of voice, speech, gestures, etc. This data is sent to the server in real time.

[1542] Example: A parent speaking to their baby in a gentle voice is recorded using a camera and microphone, and the video and audio data is sent to a server.

[1543] Step 2: Monitor your baby

[1544] Device: Monitors the baby's crying and behavior in real time. Cameras and microphones capture the crying and video and send it to the server.

[1545] Example: When a baby starts crying, the crying is recorded and the video is sent to a server.

[1546] Step 3: Parental sentiment analysis

[1547] Server: Analyzes the received parent's facial expression data and identifies emotional patterns such as "smiling," "sad," or "surprised." It also analyzes the tone of voice and the content of the speech to assign an emotional label.

[1548] Example: The server analyzes data on a parent's "smile" and "gentle voice" and assigns the emotion label "gentle" to them.

[1549] Step 4: Analyze the baby's condition

[1550] Server: Analyzes the tone of a baby's cry and body language to determine the reason for the crying (e.g., lonely, scared, or asking for a diaper change).

[1551] Example: A server analyzes a baby's cry and determines that it sounds lonely.

[1552] Step 5: Parental emotion recognition by emotion engine

[1553] Server: Recognizes the parent's real-time emotions using an emotion engine. Analyzes the parent's video and audio data to detect their current emotional state (e.g., stress, relief).

[1554] Example: The emotion engine analyzes the video and audio data of a parent and recognizes that the parent is feeling stressed.

[1555] Step 6: Parent Reaction Generation

[1556] Server: Based on the results of the parent's emotion analysis and the emotion engine, the server predicts the most appropriate parental response. Using generative AI, it generates video and audio that realistically reproduces the parent's facial expressions and voice.

[1557] Example: If a baby is judged to be "lonely" and the parent is recognized as "stressed," the AI ​​will generate a video of the parent speaking to the baby in a calm, gentle voice.

[1558] Step 7: Select and send your video

[1559] Server: Selects the most appropriate parental reaction video from the generated videos and sends it to the device.

[1560] Example: If a video of a parent speaking in a gentle voice and a video of a parent smiling and waving are generated, select the video of a baby crying and speaking in a gentle voice and send it to the device.

[1561] Step 8: Play the video

[1562] Device: The parent's video and audio are sent from the server and played in real time on the tablet or smartphone screen, with the audio played through the speaker.

[1563] Example: Play a video of a parent speaking to a baby in a gentle voice to reassure the baby.

[1564] This series of steps enables the system to respond quickly and appropriately to the baby's emotions and requests, and by reflecting the emotions of the user (parent), it can provide a more realistic and effective way to soothe a baby.

[1565] Example 2

[1566] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1567] In today's busy home environment, it is often difficult for parents to spend all their time with their babies. In particular, if a baby starts crying or feels anxious, parents' inability to respond quickly can have an impact on the baby's emotions. Another issue is that parents themselves may find it difficult to respond optimally if they are stressed.

[1568] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for monitoring the baby's condition, means for recording the words, actions, voice, and facial expressions of the parent, means for analyzing the monitored baby's condition data and the recorded parent data, means for generating video and audio of the parent's reaction based on the analysis, means for playing the generated video and audio for the baby, and means for recognizing the parent's real-time emotions and reflecting the data in the operation of the entire system. This allows the baby to be soothed effectively even when the parent is away or busy, and by reflecting the parent's emotions, a more realistic and appropriate response is possible.

[1569] A "means for monitoring the baby's condition" is a device or system that uses a camera or microphone to record the baby's movements and cries and analyzes the data.

[1570] "Means for recording parental behavior, voice, and facial expressions" refers to a device or system that uses a camera or microphone to record the words spoken by a parent, their tone of voice, and their facial expressions.

[1571] "Monitored baby status data" refers to data on the baby's movements and crying collected by monitoring means such as cameras and microphones.

[1572] "Recorded parent data" refers to data such as parental words, tone of voice, and facial expressions collected through means that record parental behavior, voice, and facial expressions.

[1573] "Means of analysis" refers to software and algorithms that analyze the collected data and identify the emotions and states of the baby and parent.

[1574] The "means for generating parental reaction video and audio" is a device or system that generates video clips and audio to reproduce the parent's actions and voice based on the analysis results.

[1575] The "means for playing the generated video and audio to the baby" is a device or system for playing the generated parent's reaction video and audio to the baby in a timely manner.

[1576] "Means for recognizing the parent's real-time emotions and reflecting that data in the operation of the entire system" refers to a device or system that analyzes the parent's current emotional state using an emotion engine and takes the results into account and reflects them in the operation of the system.

[1577] An "emotion engine" is software or an algorithm that analyzes video and audio data to identify the emotional state of parents and babies and reflect that information in the system's operation.

[1578] "Generative AI" refers to the artificial intelligence models and techniques used to generate parental reaction videos and audio.

[1579] This invention is a baby soothing system that combines an emotion engine and has a dedicated configuration for soothing a baby on behalf of the parent. This system monitors the baby's condition, records and analyzes the parent's words, actions, voice, and facial expressions, and generates and plays appropriate video and audio responses from the parent based on the analysis results. The system also aims to provide more accurate responses by recognizing the emotions of the parent (user) using the emotion engine and reflecting these in the operation of the entire system.

[1580] The system consists of the following major components:

[1581] 1. Data Collection Module

[1582] 2. Data transmission module

[1583] 3. Sentiment Analysis Module

[1584] 4. Emotion Engine Module

[1585] 5. Generative AI Module

[1586] 6. Video Playback Module

[1587] 1. Data Collection Module

[1588] The device uses a camera and microphone to record the parent's words, actions, voice, and facial expressions. This data captures the interaction between the baby and the parent and transmits it to the server in real time. The baby's crying and behavior are also recorded and transmitted to the server.

[1589] Examples:

[1590] The device records the parent speaking to the child in a gentle voice and uploads the video and audio data to a server.

[1591] 2. Data transmission module

[1592] The device transmits the collected data to the server in real time, which reflects the real-time status of the baby and parent and is analyzed in the next processing step.

[1593] Examples:

[1594] The device records the parent's facial expressions and voice data and sends it to the server.

[1595] 3. Sentiment Analysis Module

[1596] The server analyzes the received parent data and identifies emotional patterns such as "smiling," "sad," and "surprised." It also analyzes the baby's data and estimates the baby's state (e.g., loneliness, fear, or request for a diaper change) from the tone of the crying and behavior.

[1597] Examples:

[1598] The server analyzes the tone of the baby's cry to determine whether the baby is feeling lonely, and then, based on past data collected from parents, identifies a pattern of parents speaking to their babies in a gentle voice when they feel lonely.

[1599] 4. Emotion Engine Module

[1600] The server uses an emotion engine to recognize the parent's real-time emotions and reflects the results in the analysis data, thereby making it possible to understand the parent's current emotional state.

[1601] Examples:

[1602] The emotion engine analyzes the parent's facial expressions and tone of voice and determines that the parent is stressed. This information is reflected in subsequent processing.

[1603] 5. Generative AI Module

[1604] The server uses generative AI to generate the parent's reaction video and audio based on the emotion analysis data and the results of the emotion engine. At this stage, technology is used to reflect the user's emotions and realistically reproduce the parent's facial expressions and voice.

[1605] Examples:

[1606] If it determines that the baby "requests a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1607] 6. Video Playback Module

[1608] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1609] Examples:

[1610] If a baby is "scared and crying," the device will play a video of "a parent reassuring the baby with a soft smile."

[1611] Example prompts to be input to the generative AI model

[1612] "If your baby is feeling lonely, generate a video of the parent talking to them in a gentle voice."

[1613] "If a baby is scared and crying, generate a video of the parent reassuring the baby with a soft smile."

[1614] "If a baby requests a diaper change, generate a video of the parent taking the baby to the changing table."

[1615] This system allows babies to be soothed effectively even when parents are away or busy, and by reflecting the parents' emotions, it allows for more realistic and appropriate responses.

[1616] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1617] Step 1:

[1618] Data collection

[1619] The device uses a camera and microphone to record the baby's state and the parent's words, actions, voice, and facial expressions, thereby collecting scenes of interaction between the baby and the parent.

[1620] Input: Baby's cries and movements, parent's facial expressions and words.

[1621] Output: Collected video and audio data.

[1622] Specific operation: The device captures real-time footage of the baby crying and the parent's gentle response, and records this data.

[1623] Step 2:

[1624] Data transmission

[1625] The device transmits the collected data in real time to a server, which prepares the data for the next analysis step.

[1626] Input: Collected video and audio data.

[1627] Output: The data sent to the server.

[1628] Specific operation: The device records the baby's crying and the parent's reaction, and sends the video and audio data to a server via the network.

[1629] Step 3:

[1630] sentiment analysis

[1631] The server analyzes the received parent data to identify the parent's emotional patterns, and analyzes the baby data to estimate the baby's state from the tone of the crying and behavior.

[1632] Input: Video and audio data sent to the server.

[1633] Output: Parent and baby emotion pattern data.

[1634] Specific operation: The server analyzes the baby's cry and determines that the baby is "lonely," and at the same time identifies patterns from the parent's facial expressions and tone of voice that indicate "the parent has the intention to comfort the baby."

[1635] Step 4:

[1636] Emotion Engine Recognition

[1637] The server recognizes the parent's real-time emotions using an emotion engine and reflects that data in the overall system behavior, which takes the parent's current emotional state into account when proceeding to the next step.

[1638] Input: Parent video and audio data.

[1639] Output: Parent emotion data from the emotion engine.

[1640] How it works: The emotion engine analyzes the parent's facial expressions and tone of voice to determine that the parent is stressed. This information is used as input for the generative AI.

[1641] Step 5:

[1642] Generative AI video creation

[1643] The server uses generative AI to generate parental reaction videos and audio based on the emotion analysis results and emotion engine data, reflecting the user's emotions and recreating realistic facial expressions and voices.

[1644] Input: Sentiment analysis data and sentiment engine data.

[1645] Output: Generated reaction video and audio data.

[1646] Specific behavior: If it is determined that the baby "requires a diaper change," the generative AI creates a video of the parent taking the baby to the changing table.

[1647] Step 6:

[1648] Video playback

[1649] The device receives the parent's video and audio responses sent from the server and plays them back to the baby in a timely manner, allowing the baby to sense the parent's presence and feel a sense of security.

[1650] Input: Generated reaction video and audio data.

[1651] Output: Video and audio played to the baby.

[1652] What it does: The device calms a crying baby by playing a video of a parent comforting the baby with a soft smile.

[1653] (Application example 2)

[1654] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1655] In the workplace, reducing worker stress and fatigue and creating a safe and efficient work environment are important challenges. However, conventional methods make it difficult to grasp workers' emotions and physical conditions in real time and take appropriate measures. As a result, workers' health risks increase and work efficiency may decline. For this reason, there is a demand for a system that can monitor workers' emotions and physical conditions in real time and take appropriate measures.

[1656] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1657] In this invention, the server includes a means for monitoring the state of workers, a means for recording the words, actions, voices, and facial expressions of workers, a means for analyzing the monitored state data of workers and the recorded data of workers, a means for generating appropriate reaction video and audio based on the analysis, and a means for playing the generated video and audio to the workers. This makes it possible to monitor the emotions and physical condition of workers in real time and take appropriate measures.

[1658] "Means for monitoring worker status" refers to devices or systems for monitoring the physical status of workers, such as their complexion, posture, and movements, in real time.

[1659] "Means for recording the words, actions, voice, and facial expressions of workers" refers to devices such as cameras and microphones that record the content of workers' words, tone of voice, facial expressions, etc.

[1660] "Means for analyzing monitored worker condition data and recorded worker data" refers to software and hardware for analyzing collected worker physical and emotional data.

[1661] "Means for generating appropriate response videos and audio based on analysis" refers to artificial intelligence (AI) technology for automatically generating appropriate response messages and videos according to the emotional state of workers.

[1662] "Means for playing the generated video and audio to workers" refers to devices such as displays and speakers used to communicate the generated response messages and videos to workers.

[1663] The system that realizes this application example, a worker emotion recognition support robot, consists of the following main components:

[1664] System configuration and operation

[1665] 1. Means of monitoring the condition of workers

[1666] Cameras suitable for the work environment (e.g., general surveillance cameras or high-resolution cameras) are used to monitor the status of workers in real time, and highly sensitive microphones are installed to collect the voices of workers.

[1667] 2. Means of recording the words, actions, voices, and facial expressions of workers

[1668] The robots will be fitted with devices that record the workers' facial expressions and tone of voice, allowing for detailed recording of their behavior and emotions.

[1669] 3. Means of analyzing monitored worker status data and recorded worker data

[1670] The server will be installed with software that uses Google TensorFlow to analyze worker status data, analyzing facial expressions and voice to identify the worker's level of stress and fatigue.

[1671] 4. A method for generating appropriate reaction video and audio based on the analysis

[1672] The server uses OpenAI's GPT-4 to generate appropriate reaction videos and audio messages tailored to the worker's state. This generative AI model uses past data to create more natural and effective responses.

[1673] 5. A means of playing the generated video and audio to workers

[1674] The generated video and audio are then played back to the worker through the robot's on-board display and speakers, allowing the worker to receive encouragement and break suggestions at appropriate times.

[1675] Examples of the system and how to use it

[1676] How to use this system will be explained with concrete examples.

[1677] 1. Data Collection

[1678] The device is equipped with a camera and microphone that records the worker's facial expressions and voice in real time and sends the data to a server. For example, if a worker says "I'm tired" while working, the microphone will collect the voice.

[1679] 2. Sentiment analysis

[1680] The data is then analyzed on the server, and TensorFlow is used to identify the worker's fatigue and stress levels for the day based on subtle changes in facial expressions and tone of voice.

[1681] 3. Emotion Engine

[1682] The server uses Amazon Comprehend to recognize the worker's real-time emotions and incorporates that information into the next steps, allowing for more accurate responses.

[1683] 4. Generation AI

[1684] Using GPT-4, the system generates voice messages and videos based on the results of sentiment analysis, such as encouraging workers or suggesting them to take a break. For example, it generates a message saying, "It's a good idea to take a short break."

[1685] 5. Video playback

[1686] The generated messages and videos are played through the robot's displays and speakers, and if a worker shows signs of fatigue, a video message such as "Take a deep breath to relax" will be displayed at that point.

[1687] Examples of prompt statements

[1688] An example of an input prompt for the generative AI model is as follows:

[1689] If a worker looks tired, generate a message saying, "Good work. Maybe you should take a break."

[1690] If a worker is feeling stressed, generate a voice message saying, "Take a deep breath to relax. It's okay."

[1691] This system allows workers' emotions and physical condition to be monitored in real time, and appropriate measures to be taken to reduce health risks and improve work efficiency.

[1692] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1693] Step 1:

[1694] The camera and microphone installed on the terminal collect the facial expressions and voices of the workers. The system records what the workers say and their facial expressions while working and sends this data to a server in real time. The input is the workers' audio and video data, and the output is the transmission of this data to the server.

[1695] Step 2:

[1696] The server analyzes the received facial and voice data. Using Google TensorFlow, it analyzes the worker's emotional state (e.g., fatigue, stress, concentration, etc.) from their facial expressions. It also analyzes the tone and content of their voice from the audio data to determine their emotional state. The input is the worker's video and audio data, and the output is the worker's emotional state data as an analysis result.

[1697] Step 3:

[1698] The server uses an emotion engine to perform a more detailed analysis of the worker's real-time emotions and incorporates the results into the overall data. It uses Amazon Comprehend to identify the emotional state and output intermediate data for application to generative AI. The input is facial expression and voice analysis results, and the output is detailed emotional data for application to the generative AI model.

[1699] Step 4:

[1700] The server uses a generative AI model (OpenAI GPT-4) to generate appropriate response messages and videos based on the worker's emotional state. The generative AI model creates messages and videos that are optimal for the worker's situation based on the prompt text. The input is detailed emotional data and the prompt text, and the output is the generated response message and video data.

[1701] Step 5:

[1702] The generated messages and videos are sent from the server to the terminal and played to the worker through the terminal's display and speaker. The user watches the presented message or video and takes appropriate action (e.g., take a break or take a deep breath). The input is the generated response message and video data, and the output is to provide them to the worker.

[1703] This series of steps will create a system that supports worker health and work efficiency in real time.

[1704] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1705] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1706] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1707] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1708] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1709] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1710] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1711] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1712] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1713] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1714] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1715] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1716] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1717] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1718] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1719] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1720] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1721] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1722] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1723] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1724] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1725] The following is further disclosed regarding the above embodiment.

[1726] (Claim 1)

[1727] A means of monitoring the baby's condition;

[1728] A means of recording the words, actions, voices, and facial expressions of parents;

[1729] means for analyzing the monitored baby condition data and the recorded parent data;

[1730] means for generating parental reaction video and audio based on the analysis;

[1731] A means of playing the generated video and audio to the baby;

[1732] A system including:

[1733] (Claim 2)

[1734] It has the means to analyze the tone of a baby's cry and its behavior,

[1735] 10. The system of claim 1, further comprising means for identifying common parental response patterns.

[1736] (Claim 3)

[1737] Select the best parental reaction video from the generated videos,

[1738] 10. The system of claim 1, further comprising means for playing back to the baby in a timely manner.

[1739] "Example 1"

[1740] (Claim 1)

[1741] A means of monitoring the baby's condition;

[1742] A means of recording the words, actions, voices, and facial expressions of parents;

[1743] means for analyzing the monitored baby condition data and the recorded parent data;

[1744] means for generating parental reaction video and audio based on the analysis;

[1745] A means of playing the generated video and audio to the baby;

[1746] A system including:

[1747] (Claim 2)

[1748] 2. The system of claim 1, further comprising means for analyzing the tone and behavior of the baby's crying and means for identifying the parent's emotional patterns, and means for generating a parent's response using a generative AI model based on the analysis results.

[1749] (Claim 3)

[1750] 2. The system according to claim 1, further comprising means for selecting the most appropriate reaction video for the baby from among the generated reaction videos of parents and playing it back to the baby in a timely manner.

[1751] "Application Example 1"

[1752] (Claim 1)

[1753] a means for monitoring the status of the delivery personnel;

[1754] A means of recording the delivery person's actions, voice, and facial expressions;

[1755] means for analyzing the monitored delivery personnel status data and the recorded delivery personnel data;

[1756] A means for generating a response video and audio of a delivery person based on the analysis;

[1757] a means for playing the generated video and audio to the customer;

[1758] A system including:

[1759] (Claim 2)

[1760] It has a means of analyzing the actions and facial expressions of delivery personnel,

[1761] 10. The system of claim 1, further comprising means for identifying common response patterns among delivery personnel.

[1762] (Claim 3)

[1763] Select the best response video from the multiple delivery staff responses generated,

[1764] 10. The system of claim 1, further comprising means for playing the content to the customer in a timely manner.

[1765] "Example 2: Combining Emotion Engines"

[1766] (Claim 1)

[1767] A means of monitoring the baby's condition;

[1768] A means of recording the words, actions, voices, and facial expressions of parents;

[1769] means for analyzing the monitored baby condition data and the recorded parent data;

[1770] means for generating parental reaction video and audio based on the analysis;

[1771] A means of playing the generated video and audio to the baby;

[1772] A means of recognizing parents' real-time emotions and reflecting that data in the behavior of the entire system;

[1773] A system including:

[1774] (Claim 2)

[1775] It has the means to analyze the tone of a baby's cry and its behavior,

[1776] a means of identifying common parental response patterns;

[1777] 10. The system of claim 1, further comprising means for generating parental reaction video and audio using generative AI.

[1778] (Claim 3)

[1779] Select the best parental reaction video from the generated videos,

[1780] 10. The system of claim 1, further comprising means for playing back to the baby in a timely manner.

[1781] "Application example 2 when combining emotion engines"

[1782] (Claim 1)

[1783] means of monitoring the condition of workers;

[1784] a means of recording the words, actions, voices, and facial expressions of workers;

[1785] means for analyzing the monitored worker condition data and the recorded worker data;

[1786] means for generating appropriate reaction video and audio based on the analysis;

[1787] a means for playing the generated video and audio to the worker;

[1788] A system including:

[1789] (Claim 2)

[1790] Have the means to analyze the tone of voice and behavior of workers;

[1791] 10. The system of claim 1, further comprising means for identifying common patterns of responses among workers.

[1792] (Claim 3)

[1793] Select the best reaction video from the multiple generated videos,

[1794] 10. The system of claim 1, further comprising means for replaying to workers in a timely manner. [Explanation of symbols]

[1795] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of monitoring the baby's condition; A means of recording the words, actions, voices, and facial expressions of parents; means for analyzing the monitored baby condition data and the recorded parent data; means for generating parental reaction video and audio based on the analysis; A means of playing the generated video and audio to the baby; A system including:

2. It has the means to analyze the tone of a baby's cry and its behavior, 10. The system of claim 1, further comprising means for identifying common parental response patterns.

3. Select the best parental reaction video from the generated videos, 10. The system of claim 1, further comprising means for playing the baby in a timely manner.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A