system
A system using sensors and algorithms to analyze a baby's vocalizations, facial expressions, and movements provides caregivers with real-time emotional insights and guidance, addressing the challenge of parental anxiety in understanding baby emotions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Parents struggle to accurately determine their baby's emotions from initial vocalizations, expressions, and movements, leading to increased anxiety and stress in childcare.
A system that acquires a baby's vocalizations, facial expressions, and body movements using microphones, cameras, and motion sensors, preprocesses the data, and uses algorithms to determine emotions, generating specific advice for caregivers based on these determinations.
Enables accurate real-time emotion assessment of babies, providing caregivers with actionable advice to enhance childcare interactions and reduce parental stress.
Smart Images

Figure 2026071046000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is difficult to determine the emotion of a baby from its initial vocalizations, expressions, and movements, and many parents have anxiety and stress due to this. In particular, when parents cannot understand what the baby wants or what kind of emotions it has and cannot take appropriate actions, the burden on parents in child-rearing increases. To solve such problems, there is a need for a technology that determines the emotion of a baby and provides specific countermeasures based on it.
Means for Solving the Problems
[0005] This invention comprises a voice acquisition means for acquiring the baby's vocalizations, an image acquisition means for acquiring the baby's facial expressions, and a motion acquisition means for acquiring the baby's body movements. A data processing means is provided for preprocessing and analyzing the data acquired by these means, and an emotion determination means determines the baby's emotions based on the data preprocessed by this data processing means. Furthermore, an advice generation means generates specific countermeasures for the parents based on the determined emotions, and an information provision means provides the generated advice to the parents. This system makes it possible to accurately grasp the baby's emotions and support childcare.
[0006] "Voice acquisition means" refers to devices and technologies for capturing a baby's vocalizations, and includes microphones and audio sensors.
[0007] "Image acquisition means" refers to devices and technologies for photographing and acquiring a baby's facial expressions, and includes cameras, image sensors, and the like.
[0008] "Motion acquisition means" refers to devices and technologies for recording and acquiring a baby's body movements, and includes motion sensors and video cameras.
[0009] "Data processing means" refers to technologies and devices that have the function of pre-processing acquired audio, image, and motion data and then converting it into an analyzable format.
[0010] "Emotion determination means" refers to algorithms or devices used to determine a baby's emotions based on pre-processed data.
[0011] "Advice generation means" refers to devices or technologies used to construct specific strategies for dealing with parents based on the emotions that have been determined.
[0012] "Information provision means" refers to devices and technologies used to convey generated advice to parents, including displays and speakers. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. This system mainly consists of three components: a terminal, a server, and a user.
[0035] Terminal operation
[0036] The device is placed around the baby, and in this case, it would be a smartphone or tablet. The device detects the baby's facial expressions with its built-in camera, vocalizations with its microphone, and body movements with its motion sensor. These sensors collect data in real time and temporarily store it on the device. The data is then pre-processed before being sent to the server. For example, if the baby is captured laughing, the camera captures an image of that moment, and the audio captures the laughter.
[0037] Server operation
[0038] The server plays a central role in receiving and analyzing data sent from the terminal. The received data is first preprocessed, including noise reduction and standardization. Next, the emotion determination engine uses algorithms for voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, data combining a smile and a high-pitched voice would be determined to be "happy." Based on this determination, the generating AI constructs appropriate advice for the parents. For example, it might generate advice such as, "Let's play with the baby."
[0039] User actions
[0040] The user's role is to review the information received on the device and use it to guide their next actions. The emotion assessment results and advice sent from the server are displayed on the device's interface and provided in a format that is easy for the user to understand. Based on the displayed information, the user can decide on specific actions. For example, by following the displayed advice and playing with the baby, the user can facilitate smoother communication between parent and child.
[0041] As described above, the system of the present invention supports childcare by determining the baby's emotions in real time and suggesting the most appropriate response to the parents.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The device acquires audio data from the baby's surroundings using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. This data is temporarily stored in real time.
[0045] Step 2:
[0046] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0047] Step 3:
[0048] The terminal sends pre-processed data to the server. The data is securely transferred to the server via a security protocol.
[0049] Step 4:
[0050] The server receives audio, image, and motion data transmitted from the terminal and stores it in a database. This data is immediately used for analysis.
[0051] Step 5:
[0052] The server uses stored data to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, it analyzes a smiling image and a cheerful voice to determine that the baby is "happy."
[0053] Step 6:
[0054] Based on the analysis results, the server generates specific advice for parents using AI. For example, it might suggest, "Continue playing with your baby."
[0055] Step 7:
[0056] The server sends the judgment result and generated advice to the terminal.
[0057] Step 8:
[0058] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format that is easy for parents to understand.
[0059] Step 9:
[0060] The user uses the information displayed on the device as a reference to take specific actions for the baby. For example, if the device determines that the baby is sleepy, the user will take appropriate action to put the baby to sleep.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] In modern childcare, it is crucial for parents to quickly and accurately understand their baby's emotions and condition and respond appropriately. However, many parents find it difficult to accurately grasp their baby's emotions amidst their busy daily lives, which can lead to them being unable to choose appropriate childcare methods. This invention aims to solve this problem and support childcare more effectively by determining the baby's emotions in real time and providing parents with the most suitable response.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes an acquisition means for acquiring voice, an acquisition means for acquiring facial expressions, an acquisition means for acquiring body movements, a processing means for preprocessing and analyzing the acquired data, a determination means for determining emotions based on the preprocessed data, a generation means for generating specific countermeasures based on the determined emotions, a provision means for providing the generated countermeasures, a function for collecting data in real time using sensors at the terminal, temporarily storing and preprocessing it, and a function for using encryption technology in data transmission to realize secure data analysis on the server. This makes it possible to quickly analyze the emotional state of a baby and provide specific and practical advice to the guardian.
[0066] "Acquisition means" refers to a device or method for collecting data such as voice, facial expressions, or body movements.
[0067] "Processing means" refers to an apparatus or method for preprocessing acquired data and performing data analysis.
[0068] "Determination means" refers to a device or method for analyzing and determining emotions based on pre-processed data.
[0069] "Generating means" refers to a device or method for automatically creating specific countermeasures based on the determined emotions.
[0070] "Means of delivery" refers to a device or method for communicating the generated countermeasures to users such as guardians.
[0071] A "sensor" is a device that detects physical phenomena in the environment and outputs them in the form of electrical signals or other similar information.
[0072] "Encryption technology" is a technique that transforms data according to specific rules in order to protect it from unauthorized access.
[0073] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. The specific configuration and operation of the system are described below.
[0074] Terminal operation
[0075] The device takes the form of a smartphone or tablet and is placed near the baby. It has a built-in camera, microphone, and motion sensor, which are used to capture the baby's facial expressions, voice, and body movements in real time. Specifically, the camera can capture the baby's smile, the microphone can record laughter, and the motion sensor can detect arm and leg movements. The collected data is temporarily stored on the device and pre-processed before transmission.
[0076] Server operation
[0077] The server plays a central role in secure and efficient data processing. Data transmitted from the terminal undergoes preprocessing such as noise reduction and standardization, followed by analysis for emotion determination. The server uses a generative AI model to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. Specifically, if a smile or high-pitched laughter is detected, it determines that the baby is "happy" and generates advice based on this.
[0078] User actions
[0079] Users review the assessment results and advice through an interface provided on their device. Based on this information, users decide on their actual actions. For example, if they receive the advice, "Please play with your baby," they can immediately take appropriate action for playtime. As a result, richer communication between parents and children is promoted.
[0080] Example of a prompt
[0081] "We have data that detects a baby's smile and high-pitched voice. What emotions does this represent, and what kind of advice should we offer?"
[0082] This invention is a system that highly analyzes a baby's emotional state and provides parents with an easy-to-understand and effective childcare method.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The device uses a camera, microphone, and motion sensor near the baby to acquire data on facial expressions, voice, and body movements in real time. The input is the baby's natural movements and voice, which are acquired as digital data by the sensors. The output is the raw data obtained from each sensor.
[0086] Step 2:
[0087] The terminal temporarily stores the acquired raw data in its internal memory and performs preprocessing for data transmission. This preprocessing adjusts the image resolution and removes noise from audio data. The input to this process is raw data, and the output is compressed and low-noise data for transmission.
[0088] Step 3:
[0089] The terminal sends pre-processed data to the server. During this process, the data is encrypted and transmitted securely. The input is the pre-processed data, and the output is the data received on the server side. The data is transmitted using Wi-Fi or a mobile network.
[0090] Step 4:
[0091] The server preprocesses and standardizes the received data again. Here, the audio and image formats are unified, and the motion data is converted into a format that is easy to analyze. The input is the data sent from the terminal, and the output is standardized, ready-to-analyze data.
[0092] Step 5:
[0093] The server utilizes a generative AI model to perform voice analysis, image recognition, and motion analysis on standardized data to determine the baby's emotions. The input is standardized data, and the output is the determined emotion information. For example, if laughter and a smile match, the emotion "happy" is output.
[0094] Step 6:
[0095] The server uses a generative AI model to generate specific advice for parents based on the determined emotional information. The input is the determined emotional information, and the output is an advice statement. For example, the advice "Let's play with the baby" might be generated.
[0096] Step 7:
[0097] The server sends the generated advice to the terminal, which then notifies the user. The advice is displayed on the terminal screen in a format that is easy for the user to understand. The input is the advice text, and the output is the user's action based on their understanding.
[0098] (Application Example 1)
[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0100] When providing product recommendations tailored to a baby's condition in a physical store, there are challenges in accurately assessing emotions in real time and providing optimal product information. Furthermore, providing appropriate advice based on the baby's condition requires specialized knowledge from store staff, potentially leading to a lack of consistency in customer service.
[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0102] In this invention, the server includes an input acquisition means for collecting acquired baby voice data, facial expression data, and motion data; a data processing means for preprocessing and analyzing the data obtained from the input acquisition means; and an information recommendation means for determining the baby's emotions based on the preprocessed data and providing recommended product information according to the determination result. This makes it possible to provide real-time product suggestions and advice to parents in a physical store, tailored to the baby's situation.
[0103] "Input acquisition means" refers to devices and methods for collecting voice data, facial expression data, and motion data obtained from a baby.
[0104] "Data processing means" refers to devices and methods for preprocessing and analyzing data obtained from input acquisition means. This means removes noise and standardizes the data.
[0105] An "information recommendation means" is a device or method that determines a baby's emotions based on the results analyzed by a data processing means and provides product information recommended according to those results.
[0106] "Information output means" refers to devices or methods that provide information to user terminals and are used when selecting recommended products in physical stores.
[0107] The system necessary to implement this invention mainly consists of three elements: a user terminal, a server, and a physical store.
[0108] The user device is a smartphone or tablet, equipped with a camera, microphone, and motion sensor. The camera captures the baby's facial expressions, the microphone collects sound, and the motion sensor detects body movements. These sensors collect data in real time, temporarily store it on the user device, and then preprocess it before sending it to the server. Specifically, if the baby is smiling, the facial expression and sound are captured simultaneously.
[0109] The server plays a primary role in receiving and analyzing data sent from user terminals. The received data undergoes preprocessing such as noise reduction and standardization. Next, emotion determination is performed using algorithms for voice analysis, image recognition, and motion analysis. AI models such as TENSORFLOW® are used for this. For example, if a baby's smile and high-pitched voice are detected simultaneously, it is determined to be "happy." Based on this result, a generative AI model creates specific advice for the parents, such as, "Your child is happy. Playing together with toys would be ideal."
[0110] In physical stores, these assessment results and advice are displayed on tablets or dedicated terminals, and staff members make appropriate product suggestions to parents. For example, if the system detects that the baby is crying, it will suggest products such as, "This product is recommended for times when babies tend to cry."
[0111] For example, if a baby becomes fussy when visiting a store, the system will determine in real time that the baby is feeling "anxious" and automatically suggest to the staff things like "a stuffed animal to calm the child."
[0112] An example of a prompt to pass to a generative AI model might be: "Please analyze the following data: baby's voice, baby's facial expression image, and high-frequency sound. Next, determine what emotion the baby is feeling and suggest specific products to recommend to the parents."
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] The user collects the baby's facial expressions, voice, and body movements using the user's device's camera, microphone, and motion sensor. The input for this step is raw data of the baby, and the output is data temporarily stored on the device. Specifically, if the baby is laughing, the facial expression data and laughter are captured on the smartphone.
[0116] Step 2:
[0117] The terminal preprocesses the collected data by denoising and standardizing it, preparing it for transmission to the server. The input for this step is the raw data within the terminal, and the output is the preprocessed data. For example, background noise in audio may be removed, and the brightness of images may be adjusted.
[0118] Step 3:
[0119] The server receives preprocessed data sent from the user's terminal and performs speech analysis, image recognition, and motion analysis. The input for this step is preprocessed data, and the output is information about the baby's emotions. Specifically, TensorFlow is used to determine if the baby is "happy" from their smile and laughter.
[0120] Step 4:
[0121] Based on the emotion assessment results, the server uses a generative AI model to create specific advice for the parents. The input for this step is the emotion assessment results, and the output is the generated advice text. For example, advice such as, "The baby is having fun. Playing together with toys is ideal," is generated.
[0122] Step 5:
[0123] Users can choose the most suitable product for their baby based on information displayed on a terminal or tablet in a physical store. The input in this step is generated advice, and the output is product information for the parent to choose. For example, information such as "We recommend this product during times when your baby is prone to crying" is presented to help the parent make their selection.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] This invention combines a system that acquires a baby's voice, facial expressions, and body movements, and determines the baby's emotions based on that data, with an emotion engine that recognizes the emotions of the user (the caregiver). The system operates through interaction between the terminal, server, and user.
[0126] Terminal operation
[0127] The device is placed near the baby and uses a microphone, camera, and motion sensor to capture voice, facial expressions, and movements. In addition, the device is equipped with sensors to capture the user's facial expressions and voice, collecting real-time emotional data. This data is temporarily stored and pre-processed before being sent to the server.
[0128] Server operation
[0129] The server receives voice, image, and motion data from the baby and user sent from the terminal, and then performs preprocessing such as noise reduction and data format standardization. Next, it uses an emotion determination engine on the baby's data to perform voice analysis, image recognition, and motion analysis. This determines the baby's emotions, such as "happy," "anxious," or "sleepy." The user's data is separately analyzed for emotions to determine whether the caregiver is in a state of "stress" or "security."
[0130] The server uses generative AI to design specific actions parents should take, based on the emotional data of both the baby and the user. The advice is dynamically adjusted; for example, if the user is feeling stressed, relaxation advice may be added. If the baby is anxious and the user is also feeling stressed, a more comprehensive piece of advice might be generated, such as, "Hold your baby to calm them down and take some deep breaths."
[0131] User actions
[0132] Users review the emotion assessment results and advice displayed on their device. Based on the displayed information, users can take appropriate actions towards their baby. The results and advice are presented through concise and easy-to-understand notifications and a visual interface. This helps parents reduce anxiety in childcare and enables them to respond appropriately to emotions.
[0133] In this way, the present invention can improve the childcare environment by simultaneously determining the emotions of both the baby and the parent and providing appropriate countermeasures. This is expected to promote better communication between parents and children.
[0134] The following describes the processing flow.
[0135] Step 1:
[0136] The device acquires audio data from the baby's environment using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. Similarly, it also acquires the user's voice and facial expressions using sensors. This data is temporarily stored in real time.
[0137] Step 2:
[0138] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0139] Step 3:
[0140] The device sends pre-processed baby and user data to the server. A secure protocol is used for data transfer.
[0141] Step 4:
[0142] The server receives data sent from the terminal and stores it in the database. This data is immediately used for analysis.
[0143] Step 5:
[0144] The server performs voice analysis, image recognition, and motion analysis on the baby's data to determine the baby's emotions. For example, it might determine that the baby is "happy" based on a combination of a smiling image and a high-pitched voice.
[0145] Step 6:
[0146] Similarly, the server analyzes user data to determine the user's emotions. For example, it might determine that the user is stressed based on a tired tone of voice and an anxious facial expression.
[0147] Step 7:
[0148] Based on the emotional information of both the baby and the user, the server uses a generating AI to create specific advice for the parents. The user's emotional state is also taken into consideration, so advice such as, "Speak gently to your baby and also take time to calm yourself down," is generated as needed.
[0149] Step 8:
[0150] The server sends the judgment result and advice to the terminal.
[0151] Step 9:
[0152] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format, taking care to ensure that parents can easily understand and act upon the advice.
[0153] Step 10:
[0154] Based on the information displayed on the device, users take specific actions for their baby while also managing their own state of mind. For example, they might incorporate relaxation techniques to reduce their own stress while soothing their baby.
[0155] (Example 2)
[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0157] There is a need to accurately assess a baby's emotions and, based on that assessment, quickly and precisely provide parents with guidance on how to respond. Conventional systems have been insufficient in assessing babies' emotions and analyzing parents' mental states, and have limited ability to generate specific countermeasures based on those results. Therefore, these issues need to be addressed.
[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0159] In this invention, the server includes means for acquiring sound, means for acquiring images, and means for acquiring movement. This makes it possible to comprehensively determine the baby's emotions from various data such as voice, facial expressions, and movement, and further analyze the caregiver's mental state, in order to provide dynamically adjusted specific countermeasures.
[0160] "Means of acquiring sound" refers to technologies for collecting ambient sounds, particularly those used to detect sounds like a baby's voice.
[0161] "Means for acquiring images" refers to methods that use visual sensors such as cameras to capture images of the surroundings, and are used in particular to analyze the baby's facial expressions.
[0162] "Means for acquiring motion" refers to methods that use motion sensors or similar devices to detect body movements, and are particularly used to detect patterns in a baby's movements.
[0163] "Processing means" refers to a system that has the function of organizing the acquired data and converting it into a format suitable for analysis.
[0164] "Determination means" refers to algorithms and technologies used to analyze pre-processed data and evaluate specific states such as emotions.
[0165] "Additional emotion analysis means" refers to technology that evaluates a user's mental state from their speech patterns and facial expressions to gain a deeper understanding of their emotions.
[0166] "Generative means" refers to the technology used to construct and generate specific countermeasures and advice based on the analysis results.
[0167] "Information provision means" refers to screens, notification systems, etc., that present generated advice and information to users.
[0168] This invention is a system implemented through terminals placed around a baby, a server operating in the backend, and a user receiving information. Specific embodiments are described below.
[0169] Terminal configuration and operation
[0170] The device is equipped with a microphone to capture sound, a camera to capture images, and a motion sensor to capture movement. This device is placed near the baby and collects voice, facial expressions, and movements in real time. For example, if the baby is smiling, the camera captures the smile, and the motion sensor detects the baby's active movements. This data is temporarily stored for further processing and then transmitted to a server via the network.
[0171] Server configuration and operation
[0172] The server receives data transmitted from the terminal and performs preprocessing on it. For example, it performs noise reduction on audio data and standardizes the image data format. Next, the server analyzes the data using emotion determination tools. The analysis includes audio analysis, image recognition, and motion analysis, and determines the baby's emotions as "happy," "anxious," "sleepy," etc. At the same time, the server uses additional emotion analysis tools to evaluate the user's state as "stressed" or "at ease."
[0173] Furthermore, the server utilizes generative AI based on these analysis results to suggest specific actions the user should take. For example, if the user is stressed and the baby is anxious, it might recommend "hugging the baby and taking deep breaths together." The generative AI model incorporates historical data and expertise, providing sophisticated solutions.
[0174] An example of a prompt might be, "Please suggest ways to cope when a baby is crying and the caregiver is feeling stressed."
[0175] User actions
[0176] Users review the analysis results and advice displayed on their device. The information is presented in an easy-to-understand interface, enabling quick and accurate responses. This allows users to take appropriate action and reduce anxiety related to childcare.
[0177] This invention's system allows for the simultaneous evaluation of the baby's emotions and the user's mental state, contributing to better communication between parent and child.
[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0179] Step 1:
[0180] The device collects the baby's voice, facial expressions, and movements. Inputs include sounds, images, and movements from the baby's surroundings. Specifically, a microphone picks up voice, a camera captures facial expressions, and a motion sensor detects movement. This data is temporarily stored within the device.
[0181] Step 2:
[0182] The terminal preprocesses the collected data. Because the input data may contain noise or incorrect formatting, the audio data is denoised, and the image data is resized and converted before being sent to the server. The preprocessed data is then output.
[0183] Step 3:
[0184] The terminal sends pre-processed data to the server. The specific data sent includes de-noise-removed audio, facial expression images converted to an appropriate format, and motion information.
[0185] Step 4:
[0186] The server analyzes the received data. It receives pre-processed audio, image, and motion data as input and uses an emotion determination engine to determine the baby's emotions. The specific output is an emotional state such as "happy," "anxious," or "sleepy." Simultaneously, it performs additional emotion analysis to evaluate the user's state.
[0187] Step 5:
[0188] Based on the assessment results, the server generates a response plan for the parent using a generated AI model. The prompt text is a specific situation corresponding to the baby's and the user's state, and specific advice is output based on that. For example, advice such as "Hold your baby and take a deep breath together" might be generated.
[0189] Step 6:
[0190] The user reviews the analysis results and generated advice displayed on the device. The input is the information displayed on the device's screen, and the user uses this to take appropriate action for their baby. This reduces anxiety about childcare.
[0191] (Application Example 2)
[0192] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0193] Conventional baby emotion assessment systems only consider the baby's emotions, making it difficult to obtain specific countermeasures that take into account the emotions of parents or store staff. Furthermore, there was a lack of mechanisms to understand the emotions of both customers and staff and provide appropriate countermeasures, which is crucial for smooth communication with customers in physical stores.
[0194] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0195] In this invention, the server includes an acoustic acquisition means for acquiring the voices of the baby and the customer, an image acquisition means for acquiring the facial expressions of the baby and the customer, and a data processing means for pre-processing and analyzing the acquired information. This makes it possible to determine the emotions of both the baby and the customer, as well as the parent or store staff, in real time and provide specific and effective countermeasures.
[0196] "Sound acquisition means" refers to a device that has the function of detecting and recording ambient sounds.
[0197] An "image acquisition means" is a device that has the function of acquiring visual information and converting it into digital data.
[0198] A "motion sensor means" is a sensor that can detect the movement of an object or human body and record the changes as data.
[0199] A "means of recognition" is a mechanism that processes information to detect specific objects or patterns and understand their meaning.
[0200] "Data processing means" refers to a device or software that has the function of organizing acquired information and converting it into useful information.
[0201] "Analysis means" refers to a device or program that performs analysis to derive meaningful results from obtained data.
[0202] A "response generation means" is a function for constructing specific action guidelines and proposals based on the analyzed results.
[0203] "Communication means" refers to a technology or device that transmits digital information from one point to another.
[0204] The system that realizes this application example facilitates communication by determining the emotions of customers and staff in a store in real time and providing appropriate responses. The server receives customer voice and facial expression data via sound acquisition means and image acquisition means. Hardware devices such as Logitech cameras and Blue microphones are used for this. The received data is pre-processed by data processing means, such as noise reduction and format conversion, and emotion determination is performed by analysis means using software such as Amazon Rekognition or Microsoft® Azure® Face API.
[0205] The server generates specific response strategies for store staff based on the determined emotions. The response generation system combines technologies such as Google® Cloud Speech-to-Text and IBM Watson® to provide guidelines on how staff should interact with customers. The generated responses are then transmitted to staff members' smartphones via communication channels, thereby improving the customer experience.
[0206] For example, if a customer shows signs of anxiety in the store, the server analyzes the data and sends a notification to the staff member's smartphone with advice such as, "Let's recommend a calming product to this customer." This advice provides staff with quick information to help them decide what to do about the customer.
[0207] An example of a prompt would be, "Analyze the customer's emotional state from their facial expressions and voice data, and tell me specifically how to respond if the customer is feeling anxious." This establishes a process in which the generative AI model generates accurate advice.
[0208] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0209] Step 1:
[0210] The terminal acquires customer voice and facial expression data. It captures the customer's voice using an acoustic acquisition device and photographs their facial expressions using an image acquisition device. The acquired data is temporarily stored. The input is customer voice and image data, and the output is digital data that requires preprocessing.
[0211] Step 2:
[0212] The terminal sends the stored audio and image data to the server. The server performs data processing on the received data, including noise reduction and format standardization. The input is the raw data sent from the terminal, and the output is the data after noise has been removed, making it ready for analysis.
[0213] Step 3:
[0214] The server uses pre-processed data to perform emotion determination through analytical methods. It utilizes Amazon Rekognition and Google Cloud Speech-to-Text to analyze emotions from customer voice and facial expressions. The input is denoised data, and the output is the customer's emotional state (e.g., anxiety, reassurance).
[0215] Step 4:
[0216] The server uses a generative AI model to construct a response based on the customer's emotional state. It sends prompts to the generative AI model to generate an appropriate response. The input is the emotional state and the prompts to the generative AI, while the output is the specific response.
[0217] Step 5:
[0218] The server sends the generated response to the user (staff) via a communication method. The staff checks the response displayed on their smartphone and handles the customer interaction. The input is the generated response, and the output is the specific action plan that the staff should take.
[0219] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0220] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0221] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0222] [Second Embodiment]
[0223] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0224] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0225] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0226] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0227] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0228] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0229] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0230] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0231] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0232] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0233] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0234] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0235] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. This system mainly consists of three components: a terminal, a server, and a user.
[0236] Terminal operation
[0237] The device is placed around the baby, and in this case, it would be a smartphone or tablet. The device detects the baby's facial expressions with its built-in camera, vocalizations with its microphone, and body movements with its motion sensor. These sensors collect data in real time and temporarily store it on the device. The data is then pre-processed before being sent to the server. For example, if the baby is captured laughing, the camera captures an image of that moment, and the audio captures the laughter.
[0238] Server operation
[0239] The server plays a central role in receiving and analyzing data sent from the terminal. The received data is first preprocessed, including noise reduction and standardization. Next, the emotion determination engine uses algorithms for voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, data combining a smile and a high-pitched voice would be determined to be "happy." Based on this determination, the generating AI constructs appropriate advice for the parents. For example, it might generate advice such as, "Let's play with the baby."
[0240] User actions
[0241] The user's role is to review the information received on the device and use it to guide their next actions. The emotion assessment results and advice sent from the server are displayed on the device's interface and provided in a format that is easy for the user to understand. Based on the displayed information, the user can decide on specific actions. For example, by following the displayed advice and playing with the baby, the user can facilitate smoother communication between parent and child.
[0242] As described above, the system of the present invention supports childcare by determining the baby's emotions in real time and suggesting the most appropriate response to the parents.
[0243] The following describes the processing flow.
[0244] Step 1:
[0245] The device acquires audio data from the baby's surroundings using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. This data is temporarily stored in real time.
[0246] Step 2:
[0247] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0248] Step 3:
[0249] The terminal sends pre-processed data to the server. The data is securely transferred to the server via a security protocol.
[0250] Step 4:
[0251] The server receives audio, image, and motion data transmitted from the terminal and stores it in a database. This data is immediately used for analysis.
[0252] Step 5:
[0253] The server uses stored data to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, it analyzes a smiling image and a cheerful voice to determine that the baby is "happy."
[0254] Step 6:
[0255] Based on the analysis results, the server generates specific advice for parents using AI. For example, it might suggest, "Continue playing with your baby."
[0256] Step 7:
[0257] The server sends the judgment result and generated advice to the terminal.
[0258] Step 8:
[0259] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format that is easy for parents to understand.
[0260] Step 9:
[0261] The user uses the information displayed on the device as a reference to take specific actions for the baby. For example, if the device determines that the baby is sleepy, the user will take appropriate action to put the baby to sleep.
[0262] (Example 1)
[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0264] In modern childcare, it is crucial for parents to quickly and accurately understand their baby's emotions and condition and respond appropriately. However, many parents find it difficult to accurately grasp their baby's emotions amidst their busy daily lives, which can lead to them being unable to choose appropriate childcare methods. This invention aims to solve this problem and support childcare more effectively by determining the baby's emotions in real time and providing parents with the most suitable response.
[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0266] In this invention, the server includes an acquisition means for acquiring voice, an acquisition means for acquiring facial expressions, an acquisition means for acquiring body movements, a processing means for preprocessing and analyzing the acquired data, a determination means for determining emotions based on the preprocessed data, a generation means for generating specific countermeasures based on the determined emotions, a provision means for providing the generated countermeasures, a function for collecting data in real time using sensors at the terminal, temporarily storing and preprocessing it, and a function for using encryption technology in data transmission to realize secure data analysis on the server. This makes it possible to quickly analyze the emotional state of a baby and provide specific and practical advice to the guardian.
[0267] "Acquisition means" refers to a device or method for collecting data such as voice, facial expressions, or body movements.
[0268] "Processing means" refers to an apparatus or method for preprocessing acquired data and performing data analysis.
[0269] "Determination means" refers to a device or method for analyzing and determining emotions based on pre-processed data.
[0270] "Generating means" refers to a device or method for automatically creating specific countermeasures based on the determined emotions.
[0271] "Means of delivery" refers to a device or method for communicating the generated countermeasures to users such as guardians.
[0272] A "sensor" is a device that detects physical phenomena in the environment and outputs them in the form of electrical signals or other similar information.
[0273] "Encryption technology" is a technique that transforms data according to specific rules in order to protect it from unauthorized access.
[0274] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. The specific configuration and operation of the system are described below.
[0275] Terminal operation
[0276] The device takes the form of a smartphone or tablet and is placed near the baby. It has a built-in camera, microphone, and motion sensor, which are used to capture the baby's facial expressions, voice, and body movements in real time. Specifically, the camera can capture the baby's smile, the microphone can record laughter, and the motion sensor can detect arm and leg movements. The collected data is temporarily stored on the device and pre-processed before transmission.
[0277] Server operation
[0278] The server plays a central role in secure and efficient data processing. Data transmitted from the terminal undergoes preprocessing such as noise reduction and standardization, followed by analysis for emotion determination. The server uses a generative AI model to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. Specifically, if a smile or high-pitched laughter is detected, it determines that the baby is "happy" and generates advice based on this.
[0279] User actions
[0280] Users review the assessment results and advice through an interface provided on their device. Based on this information, users decide on their actual actions. For example, if they receive the advice, "Please play with your baby," they can immediately take appropriate action for playtime. As a result, richer communication between parents and children is promoted.
[0281] Example of a prompt
[0282] "There is data detecting a baby's smile and loud voice. What kind of emotions and what kind of advice should be provided?"
[0283] This invention is a system that highly analyzes a baby's emotional state and provides an easy-to-understand and effective child-rearing method for parents.
[0284] The flow of the specific process in Example 1 will be described using FIG. 11.
[0285] Step 1:
[0286] The terminal uses a camera, a microphone, and a motion sensor near the baby to acquire data on expressions, voices, and body movements in real time. The input is the natural movements and voices of the baby, which are acquired as digital data by the sensors. The output is the raw data obtained by each sensor.
[0287] Step 2:
[0288] The terminal temporarily stores the acquired raw data in internal memory and performs preprocessing for data transmission. In the preprocessing, the resolution of the image is adjusted and the noise of the voice data is removed. The input of this process is the raw data, and the output is the data for transmission with compression and less noise.
[0289] Step 3:
[0290] The terminal transmits the preprocessed data to the server. At this time, the data is encrypted and transferred securely. The input is the preprocessed data, and the output is the data reception on the server side. The data is transmitted using Wi-Fi or a mobile network.
[0291] Step 4:
[0292] The server preprocesses and standardizes the received data again. Here, the audio and image formats are unified, and the motion data is converted into a format that is easy to analyze. The input is the data sent from the terminal, and the output is standardized, ready-to-analyze data.
[0293] Step 5:
[0294] The server utilizes a generative AI model to perform voice analysis, image recognition, and motion analysis on standardized data to determine the baby's emotions. The input is standardized data, and the output is the determined emotion information. For example, if laughter and a smile match, the emotion "happy" is output.
[0295] Step 6:
[0296] The server uses a generative AI model to generate specific advice for parents based on the determined emotional information. The input is the determined emotional information, and the output is an advice statement. For example, the advice "Let's play with the baby" might be generated.
[0297] Step 7:
[0298] The server sends the generated advice to the terminal, which then notifies the user. The advice is displayed on the terminal screen in a format that is easy for the user to understand. The input is the advice text, and the output is the user's action based on their understanding.
[0299] (Application Example 1)
[0300] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0301] When making product recommendations according to the baby's situation in a physical store, there is a problem that it is difficult to perform accurate emotion judgment and provide optimal product information in real time. In addition, in order for store staff to provide appropriate advice based on the baby's condition, specialized knowledge is required, and there is a possibility of lacking consistency in customer service.
[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0303] In this invention, the server includes an input acquisition means for collecting the acquired voice data, expression data, and motion data of the baby, a data processing means for preprocessing and analyzing the data obtained from the input acquisition means, and an information recommendation means for determining the emotion of the baby based on the preprocessed data and providing product information recommended according to the determination result. Thereby, it becomes possible to provide real-time product recommendations according to the baby's situation and advice to parents in a physical store.
[0304] The "input acquisition means" is a device or method for collecting voice data, expression data, and motion data acquired from the baby.
[0305] The "data processing means" is a device or method for preprocessing and analyzing the data obtained from the input acquisition means. By this means, noise is removed and the data is standardized.
[0306] The "information recommendation means" is a device or method for determining the emotion of the baby based on the result analyzed by the data processing means and providing product information recommended according to the result.
[0307] The "information output means" is a device or method for providing information to the user terminal and used when selecting products recommended in a physical store.
[0308] The system necessary for implementing this invention is mainly composed of three elements: a user terminal, a server, and a physical store.
[0309] The user device is a smartphone or tablet, equipped with a camera, microphone, and motion sensor. The camera captures the baby's facial expressions, the microphone collects sound, and the motion sensor detects body movements. These sensors collect data in real time, temporarily store it on the user device, and then preprocess it before sending it to the server. Specifically, if the baby is smiling, the facial expression and sound are captured simultaneously.
[0310] The server plays a primary role in receiving and analyzing data sent from user terminals. The received data undergoes preprocessing such as noise reduction and standardization. Next, emotion determination is performed using algorithms for speech analysis, image recognition, and motion analysis. AI models such as TensorFlow are used for this. For example, if a baby's smile and high-pitched voice are detected simultaneously, it is determined to be "happy." Based on this result, a generative AI model creates specific advice for the parents, such as, "Your child is happy. Playing together with toys would be ideal."
[0311] In physical stores, these assessment results and advice are displayed on tablets or dedicated terminals, and staff members make appropriate product suggestions to parents. For example, if the system detects that the baby is crying, it will suggest products such as, "This product is recommended for times when babies tend to cry."
[0312] For example, if a baby becomes fussy when visiting a store, the system will determine in real time that the baby is feeling "anxious" and automatically suggest to the staff things like "a stuffed animal to calm the child."
[0313] An example of a prompt to pass to a generative AI model might be: "Please analyze the following data: baby's voice, baby's facial expression image, and high-frequency sound. Next, determine what emotion the baby is feeling and suggest specific products to recommend to the parents."
[0314] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0315] Step 1:
[0316] The user collects the baby's facial expressions, voice, and body movements using the user's device's camera, microphone, and motion sensor. The input for this step is raw data of the baby, and the output is data temporarily stored on the device. Specifically, if the baby is laughing, the facial expression data and laughter are captured on the smartphone.
[0317] Step 2:
[0318] The terminal preprocesses the collected data by denoising and standardizing it, preparing it for transmission to the server. The input for this step is the raw data within the terminal, and the output is the preprocessed data. For example, background noise in audio may be removed, and the brightness of images may be adjusted.
[0319] Step 3:
[0320] The server receives preprocessed data sent from the user's terminal and performs speech analysis, image recognition, and motion analysis. The input for this step is preprocessed data, and the output is information about the baby's emotions. Specifically, TensorFlow is used to determine if the baby is "happy" from their smile and laughter.
[0321] Step 4:
[0322] Based on the emotion assessment results, the server uses a generative AI model to create specific advice for the parents. The input for this step is the emotion assessment results, and the output is the generated advice text. For example, advice such as, "The baby is having fun. Playing together with toys is ideal," is generated.
[0323] Step 5:
[0324] Users can choose the most suitable product for their baby based on information displayed on a terminal or tablet in a physical store. The input in this step is generated advice, and the output is product information for the parent to choose. For example, information such as "We recommend this product during times when your baby is prone to crying" is presented to help the parent make their selection.
[0325] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0326] This invention combines a system that acquires a baby's voice, facial expressions, and body movements, and determines the baby's emotions based on that data, with an emotion engine that recognizes the emotions of the user (the caregiver). The system operates through interaction between the terminal, server, and user.
[0327] Terminal operation
[0328] The device is placed near the baby and uses a microphone, camera, and motion sensor to capture voice, facial expressions, and movements. In addition, the device is equipped with sensors to capture the user's facial expressions and voice, collecting real-time emotional data. This data is temporarily stored and pre-processed before being sent to the server.
[0329] Server operation
[0330] The server receives voice, image, and motion data from the baby and user sent from the terminal, and then performs preprocessing such as noise reduction and data format standardization. Next, it uses an emotion determination engine on the baby's data to perform voice analysis, image recognition, and motion analysis. This determines the baby's emotions, such as "happy," "anxious," or "sleepy." The user's data is separately analyzed for emotions to determine whether the caregiver is in a state of "stress" or "security."
[0331] The server uses generative AI to design specific actions parents should take, based on the emotional data of both the baby and the user. The advice is dynamically adjusted; for example, if the user is feeling stressed, relaxation advice may be added. If the baby is anxious and the user is also feeling stressed, a more comprehensive piece of advice might be generated, such as, "Hold your baby to calm them down and take some deep breaths."
[0332] User actions
[0333] Users review the emotion assessment results and advice displayed on their device. Based on the displayed information, users can take appropriate actions towards their baby. The results and advice are presented through concise and easy-to-understand notifications and a visual interface. This helps parents reduce anxiety in childcare and enables them to respond appropriately to emotions.
[0334] In this way, the present invention can improve the childcare environment by simultaneously determining the emotions of both the baby and the parent and providing appropriate countermeasures. This is expected to promote better communication between parents and children.
[0335] The following describes the processing flow.
[0336] Step 1:
[0337] The device acquires audio data from the baby's environment using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. Similarly, it also acquires the user's voice and facial expressions using sensors. This data is temporarily stored in real time.
[0338] Step 2:
[0339] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0340] Step 3:
[0341] The device sends pre-processed baby and user data to the server. A secure protocol is used for data transfer.
[0342] Step 4:
[0343] The server receives data sent from the terminal and stores it in the database. This data is immediately used for analysis.
[0344] Step 5:
[0345] The server performs voice analysis, image recognition, and motion analysis on the baby's data to determine the baby's emotions. For example, it might determine that the baby is "happy" based on a combination of a smiling image and a high-pitched voice.
[0346] Step 6:
[0347] Similarly, the server analyzes user data to determine the user's emotions. For example, it might determine that the user is stressed based on a tired tone of voice and an anxious facial expression.
[0348] Step 7:
[0349] Based on the emotional information of both the baby and the user, the server uses a generating AI to create specific advice for the parents. The user's emotional state is also taken into consideration, so advice such as, "Speak gently to your baby and also take time to calm yourself down," is generated as needed.
[0350] Step 8:
[0351] The server sends the judgment result and advice to the terminal.
[0352] Step 9:
[0353] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format, taking care to ensure that parents can easily understand and act upon the advice.
[0354] Step 10:
[0355] Based on the information displayed on the device, users take specific actions for their baby while also managing their own state of mind. For example, they might incorporate relaxation techniques to reduce their own stress while soothing their baby.
[0356] (Example 2)
[0357] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0358] There is a need to accurately assess a baby's emotions and, based on that assessment, quickly and precisely provide parents with guidance on how to respond. Conventional systems have been insufficient in assessing babies' emotions and analyzing parents' mental states, and have limited ability to generate specific countermeasures based on those results. Therefore, these issues need to be addressed.
[0359] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0360] In this invention, the server includes means for acquiring sound, means for acquiring images, and means for acquiring movement. This makes it possible to comprehensively determine the baby's emotions from various data such as voice, facial expressions, and movement, and further analyze the caregiver's mental state, in order to provide dynamically adjusted specific countermeasures.
[0361] "Means of acquiring sound" refers to technologies for collecting ambient sounds, particularly those used to detect sounds like a baby's voice.
[0362] "Means for acquiring images" refers to methods that use visual sensors such as cameras to capture images of the surroundings, and are used in particular to analyze the baby's facial expressions.
[0363] "Means for acquiring motion" refers to methods that use motion sensors or similar devices to detect body movements, and are particularly used to detect patterns in a baby's movements.
[0364] "Processing means" refers to a system that has the function of organizing the acquired data and converting it into a format suitable for analysis.
[0365] "Determination means" refers to algorithms and technologies used to analyze pre-processed data and evaluate specific states such as emotions.
[0366] "Additional emotion analysis means" refers to technology that evaluates a user's mental state from their speech patterns and facial expressions to gain a deeper understanding of their emotions.
[0367] "Generative means" refers to the technology used to construct and generate specific countermeasures and advice based on the analysis results.
[0368] "Information provision means" refers to screens, notification systems, etc., that present generated advice and information to users.
[0369] This invention is a system implemented through terminals placed around a baby, a server operating in the backend, and a user receiving information. Specific embodiments are described below.
[0370] Terminal configuration and operation
[0371] The device is equipped with a microphone to capture sound, a camera to capture images, and a motion sensor to capture movement. This device is placed near the baby and collects voice, facial expressions, and movements in real time. For example, if the baby is smiling, the camera captures the smile, and the motion sensor detects the baby's active movements. This data is temporarily stored for further processing and then transmitted to a server via the network.
[0372] Server configuration and operation
[0373] The server receives data transmitted from the terminal and performs preprocessing on it. For example, it performs noise reduction on audio data and standardizes the image data format. Next, the server analyzes the data using emotion determination tools. The analysis includes audio analysis, image recognition, and motion analysis, and determines the baby's emotions as "happy," "anxious," "sleepy," etc. At the same time, the server uses additional emotion analysis tools to evaluate the user's state as "stressed" or "at ease."
[0374] Furthermore, the server utilizes generative AI based on these analysis results to suggest specific actions the user should take. For example, if the user is stressed and the baby is anxious, it might recommend "hugging the baby and taking deep breaths together." The generative AI model incorporates historical data and expertise, providing sophisticated solutions.
[0375] An example of a prompt might be, "Please suggest ways to cope when a baby is crying and the caregiver is feeling stressed."
[0376] User actions
[0377] Users review the analysis results and advice displayed on their device. The information is presented in an easy-to-understand interface, enabling quick and accurate responses. This allows users to take appropriate action and reduce anxiety related to childcare.
[0378] This invention's system allows for the simultaneous evaluation of the baby's emotions and the user's mental state, contributing to better communication between parent and child.
[0379] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0380] Step 1:
[0381] The device collects the baby's voice, facial expressions, and movements. Inputs include sounds, images, and movements from the baby's surroundings. Specifically, a microphone picks up voice, a camera captures facial expressions, and a motion sensor detects movement. This data is temporarily stored within the device.
[0382] Step 2:
[0383] The terminal preprocesses the collected data. Because the input data may contain noise or incorrect formatting, the audio data is denoised, and the image data is resized and converted before being sent to the server. The preprocessed data is then output.
[0384] Step 3:
[0385] The terminal sends pre-processed data to the server. The specific data sent includes de-noise-removed audio, facial expression images converted to an appropriate format, and motion information.
[0386] Step 4:
[0387] The server analyzes the received data. It receives pre-processed audio, image, and motion data as input and uses an emotion determination engine to determine the baby's emotions. The specific output is an emotional state such as "happy," "anxious," or "sleepy." Simultaneously, it performs additional emotion analysis to evaluate the user's state.
[0388] Step 5:
[0389] Based on the assessment results, the server generates a response plan for the parent using a generated AI model. The prompt text is a specific situation corresponding to the baby's and the user's state, and specific advice is output based on that. For example, advice such as "Hold your baby and take a deep breath together" might be generated.
[0390] Step 6:
[0391] The user reviews the analysis results and generated advice displayed on the device. The input is the information displayed on the device's screen, and the user uses this to take appropriate action for their baby. This reduces anxiety about childcare.
[0392] (Application Example 2)
[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0394] Conventional baby emotion assessment systems only consider the baby's emotions, making it difficult to obtain specific countermeasures that take into account the emotions of parents or store staff. Furthermore, there was a lack of mechanisms to understand the emotions of both customers and staff and provide appropriate countermeasures, which is crucial for smooth communication with customers in physical stores.
[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0396] In this invention, the server includes an acoustic acquisition means for acquiring the voices of the baby and the customer, an image acquisition means for acquiring the facial expressions of the baby and the customer, and a data processing means for pre-processing and analyzing the acquired information. This makes it possible to determine the emotions of both the baby and the customer, as well as the parent or store staff, in real time and provide specific and effective countermeasures.
[0397] "Sound acquisition means" refers to a device that has the function of detecting and recording ambient sounds.
[0398] An "image acquisition means" is a device that has the function of acquiring visual information and converting it into digital data.
[0399] A "motion sensor means" is a sensor that can detect the movement of an object or human body and record the changes as data.
[0400] A "means of recognition" is a mechanism that processes information to detect specific objects or patterns and understand their meaning.
[0401] "Data processing means" refers to a device or software that has the function of organizing acquired information and converting it into useful information.
[0402] "Analysis means" refers to a device or program that performs analysis to derive meaningful results from obtained data.
[0403] A "response generation means" is a function for constructing specific action guidelines and proposals based on the analyzed results.
[0404] "Communication means" refers to a technology or device that transmits digital information from one point to another.
[0405] The system that implements this application facilitates communication by determining the emotions of customers and staff in a store in real time and providing appropriate responses. The server receives customer voice and facial expression data via acoustic and image acquisition means. Hardware devices such as Logitech cameras and Blue microphones are used for this purpose. The received data is pre-processed by data processing means, such as noise reduction and format conversion, and emotion determination is performed by analysis means using software such as Amazon Rekognition or Microsoft Azure Face API.
[0406] The server generates specific response strategies for store staff based on the determined emotions. The response generation system combines technologies such as Google Cloud Speech-to-Text and IBM Watson to provide guidelines on how staff should interact with customers. The generated responses are then transmitted to staff members' smartphones via communication channels, thereby improving the customer experience.
[0407] For example, if a customer shows signs of anxiety in the store, the server analyzes the data and sends a notification to the staff member's smartphone with advice such as, "Let's recommend a calming product to this customer." This advice provides staff with quick information to help them decide what to do about the customer.
[0408] An example of a prompt would be, "Analyze the customer's emotional state from their facial expressions and voice data, and tell me specifically how to respond if the customer is feeling anxious." This establishes a process in which the generative AI model generates accurate advice.
[0409] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0410] Step 1:
[0411] The terminal acquires customer voice and facial expression data. It captures the customer's voice using an acoustic acquisition device and photographs their facial expressions using an image acquisition device. The acquired data is temporarily stored. The input is customer voice and image data, and the output is digital data that requires preprocessing.
[0412] Step 2:
[0413] The terminal sends the stored audio and image data to the server. The server performs data processing on the received data, including noise reduction and format standardization. The input is the raw data sent from the terminal, and the output is the data after noise has been removed, making it ready for analysis.
[0414] Step 3:
[0415] The server uses pre-processed data to perform emotion determination through analytical methods. It utilizes Amazon Rekognition and Google Cloud Speech-to-Text to analyze emotions from customer voice and facial expressions. The input is denoised data, and the output is the customer's emotional state (e.g., anxiety, reassurance).
[0416] Step 4:
[0417] The server uses a generative AI model to construct a response based on the customer's emotional state. It sends prompts to the generative AI model to generate an appropriate response. The input is the emotional state and the prompts to the generative AI, while the output is the specific response.
[0418] Step 5:
[0419] The server sends the generated response to the user (staff) via a communication method. The staff checks the response displayed on their smartphone and handles the customer interaction. The input is the generated response, and the output is the specific action plan that the staff should take.
[0420] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0421] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0422] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0423] [Third Embodiment]
[0424] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0425] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0426] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0427] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0428] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0430] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0431] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0432] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0433] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0434] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0435] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0436] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. This system mainly consists of three components: a terminal, a server, and a user.
[0437] Terminal operation
[0438] The device is placed around the baby, and in this case, it would be a smartphone or tablet. The device detects the baby's facial expressions with its built-in camera, vocalizations with its microphone, and body movements with its motion sensor. These sensors collect data in real time and temporarily store it on the device. The data is then pre-processed before being sent to the server. For example, if the baby is captured laughing, the camera captures an image of that moment, and the audio captures the laughter.
[0439] Server operation
[0440] The server plays a central role in receiving and analyzing data sent from the terminal. The received data is first preprocessed, including noise reduction and standardization. Next, the emotion determination engine uses algorithms for voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, data combining a smile and a high-pitched voice would be determined to be "happy." Based on this determination, the generating AI constructs appropriate advice for the parents. For example, it might generate advice such as, "Let's play with the baby."
[0441] User actions
[0442] The user's role is to review the information received on the device and use it to guide their next actions. The emotion assessment results and advice sent from the server are displayed on the device's interface and provided in a format that is easy for the user to understand. Based on the displayed information, the user can decide on specific actions. For example, by following the displayed advice and playing with the baby, the user can facilitate smoother communication between parent and child.
[0443] As described above, the system of the present invention supports childcare by determining the baby's emotions in real time and suggesting the most appropriate response to the parents.
[0444] The following describes the processing flow.
[0445] Step 1:
[0446] The device acquires audio data from the baby's surroundings using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. This data is temporarily stored in real time.
[0447] Step 2:
[0448] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0449] Step 3:
[0450] The terminal sends pre-processed data to the server. The data is securely transferred to the server via a security protocol.
[0451] Step 4:
[0452] The server receives audio, image, and motion data transmitted from the terminal and stores it in a database. This data is immediately used for analysis.
[0453] Step 5:
[0454] The server uses stored data to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, it analyzes a smiling image and a cheerful voice to determine that the baby is "happy."
[0455] Step 6:
[0456] Based on the analysis results, the server generates specific advice for parents using AI. For example, it might suggest, "Continue playing with your baby."
[0457] Step 7:
[0458] The server sends the judgment result and generated advice to the terminal.
[0459] Step 8:
[0460] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format that is easy for parents to understand.
[0461] Step 9:
[0462] The user uses the information displayed on the device as a reference to take specific actions for the baby. For example, if the device determines that the baby is sleepy, the user will take appropriate action to put the baby to sleep.
[0463] (Example 1)
[0464] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] In modern childcare, it is crucial for parents to quickly and accurately understand their baby's emotions and condition and respond appropriately. However, many parents find it difficult to accurately grasp their baby's emotions amidst their busy daily lives, which can lead to them being unable to choose appropriate childcare methods. This invention aims to solve this problem and support childcare more effectively by determining the baby's emotions in real time and providing parents with the most suitable response.
[0466] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0467] In this invention, the server includes an acquisition means for acquiring voice, an acquisition means for acquiring facial expressions, an acquisition means for acquiring body movements, a processing means for preprocessing and analyzing the acquired data, a determination means for determining emotions based on the preprocessed data, a generation means for generating specific countermeasures based on the determined emotions, a provision means for providing the generated countermeasures, a function for collecting data in real time using sensors at the terminal, temporarily storing and preprocessing it, and a function for using encryption technology in data transmission to realize secure data analysis on the server. This makes it possible to quickly analyze the emotional state of a baby and provide specific and practical advice to the guardian.
[0468] "Acquisition means" refers to a device or method for collecting data such as voice, facial expressions, or body movements.
[0469] "Processing means" refers to an apparatus or method for preprocessing acquired data and performing data analysis.
[0470] "Determination means" refers to a device or method for analyzing and determining emotions based on pre-processed data.
[0471] "Generating means" refers to a device or method for automatically creating specific countermeasures based on the determined emotions.
[0472] "Means of delivery" refers to a device or method for communicating the generated countermeasures to users such as guardians.
[0473] A "sensor" is a device that detects physical phenomena in the environment and outputs them in the form of electrical signals or other similar information.
[0474] "Encryption technology" is a technique that transforms data according to specific rules in order to protect it from unauthorized access.
[0475] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. The specific configuration and operation of the system are described below.
[0476] Terminal operation
[0477] The device takes the form of a smartphone or tablet and is placed near the baby. It has a built-in camera, microphone, and motion sensor, which are used to capture the baby's facial expressions, voice, and body movements in real time. Specifically, the camera can capture the baby's smile, the microphone can record laughter, and the motion sensor can detect arm and leg movements. The collected data is temporarily stored on the device and pre-processed before transmission.
[0478] Server operation
[0479] The server plays a central role in secure and efficient data processing. Data transmitted from the terminal undergoes preprocessing such as noise reduction and standardization, followed by analysis for emotion determination. The server uses a generative AI model to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. Specifically, if a smile or high-pitched laughter is detected, it determines that the baby is "happy" and generates advice based on this.
[0480] User actions
[0481] Users review the assessment results and advice through an interface provided on their device. Based on this information, users decide on their actual actions. For example, if they receive the advice, "Please play with your baby," they can immediately take appropriate action for playtime. As a result, richer communication between parents and children is promoted.
[0482] Example of a prompt
[0483] "We have data that detects a baby's smile and high-pitched voice. What emotions does this represent, and what kind of advice should we offer?"
[0484] This invention is a system that highly analyzes a baby's emotional state and provides parents with an easy-to-understand and effective childcare method.
[0485] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0486] Step 1:
[0487] The device uses a camera, microphone, and motion sensor near the baby to acquire data on facial expressions, voice, and body movements in real time. The input is the baby's natural movements and voice, which are acquired as digital data by the sensors. The output is the raw data obtained from each sensor.
[0488] Step 2:
[0489] The terminal temporarily stores the acquired raw data in its internal memory and performs preprocessing for data transmission. This preprocessing adjusts the image resolution and removes noise from audio data. The input to this process is raw data, and the output is compressed and low-noise data for transmission.
[0490] Step 3:
[0491] The terminal sends pre-processed data to the server. During this process, the data is encrypted and transmitted securely. The input is the pre-processed data, and the output is the data received on the server side. The data is transmitted using Wi-Fi or a mobile network.
[0492] Step 4:
[0493] The server preprocesses and standardizes the received data again. Here, the audio and image formats are unified, and the motion data is converted into a format that is easy to analyze. The input is the data sent from the terminal, and the output is standardized, ready-to-analyze data.
[0494] Step 5:
[0495] The server utilizes a generative AI model to perform voice analysis, image recognition, and motion analysis on standardized data to determine the baby's emotions. The input is standardized data, and the output is the determined emotion information. For example, if laughter and a smile match, the emotion "happy" is output.
[0496] Step 6:
[0497] The server uses a generative AI model to generate specific advice for parents based on the determined emotional information. The input is the determined emotional information, and the output is an advice statement. For example, the advice "Let's play with the baby" might be generated.
[0498] Step 7:
[0499] The server sends the generated advice to the terminal, which then notifies the user. The advice is displayed on the terminal screen in a format that is easy for the user to understand. The input is the advice text, and the output is the user's action based on their understanding.
[0500] (Application Example 1)
[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] When providing product recommendations tailored to a baby's condition in a physical store, there are challenges in accurately assessing emotions in real time and providing optimal product information. Furthermore, providing appropriate advice based on the baby's condition requires specialized knowledge from store staff, potentially leading to a lack of consistency in customer service.
[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0504] In this invention, the server includes an input acquisition means for collecting acquired baby voice data, facial expression data, and motion data; a data processing means for preprocessing and analyzing the data obtained from the input acquisition means; and an information recommendation means for determining the baby's emotions based on the preprocessed data and providing recommended product information according to the determination result. This makes it possible to provide real-time product suggestions and advice to parents in a physical store, tailored to the baby's situation.
[0505] "Input acquisition means" refers to devices and methods for collecting voice data, facial expression data, and motion data obtained from a baby.
[0506] "Data processing means" refers to devices and methods for preprocessing and analyzing data obtained from input acquisition means. This means removes noise and standardizes the data.
[0507] An "information recommendation means" is a device or method that determines a baby's emotions based on the results analyzed by a data processing means and provides product information recommended according to those results.
[0508] "Information output means" refers to devices or methods that provide information to user terminals and are used when selecting recommended products in physical stores.
[0509] The system necessary to implement this invention mainly consists of three elements: a user terminal, a server, and a physical store.
[0510] The user device is a smartphone or tablet, equipped with a camera, microphone, and motion sensor. The camera captures the baby's facial expressions, the microphone collects sound, and the motion sensor detects body movements. These sensors collect data in real time, temporarily store it on the user device, and then preprocess it before sending it to the server. Specifically, if the baby is smiling, the facial expression and sound are captured simultaneously.
[0511] The server plays a primary role in receiving and analyzing data sent from user terminals. The received data undergoes preprocessing such as noise reduction and standardization. Next, emotion determination is performed using algorithms for speech analysis, image recognition, and motion analysis. AI models such as TensorFlow are used for this. For example, if a baby's smile and high-pitched voice are detected simultaneously, it is determined to be "happy." Based on this result, a generative AI model creates specific advice for the parents, such as, "Your child is happy. Playing together with toys would be ideal."
[0512] In physical stores, these assessment results and advice are displayed on tablets or dedicated terminals, and staff members make appropriate product suggestions to parents. For example, if the system detects that the baby is crying, it will suggest products such as, "This product is recommended for times when babies tend to cry."
[0513] For example, if a baby becomes fussy when visiting a store, the system will determine in real time that the baby is feeling "anxious" and automatically suggest to the staff things like "a stuffed animal to calm the child."
[0514] An example of a prompt to pass to a generative AI model might be: "Please analyze the following data: baby's voice, baby's facial expression image, and high-frequency sound. Next, determine what emotion the baby is feeling and suggest specific products to recommend to the parents."
[0515] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0516] Step 1:
[0517] The user collects the baby's facial expressions, voice, and body movements using the user's device's camera, microphone, and motion sensor. The input for this step is raw data of the baby, and the output is data temporarily stored on the device. Specifically, if the baby is laughing, the facial expression data and laughter are captured on the smartphone.
[0518] Step 2:
[0519] The terminal preprocesses the collected data by denoising and standardizing it, preparing it for transmission to the server. The input for this step is the raw data within the terminal, and the output is the preprocessed data. For example, background noise in audio may be removed, and the brightness of images may be adjusted.
[0520] Step 3:
[0521] The server receives preprocessed data sent from the user's terminal and performs speech analysis, image recognition, and motion analysis. The input for this step is preprocessed data, and the output is information about the baby's emotions. Specifically, TensorFlow is used to determine if the baby is "happy" from their smile and laughter.
[0522] Step 4:
[0523] Based on the emotion assessment results, the server uses a generative AI model to create specific advice for the parents. The input for this step is the emotion assessment results, and the output is the generated advice text. For example, advice such as, "The baby is having fun. Playing together with toys is ideal," is generated.
[0524] Step 5:
[0525] Users can choose the most suitable product for their baby based on information displayed on a terminal or tablet in a physical store. The input in this step is generated advice, and the output is product information for the parent to choose. For example, information such as "We recommend this product during times when your baby is prone to crying" is presented to help the parent make their selection.
[0526] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0527] This invention combines a system that acquires a baby's voice, facial expressions, and body movements, and determines the baby's emotions based on that data, with an emotion engine that recognizes the emotions of the user (the caregiver). The system operates through interaction between the terminal, server, and user.
[0528] Terminal operation
[0529] The device is placed near the baby and uses a microphone, camera, and motion sensor to capture voice, facial expressions, and movements. In addition, the device is equipped with sensors to capture the user's facial expressions and voice, collecting real-time emotional data. This data is temporarily stored and pre-processed before being sent to the server.
[0530] Server operation
[0531] The server receives voice, image, and motion data from the baby and user sent from the terminal, and then performs preprocessing such as noise reduction and data format standardization. Next, it uses an emotion determination engine on the baby's data to perform voice analysis, image recognition, and motion analysis. This determines the baby's emotions, such as "happy," "anxious," or "sleepy." The user's data is separately analyzed for emotions to determine whether the caregiver is in a state of "stress" or "security."
[0532] The server uses generative AI to design specific actions parents should take, based on the emotional data of both the baby and the user. The advice is dynamically adjusted; for example, if the user is feeling stressed, relaxation advice may be added. If the baby is anxious and the user is also feeling stressed, a more comprehensive piece of advice might be generated, such as, "Hold your baby to calm them down and take some deep breaths."
[0533] User actions
[0534] Users review the emotion assessment results and advice displayed on their device. Based on the displayed information, users can take appropriate actions towards their baby. The results and advice are presented through concise and easy-to-understand notifications and a visual interface. This helps parents reduce anxiety in childcare and enables them to respond appropriately to emotions.
[0535] In this way, the present invention can improve the childcare environment by simultaneously determining the emotions of both the baby and the parent and providing appropriate countermeasures. This is expected to promote better communication between parents and children.
[0536] The following describes the processing flow.
[0537] Step 1:
[0538] The device acquires audio data from the baby's environment using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. Similarly, it also acquires the user's voice and facial expressions using sensors. This data is temporarily stored in real time.
[0539] Step 2:
[0540] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0541] Step 3:
[0542] The device sends pre-processed baby and user data to the server. A secure protocol is used for data transfer.
[0543] Step 4:
[0544] The server receives data sent from the terminal and stores it in the database. This data is immediately used for analysis.
[0545] Step 5:
[0546] The server performs voice analysis, image recognition, and motion analysis on the baby's data to determine the baby's emotions. For example, it might determine that the baby is "happy" based on a combination of a smiling image and a high-pitched voice.
[0547] Step 6:
[0548] Similarly, the server analyzes user data to determine the user's emotions. For example, it might determine that the user is stressed based on a tired tone of voice and an anxious facial expression.
[0549] Step 7:
[0550] Based on the emotional information of both the baby and the user, the server uses a generating AI to create specific advice for the parents. The user's emotional state is also taken into consideration, so advice such as, "Speak gently to your baby and also take time to calm yourself down," is generated as needed.
[0551] Step 8:
[0552] The server sends the judgment result and advice to the terminal.
[0553] Step 9:
[0554] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format, taking care to ensure that parents can easily understand and act upon the advice.
[0555] Step 10:
[0556] Based on the information displayed on the device, users take specific actions for their baby while also managing their own state of mind. For example, they might incorporate relaxation techniques to reduce their own stress while soothing their baby.
[0557] (Example 2)
[0558] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0559] There is a need to accurately assess a baby's emotions and, based on that assessment, quickly and precisely provide parents with guidance on how to respond. Conventional systems have been insufficient in assessing babies' emotions and analyzing parents' mental states, and have limited ability to generate specific countermeasures based on those results. Therefore, these issues need to be addressed.
[0560] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0561] In this invention, the server includes means for acquiring sound, means for acquiring images, and means for acquiring movement. This makes it possible to comprehensively determine the baby's emotions from various data such as voice, facial expressions, and movement, and further analyze the caregiver's mental state, in order to provide dynamically adjusted specific countermeasures.
[0562] "Means of acquiring sound" refers to technologies for collecting ambient sounds, particularly those used to detect sounds like a baby's voice.
[0563] "Means for acquiring images" refers to methods that use visual sensors such as cameras to capture images of the surroundings, and are used in particular to analyze the baby's facial expressions.
[0564] "Means for acquiring motion" refers to methods that use motion sensors or similar devices to detect body movements, and are particularly used to detect patterns in a baby's movements.
[0565] "Processing means" refers to a system that has the function of organizing the acquired data and converting it into a format suitable for analysis.
[0566] "Determination means" refers to algorithms and technologies used to analyze pre-processed data and evaluate specific states such as emotions.
[0567] "Additional emotion analysis means" refers to technology that evaluates a user's mental state from their speech patterns and facial expressions to gain a deeper understanding of their emotions.
[0568] "Generative means" refers to the technology used to construct and generate specific countermeasures and advice based on the analysis results.
[0569] "Information provision means" refers to screens, notification systems, etc., that present generated advice and information to users.
[0570] This invention is a system implemented through terminals placed around a baby, a server operating in the backend, and a user receiving information. Specific embodiments are described below.
[0571] Terminal configuration and operation
[0572] The device is equipped with a microphone to capture sound, a camera to capture images, and a motion sensor to capture movement. This device is placed near the baby and collects voice, facial expressions, and movements in real time. For example, if the baby is smiling, the camera captures the smile, and the motion sensor detects the baby's active movements. This data is temporarily stored for further processing and then transmitted to a server via the network.
[0573] Server configuration and operation
[0574] The server receives data transmitted from the terminal and performs preprocessing on it. For example, it performs noise reduction on audio data and standardizes the image data format. Next, the server analyzes the data using emotion determination tools. The analysis includes audio analysis, image recognition, and motion analysis, and determines the baby's emotions as "happy," "anxious," "sleepy," etc. At the same time, the server uses additional emotion analysis tools to evaluate the user's state as "stressed" or "at ease."
[0575] Furthermore, the server utilizes generative AI based on these analysis results to suggest specific actions the user should take. For example, if the user is stressed and the baby is anxious, it might recommend "hugging the baby and taking deep breaths together." The generative AI model incorporates historical data and expertise, providing sophisticated solutions.
[0576] An example of a prompt might be, "Please suggest ways to cope when a baby is crying and the caregiver is feeling stressed."
[0577] User actions
[0578] Users review the analysis results and advice displayed on their device. The information is presented in an easy-to-understand interface, enabling quick and accurate responses. This allows users to take appropriate action and reduce anxiety related to childcare.
[0579] This invention's system allows for the simultaneous evaluation of the baby's emotions and the user's mental state, contributing to better communication between parent and child.
[0580] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0581] Step 1:
[0582] The device collects the baby's voice, facial expressions, and movements. Inputs include sounds, images, and movements from the baby's surroundings. Specifically, a microphone picks up voice, a camera captures facial expressions, and a motion sensor detects movement. This data is temporarily stored within the device.
[0583] Step 2:
[0584] The terminal preprocesses the collected data. Because the input data may contain noise or incorrect formatting, the audio data is denoised, and the image data is resized and converted before being sent to the server. The preprocessed data is then output.
[0585] Step 3:
[0586] The terminal sends pre-processed data to the server. The specific data sent includes de-noise-removed audio, facial expression images converted to an appropriate format, and motion information.
[0587] Step 4:
[0588] The server analyzes the received data. It receives pre-processed audio, image, and motion data as input and uses an emotion determination engine to determine the baby's emotions. The specific output is an emotional state such as "happy," "anxious," or "sleepy." Simultaneously, it performs additional emotion analysis to evaluate the user's state.
[0589] Step 5:
[0590] Based on the assessment results, the server generates a response plan for the parent using a generated AI model. The prompt text is a specific situation corresponding to the baby's and the user's state, and specific advice is output based on that. For example, advice such as "Hold your baby and take a deep breath together" might be generated.
[0591] Step 6:
[0592] The user reviews the analysis results and generated advice displayed on the device. The input is the information displayed on the device's screen, and the user uses this to take appropriate action for their baby. This reduces anxiety about childcare.
[0593] (Application Example 2)
[0594] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0595] Conventional baby emotion assessment systems only consider the baby's emotions, making it difficult to obtain specific countermeasures that take into account the emotions of parents or store staff. Furthermore, there was a lack of mechanisms to understand the emotions of both customers and staff and provide appropriate countermeasures, which is crucial for smooth communication with customers in physical stores.
[0596] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0597] In this invention, the server includes an acoustic acquisition means for acquiring the voices of the baby and the customer, an image acquisition means for acquiring the facial expressions of the baby and the customer, and a data processing means for pre-processing and analyzing the acquired information. This makes it possible to determine the emotions of both the baby and the customer, as well as the parent or store staff, in real time and provide specific and effective countermeasures.
[0598] "Sound acquisition means" refers to a device that has the function of detecting and recording ambient sounds.
[0599] An "image acquisition means" is a device that has the function of acquiring visual information and converting it into digital data.
[0600] A "motion sensor means" is a sensor that can detect the movement of an object or human body and record the changes as data.
[0601] A "means of recognition" is a mechanism that processes information to detect specific objects or patterns and understand their meaning.
[0602] "Data processing means" refers to a device or software that has the function of organizing acquired information and converting it into useful information.
[0603] "Analysis means" refers to a device or program that performs analysis to derive meaningful results from obtained data.
[0604] A "response generation means" is a function for constructing specific action guidelines and proposals based on the analyzed results.
[0605] "Communication means" refers to a technology or device that transmits digital information from one point to another.
[0606] The system that implements this application facilitates communication by determining the emotions of customers and staff in a store in real time and providing appropriate responses. The server receives customer voice and facial expression data via acoustic and image acquisition means. Hardware devices such as Logitech cameras and Blue microphones are used for this purpose. The received data is pre-processed by data processing means, such as noise reduction and format conversion, and emotion determination is performed by analysis means using software such as Amazon Rekognition or Microsoft Azure Face API.
[0607] The server generates specific response strategies for store staff based on the determined emotions. The response generation system combines technologies such as Google Cloud Speech-to-Text and IBM Watson to provide guidelines on how staff should interact with customers. The generated responses are then transmitted to staff members' smartphones via communication channels, thereby improving the customer experience.
[0608] For example, if a customer shows signs of anxiety in the store, the server analyzes the data and sends a notification to the staff member's smartphone with advice such as, "Let's recommend a calming product to this customer." This advice provides staff with quick information to help them decide what to do about the customer.
[0609] An example of a prompt would be, "Analyze the customer's emotional state from their facial expressions and voice data, and tell me specifically how to respond if the customer is feeling anxious." This establishes a process in which the generative AI model generates accurate advice.
[0610] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0611] Step 1:
[0612] The terminal acquires customer voice and facial expression data. It captures the customer's voice using an acoustic acquisition device and photographs their facial expressions using an image acquisition device. The acquired data is temporarily stored. The input is customer voice and image data, and the output is digital data that requires preprocessing.
[0613] Step 2:
[0614] The terminal sends the stored audio and image data to the server. The server performs data processing on the received data, including noise reduction and format standardization. The input is the raw data sent from the terminal, and the output is the data after noise has been removed, making it ready for analysis.
[0615] Step 3:
[0616] The server uses pre-processed data to perform emotion determination through analytical methods. It utilizes Amazon Rekognition and Google Cloud Speech-to-Text to analyze emotions from customer voice and facial expressions. The input is denoised data, and the output is the customer's emotional state (e.g., anxiety, reassurance).
[0617] Step 4:
[0618] The server uses a generative AI model to construct a response based on the customer's emotional state. It sends prompts to the generative AI model to generate an appropriate response. The input is the emotional state and the prompts to the generative AI, while the output is the specific response.
[0619] Step 5:
[0620] The server sends the generated response to the user (staff) via a communication method. The staff checks the response displayed on their smartphone and handles the customer interaction. The input is the generated response, and the output is the specific action plan that the staff should take.
[0621] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0622] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0623] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0624] [Fourth Embodiment]
[0625] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0626] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0627] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0628] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0629] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0630] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0631] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0632] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0633] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0634] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0635] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0636] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0637] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0638] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. This system mainly consists of three components: a terminal, a server, and a user.
[0639] Terminal operation
[0640] The device is placed around the baby, and in this case, it would be a smartphone or tablet. The device detects the baby's facial expressions with its built-in camera, vocalizations with its microphone, and body movements with its motion sensor. These sensors collect data in real time and temporarily store it on the device. The data is then pre-processed before being sent to the server. For example, if the baby is captured laughing, the camera captures an image of that moment, and the audio captures the laughter.
[0641] Server operation
[0642] The server plays a central role in receiving and analyzing data sent from the terminal. The received data is first preprocessed, including noise reduction and standardization. Next, the emotion determination engine uses algorithms for voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, data combining a smile and a high-pitched voice would be determined to be "happy." Based on this determination, the generating AI constructs appropriate advice for the parents. For example, it might generate advice such as, "Let's play with the baby."
[0643] User actions
[0644] The user's role is to review the information received on the device and use it to guide their next actions. The emotion assessment results and advice sent from the server are displayed on the device's interface and provided in a format that is easy for the user to understand. Based on the displayed information, the user can decide on specific actions. For example, by following the displayed advice and playing with the baby, the user can facilitate smoother communication between parent and child.
[0645] As described above, the system of the present invention supports childcare by determining the baby's emotions in real time and suggesting the most appropriate response to the parents.
[0646] The following describes the processing flow.
[0647] Step 1:
[0648] The device acquires audio data from the baby's surroundings using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. This data is temporarily stored in real time.
[0649] Step 2:
[0650] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0651] Step 3:
[0652] The terminal sends pre-processed data to the server. The data is securely transferred to the server via a security protocol.
[0653] Step 4:
[0654] The server receives audio, image, and motion data transmitted from the terminal and stores it in a database. This data is immediately used for analysis.
[0655] Step 5:
[0656] The server uses stored data to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. For example, it analyzes a smiling image and a cheerful voice to determine that the baby is "happy."
[0657] Step 6:
[0658] Based on the analysis results, the server generates specific advice for parents using AI. For example, it might suggest, "Continue playing with your baby."
[0659] Step 7:
[0660] The server sends the judgment result and generated advice to the terminal.
[0661] Step 8:
[0662] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format that is easy for parents to understand.
[0663] Step 9:
[0664] The user uses the information displayed on the device as a reference to take specific actions for the baby. For example, if the device determines that the baby is sleepy, the user will take appropriate action to put the baby to sleep.
[0665] (Example 1)
[0666] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0667] In modern childcare, it is crucial for parents to quickly and accurately understand their baby's emotions and condition and respond appropriately. However, many parents find it difficult to accurately grasp their baby's emotions amidst their busy daily lives, which can lead to them being unable to choose appropriate childcare methods. This invention aims to solve this problem and support childcare more effectively by determining the baby's emotions in real time and providing parents with the most suitable response.
[0668] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0669] In this invention, the server includes an acquisition means for acquiring voice, an acquisition means for acquiring facial expressions, an acquisition means for acquiring body movements, a processing means for preprocessing and analyzing the acquired data, a determination means for determining emotions based on the preprocessed data, a generation means for generating specific countermeasures based on the determined emotions, a provision means for providing the generated countermeasures, a function for collecting data in real time using sensors at the terminal, temporarily storing and preprocessing it, and a function for using encryption technology in data transmission to realize secure data analysis on the server. This makes it possible to quickly analyze the emotional state of a baby and provide specific and practical advice to the guardian.
[0670] "Acquisition means" refers to a device or method for collecting data such as voice, facial expressions, or body movements.
[0671] "Processing means" refers to an apparatus or method for preprocessing acquired data and performing data analysis.
[0672] "Determination means" refers to a device or method for analyzing and determining emotions based on pre-processed data.
[0673] "Generating means" refers to a device or method for automatically creating specific countermeasures based on the determined emotions.
[0674] "Means of delivery" refers to a device or method for communicating the generated countermeasures to users such as guardians.
[0675] A "sensor" is a device that detects physical phenomena in the environment and outputs them in the form of electrical signals or other similar information.
[0676] "Encryption technology" is a technique that transforms data according to specific rules in order to protect it from unauthorized access.
[0677] This invention is a system that acquires a baby's voice, facial expressions, and body movements, determines their emotions based on that data, and provides appropriate advice to the caregiver. The specific configuration and operation of the system are described below.
[0678] Terminal operation
[0679] The device takes the form of a smartphone or tablet and is placed near the baby. It has a built-in camera, microphone, and motion sensor, which are used to capture the baby's facial expressions, voice, and body movements in real time. Specifically, the camera can capture the baby's smile, the microphone can record laughter, and the motion sensor can detect arm and leg movements. The collected data is temporarily stored on the device and pre-processed before transmission.
[0680] Server operation
[0681] The server plays a central role in secure and efficient data processing. Data transmitted from the terminal undergoes preprocessing such as noise reduction and standardization, followed by analysis for emotion determination. The server uses a generative AI model to perform voice analysis, image recognition, and motion analysis to determine the baby's emotions. Specifically, if a smile or high-pitched laughter is detected, it determines that the baby is "happy" and generates advice based on this.
[0682] User actions
[0683] Users review the assessment results and advice through an interface provided on their device. Based on this information, users decide on their actual actions. For example, if they receive the advice, "Please play with your baby," they can immediately take appropriate action for playtime. As a result, richer communication between parents and children is promoted.
[0684] Example of a prompt
[0685] "We have data that detects a baby's smile and high-pitched voice. What emotions does this represent, and what kind of advice should we offer?"
[0686] This invention is a system that highly analyzes a baby's emotional state and provides parents with an easy-to-understand and effective childcare method.
[0687] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0688] Step 1:
[0689] The device uses a camera, microphone, and motion sensor near the baby to acquire data on facial expressions, voice, and body movements in real time. The input is the baby's natural movements and voice, which are acquired as digital data by the sensors. The output is the raw data obtained from each sensor.
[0690] Step 2:
[0691] The terminal temporarily stores the acquired raw data in its internal memory and performs preprocessing for data transmission. This preprocessing adjusts the image resolution and removes noise from audio data. The input to this process is raw data, and the output is compressed and low-noise data for transmission.
[0692] Step 3:
[0693] The terminal sends pre-processed data to the server. During this process, the data is encrypted and transmitted securely. The input is the pre-processed data, and the output is the data received on the server side. The data is transmitted using Wi-Fi or a mobile network.
[0694] Step 4:
[0695] The server preprocesses and standardizes the received data again. Here, the audio and image formats are unified, and the motion data is converted into a format that is easy to analyze. The input is the data sent from the terminal, and the output is standardized, ready-to-analyze data.
[0696] Step 5:
[0697] The server utilizes a generative AI model to perform voice analysis, image recognition, and motion analysis on standardized data to determine the baby's emotions. The input is standardized data, and the output is the determined emotion information. For example, if laughter and a smile match, the emotion "happy" is output.
[0698] Step 6:
[0699] The server uses a generative AI model to generate specific advice for parents based on the determined emotional information. The input is the determined emotional information, and the output is an advice statement. For example, the advice "Let's play with the baby" might be generated.
[0700] Step 7:
[0701] The server sends the generated advice to the terminal, which then notifies the user. The advice is displayed on the terminal screen in a format that is easy for the user to understand. The input is the advice text, and the output is the user's action based on their understanding.
[0702] (Application Example 1)
[0703] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0704] When providing product recommendations tailored to a baby's condition in a physical store, there are challenges in accurately assessing emotions in real time and providing optimal product information. Furthermore, providing appropriate advice based on the baby's condition requires specialized knowledge from store staff, potentially leading to a lack of consistency in customer service.
[0705] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0706] In this invention, the server includes an input acquisition means for collecting acquired baby voice data, facial expression data, and motion data; a data processing means for preprocessing and analyzing the data obtained from the input acquisition means; and an information recommendation means for determining the baby's emotions based on the preprocessed data and providing recommended product information according to the determination result. This makes it possible to provide real-time product suggestions and advice to parents in a physical store, tailored to the baby's situation.
[0707] "Input acquisition means" refers to devices and methods for collecting voice data, facial expression data, and motion data obtained from a baby.
[0708] "Data processing means" refers to devices and methods for preprocessing and analyzing data obtained from input acquisition means. This means removes noise and standardizes the data.
[0709] An "information recommendation means" is a device or method that determines a baby's emotions based on the results analyzed by a data processing means and provides product information recommended according to those results.
[0710] "Information output means" refers to devices or methods that provide information to user terminals and are used when selecting recommended products in physical stores.
[0711] The system necessary to implement this invention mainly consists of three elements: a user terminal, a server, and a physical store.
[0712] The user device is a smartphone or tablet, equipped with a camera, microphone, and motion sensor. The camera captures the baby's facial expressions, the microphone collects sound, and the motion sensor detects body movements. These sensors collect data in real time, temporarily store it on the user device, and then preprocess it before sending it to the server. Specifically, if the baby is smiling, the facial expression and sound are captured simultaneously.
[0713] The server plays a primary role in receiving and analyzing data sent from user terminals. The received data undergoes preprocessing such as noise reduction and standardization. Next, emotion determination is performed using algorithms for speech analysis, image recognition, and motion analysis. AI models such as TensorFlow are used for this. For example, if a baby's smile and high-pitched voice are detected simultaneously, it is determined to be "happy." Based on this result, a generative AI model creates specific advice for the parents, such as, "Your child is happy. Playing together with toys would be ideal."
[0714] In physical stores, these assessment results and advice are displayed on tablets or dedicated terminals, and staff members make appropriate product suggestions to parents. For example, if the system detects that the baby is crying, it will suggest products such as, "This product is recommended for times when babies tend to cry."
[0715] For example, if a baby becomes fussy when visiting a store, the system will determine in real time that the baby is feeling "anxious" and automatically suggest to the staff things like "a stuffed animal to calm the child."
[0716] An example of a prompt to pass to a generative AI model might be: "Please analyze the following data: baby's voice, baby's facial expression image, and high-frequency sound. Next, determine what emotion the baby is feeling and suggest specific products to recommend to the parents."
[0717] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0718] Step 1:
[0719] The user collects the baby's facial expressions, voice, and body movements using the user's device's camera, microphone, and motion sensor. The input for this step is raw data of the baby, and the output is data temporarily stored on the device. Specifically, if the baby is laughing, the facial expression data and laughter are captured on the smartphone.
[0720] Step 2:
[0721] The terminal preprocesses the collected data by denoising and standardizing it, preparing it for transmission to the server. The input for this step is the raw data within the terminal, and the output is the preprocessed data. For example, background noise in audio may be removed, and the brightness of images may be adjusted.
[0722] Step 3:
[0723] The server receives preprocessed data sent from the user's terminal and performs speech analysis, image recognition, and motion analysis. The input for this step is preprocessed data, and the output is information about the baby's emotions. Specifically, TensorFlow is used to determine if the baby is "happy" from their smile and laughter.
[0724] Step 4:
[0725] Based on the emotion assessment results, the server uses a generative AI model to create specific advice for the parents. The input for this step is the emotion assessment results, and the output is the generated advice text. For example, advice such as, "The baby is having fun. Playing together with toys is ideal," is generated.
[0726] Step 5:
[0727] Users can choose the most suitable product for their baby based on information displayed on a terminal or tablet in a physical store. The input in this step is generated advice, and the output is product information for the parent to choose. For example, information such as "We recommend this product during times when your baby is prone to crying" is presented to help the parent make their selection.
[0728] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0729] This invention combines a system that acquires a baby's voice, facial expressions, and body movements, and determines the baby's emotions based on that data, with an emotion engine that recognizes the emotions of the user (the caregiver). The system operates through interaction between the terminal, server, and user.
[0730] Terminal operation
[0731] The device is placed near the baby and uses a microphone, camera, and motion sensor to capture voice, facial expressions, and movements. In addition, the device is equipped with sensors to capture the user's facial expressions and voice, collecting real-time emotional data. This data is temporarily stored and pre-processed before being sent to the server.
[0732] Server operation
[0733] The server receives voice, image, and motion data from the baby and user sent from the terminal, and then performs preprocessing such as noise reduction and data format standardization. Next, it uses an emotion determination engine on the baby's data to perform voice analysis, image recognition, and motion analysis. This determines the baby's emotions, such as "happy," "anxious," or "sleepy." The user's data is separately analyzed for emotions to determine whether the caregiver is in a state of "stress" or "security."
[0734] The server uses generative AI to design specific actions parents should take, based on the emotional data of both the baby and the user. The advice is dynamically adjusted; for example, if the user is feeling stressed, relaxation advice may be added. If the baby is anxious and the user is also feeling stressed, a more comprehensive piece of advice might be generated, such as, "Hold your baby to calm them down and take some deep breaths."
[0735] User actions
[0736] Users review the emotion assessment results and advice displayed on their device. Based on the displayed information, users can take appropriate actions towards their baby. The results and advice are presented through concise and easy-to-understand notifications and a visual interface. This helps parents reduce anxiety in childcare and enables them to respond appropriately to emotions.
[0737] In this way, the present invention can improve the childcare environment by simultaneously determining the emotions of both the baby and the parent and providing appropriate countermeasures. This is expected to promote better communication between parents and children.
[0738] The following describes the processing flow.
[0739] Step 1:
[0740] The device acquires audio data from the baby's environment using a microphone, captures facial expression data with a camera, and records body movements with a motion sensor. Similarly, it also acquires the user's voice and facial expressions using sensors. This data is temporarily stored in real time.
[0741] Step 2:
[0742] The terminal performs preprocessing on the collected audio, image, and motion data, such as noise reduction and format conversion. This processing makes the data suitable for analysis on the server.
[0743] Step 3:
[0744] The device sends pre-processed baby and user data to the server. A secure protocol is used for data transfer.
[0745] Step 4:
[0746] The server receives data sent from the terminal and stores it in the database. This data is immediately used for analysis.
[0747] Step 5:
[0748] The server performs voice analysis, image recognition, and motion analysis on the baby's data to determine the baby's emotions. For example, it might determine that the baby is "happy" based on a combination of a smiling image and a high-pitched voice.
[0749] Step 6:
[0750] Similarly, the server analyzes user data to determine the user's emotions. For example, it might determine that the user is stressed based on a tired tone of voice and an anxious facial expression.
[0751] Step 7:
[0752] Based on the emotional information of both the baby and the user, the server uses a generating AI to create specific advice for the parents. The user's emotional state is also taken into consideration, so advice such as, "Speak gently to your baby and also take time to calm yourself down," is generated as needed.
[0753] Step 8:
[0754] The server sends the judgment result and advice to the terminal.
[0755] Step 9:
[0756] The device displays the received emotion assessment results and advice to the user. The display is presented in a visually organized format, taking care to ensure that parents can easily understand and act upon the advice.
[0757] Step 10:
[0758] Based on the information displayed on the device, users take specific actions for their baby while also managing their own state of mind. For example, they might incorporate relaxation techniques to reduce their own stress while soothing their baby.
[0759] (Example 2)
[0760] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] There is a need to accurately assess a baby's emotions and, based on that assessment, quickly and precisely provide parents with guidance on how to respond. Conventional systems have been insufficient in assessing babies' emotions and analyzing parents' mental states, and have limited ability to generate specific countermeasures based on those results. Therefore, these issues need to be addressed.
[0762] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0763] In this invention, the server includes means for acquiring sound, means for acquiring images, and means for acquiring movement. This makes it possible to comprehensively determine the baby's emotions from various data such as voice, facial expressions, and movement, and further analyze the caregiver's mental state, in order to provide dynamically adjusted specific countermeasures.
[0764] "Means of acquiring sound" refers to technologies for collecting ambient sounds, particularly those used to detect sounds like a baby's voice.
[0765] "Means for acquiring images" refers to methods that use visual sensors such as cameras to capture images of the surroundings, and are used in particular to analyze the baby's facial expressions.
[0766] "Means for acquiring motion" refers to methods that use motion sensors or similar devices to detect body movements, and are particularly used to detect patterns in a baby's movements.
[0767] "Processing means" refers to a system that has the function of organizing the acquired data and converting it into a format suitable for analysis.
[0768] "Determination means" refers to algorithms and technologies used to analyze pre-processed data and evaluate specific states such as emotions.
[0769] "Additional emotion analysis means" refers to technology that evaluates a user's mental state from their speech patterns and facial expressions to gain a deeper understanding of their emotions.
[0770] "Generative means" refers to the technology used to construct and generate specific countermeasures and advice based on the analysis results.
[0771] "Information provision means" refers to screens, notification systems, etc., that present generated advice and information to users.
[0772] This invention is a system implemented through terminals placed around a baby, a server operating in the backend, and a user receiving information. Specific embodiments are described below.
[0773] Terminal configuration and operation
[0774] The device is equipped with a microphone to capture sound, a camera to capture images, and a motion sensor to capture movement. This device is placed near the baby and collects voice, facial expressions, and movements in real time. For example, if the baby is smiling, the camera captures the smile, and the motion sensor detects the baby's active movements. This data is temporarily stored for further processing and then transmitted to a server via the network.
[0775] Server configuration and operation
[0776] The server receives data transmitted from the terminal and performs preprocessing on it. For example, it performs noise reduction on audio data and standardizes the image data format. Next, the server analyzes the data using emotion determination tools. The analysis includes audio analysis, image recognition, and motion analysis, and determines the baby's emotions as "happy," "anxious," "sleepy," etc. At the same time, the server uses additional emotion analysis tools to evaluate the user's state as "stressed" or "at ease."
[0777] Furthermore, the server utilizes generative AI based on these analysis results to suggest specific actions the user should take. For example, if the user is stressed and the baby is anxious, it might recommend "hugging the baby and taking deep breaths together." The generative AI model incorporates historical data and expertise, providing sophisticated solutions.
[0778] An example of a prompt might be, "Please suggest ways to cope when a baby is crying and the caregiver is feeling stressed."
[0779] User actions
[0780] Users review the analysis results and advice displayed on their device. The information is presented in an easy-to-understand interface, enabling quick and accurate responses. This allows users to take appropriate action and reduce anxiety related to childcare.
[0781] This invention's system allows for the simultaneous evaluation of the baby's emotions and the user's mental state, contributing to better communication between parent and child.
[0782] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0783] Step 1:
[0784] The device collects the baby's voice, facial expressions, and movements. Inputs include sounds, images, and movements from the baby's surroundings. Specifically, a microphone picks up voice, a camera captures facial expressions, and a motion sensor detects movement. This data is temporarily stored within the device.
[0785] Step 2:
[0786] The terminal preprocesses the collected data. Because the input data may contain noise or incorrect formatting, the audio data is denoised, and the image data is resized and converted before being sent to the server. The preprocessed data is then output.
[0787] Step 3:
[0788] The terminal sends pre-processed data to the server. The specific data sent includes de-noise-removed audio, facial expression images converted to an appropriate format, and motion information.
[0789] Step 4:
[0790] The server analyzes the received data. It receives pre-processed audio, image, and motion data as input and uses an emotion determination engine to determine the baby's emotions. The specific output is an emotional state such as "happy," "anxious," or "sleepy." Simultaneously, it performs additional emotion analysis to evaluate the user's state.
[0791] Step 5:
[0792] Based on the assessment results, the server generates a response plan for the parent using a generated AI model. The prompt text is a specific situation corresponding to the baby's and the user's state, and specific advice is output based on that. For example, advice such as "Hold your baby and take a deep breath together" might be generated.
[0793] Step 6:
[0794] The user reviews the analysis results and generated advice displayed on the device. The input is the information displayed on the device's screen, and the user uses this to take appropriate action for their baby. This reduces anxiety about childcare.
[0795] (Application Example 2)
[0796] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0797] Conventional baby emotion assessment systems only consider the baby's emotions, making it difficult to obtain specific countermeasures that take into account the emotions of parents or store staff. Furthermore, there was a lack of mechanisms to understand the emotions of both customers and staff and provide appropriate countermeasures, which is crucial for smooth communication with customers in physical stores.
[0798] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0799] In this invention, the server includes an acoustic acquisition means for acquiring the voices of the baby and the customer, an image acquisition means for acquiring the facial expressions of the baby and the customer, and a data processing means for pre-processing and analyzing the acquired information. This makes it possible to determine the emotions of both the baby and the customer, as well as the parent or store staff, in real time and provide specific and effective countermeasures.
[0800] "Sound acquisition means" refers to a device that has the function of detecting and recording ambient sounds.
[0801] An "image acquisition means" is a device that has the function of acquiring visual information and converting it into digital data.
[0802] A "motion sensor means" is a sensor that can detect the movement of an object or human body and record the changes as data.
[0803] A "means of recognition" is a mechanism that processes information to detect specific objects or patterns and understand their meaning.
[0804] "Data processing means" refers to a device or software that has the function of organizing acquired information and converting it into useful information.
[0805] "Analysis means" refers to a device or program that performs analysis to derive meaningful results from obtained data.
[0806] A "response generation means" is a function for constructing specific action guidelines and proposals based on the analyzed results.
[0807] "Communication means" refers to a technology or device that transmits digital information from one point to another.
[0808] The system that implements this application facilitates communication by determining the emotions of customers and staff in a store in real time and providing appropriate responses. The server receives customer voice and facial expression data via acoustic and image acquisition means. Hardware devices such as Logitech cameras and Blue microphones are used for this purpose. The received data is pre-processed by data processing means, such as noise reduction and format conversion, and emotion determination is performed by analysis means using software such as Amazon Rekognition or Microsoft Azure Face API.
[0809] The server generates specific response strategies for store staff based on the determined emotions. The response generation system combines technologies such as Google Cloud Speech-to-Text and IBM Watson to provide guidelines on how staff should interact with customers. The generated responses are then transmitted to staff members' smartphones via communication channels, thereby improving the customer experience.
[0810] For example, if a customer shows signs of anxiety in the store, the server analyzes the data and sends a notification to the staff member's smartphone with advice such as, "Let's recommend a calming product to this customer." This advice provides staff with quick information to help them decide what to do about the customer.
[0811] An example of a prompt would be, "Analyze the customer's emotional state from their facial expressions and voice data, and tell me specifically how to respond if the customer is feeling anxious." This establishes a process in which the generative AI model generates accurate advice.
[0812] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0813] Step 1:
[0814] The terminal acquires customer voice and facial expression data. It captures the customer's voice using an acoustic acquisition device and photographs their facial expressions using an image acquisition device. The acquired data is temporarily stored. The input is customer voice and image data, and the output is digital data that requires preprocessing.
[0815] Step 2:
[0816] The terminal sends the stored audio and image data to the server. The server performs data processing on the received data, including noise reduction and format standardization. The input is the raw data sent from the terminal, and the output is the data after noise has been removed, making it ready for analysis.
[0817] Step 3:
[0818] The server uses pre-processed data to perform emotion determination through analytical methods. It utilizes Amazon Rekognition and Google Cloud Speech-to-Text to analyze emotions from customer voice and facial expressions. The input is denoised data, and the output is the customer's emotional state (e.g., anxiety, reassurance).
[0819] Step 4:
[0820] The server uses a generative AI model to construct a response based on the customer's emotional state. It sends prompts to the generative AI model to generate an appropriate response. The input is the emotional state and the prompts to the generative AI, while the output is the specific response.
[0821] Step 5:
[0822] The server sends the generated response to the user (staff) via a communication method. The staff checks the response displayed on their smartphone and handles the customer interaction. The input is the generated response, and the output is the specific action plan that the staff should take.
[0823] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0824] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0825] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0826] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0827] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0828] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0829] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0830] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0831] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0832] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0833] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0834] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0835] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0836] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0837] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0838] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0839] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0840] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0841] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0842] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0843] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0844] The following is further disclosed regarding the embodiments described above.
[0845] (Claim 1)
[0846] A means of acquiring sounds to capture a baby's voice,
[0847] A means of acquiring images to capture a baby's facial expressions,
[0848] A means of acquiring the body movements of a baby,
[0849] A data processing means for preprocessing and analyzing acquired data,
[0850] An emotion determination means for determining a baby's emotions based on pre-processed data,
[0851] An advice generation method that generates specific countermeasures for dealing with parents based on the determined emotions,
[0852] Information provision means to provide the generated advice to parents,
[0853] A system that includes this.
[0854] (Claim 2)
[0855] The system according to claim 1, characterized in that the emotion determination means determines emotion using speech analysis, image recognition, and motion analysis algorithms.
[0856] (Claim 3)
[0857] The system according to claim 1, characterized in that the advice generation means constructs advice for parents based on past data and expert knowledge.
[0858] "Example 1"
[0859] (Claim 1)
[0860] A means of acquiring sound,
[0861] A means of acquiring facial expressions,
[0862] Means for acquiring body movements,
[0863] Processing means for preprocessing and analyzing acquired data,
[0864] A determination means for determining emotions based on pre-processed data,
[0865] A generation means for generating specific countermeasures based on the determined emotions,
[0866] A means of providing the generated countermeasures,
[0867] The terminal has a function to collect data in real time using sensors, and to temporarily store and preprocess that data.
[0868] To enable secure data analysis on the server, the system includes a function that uses encryption technology during data transmission,
[0869] A system that includes this.
[0870] (Claim 2)
[0871] The system according to claim 1, which determines emotions by using speech analysis, image recognition, and motion analysis algorithms.
[0872] (Claim 3)
[0873] The system according to claim 1, which constructs advice based on past data and knowledge.
[0874] "Application Example 1"
[0875] (Claim 1)
[0876] An input acquisition means for collecting acquired baby voice data, facial expression data, and movement data,
[0877] A data processing means for preprocessing and analyzing the data obtained from the input acquisition means,
[0878] An information recommendation system that determines the baby's emotions based on pre-processed data and provides recommended product information according to the determination result,
[0879] A means of providing information to user terminals and outputting information for selecting recommended products in physical stores,
[0880] A system that includes this.
[0881] (Claim 2)
[0882] The information recommendation means is characterized by determining the baby's emotions using voice analysis, image recognition, and motion analysis algorithms, and selecting recommended products in a physical store based on the results.
[0883] (Claim 3)
[0884] The information recommendation means is characterized by constructing and providing recommended product information and advice to users based on past analysis results and expert knowledge, as described in claim 1.
[0885] "Example 2 of combining an emotion engine"
[0886] (Claim 1)
[0887] Means of acquiring sound,
[0888] Means of acquiring images,
[0889] Means for obtaining the action,
[0890] Processing means for preprocessing and analyzing acquired data,
[0891] A determination means for determining emotions based on pre-processed data,
[0892] An additional emotion analysis method for analyzing the user's mental state,
[0893] A generation means for generating specific countermeasures based on the determined emotions,
[0894] Information provision means that provide the generated countermeasures,
[0895] A system that includes this.
[0896] (Claim 2)
[0897] The system according to claim 1, characterized in that it dynamically adjusts countermeasures based on analyzed emotions and the user's mental state.
[0898] (Claim 3)
[0899] The system according to claim 1, characterized by constructing advice for users using historical data, expert knowledge, and generative models.
[0900] "Application example 2 when combining with an emotional engine"
[0901] (Claim 1)
[0902] A means of acquiring sounds to capture a baby's voice,
[0903] A means of acquiring images to capture a baby's facial expressions,
[0904] A motion sensor means for acquiring the movements of a baby's body,
[0905] A recognition means for indexing and acquiring customer voice and facial expressions,
[0906] A data processing means for preprocessing and analyzing acquired information,
[0907] An analytical means for determining emotions based on pre-processed information,
[0908] A response generation means that generates a response to the parent or store staff based on the determined emotion,
[0909] A means of communication to provide the generated countermeasures to the parent or store staff,
[0910] A system that includes this.
[0911] (Claim 2)
[0912] The system according to claim 1, characterized in that the analysis means determines emotion using acoustic analysis, image recognition, and motion analysis algorithms.
[0913] (Claim 3)
[0914] The system according to claim 1, characterized in that the response generation means designs a response to a parent or store staff based on past information and cases. [Explanation of Symbols]
[0915] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring sounds to capture a baby's voice, A means of acquiring images to capture a baby's facial expressions, A means of acquiring the body movements of a baby, A data processing means for preprocessing and analyzing acquired data, An emotion determination means for determining a baby's emotions based on pre-processed data, An advice generation method that generates specific countermeasures for dealing with parents based on the determined emotions, Information provision means to provide the generated advice to parents, A system that includes this.
2. The system according to claim 1, characterized in that the emotion determination means determines emotion using speech analysis, image recognition, and motion analysis algorithms.
3. The system according to claim 1, characterized in that the advice generation means constructs advice for parents based on past data and expert knowledge.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A