System

The system addresses the challenge of understanding baby feelings and needs by analyzing voice data and providing appropriate reactions, enhancing communication and bonding through a database construction, analysis, and feedback mechanism.

JP2026018645APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119967
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional technologies face challenges in accurately understanding a baby's feelings and needs, making it difficult to provide appropriate reactions.

Method used

A system comprising a database construction unit, analysis unit, and feedback unit that collects, analyzes, and responds to baby voice data, incorporating emotion identification and reaction generation to provide appropriate feedback to caregivers.

Benefits of technology

The system effectively understands and responds to a baby's feelings and needs, reducing stress and promoting smoother communication and deeper bonding between parents and caregivers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018645000001_ABST
    Figure 2026018645000001_ABST
Patent Text Reader

Abstract

An object of a system according to an embodiment is to understand the feelings and needs of a baby and provide an appropriate reaction.SOLUTION: A system includes a database construction unit, an analysis unit, a reaction generation unit, and a feedback unit. The database construction unit collects voice data of the baby. The analysis unit analyzes the voice data of the baby collected by the database construction unit. The reaction generation unit generates an appropriate reaction on the basis of the result analyzed by the analysis unit. The feedback unit feeds back the reaction generated by the reaction generation unit to the parent or the nursery teacher.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional technology has had the challenge of making it difficult to accurately understand a baby's feelings and needs and provide appropriate reactions.

[0005] The system according to the embodiment aims to understand the feelings and needs of the baby and provide appropriate reactions. [Means for solving the problem]

[0006] The system according to the embodiment includes a database construction unit, an analysis unit, a reaction generation unit, and a feedback unit. The database construction unit collects baby voice data. The analysis unit analyzes the baby voice data collected by the database construction unit. The reaction generation unit generates an appropriate reaction based on the analysis result by the analysis unit. The feedback unit feeds back the reaction generated by the reaction generation unit to the parent or caregiver. [Effects of the Invention]

[0007] The system according to the embodiment can understand the feelings and needs of the baby and provide appropriate reactions. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION

[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0010] First, the terms used in the following description will be explained.

[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).

[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0013] In the following embodiments, the coded storage is one or more nonvolatile storage devices that store various programs, various parameters, etc. Examples of nonvolatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.

[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example 1) The AI-based service according to an embodiment of the present invention is a system that provides a deeper understanding of a baby's feelings and promotes communication with parents and caregivers. This system is built on a database of baby talk, having learned the languages ​​of tens of thousands of babies. This allows the AI-based service to provide a deeper understanding of a baby's feelings and promote communication with parents and caregivers.

[0029] The AI-based service according to the embodiment includes a database construction unit, an analysis unit, a reaction generation unit, and a feedback unit. The database construction unit collects baby voice data. For example, the database construction unit collects voice data such as baby crying, laughing, and babbling. The database construction unit also stores the baby voice data in a digital format. For example, the database construction unit stores the voice data in MP3 or WAV format. The database construction unit also adds the collected voice data to a baby talk database. For example, the collected voice data is indexed and stored in the database. The analysis unit analyzes the baby voice data collected by the database construction unit. For example, the analysis unit converts the baby voice data into text using voice recognition technology. The analysis unit also estimates the baby's emotion using an emotion analysis algorithm. For example, the analysis unit analyzes the baby's crying pattern to estimate "sad." The analysis unit also analyzes the voice data to identify the baby's needs. For example, the analysis unit analyzes the baby's crying to identify "my diaper is wet." The reaction generation unit generates an appropriate reaction based on the analysis results obtained by the analysis unit. For example, if a baby expresses "I'm hungry," the reaction generation unit responds by saying, "You're hungry, I'll prepare dinner now." The reaction generation unit also generates visual feedback. For example, when a baby smiles, a smiley face icon is displayed. The reaction generation unit also generates audio responses. For example, when a baby is crying, the reaction generation unit asks, "What's wrong?" The feedback unit feeds back the reactions generated by the reaction generation unit to a parent or a caregiver. For example, the feedback unit notifies the parent of the reason why the baby is crying. The feedback unit also reports the baby's condition to the parent or caregiver. For example, the feedback unit notifies the parent or caregiver that "The baby is crying because his diaper is wet." The feedback unit also provides advice regarding the baby's needs. For example, the feedback unit advises, "When the baby smiles, it's a good idea to respond by smiling along with him." This enables the AI-based service according to the embodiment to better understand the baby's feelings and promote communication with the parent or caregiver.For example, accurately understanding why a baby is crying and responding quickly can reduce the baby's stress. Furthermore, smoother communication with the baby can deepen the bond between the parent and the childcare worker, which is expected to promote the baby's healthy development.

[0030] The database construction unit can collect facial expression or movement data in addition to the baby's voice data and construct a multimodal database. For example, the database construction unit simultaneously collects not only the baby's voice data but also facial expression and movement data and constructs a multimodal database that integrates these. For example, the voice, facial expression, and hand movements of a baby when laughing are simultaneously recorded. The database construction unit also captures the baby's facial expression data with a camera and collects movement data with a sensor. For example, the sensor detects the baby's hand waving movement and saves that data. Furthermore, the database construction unit adds the collected facial expression and movement data to a baby talk database. For example, the collected facial expression data is indexed and saved in the database. In this way, by collecting facial expression and movement data in addition to the baby's voice data, more accurate analysis is possible.

[0031] The database construction unit can learn different language patterns for each baby's developmental stage and create databases for each age. For example, the database construction unit learns different language patterns for each baby's developmental stage and creates databases for each age. For example, different language patterns are recorded for each stage, such as 3 months, 6 months, and 1 year old. The database construction unit also collects voice data according to the baby's developmental stage. For example, it collects crying data from a 3-month-old baby and adds it to the database. Furthermore, the database construction unit classifies and saves the collected voice data by age. For example, it classifies and saves voice data from a 6-month-old baby in a different category. In this way, by learning different language patterns for each baby's developmental stage, appropriate analysis according to age becomes possible.

[0032] The database construction unit can collect data on babies from different cultural or linguistic regions and construct a global database. The database construction unit, for example, collects voice data of babies from different cultural or linguistic regions and constructs a global database. For example, it collects voice data of babies from Asia, Europe, America, etc. The database construction unit also collects voice data of babies from different linguistic regions. For example, it collects voice data of babies from English-speaking regions, Spanish-speaking regions, etc. The database construction unit then adds the collected voice data to the global database. For example, it classifies and stores the collected voice data by region. This allows global analysis to be performed by collecting data from different cultural or linguistic regions.

[0033] The database construction unit can collect not only baby voice data but also reaction data of parents or childcare workers to create a two-way communication database. The database construction unit, for example, collects reaction data of parents and childcare workers along with baby voice data to create a two-way communication database. For example, it records parent reactions when a baby cries. The database construction unit also collects voice response data of parents and childcare workers. For example, it collects voice data of parents speaking to their babies. The database construction unit then adds the collected reaction data to a baby talk database. For example, it indexes and stores the collected reaction data in the database. This allows two-way communication by collecting reaction data of parents and childcare workers as well.

[0034] The analysis unit can analyze the baby's voice data in real time and instantly identify the meaning. The analysis unit, for example, develops a system that analyzes the baby's voice data in real time and instantly identifies the meaning. For example, when a baby says "woo-woo," it immediately identifies that the baby is "hungry." The analysis unit also develops an algorithm for analyzing voice data in real time. For example, it converts voice data into text in real time using voice recognition technology. Furthermore, the analysis unit builds a system that outputs the analysis results in real time. For example, it analyzes the baby's voice data and instantly notifies the parents or childcare workers. This makes it possible to analyze the baby's voice data in real time and instantly identify the meaning, enabling a rapid response.

[0035] The analysis unit can employ filtering technology to remove background or environmental sounds when analyzing the baby's voice pattern. For example, the analysis unit employs filtering technology to remove background or environmental sounds when analyzing the baby's voice pattern. For example, television sounds and outside noises are removed to analyze only the baby's voice. The analysis unit also employs noise canceling technology to remove background sounds. For example, ambient noise is removed when analyzing a baby's crying. Furthermore, the analysis unit employs digital filtering technology to remove environmental sounds. For example, wind sounds and car sounds are removed to analyze the baby's voice. In this way, the accuracy of analyzing the baby's voice data is improved by removing background or environmental sounds.

[0036] The analysis unit can identify meaning from both the baby's voice and facial expression by analyzing their facial expressions. For example, the analysis unit can identify meaning from both the baby's voice and facial expression by analyzing their facial expressions in addition to analyzing their voice. For example, if a baby is laughing while saying "woo-woo," it can identify the baby as "happy." The analysis unit also uses facial recognition technology to analyze the baby's facial expression data. For example, it can detect the baby's smile and identify the baby as "happy." The analysis unit can also develop algorithms that integrate and analyze voice data and facial expression data. For example, it can build an algorithm that combines voice data and facial expression data for analysis. This enables more accurate analysis by identifying meaning from both voice and facial expression.

[0037] When analyzing the baby's voice data, the analysis unit can simultaneously analyze the reaction data of the parent or childcare worker, taking interactions into consideration. For example, when analyzing the baby's voice data, the analysis unit simultaneously analyzes the reaction data of the parent or childcare worker, taking interactions into consideration. For example, the analysis unit analyzes the parent's reaction when the baby cries, taking interactions into consideration. The analysis unit also analyzes the voice response data of the parent or childcare worker. For example, it analyzes voice data of the parent speaking to the baby. Furthermore, the analysis unit develops an algorithm for analyzing interactions between the baby and the parent or childcare worker. For example, it constructs an algorithm for analyzing a combination of the baby's cry and the parent's reaction. This makes it possible to simultaneously analyze the reaction data of the parent and childcare worker, thereby taking interactions into consideration in the analysis.

[0038] The analysis unit can develop an algorithm that identifies patterns by comparing with past data. For example, the analysis unit develops an algorithm that identifies patterns by comparing with past data in order to understand a baby's expressions and needs. For example, it compares past crying data with current crying to identify specific needs. The analysis unit also collects past audio data and stores it in a database. For example, it collects baby crying data from the past year. Furthermore, the analysis unit builds an algorithm that identifies patterns based on the collected past data. For example, it analyzes past data and extracts common patterns. This allows for more accurate analysis by comparing with past data to identify patterns.

[0039] The analysis unit can incorporate feedback from parents or childcare workers and correct the analysis results. For example, when understanding a baby's expressions and needs, the analysis unit incorporates feedback from parents or childcare workers and corrects the analysis results. For example, if a parent says, "I'm hungry," the analysis results are corrected based on that feedback. The analysis unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on their baby's condition. Furthermore, the analysis unit builds an algorithm that corrects the analysis results based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and corrects the analysis results. In this way, the accuracy of the analysis results can be improved by incorporating feedback from parents and childcare workers.

[0040] The analysis unit can compare data from different cultural or linguistic regions and identify common patterns. For example, the analysis unit compares data from babies from different cultural or linguistic regions and identifies common patterns. For example, it compares crying data from babies in Asia, Europe, and America to identify common needs. The analysis unit also collects voice data from babies from different linguistic regions and stores it in a database. For example, it collects voice data from babies in English-speaking and Spanish-speaking regions. The analysis unit then builds an algorithm to identify common patterns based on the collected data. For example, it analyzes data from different cultural or linguistic regions and extracts common patterns. This makes it possible to identify common patterns by comparing data from different cultural or linguistic regions.

[0041] The analysis unit can also analyze reaction data of parents or childcare workers and take interactions into account. For example, when understanding a baby's expressions and needs, the analysis unit also analyzes reaction data of parents and childcare workers and takes interactions into account. For example, it analyzes the parent's reaction when the baby cries and takes interactions into account. The analysis unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents talking to their babies. Furthermore, the analysis unit develops algorithms to analyze interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to analyze parent and childcare worker reaction data and take interactions into account.

[0042] The reaction generation unit can develop an algorithm that refers to past success stories. For example, when providing a reaction that meets the baby's needs, the reaction generation unit develops an algorithm that refers to past success stories. For example, it learns reaction patterns that have been successful in the past and applies them in similar situations. The reaction generation unit also collects past success stories and stores them in a database. For example, it collects data on reactions that have been effective in the past. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected success stories. For example, it analyzes success stories and extracts reaction patterns. In this way, by developing an algorithm that refers to past success stories, more effective reactions become possible.

[0043] The reaction generation unit can incorporate feedback from parents or childcare workers and optimize reactions. For example, when providing a reaction according to the baby's needs, the reaction generation unit incorporates feedback from parents or childcare workers and optimizes the reaction. For example, if a parent says, "This reaction was effective," the reaction is optimized based on that feedback. The reaction generation unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the reaction generation unit builds an algorithm that optimizes reactions based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies reactions. In this way, the accuracy of reactions is improved by incorporating feedback from parents and childcare workers.

[0044] The reaction generation unit can refer to reaction patterns from different cultural or linguistic regions. For example, the reaction generation unit refers to reaction patterns from different cultural or linguistic regions when providing reactions that meet the needs of a baby. For example, the reaction generation unit learns and applies reaction patterns from Asia, Europe, and America. The reaction generation unit also collects reaction data from different cultural or linguistic regions and stores it in a database. For example, it collects reaction data from babies in each region. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected reaction data. For example, it develops an algorithm that analyzes reaction patterns from different cultural or linguistic regions and generates reactions. This makes it possible to provide more diverse reactions by referring to reaction patterns from different cultural or linguistic regions.

[0045] The reaction generation unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, when providing a reaction according to the baby's needs, the reaction generation unit also analyzes reaction data of parents and childcare workers and takes interactions into consideration. For example, it analyzes the parent's reaction when the baby cries and takes interactions into consideration. The reaction generation unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents talking to their babies. Furthermore, the reaction generation unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to analyze reaction data of parents and childcare workers and take interactions into consideration.

[0046] The feedback unit can show trends by comparing with past data. For example, when feeding back the results of understanding a baby's needs and expressions, the feedback unit shows trends by comparing with past data. For example, it shows trends by comparing crying data from the past month with current crying data. The feedback unit also collects past data and stores it in a database. For example, it collects crying data from babies over the past year. Furthermore, the feedback unit builds an algorithm that shows trends based on the collected past data. For example, it develops an algorithm that analyzes past data and extracts trends of change. This makes it easier to understand changes in the baby's condition by showing trends by comparing with past data.

[0047] The feedback unit can incorporate feedback from parents or childcare workers and optimize the feedback content. For example, when providing feedback on the results of understanding the baby's needs and expressions, the feedback unit incorporates feedback from parents or childcare workers and optimizes the feedback content. For example, if a parent says, "This feedback was helpful," the content is optimized based on that feedback. The feedback unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the feedback unit builds an algorithm that optimizes the feedback content based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies the content. In this way, the accuracy of the feedback content is improved by incorporating feedback from parents and childcare workers.

[0048] The feedback unit can refer to data from different cultural or linguistic regions. For example, the feedback unit refers to data from different cultural or linguistic regions when providing feedback on the results of understanding a baby's needs and expressions. For example, the feedback unit refers to data from Asia, Europe, and America. The feedback unit also collects data from different cultural or linguistic regions and stores it in a database. For example, it collects data on babies in each region. Furthermore, the feedback unit builds an algorithm that provides feedback based on the collected data. For example, it develops an algorithm that analyzes data from different cultural or linguistic regions and provides feedback. This makes it possible to provide more diverse feedback by referring to data from different cultural or linguistic regions.

[0049] The feedback unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, when providing feedback on the results of understanding a baby's needs and expressions, the feedback unit also analyzes reaction data of parents and childcare workers and takes interactions into consideration. For example, it analyzes the parent's reaction when the baby cries and takes interactions into consideration. The feedback unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents speaking to their babies. Furthermore, the feedback unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to provide feedback that takes interactions into consideration by analyzing reaction data of parents and childcare workers.

[0050] The feedback unit can develop an algorithm that provides advice by referring to past success stories. For example, the feedback unit develops an algorithm that provides advice by referring to past success stories in order to promote communication with a baby. For example, it learns communication patterns that have been successful in the past and applies them in similar situations. The feedback unit also collects past success stories and stores them in a database. For example, it collects data on communication that has been effective in the past. Furthermore, the feedback unit builds an algorithm that provides advice based on the collected success stories. For example, it develops an algorithm that analyzes success stories and extracts advice patterns. This enables more effective feedback by providing advice by referring to past success stories.

[0051] The feedback unit can incorporate feedback from parents or childcare workers and optimize the advice content. The feedback unit incorporates feedback from parents or childcare workers and optimizes the advice content, for example, to promote communication with the baby. For example, if a parent says, "This advice was helpful," the content is optimized based on that feedback. The feedback unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the feedback unit builds an algorithm that optimizes the advice content based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies the content. In this way, the accuracy of the advice content is improved by incorporating feedback from parents and childcare workers.

[0052] The feedback unit can refer to communication patterns in different cultural or linguistic regions. For example, the feedback unit refers to communication patterns in different cultural or linguistic regions to promote communication with babies. For example, it learns and applies communication patterns from Asia, Europe, and America. The feedback unit also collects communication data from different cultural or linguistic regions and stores it in a database. For example, it collects communication data from babies in each region. Furthermore, the feedback unit builds an algorithm that provides advice based on the collected data. For example, it develops an algorithm that analyzes communication patterns in different cultural or linguistic regions and provides advice. This makes it possible to provide more diverse advice by referring to communication patterns from different cultural or linguistic regions.

[0053] The feedback unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, in order to promote communication with babies, the feedback unit can also analyze reaction data of parents and childcare workers and take interactions into consideration. For example, it can analyze the parent's reaction when the baby cries and take interactions into consideration. The feedback unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it can collect voice data of parents talking to their babies. Furthermore, the feedback unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it can build an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to provide feedback that takes interactions into consideration by also analyzing reaction data of parents and childcare workers.

[0054] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0055] The database construction unit can collect reaction data of parents or childcare workers in addition to baby's voice data, and create a two-way communication database. For example, it records the parent's reaction when the baby cries. The database construction unit also collects voice response data of parents and childcare workers. For example, it collects voice data of parents talking to their babies. The database construction unit then adds the collected reaction data to a baby talk database. For example, it indexes and stores the collected reaction data in the database. This allows two-way communication by collecting reaction data of parents and childcare workers.

[0056] The analysis unit can analyze the baby's voice data in real time and instantly identify the meaning. For example, when a baby says "uuuuu," it can immediately identify that the baby is "hungry." The analysis unit also develops algorithms for analyzing voice data in real time. For example, it can convert voice data into text in real time using voice recognition technology. The analysis unit then builds a system that outputs the analysis results in real time. For example, it can analyze the baby's voice data and instantly notify the parents or childcare workers. This allows for real-time analysis of the baby's voice data and instantaneous identification of the meaning, enabling a rapid response.

[0057] The reaction generation unit can develop an algorithm that takes past success cases into consideration. For example, it learns reaction patterns that have been successful in the past and applies them in similar situations. The reaction generation unit also collects past success cases and stores them in a database. For example, it collects data on reactions that have been effective in the past. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected success cases. For example, it analyzes success cases and extracts reaction patterns. In this way, by developing an algorithm that takes past success cases into consideration, more effective reactions become possible.

[0058] The feedback unit can compare data with past data to show trends. For example, it can compare crying data from the past month with current crying data to show trends. The feedback unit also collects past data and stores it in a database. For example, it can collect data on babies' crying over the past year. The feedback unit then builds an algorithm that shows trends based on the collected past data. For example, it develops an algorithm that analyzes past data and extracts trends of change. This makes it easier to understand changes in the baby's condition by comparing it with past data to show trends.

[0059] The database construction unit can collect data on babies from different cultural or linguistic regions to construct a global database. For example, it collects voice data from babies from Asia, Europe, America, etc. The database construction unit also collects voice data from babies from different linguistic regions. For example, it collects voice data from babies from English-speaking countries, Spanish-speaking countries, etc. The database construction unit then adds the collected voice data to the global database. For example, it classifies and stores the collected voice data by region. This allows global analysis to be performed by collecting data from different cultural or linguistic regions.

[0060] The processing flow of the first embodiment will be briefly explained below.

[0061] Step 1: The database construction unit collects baby voice data and stores it digitally. For example, it collects baby crying, laughing, babbling, and other sounds and stores them in MP3 or WAV format. It also indexes and adds the collected voice data to a baby talk database. Step 2: The analysis unit analyzes the baby's voice data collected by the database construction unit. For example, it converts the baby's voice data into text using voice recognition technology and estimates the baby's emotions using an emotion analysis algorithm. It also analyzes the voice data to identify the baby's needs. Step 3: The reaction generator generates an appropriate reaction based on the results of the analysis. For example, if the baby says, "I'm hungry," the reaction generator responds by saying, "I see you're hungry. I'll prepare some food now." It also generates visual feedback and audio responses. Step 4: The feedback unit feeds back the reactions generated by the reaction generation unit to the parent or caregiver, for example, informing the parent why the baby is crying, reporting the baby's condition, or providing advice regarding the baby's needs.

[0062] (Example 2) The AI-based service according to an embodiment of the present invention is a system that provides a deeper understanding of a baby's feelings and promotes communication with parents and caregivers. This system is built on a database of baby talk, having learned the languages ​​of tens of thousands of babies. This allows the AI-based service to provide a deeper understanding of a baby's feelings and promote communication with parents and caregivers.

[0063] The AI-based service according to the embodiment includes a database construction unit, an analysis unit, a reaction generation unit, and a feedback unit. The database construction unit collects baby voice data. For example, the database construction unit collects voice data such as baby crying, laughing, and babbling. The database construction unit also stores the baby voice data in a digital format. For example, the database construction unit stores the voice data in MP3 or WAV format. The database construction unit also adds the collected voice data to a baby talk database. For example, the collected voice data is indexed and stored in the database. The analysis unit analyzes the baby voice data collected by the database construction unit. For example, the analysis unit converts the baby voice data into text using voice recognition technology. The analysis unit also estimates the baby's emotion using an emotion analysis algorithm. For example, the analysis unit analyzes the baby's crying pattern to estimate "sad." The analysis unit also analyzes the voice data to identify the baby's needs. For example, the analysis unit analyzes the baby's crying to identify "my diaper is wet." The reaction generation unit generates an appropriate reaction based on the analysis results obtained by the analysis unit. For example, if a baby expresses "I'm hungry," the reaction generation unit responds by saying, "You're hungry, I'll prepare dinner now." The reaction generation unit also generates visual feedback. For example, when a baby smiles, a smiley face icon is displayed. The reaction generation unit also generates audio responses. For example, when a baby is crying, the reaction generation unit asks, "What's wrong?" The feedback unit feeds back the reactions generated by the reaction generation unit to a parent or a caregiver. For example, the feedback unit notifies the parent of the reason why the baby is crying. The feedback unit also reports the baby's condition to the parent or caregiver. For example, the feedback unit notifies the parent or caregiver that "The baby is crying because his diaper is wet." The feedback unit also provides advice regarding the baby's needs. For example, the feedback unit advises, "When the baby smiles, it's a good idea to respond by smiling along with him." This enables the AI-based service according to the embodiment to better understand the baby's feelings and promote communication with the parent or caregiver.For example, accurately understanding why a baby is crying and responding quickly can reduce the baby's stress. Furthermore, smoother communication with the baby can deepen the bond between the parent and the childcare worker, which is expected to promote the baby's healthy development.

[0064] The database construction unit can collect facial expression or movement data in addition to the baby's voice data and construct a multimodal database. For example, the database construction unit simultaneously collects not only the baby's voice data but also facial expression and movement data and constructs a multimodal database that integrates these. For example, the voice, facial expression, and hand movements of a baby when laughing are simultaneously recorded. The database construction unit also captures the baby's facial expression data with a camera and collects movement data with a sensor. For example, the sensor detects the baby's hand waving movement and saves that data. Furthermore, the database construction unit adds the collected facial expression and movement data to a baby talk database. For example, the collected facial expression data is indexed and saved in the database. In this way, by collecting facial expression and movement data in addition to the baby's voice data, more accurate analysis is possible.

[0065] The database construction unit can learn different language patterns for each baby's developmental stage and create databases for each age. For example, the database construction unit learns different language patterns for each baby's developmental stage and creates databases for each age. For example, different language patterns are recorded for each stage, such as 3 months, 6 months, and 1 year old. The database construction unit also collects voice data according to the baby's developmental stage. For example, it collects crying data from a 3-month-old baby and adds it to the database. Furthermore, the database construction unit classifies and saves the collected voice data by age. For example, it classifies and saves voice data from a 6-month-old baby in a different category. In this way, by learning different language patterns for each baby's developmental stage, appropriate analysis according to age becomes possible.

[0066] The database construction unit can use the emotion estimation function to add the baby's emotional state to the database and perform emotion-based analysis. The database construction unit, for example, uses the emotion estimation function to estimate the baby's emotional state from voice data and add the data to the database. For example, it analyzes voice data when the baby is crying and records the emotional state as "sad." The database construction unit also estimates the baby's emotional state from facial expression data and adds the data to the database. For example, it analyzes facial expression data when the baby is smiling and records the emotional state as "happy." The database construction unit also adds the collected emotional data to a baby talk database. For example, it indexes and stores the collected emotional data in the database. This makes it possible to use the emotion estimation function to perform analysis that takes the baby's emotional state into consideration.

[0067] The database construction unit can collect data on babies from different cultural or linguistic regions and construct a global database. The database construction unit, for example, collects voice data of babies from different cultural or linguistic regions and constructs a global database. For example, it collects voice data of babies from Asia, Europe, America, etc. The database construction unit also collects voice data of babies from different linguistic regions. For example, it collects voice data of babies from English-speaking regions, Spanish-speaking regions, etc. The database construction unit then adds the collected voice data to the global database. For example, it classifies and stores the collected voice data by region. This allows global analysis to be performed by collecting data from different cultural or linguistic regions.

[0068] The database construction unit can collect not only baby voice data but also reaction data of parents or childcare workers to create a two-way communication database. The database construction unit, for example, collects reaction data of parents and childcare workers along with baby voice data to create a two-way communication database. For example, it records parent reactions when a baby cries. The database construction unit also collects voice response data of parents and childcare workers. For example, it collects voice data of parents speaking to their babies. The database construction unit then adds the collected reaction data to a baby talk database. For example, it indexes and stores the collected reaction data in the database. This allows two-way communication by collecting reaction data of parents and childcare workers as well.

[0069] The database construction unit uses the emotion estimation function to construct a database based on the baby's emotions and can perform analysis according to the emotions. The database construction unit, for example, uses the emotion estimation function to estimate the emotional state from the baby's voice data and add the data to the database. For example, it analyzes voice data when the baby is crying and records the emotional state as "sad." The database construction unit also estimates the emotional state from the baby's facial expression data and adds the data to the database. For example, it analyzes facial expression data when the baby is smiling and records the emotional state as "happy." The database construction unit also adds the collected emotional data to a baby talk database. For example, it indexes and stores the collected emotional data in the database. This makes it possible to use the emotion estimation function to perform analysis based on the baby's emotions.

[0070] The analysis unit can analyze the baby's voice data in real time and instantly identify the meaning. The analysis unit, for example, develops a system that analyzes the baby's voice data in real time and instantly identifies the meaning. For example, when a baby says "woo-woo," it immediately identifies that the baby is "hungry." The analysis unit also develops an algorithm for analyzing voice data in real time. For example, it converts voice data into text in real time using voice recognition technology. Furthermore, the analysis unit builds a system that outputs the analysis results in real time. For example, it analyzes the baby's voice data and instantly notifies the parents or childcare workers. This makes it possible to analyze the baby's voice data in real time and instantly identify the meaning, enabling a rapid response.

[0071] The analysis unit can employ filtering technology to remove background or environmental sounds when analyzing the baby's voice pattern. For example, the analysis unit employs filtering technology to remove background or environmental sounds when analyzing the baby's voice pattern. For example, television sounds and outside noises are removed to analyze only the baby's voice. The analysis unit also employs noise canceling technology to remove background sounds. For example, ambient noise is removed when analyzing a baby's crying. Furthermore, the analysis unit employs digital filtering technology to remove environmental sounds. For example, wind sounds and car sounds are removed to analyze the baby's voice. In this way, the accuracy of analyzing the baby's voice data is improved by removing background or environmental sounds.

[0072] The analysis unit can use the emotion estimation function to estimate emotions from the baby's voice and correct the analysis results based on the estimated emotions. For example, the analysis unit uses the emotion estimation function to estimate the emotional state from the baby's voice and reflects the data in the analysis results. For example, it analyzes voice data when the baby is crying and corrects the analysis results by determining the emotional state as "sad." The analysis unit also estimates the emotional state from the baby's facial expression data and reflects the data in the analysis results. For example, it analyzes facial expression data when the baby is smiling and corrects the analysis results by determining the emotional state as "happy." The analysis unit also develops an algorithm to correct the analysis results based on the collected emotional data. For example, it builds an algorithm to modify the analysis results depending on the emotional state. This makes it possible to use the emotion estimation function to correct the analysis results based on the baby's emotions.

[0073] The analysis unit can identify meaning from both the baby's voice and facial expression by analyzing their facial expressions. For example, the analysis unit can identify meaning from both the baby's voice and facial expression by analyzing their facial expressions in addition to analyzing their voice. For example, if a baby is laughing while saying "woo-woo," it can identify the baby as "happy." The analysis unit also uses facial recognition technology to analyze the baby's facial expression data. For example, it can detect the baby's smile and identify the baby as "happy." The analysis unit can also develop algorithms that integrate and analyze voice data and facial expression data. For example, it can build an algorithm that combines voice data and facial expression data for analysis. This enables more accurate analysis by identifying meaning from both voice and facial expression.

[0074] When analyzing the baby's voice data, the analysis unit can simultaneously analyze the reaction data of the parent or childcare worker, taking interactions into consideration. For example, when analyzing the baby's voice data, the analysis unit simultaneously analyzes the reaction data of the parent or childcare worker, taking interactions into consideration. For example, the analysis unit analyzes the parent's reaction when the baby cries, taking interactions into consideration. The analysis unit also analyzes the voice response data of the parent or childcare worker. For example, it analyzes voice data of the parent speaking to the baby. Furthermore, the analysis unit develops an algorithm for analyzing interactions between the baby and the parent or childcare worker. For example, it constructs an algorithm for analyzing a combination of the baby's cry and the parent's reaction. This makes it possible to simultaneously analyze the reaction data of the parent and childcare worker, thereby taking interactions into consideration in the analysis.

[0075] The analysis unit can use the emotion estimation function to estimate the emotion from both the baby's voice and facial expression, and correct the analysis results based on that emotion. For example, the analysis unit uses the emotion estimation function to estimate the emotional state from both the baby's voice and facial expression, and reflects that data in the analysis results. For example, it analyzes the voice and facial expression data when the baby is crying, and corrects the analysis results by determining the emotional state as "sad." The analysis unit also estimates the emotional state from the baby's facial expression data, and reflects that data in the analysis results. For example, it analyzes facial expression data when the baby is smiling, and corrects the analysis results by determining the emotional state as "happy." The analysis unit also develops an algorithm to correct the analysis results based on the collected emotional data. For example, it builds an algorithm to correct the analysis results depending on the emotional state. This enables more accurate analysis by estimating emotions from both the voice and facial expression, and correcting the analysis results based on that emotion.

[0076] The analysis unit can develop an algorithm that identifies patterns by comparing with past data. For example, the analysis unit develops an algorithm that identifies patterns by comparing with past data in order to understand a baby's expressions and needs. For example, it compares past crying data with current crying to identify specific needs. The analysis unit also collects past audio data and stores it in a database. For example, it collects baby crying data from the past year. Furthermore, the analysis unit builds an algorithm that identifies patterns based on the collected past data. For example, it analyzes past data and extracts common patterns. This allows for more accurate analysis by comparing with past data to identify patterns.

[0077] The analysis unit can incorporate feedback from parents or childcare workers and correct the analysis results. For example, when understanding a baby's expressions and needs, the analysis unit incorporates feedback from parents or childcare workers and corrects the analysis results. For example, if a parent says, "I'm hungry," the analysis results are corrected based on that feedback. The analysis unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on their baby's condition. Furthermore, the analysis unit builds an algorithm that corrects the analysis results based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and corrects the analysis results. In this way, the accuracy of the analysis results can be improved by incorporating feedback from parents and childcare workers.

[0078] The analysis unit can use the emotion estimation function to consider the emotional state of the baby and identify needs based on the emotion. The analysis unit, for example, uses the emotion estimation function to consider the emotional state of the baby and identify needs based on the emotion. For example, when a baby is crying, the emotional state is determined to be "sad" and needs are identified. The analysis unit also estimates the emotional state from facial expression data of the baby and identifies needs based on that data. For example, by analyzing facial expression data when the baby is smiling, the emotional state is determined to be "happy" and needs are identified. The analysis unit also develops an algorithm to identify needs based on the collected emotional data. For example, it builds an algorithm to classify needs according to the emotional state. As a result, by using the emotion estimation function, it becomes possible to identify needs taking into account the emotional state of the baby.

[0079] The analysis unit can compare data from different cultural or linguistic regions and identify common patterns. For example, the analysis unit compares data from babies from different cultural or linguistic regions and identifies common patterns. For example, it compares crying data from babies in Asia, Europe, and America to identify common needs. The analysis unit also collects voice data from babies from different linguistic regions and stores it in a database. For example, it collects voice data from babies in English-speaking and Spanish-speaking regions. The analysis unit then builds an algorithm to identify common patterns based on the collected data. For example, it analyzes data from different cultural or linguistic regions and extracts common patterns. This makes it possible to identify common patterns by comparing data from different cultural or linguistic regions.

[0080] The analysis unit can also analyze reaction data of parents or childcare workers and take interactions into account. For example, when understanding a baby's expressions and needs, the analysis unit also analyzes reaction data of parents and childcare workers and takes interactions into account. For example, it analyzes the parent's reaction when the baby cries and takes interactions into account. The analysis unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents talking to their babies. Furthermore, the analysis unit develops algorithms to analyze interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to analyze parent and childcare worker reaction data and take interactions into account.

[0081] The analysis unit can use the emotion estimation function to identify needs based on the baby's emotions and propose a response according to the emotions. For example, the analysis unit uses the emotion estimation function to identify needs based on the baby's emotions and propose a response according to the emotions. For example, when a baby is crying, the emotional state is determined to be "sad" and a response is proposed. The analysis unit also estimates the emotional state from the baby's facial expression data and proposes a response based on that data. For example, it analyzes facial expression data when a baby is smiling and proposes a response according to the emotional state as "happy." Furthermore, the analysis unit develops an algorithm that proposes a response based on the collected emotional data. For example, it builds an algorithm that proposes a response according to the emotional state. As a result, by using the emotion estimation function, it is possible to identify needs based on the baby's emotions and propose a response.

[0082] The reaction generation unit can develop an algorithm that refers to past success stories. For example, when providing a reaction that meets the baby's needs, the reaction generation unit develops an algorithm that refers to past success stories. For example, it learns reaction patterns that have been successful in the past and applies them in similar situations. The reaction generation unit also collects past success stories and stores them in a database. For example, it collects data on reactions that have been effective in the past. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected success stories. For example, it analyzes success stories and extracts reaction patterns. In this way, by developing an algorithm that refers to past success stories, more effective reactions become possible.

[0083] The reaction generation unit can incorporate feedback from parents or childcare workers and optimize reactions. For example, when providing a reaction according to the baby's needs, the reaction generation unit incorporates feedback from parents or childcare workers and optimizes the reaction. For example, if a parent says, "This reaction was effective," the reaction is optimized based on that feedback. The reaction generation unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the reaction generation unit builds an algorithm that optimizes reactions based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies reactions. In this way, the accuracy of reactions is improved by incorporating feedback from parents and childcare workers.

[0084] The reaction generation unit uses the emotion estimation function to provide a reaction according to the baby's emotional state, and can conduct a dialogue based on emotions. The reaction generation unit, for example, uses the emotion estimation function to provide a reaction according to the baby's emotional state and conduct a dialogue based on emotions. For example, when a baby is crying, the emotional state is determined to be "sad" and an appropriate reaction is provided. The reaction generation unit also estimates the emotional state from the baby's facial expression data and conducts a dialogue based on that data. For example, it analyzes facial expression data when the baby is smiling and conducts a dialogue based on the emotional state as "happy." Furthermore, the reaction generation unit develops an algorithm that generates a reaction based on the collected emotional data. For example, it builds an algorithm that generates a reaction according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide a reaction and dialogue according to the baby's emotional state.

[0085] The reaction generation unit can refer to reaction patterns from different cultural or linguistic regions. For example, the reaction generation unit refers to reaction patterns from different cultural or linguistic regions when providing reactions that meet the needs of a baby. For example, the reaction generation unit learns and applies reaction patterns from Asia, Europe, and America. The reaction generation unit also collects reaction data from different cultural or linguistic regions and stores it in a database. For example, it collects reaction data from babies in each region. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected reaction data. For example, it develops an algorithm that analyzes reaction patterns from different cultural or linguistic regions and generates reactions. This makes it possible to provide more diverse reactions by referring to reaction patterns from different cultural or linguistic regions.

[0086] The reaction generation unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, when providing a reaction according to the baby's needs, the reaction generation unit also analyzes reaction data of parents and childcare workers and takes interactions into consideration. For example, it analyzes the parent's reaction when the baby cries and takes interactions into consideration. The reaction generation unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents talking to their babies. Furthermore, the reaction generation unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to analyze reaction data of parents and childcare workers and take interactions into consideration.

[0087] The reaction generation unit uses the emotion estimation function to provide a reaction based on the baby's emotion and can conduct a dialogue according to the emotion. The reaction generation unit, for example, uses the emotion estimation function to provide a reaction based on the baby's emotion and conduct a dialogue according to the emotion. For example, when a baby is crying, the emotional state is determined to be "sad" and an appropriate reaction is provided. The reaction generation unit also estimates the emotional state from the baby's facial expression data and conducts a dialogue based on that data. For example, it analyzes facial expression data when the baby is smiling and conducts a dialogue based on the emotional state as "happy". Furthermore, the reaction generation unit develops an algorithm that generates a reaction based on the collected emotional data. For example, it builds an algorithm that generates a reaction according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide reactions and conduct a dialogue based on the baby's emotion.

[0088] The feedback unit can show trends by comparing with past data. For example, when feeding back the results of understanding a baby's needs and expressions, the feedback unit shows trends by comparing with past data. For example, it shows trends by comparing crying data from the past month with current crying data. The feedback unit also collects past data and stores it in a database. For example, it collects crying data from babies over the past year. Furthermore, the feedback unit builds an algorithm that shows trends based on the collected past data. For example, it develops an algorithm that analyzes past data and extracts trends of change. This makes it easier to understand changes in the baby's condition by showing trends by comparing with past data.

[0089] The feedback unit can incorporate feedback from parents or childcare workers and optimize the feedback content. For example, when providing feedback on the results of understanding the baby's needs and expressions, the feedback unit incorporates feedback from parents or childcare workers and optimizes the feedback content. For example, if a parent says, "This feedback was helpful," the content is optimized based on that feedback. The feedback unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the feedback unit builds an algorithm that optimizes the feedback content based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies the content. In this way, the accuracy of the feedback content is improved by incorporating feedback from parents and childcare workers.

[0090] The feedback unit can use the emotion estimation function to provide feedback on the baby's emotional state and suggest a response based on the emotion. For example, the feedback unit uses the emotion estimation function to provide feedback on the baby's emotional state and suggest a response based on the emotion. For example, when a baby is crying, the emotional state is determined to be "sad" and a response is suggested. The feedback unit also estimates the emotional state from the baby's facial expression data and suggests a response based on that data. For example, it analyzes facial expression data when a baby is smiling and suggests a response based on the emotional state as "happy." Furthermore, the feedback unit develops an algorithm that suggests a response based on the collected emotional data. For example, it builds an algorithm that presents a response depending on the emotional state. As a result, by using the emotion estimation function, it is possible to provide feedback and suggest a response based on the baby's emotional state.

[0091] The feedback unit can refer to data from different cultural or linguistic regions. For example, the feedback unit refers to data from different cultural or linguistic regions when providing feedback on the results of understanding a baby's needs and expressions. For example, the feedback unit refers to data from Asia, Europe, and America. The feedback unit also collects data from different cultural or linguistic regions and stores it in a database. For example, it collects data on babies in each region. Furthermore, the feedback unit builds an algorithm that provides feedback based on the collected data. For example, it develops an algorithm that analyzes data from different cultural or linguistic regions and provides feedback. This makes it possible to provide more diverse feedback by referring to data from different cultural or linguistic regions.

[0092] The feedback unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, when providing feedback on the results of understanding a baby's needs and expressions, the feedback unit also analyzes reaction data of parents and childcare workers and takes interactions into consideration. For example, it analyzes the parent's reaction when the baby cries and takes interactions into consideration. The feedback unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it collects voice data of parents speaking to their babies. Furthermore, the feedback unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it builds an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to provide feedback that takes interactions into consideration by analyzing reaction data of parents and childcare workers.

[0093] The feedback unit can use the emotion estimation function to provide feedback based on the baby's emotions and suggest a response according to the emotions. For example, the feedback unit uses the emotion estimation function to provide feedback based on the baby's emotions and suggest a response according to the emotions. For example, when a baby is crying, the emotional state is determined to be "sad" and a response is suggested. The feedback unit also estimates the emotional state from the baby's facial expression data and suggests a response based on that data. For example, it analyzes facial expression data when the baby is smiling and suggests a response according to the emotional state as "happy." Furthermore, the feedback unit develops an algorithm that suggests a response based on the collected emotional data. For example, it builds an algorithm that presents a response according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide feedback based on the baby's emotions and suggest a response.

[0094] The feedback unit can develop an algorithm that provides advice by referring to past success stories. For example, the feedback unit develops an algorithm that provides advice by referring to past success stories in order to promote communication with a baby. For example, it learns communication patterns that have been successful in the past and applies them in similar situations. The feedback unit also collects past success stories and stores them in a database. For example, it collects data on communication that has been effective in the past. Furthermore, the feedback unit builds an algorithm that provides advice based on the collected success stories. For example, it develops an algorithm that analyzes success stories and extracts advice patterns. This enables more effective feedback by providing advice by referring to past success stories.

[0095] The feedback unit can incorporate feedback from parents or childcare workers and optimize the advice content. The feedback unit incorporates feedback from parents or childcare workers and optimizes the advice content, for example, to promote communication with the baby. For example, if a parent says, "This advice was helpful," the content is optimized based on that feedback. The feedback unit also collects feedback data from parents and childcare workers and stores it in a database. For example, it collects audio data in which parents report on the baby's condition. Furthermore, the feedback unit builds an algorithm that optimizes the advice content based on the collected feedback data. For example, it develops an algorithm that analyzes the feedback data and modifies the content. In this way, the accuracy of the advice content is improved by incorporating feedback from parents and childcare workers.

[0096] The feedback unit can use the emotion estimation function to provide advice according to the baby's emotional state, thereby promoting emotion-based communication. The feedback unit, for example, uses the emotion estimation function to provide advice according to the baby's emotional state, thereby promoting emotion-based communication. For example, when a baby is crying, the emotional state is determined to be "sad" and appropriate advice is provided. The feedback unit also estimates the emotional state from the baby's facial expression data and provides advice based on that data. For example, the feedback unit analyzes facial expression data when the baby is smiling and provides advice based on the emotional state as "happy." The feedback unit also develops an algorithm that provides advice based on the collected emotional data. For example, it builds an algorithm that generates advice according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide advice according to the baby's emotional state and promote communication.

[0097] The feedback unit can refer to communication patterns in different cultural or linguistic regions. For example, the feedback unit refers to communication patterns in different cultural or linguistic regions to promote communication with babies. For example, it learns and applies communication patterns from Asia, Europe, and America. The feedback unit also collects communication data from different cultural or linguistic regions and stores it in a database. For example, it collects communication data from babies in each region. Furthermore, the feedback unit builds an algorithm that provides advice based on the collected data. For example, it develops an algorithm that analyzes communication patterns in different cultural or linguistic regions and provides advice. This makes it possible to provide more diverse advice by referring to communication patterns from different cultural or linguistic regions.

[0098] The feedback unit can also analyze reaction data of parents or childcare workers and take interactions into consideration. For example, in order to promote communication with babies, the feedback unit can also analyze reaction data of parents and childcare workers and take interactions into consideration. For example, it can analyze the parent's reaction when the baby cries and take interactions into consideration. The feedback unit also collects voice response data of parents and childcare workers and stores it in a database. For example, it can collect voice data of parents talking to their babies. Furthermore, the feedback unit develops an algorithm that analyzes interactions between babies and parents or childcare workers. For example, it can build an algorithm that combines and analyzes the baby's crying and the parent's reaction. This makes it possible to provide feedback that takes interactions into consideration by also analyzing reaction data of parents and childcare workers.

[0099] The feedback unit can use the emotion estimation function to provide advice based on the baby's emotions and promote emotion-appropriate communication. The feedback unit, for example, uses the emotion estimation function to provide advice based on the baby's emotions and promote emotion-appropriate communication. For example, when a baby is crying, the emotional state is determined to be "sad" and appropriate advice is provided. The feedback unit also estimates the emotional state from the baby's facial expression data and provides advice based on that data. For example, the feedback unit analyzes facial expression data when the baby is smiling and provides advice based on the emotional state as "happy." The feedback unit also develops an algorithm that provides advice based on the collected emotional data. For example, it builds an algorithm that generates advice according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide advice based on the baby's emotions and promote communication.

[0100] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0101] The database construction unit can collect reaction data of parents or childcare workers in addition to baby's voice data, and create a two-way communication database. For example, it records the parent's reaction when the baby cries. The database construction unit also collects voice response data of parents and childcare workers. For example, it collects voice data of parents talking to their babies. The database construction unit then adds the collected reaction data to a baby talk database. For example, it indexes and stores the collected reaction data in the database. This allows two-way communication by collecting reaction data of parents and childcare workers.

[0102] The analysis unit can analyze the baby's voice data in real time and instantly identify the meaning. For example, when a baby says "uuuuu," it can immediately identify that the baby is "hungry." The analysis unit also develops algorithms for analyzing voice data in real time. For example, it can convert voice data into text in real time using voice recognition technology. The analysis unit then builds a system that outputs the analysis results in real time. For example, it can analyze the baby's voice data and instantly notify the parents or childcare workers. This allows for real-time analysis of the baby's voice data and instantaneous identification of the meaning, enabling a rapid response.

[0103] The reaction generation unit can develop an algorithm that takes past success cases into consideration. For example, it learns reaction patterns that have been successful in the past and applies them in similar situations. The reaction generation unit also collects past success cases and stores them in a database. For example, it collects data on reactions that have been effective in the past. Furthermore, the reaction generation unit builds an algorithm that generates reactions based on the collected success cases. For example, it analyzes success cases and extracts reaction patterns. In this way, by developing an algorithm that takes past success cases into consideration, more effective reactions become possible.

[0104] The feedback unit can compare data with past data to show trends. For example, it can compare crying data from the past month with current crying data to show trends. The feedback unit also collects past data and stores it in a database. For example, it can collect data on babies' crying over the past year. The feedback unit then builds an algorithm that shows trends based on the collected past data. For example, it develops an algorithm that analyzes past data and extracts trends of change. This makes it easier to understand changes in the baby's condition by comparing it with past data to show trends.

[0105] The database construction unit can collect data on babies from different cultural or linguistic regions to construct a global database. For example, it collects voice data from babies from Asia, Europe, America, etc. The database construction unit also collects voice data from babies from different linguistic regions. For example, it collects voice data from babies from English-speaking countries, Spanish-speaking countries, etc. The database construction unit then adds the collected voice data to the global database. For example, it classifies and stores the collected voice data by region. This allows global analysis to be performed by collecting data from different cultural or linguistic regions.

[0106] The analysis unit can use the emotion estimation function to estimate emotions from the baby's voice and correct the analysis results based on those emotions. For example, it analyzes audio data when the baby is crying and corrects the analysis results by determining the emotional state as "sad." The analysis unit also estimates the emotional state from the baby's facial expression data and reflects that data in the analysis results. For example, it analyzes facial expression data when the baby is smiling and corrects the analysis results by determining the emotional state as "happy." The analysis unit also develops an algorithm to correct the analysis results based on the collected emotional data. For example, it builds an algorithm to modify the analysis results depending on the emotional state. This makes it possible to use the emotion estimation function to correct the analysis results based on the baby's emotions.

[0107] The reaction generation unit uses the emotion estimation function to provide a reaction according to the baby's emotional state, enabling emotion-based dialogue. For example, when a baby is crying, it determines the emotional state as "sad" and provides an appropriate reaction. The reaction generation unit also estimates the baby's emotional state from facial expression data and engages in dialogue based on that data. For example, it analyzes facial expression data when the baby is smiling and determines the emotional state as "happy" and engages in dialogue. The reaction generation unit also develops an algorithm that generates reactions based on the collected emotional data. For example, it builds an algorithm that generates reactions according to the emotional state. As a result, by using the emotion estimation function, it becomes possible to provide reactions and dialogue according to the baby's emotional state.

[0108] The feedback unit can use the emotion estimation function to provide feedback on the baby's emotional state and suggest responses based on the emotion. For example, when a baby is crying, the emotional state is considered to be "sad" and a response is suggested. The feedback unit can also estimate the emotional state from the baby's facial expression data and suggest a response based on that data. For example, it can analyze facial expression data when a baby is smiling, and suggest a response based on the emotional state as "happy." The feedback unit can also develop an algorithm that suggests a response based on the collected emotional data. For example, it can build an algorithm that presents a response plan depending on the emotional state. As a result, using the emotion estimation function makes it possible to provide feedback and suggest a response based on the baby's emotional state.

[0109] The analysis unit can use the emotion estimation function to consider the emotional state of the baby and identify needs based on the emotion. For example, when a baby is crying, the emotional state is determined to be "sad" and the needs are identified. The analysis unit can also estimate the emotional state from the baby's facial expression data and identify needs based on that data. For example, it can analyze facial expression data when the baby is smiling and identify the emotional state as "happy" and the needs are identified. The analysis unit can also develop an algorithm to identify needs based on the collected emotional data. For example, it can build an algorithm to classify needs according to the emotional state. As a result, the emotion estimation function can be used to identify needs taking into account the baby's emotional state.

[0110] The feedback unit uses the emotion estimation function to provide advice according to the baby's emotional state, thereby promoting emotion-based communication. For example, when a baby is crying, the emotional state is determined to be "sad" and appropriate advice is provided. The feedback unit also estimates the baby's emotional state from facial expression data and provides advice based on that data. For example, it analyzes facial expression data when a baby is smiling and provides advice based on the emotional state as "happy." The feedback unit also develops an algorithm that provides advice based on the collected emotional data. For example, it builds an algorithm that generates advice according to the emotional state. As a result, by using the emotion estimation function, it is possible to provide advice according to the baby's emotional state and promote communication.

[0111] The processing flow of the second embodiment will be briefly explained below.

[0112] Step 1: The database construction unit collects baby voice data and stores it digitally. For example, it collects baby crying, laughing, babbling, and other sounds and stores them in MP3 or WAV format. It also indexes and adds the collected voice data to a baby talk database. Step 2: The analysis unit analyzes the baby's voice data collected by the database construction unit. For example, it converts the baby's voice data into text using voice recognition technology and estimates the baby's emotions using an emotion analysis algorithm. It also analyzes the voice data to identify the baby's needs. Step 3: The reaction generator generates an appropriate reaction based on the results of the analysis. For example, if the baby says, "I'm hungry," the reaction generator responds by saying, "I see you're hungry. I'll prepare some food now." It also generates visual feedback and audio responses. Step 4: The feedback unit feeds back the reactions generated by the reaction generation unit to the parent or caregiver, for example, informing the parent why the baby is crying, reporting the baby's condition, or providing advice regarding the baby's needs.

[0113] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0114] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0115] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0116] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0117] 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0118] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0119] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0120] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0121] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0122] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0123] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0124] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0125] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0126] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart glasses 214 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0127] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0128] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0129] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0130] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0131] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0132] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0133] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0134] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0135] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0137] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0138] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0139] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0140] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0141] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 may also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0142] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0143] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0144] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0145] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0146] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0147] 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0148] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0149] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0150] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0151] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0152] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0153] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0154] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0155] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0156] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0157] In the robot 414, the processor 46 performs the identification process. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0158] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0159] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0160] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0161] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0162] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0163] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0164] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0165] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0166] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.

[0167] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[0169] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.

[0170] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0172] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0173] The hardware resource for executing a specific process can be any of the following processors: A CPU is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A dedicated electrical circuit, such as a field-programmable gate array (FPGA), a programmable logic device (PLD), or an application-specific integrated circuit (ASIC), is a processor with a circuit configuration specifically designed to execute a specific process. Each processor has built-in or connected memory, and uses the memory to execute the specific process.

[0174] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0175] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0176] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0177] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.

[0178] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0179] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]

[0180] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot

Claims

1. A database construction department that collects baby voice data; an analysis unit that analyzes the baby's voice data collected by the database construction unit; a reaction generation unit that generates an appropriate reaction based on the result of the analysis by the analysis unit; a feedback unit that feeds back the reaction generated by the reaction generation unit to a parent or a childcare worker. A system characterized by:

2. The database construction unit In addition to the baby's voice data, facial expression and movement data will also be collected to build a multimodal database. The system of claim 1 .

3. The analysis unit Analyze the baby's voice data in real time and instantly identify its meaning The system of claim 1 .

4. The reaction generation unit Developing algorithms based on past success stories The system of claim 1 .

5. The feedback unit Compare with historical data to show trends The system of claim 1 .

6. The database construction unit Using emotion estimation function, add the baby's emotional state to the database and perform emotion-based analysis. The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A