system

The system addresses the challenge of recording childcare events and emotional support by converting voice inputs to text, summarizing, and offering advice, enhancing caregiver experience through easy data storage and physical book creation.

JP2026073360APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Caregivers face challenges in recording child-rearing events and obtaining mental support due to their busy schedules, leading to potential interruptions in diary creation and increased isolation with unmet emotional needs.

Method used

A system that records childcare events via voice input, converts them into text data, summarizes them using a generative model, and provides emotional support, with the option to store and bind the data into a physical book.

Benefits of technology

Facilitates easy recording and emotional support for caregivers, reducing their burden by providing accessible and personalized advice, and creating a tangible childcare record.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073360000001_ABST
    Figure 2026073360000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An input device for users to voice-input events related to childcare, A speech recognition device that converts the aforementioned voice input into text data, A summarization device that uses a generation model to summarize the aforementioned text data into a predetermined number of characters, A storage means for storing the summarized text data in a storage device, An output device that converts the summary text stored in the aforementioned storage device into a format suitable for bookbinding, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] During child-rearing, caregivers face the problem that it is difficult to record events related to the growth and upbringing of their children due to their daily busyness. In addition, there are limited means to easily obtain mental support and advice regarding the burden of child-rearing. As a result, there is a possibility that the creation of a child-rearing diary will be interrupted, or that caregivers may become isolated while still having concerns.

Means for Solving the Problems

[0005] This invention solves these problems by providing a system that easily records events and concerns related to childcare via voice input, converts them into text data, and summarizes them. Furthermore, it provides an easy way to obtain emotional support by generating and presenting childcare advice and encouraging messages using a generative model. The recorded summary data is stored in a storage device and, if necessary, can be bound into a book for later review.

[0006] An "input device" is a device that allows users to communicate events and concerns related to childcare to the system through voice.

[0007] A "speech recognition device" is a technology or system for converting speech input received from a user into text data.

[0008] A "summarization device" is a device or program that has the function of summarizing converted text data using a generative model based on a predetermined number of characters.

[0009] A "generative model" is an algorithm or artificial intelligence technique that uses natural language processing to generate text summaries, advice, and messages of encouragement.

[0010] "Memory means" refers to a technique or device for storing summarized text data or other information and making it accessible at a later date.

[0011] An "output device" is a device or technology that displays or converts stored summary data into a format suitable for binding and presents it to the user.

[0012] A "response generation device" is a device or program that has the function of creating advice and empathetic messages using a generation model based on the content of the user's consultation.

[0013] "Binding means" refers to a function or procedure for converting digital summary data into a physical printed object. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides technology for implementing a system that allows parents raising children to easily record daily events and receive support for their childcare concerns. The system begins with the user inputting childcare events via voice through a device such as a smartphone or tablet. First, the device records the voice in real time and converts it into text data using speech recognition. The converted text data is sent to a server, where a generative model summarizes it to a user-specified number of characters. This summarized text data is stored in a storage device and can be accessed later.

[0036] Furthermore, when a user enters a concern related to childcare, the server analyzes that concern and generates appropriate advice and encouraging messages. These messages utilize a generation AI model and are designed to provide users with appropriate and heartwarming content. The generated advice is sent back to the device and presented to the user visually or audibly.

[0037] Furthermore, if a user wishes to have their saved childcare diary bound into a physical book, the server formats this summarized data and sends it to an output device to convert it into printable media. In this way, information stored as electronic data can later be used as a physical album or childcare record.

[0038] For example, if a user inputs an event such as "Today I went down the slide by myself at the park for the first time," the device converts the audio into text, and the server summarizes it and saves it as "First slide success!" Also, if a user inputs a concern such as "Raising children is hard and tiring," the server generates a message of encouragement such as "Thank you for your hard work raising children every day. You're doing a great job," which the device displays.

[0039] In this way, users can easily record childcare information and receive support.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The user opens the smartphone app and speaks about events or concerns related to childcare. At this time, they tap the voice input button to start recording.

[0043] Step 2:

[0044] The device records the user's voice and uses a speech recognition engine to convert the voice into text data. This process takes place in real time within the device.

[0045] Step 3:

[0046] The terminal sends the converted text data to the server via an HTTP request. This request includes metadata such as the date and time the audio was recorded and the user ID.

[0047] Step 4:

[0048] The server provides the received text data to an AI model that generates a summary of a specified length. During this process, the most important content of the text is condensed.

[0049] Step 5:

[0050] The server stores the summarized text in a database. This stored data can later be accessed by users or used for bookbinding purposes.

[0051] Step 6:

[0052] If the input is a user's problem or concern, the server analyzes the text data and uses a generative AI model to generate advice and encouraging messages. The generated messages will be tailored to the user's situation.

[0053] Step 7:

[0054] The server sends the generated summary or advice back to the device. The returned data is visually displayed to the user through the app's interface, and can also be presented audibly if voice guidance is available.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] There is a need for a way to easily record daily events and parental concerns related to childcare, making them easily accessible for later reference, and for providing advice and support based on this information. However, parents raising children often have limited time and knowledge, making it difficult to do this using conventional methods. Furthermore, there is the problem of the time and effort required to physically store and bind the recorded information.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes an input means for a user to input childcare-related activities via voice input through an information terminal, a voice recognition means for converting the voice input into text format, and a summarization means that uses generation technology to compress the text data to a specified number of characters. This makes it possible to quickly and efficiently record childcare events and store them in a format that can be easily referenced later. Furthermore, since advice and support generated using generation technology can be provided in real time, the burden on parents can be reduced and emotional support can be provided.

[0060] An "information terminal" is a device used by a user for voice input, and includes smartphones, tablets, and personal computers.

[0061] "Voice input" refers to the act of users communicating events or questions related to childcare to the system using voice.

[0062] "Speech recognition" is a technology that analyzes input speech data and converts it into text format.

[0063] "Generative technology" refers to techniques that primarily use generative AI models to summarize data and create appropriate messages.

[0064] A "summarization method" is a technique that compresses text data to a specified number of characters and processes it to simplify the information.

[0065] "Memory device" refers to the function of storing summarized text data in a database so that it can be referenced later.

[0066] "Output means" refers to a device that physically or digitally displays or outputs data in order to provide information to users.

[0067] "Bookbinding methods" refer to technologies for formatting digital data into a printable format and providing it as a physical book or album.

[0068] This invention presents a specific embodiment of a system that allows parents raising children to easily record everyday events and concerns related to childcare and receive appropriate advice. The details are described below.

[0069] The term "device" refers to information terminals such as smartphones and tablets, which users use to access the system. Users input information about childcare-related events or topics they wish to discuss using voice input. The device utilizes speech recognition software (e.g., Google® Speech-to-Text API) to convert the voice data into text data.

[0070] The server receives the converted text data and utilizes generation technology. This generation technology includes a generation AI model (e.g., GPT-3® or similar models), which enables a summarization method that compresses the input text data to a specified number of characters. It also has a function to generate advice and empathetic messages according to the user's concerns.

[0071] The generated summaries and messages are stored in a database by the server. This storage mechanism allows users to refer to past records at any time. Furthermore, the stored text data can be sent to the terminal as needed and presented to the user visually or audibly.

[0072] Furthermore, if a user wishes to save their childcare diary in a physical form, the server will format the digital data and convert it into a printable format such as PDF. This allows it to be bound into a physical album or diary via an output device.

[0073] As a concrete example of its use, if a user records an event such as, "Today I went down the slide by myself at the park for the first time," the device will transcribe it into text, and the server will summarize it as, "First slide success!" Also, if a user inputs a complaint such as, "Raising children is tough and exhausting," the server will generate an encouraging message such as, "You're doing a great job with childcare every day. You're doing a wonderful job," and the device will display this message.

[0074] A concrete example of a prompt sentence is, "My child experienced a new game at the park. I would like to briefly record it." This is used when the generative AI model performs summarization.

[0075] In this way, the system provides convenient and effective record-keeping and support for parents raising children.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] Users input events and concerns related to childcare using voice commands via an information terminal. The terminal captures the user's voice using a microphone. This input data is stored as raw audio data.

[0079] Step 2:

[0080] The device uses speech recognition software to convert captured audio data into text data. Specifically, the audio waveform data is converted into text format, and noise reduction and normalization are performed accordingly. As a result, the data is output in the form of text sentences.

[0081] Step 3:

[0082] The terminal sends the converted text data to the server. For security reasons, the data is encrypted and transmitted over the network. The transmitted data is received by the server and ready for analysis.

[0083] Step 4:

[0084] The server uses generation technology to compress the received text data to a specified number of characters. Summarization is performed by inputting a prompt into the generation AI model. This prompt includes specific instructions such as, "Please summarize this text briefly." The summarization result is obtained as shortened text data.

[0085] Step 5:

[0086] The server saves the summarized text data to a database. This saving process involves organizing the data and writing it to the database. The saved data can then be searched later using specific dates, keywords, or other criteria.

[0087] Step 6:

[0088] When a user submits a parenting consultation request, the server analyzes the text data and, if necessary, uses a generative AI model to generate advice and encouraging messages. This process creates a customized response tailored to the user's concerns. The generated messages are output in text or audio format.

[0089] Step 7:

[0090] The server sends the generated message to the terminal. The received message is displayed on the terminal's screen or read aloud using text-to-speech functionality. This allows the user to confirm the message visually or audibly.

[0091] Step 8:

[0092] If a user requests a physical binding of their saved baby journal, the server formats this data and converts it into a printable format. In this process, the text data is converted into a document format such as PDF, sent to an output device, and obtained as a physical album.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] In modern parenting, parents lack the time to record daily events and are also required to respond quickly to parenting concerns. However, it is difficult to meet these needs simultaneously using traditional methods. Furthermore, the lack of adequate systems for efficiently recording parenting events and obtaining appropriate support is increasing the mental burden on parents. In addition, there is a need for improved services that comfortably support parenting, such as recommendations for relevant products based on recorded events.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes information processing means for users to input life events by voice, language processing means for converting the voice input into information data, and information compression means that uses a generation method for summarizing the information data into a predetermined number of symbols. This allows users to easily record events related to childcare and receive efficient and comprehensive support by being recommended related products based on those events.

[0098] "Information processing means" refers to a device that has the function of allowing users to input life events and consultations via voice.

[0099] "Language processing means" refers to technology that converts speech input into information data, specifically by converting it into text data through speech recognition.

[0100] "Information compression means" refers to a function that uses a generation method to summarize acquired information data into a predetermined number of symbols.

[0101] An "information storage device" is a memory device that has the function of accumulating summarized information data.

[0102] A "media control means" is a system that controls summary information stored in an information storage means and provides it to users as needed.

[0103] "Media output means" refers to a device that has the function of converting summary information extracted from information storage means into a recordable format and outputting it.

[0104] A "product recommendation tool" is a function that recommends relevant products based on the user's life events.

[0105] A "response generation means" is a device that has a generation method for generating advice or emotional messages based on information data and presenting them to the user.

[0106] This invention provides a system that allows users to easily record life events and receive childcare-related support. The system mainly consists of smart devices and a server.

[0107] First, users input childcare-related information using voice commands via the input device of their smart device. The device receives the voice input and converts the voice signal into text data using a speech recognition API. Google Speech-to-Text is one of the speech recognition APIs used.

[0108] Next, the server receives the converted text data and uses a generative AI model to summarize the information data into a predetermined number of symbols. This summarized data is stored in a cloud database via an information storage system. OpenAI's GPT model is used as the generative AI model.

[0109] Furthermore, when a user seeks advice regarding childcare, the server uses a generative AI model to generate appropriate advice messages and empathetic messages based on the converted text data, and sends them back to the device. The device then presents these messages to the user visually or audibly.

[0110] In addition, the server can select and recommend relevant childcare products from online stores based on the user's life events. This recommendation function can meet a variety of needs in childcare life.

[0111] For example, if a user inputs the voice message, "My child laughed for the first time today," the AI ​​model will summarize this as "First laugh successful!" and save it. Related products that might be recommended include toys that help elicit smiles. An example of a prompt based on this example would be, "Please record your baby's first smile and recommend related products."

[0112] This invention will improve the efficiency of childcare record-keeping and promote support, allowing users to easily obtain useful information.

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] Users input childcare-related information by voice via a smart device. The input voice data is captured by the microphone in the smart device. This voice data serves as the starting point for determining what processing will be performed next.

[0116] Step 2:

[0117] The device receives the captured audio data and converts it into text data using a speech recognition API. Google Speech-to-Text is used as the speech recognition API in this case. The audio data is processed as input, and text data is generated as output. This text data semantically preserves the information provided by the user via voice.

[0118] Step 3:

[0119] The server receives text data sent from the terminal. Based on the received text data, it uses a generative AI model (e.g., the OpenAI GPT model) to summarize the information data into a predetermined number of symbols. The goal is to compress the input text data and summarize it into language that is easy for the user to understand.

[0120] Step 4:

[0121] The server saves the summarized text data to an information storage system. A cloud database is used here to securely store the summarized information. The saved information can be accessed later and used like an electronic diary.

[0122] Step 5:

[0123] When a user asks for advice related to childcare, they input the content into the system again via voice. The terminal captures the voice data as described above and converts it into text data using a speech recognition API.

[0124] Step 6:

[0125] The server receives text data related to the consultation and uses a generative AI model to generate advice messages and empathetic messages tailored to the content of the consultation. Here, the generative AI model processes the prompt text and outputs information that is useful to the user.

[0126] Step 7:

[0127] The server sends the generated message back to the terminal, which then presents it to the user visually or audibly. Here, the feedback the user receives greatly influences the user experience.

[0128] Step 8:

[0129] The server complementarily selects and recommends relevant childcare products from a virtual store based on the user's childcare information. This recommendation information aims to provide highly relevant products based on the information entered by the user.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention relates to a system that records events and concerns related to childcare and provides support combined with emotion analysis. Users input childcare-related events via voice using a device such as a smartphone or tablet. This voice data is recorded on the device and converted into text data by a voice recognition device.

[0132] The terminal then sends the converted text data to the server. The server is equipped with a generative model and an emotion engine, which analyzes the user's emotions based on the content of the text data. The analyzed emotion data is input into the generative model, which generates a summarized text that reflects the emotional information. This results in a summary of the child's growth record that includes emotional nuances.

[0133] Furthermore, for inquiries regarding childcare, advice and empathetic messages are generated that are tailored to the user's emotions. Based on the analysis results of the emotion engine, the generation model responds in a way that is sensitive to the user's feelings. This provides users with more appropriate and personalized support. For example, if a user is feeling anxious, a message such as "I understand your anxiety. Thank you for your daily efforts" is provided.

[0134] The saved summary data retains emotional information even when converted into a format suitable for printing, as instructed by the user. This results in richer records as a parenting diary, allowing users to relive the emotions of the time when looking back. This system enables users to easily record their emotional experiences related to parenting and receive support.

[0135] The following describes the processing flow.

[0136] Step 1:

[0137] The user launches the smartphone app and speaks about events or concerns related to childcare. They then press the voice input button to start recording.

[0138] Step 2:

[0139] The device records the user's voice and uses its built-in speech recognition device to convert the voice data into text data in real time.

[0140] Step 3:

[0141] The terminal sends the converted text data to the sentiment engine for sentiment analysis. This analysis yields user sentiment data associated with the input text.

[0142] Step 4:

[0143] The device sends text data and the resulting sentiment data to the server. This data is sent in HTTP request format and includes metadata such as user information and timestamps.

[0144] Step 5:

[0145] The server inputs the received text data and sentiment data into a generative model and generates a summary of the specified length, reflecting the sentiment. The summary will reflect the nuances of the emotions.

[0146] Step 6:

[0147] The server saves the generated sentiment-reflecting summaries to a database. This saved summary data is then used by users for later review.

[0148] Step 7:

[0149] The server analyzes text data entered by the user that is categorized as a consultation or problem, and generates appropriate advice and encouraging messages. Based on the sentiment analysis results, the generative model provides responses that take the user's emotions into consideration.

[0150] Step 8:

[0151] The server sends the generated summary and advice messages to the terminal. The terminal displays the received data in its user interface, providing visual and auditory feedback to the user.

[0152] Step 9:

[0153] Users can instruct the system to bind their saved summary data into a book. The server then formats the summary data into a printable format and sends it to a printing company. During this process, emotional information is also incorporated.

[0154] (Example 2)

[0155] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0156] In recording and seeking advice regarding childcare-related events, there are limited ways to receive support that reflects the user's emotions. Traditional systems struggle to automatically generate emotionally balanced summaries and advice, preventing users from receiving personalized support tailored to their specific situations. This has led to increased childcare burdens and a lack of emotional support.

[0157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0158] In this invention, the server includes an input means for the user to input childcare-related events and consultations via an information device using voice input, a voice recognition means for converting the voice input into text information, and an emotion analysis means for analyzing the text information to identify emotions. This enables the automatic generation of summaries and advice that reflect the user's emotions.

[0159] "Input means" refers to a device or method by which a user inputs voice via an information device.

[0160] "Speech recognition means" refers to a technology or device that converts speech data into text information.

[0161] "Emotional analysis means" refers to a technology or method that analyzes textual information and identifies emotions from its content.

[0162] A "generative model" is an algorithm or computer program that generates information based on given data.

[0163] "Summarization techniques" refer to techniques or methods for shortening information and extracting key points.

[0164] A "recording medium" is a physical or electronic medium used to store data.

[0165] "Output means" refers to a device or method for presenting processed information to a user.

[0166] "Response generation means" refers to a technology or method for generating advice or messages based on input information.

[0167] "Conversion means" refers to a technique or method for changing data format to another format.

[0168] This invention is a system that records events related to childcare, performs emotional analysis, and provides appropriate support to the user. The following describes how this system can be implemented.

[0169] Users input childcare-related events and questions via voice using devices such as smartphones and tablets. These devices are equipped with microphones to convert voice data into digital data. This voice data is then converted into text data on the device using speech recognition software such as Google Speech-to-Text.

[0170] The converted text data is sent to a server via the internet. The server is equipped with an emotion engine that analyzes emotions based on the text information. For example, IBM Watson® Natural Language Understanding performs this process. Once the server receives the results of the emotion analysis, it uses a generative AI model such as OpenAI's GPT to generate a summary or response message that reflects those emotions.

[0171] In this system, for example, if a user makes a voice input such as, "My child went to kindergarten for the first time today, and I felt lonely," the emotion engine will extract the emotion of "loneliness." Then, a generative AI model will create a summary such as, "As you watch your child grow, you may experience changes in your feelings, but this is a wonderful step."

[0172] Ultimately, the device presents these summaries and messages received from the server to the user. Furthermore, if requested by the user, the saved data can be converted into a printable format, such as PDF. This allows for a more emotionally rich record of childcare experiences, while also providing access to appropriate advice.

[0173] Through this system, users can receive more personalized support in childcare. An example of a prompt message is: "Please input your childcare experiences or questions via voice. The system will analyze your emotions and provide appropriate summaries and messages."

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] The user uses the device to input childcare-related events and questions via voice. The input voice is recorded by the device's microphone and saved as digital audio data. The input data is the user's voice, and the output is that audio data.

[0177] Step 2:

[0178] The device passes the recorded audio data to speech recognition software, which converts it into text data. At this stage, audio data is input and text data is output as written information. Specifically, the device uses the internet to send the audio data to a cloud-based speech recognition service and receives the text data from there.

[0179] Step 3:

[0180] The terminal sends the converted text data to the server. In this process, text data is input and sent to the server. Specifically, the terminal sends the text data as an HTTP POST request, and communication is securely performed using the SSL / TLS protocol.

[0181] Step 4:

[0182] The server inputs the received text data into the sentiment engine for sentiment analysis. At this stage, the text data is provided as input, and sentiment scores and labels are output. Specifically, the sentiment engine analyzes the content of the text and identifies emotions such as positive, negative, and neutral.

[0183] Step 5:

[0184] The server uses the sentiment analysis results to input information into a generative AI model, which then generates summaries and messages that reflect the emotions. The input here is text data and its sentiment score, while the output is summarized text and response messages. Specifically, the generative AI model receives text and sentiment information and forms a response based on it.

[0185] Step 6:

[0186] The server returns the generated summary and response message to the terminal. In this process, the summary text and message are input and sent as output to the terminal. Specifically, the server generates an HTTP response, the terminal receives it, and displays it to the user.

[0187] Step 7:

[0188] The user operates the terminal to convert the saved summary data into a format suitable for printing, as needed. In this step, the summary data is input, and the output is in a format such as PDF. Specifically, the software on the terminal formats the data, generates a PDF, and saves it to local storage.

[0189] In this way, the system efficiently records events and emotions related to childcare and provides support.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] As there is a growing need for real-time and appropriate support for childcare-related events and concerns, there is a need to strengthen childcare-related support at physical stores and provide users with personalized, immediate advice. However, conventional systems have made it difficult to provide emotionally resonant support and to create printed information in a physical store environment.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes an input means for recognizing information input via voice, a voice recognition means for converting it into text information, and a summarization means using sentiment analysis and a generative model. This makes it possible to provide real-time advice tailored to the user's emotions, as well as compress and compile information, even in a physical store environment.

[0195] "Information input via voice" refers to data and messages that users send to a device using spoken language.

[0196] "Speech recognition means" refers to technologies and devices that analyze speech information and convert it into text information.

[0197] A "generative model" is a type of artificial intelligence technology that generates new data or information based on given data.

[0198] A "summarization tool" refers to a process or device that extracts important information from a large amount of data and summarizes it concisely.

[0199] A "memory device" refers to a device or method for holding and storing data or information.

[0200] "Output means" refers to devices or technologies for displaying or physically outputting processed information to an external source.

[0201] An "interface means" refers to a device or method that serves as a window for information exchange between people and systems.

[0202] A "digital visual device" refers to a wearable device or screen that displays visual information using digital technology.

[0203] "Personalized responses" refer to feedback and messages that are customized based on the user's individual circumstances and emotions.

[0204] "Devices for converting to print media" refer to equipment and methods used to convert digital data into physical printed materials.

[0205] The system realizing this invention works in conjunction with smart glasses or digital visual devices to process the user's childcare-related voice input. When a user asks for advice on childcare while selecting products, the system first collects the voice using the microphone on the smart glasses. The voice is recorded on the device and transmitted to a server via the network.

[0206] The server converts speech data into text using speech recognition software (e.g., Google Speech-to-Text API). This text is then analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer) for sentiment analysis. The analyzed sentiment data is then input into a generative AI model (e.g., OpenAI's GPT) to generate personalized advice and responses based on the user's emotions.

[0207] The generated response is displayed in real time on the smart glasses' screen. This allows users to receive immediate and appropriate support regarding childcare within the store. For example, if a user is unsure about choosing diapers, the smart glasses will display a message such as, "We understand your concerns. Here are some suggestions for improvement..."

[0208] Examples of prompt messages include the following:

[0209] "Users are feeling anxious about choosing diapers. To alleviate their anxiety, please recommend a brand in 10 characters or less."

[0210] This system allows for real-time consultation and information exchange regarding childcare, enabling a more effective shopping experience.

[0211] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0212] Step 1:

[0213] The user inputs childcare-related information by voice through the microphone on their smart glasses. The input voice data is temporarily stored on the device, and then the device transmits this voice data to a server via the network.

[0214] Step 2:

[0215] The server converts the received audio data into text data using speech recognition software. In this process, the input audio data is analyzed, and a corresponding string of characters is generated. The output text data is then used for subsequent processing.

[0216] Step 3:

[0217] The server inputs text data into the emotion engine, which analyzes the user's emotions. The emotion engine detects emotional cues contained in the text data and extracts emotional indicators such as joy, sadness, and anxiety. This analysis result is used as data necessary for the next generation process.

[0218] Step 4:

[0219] The analyzed emotional and textual data are input into a generative AI model. Based on this data, the generative AI model generates personalized advice and empathetic messages that resonate with the user's emotions. These output messages play a crucial role in providing instantaneous support to the user.

[0220] Step 5:

[0221] The server sends the generated advice and messages back to the device. The device displays this information on the smart glasses' screen. Users can immediately receive this feedback visually and use it as a reference for shopping or childcare advice.

[0222] This entire process allows users to seek real-time advice on childcare in physical stores and receive empathetic feedback.

[0223] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0224] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0225] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0226] [Second Embodiment]

[0227] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0228] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0229] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0230] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0231] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0232] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0233] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0234] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0235] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0236] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0237] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0238] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0239] This invention provides technology for implementing a system that allows parents raising children to easily record daily events and receive support for their childcare concerns. The system begins with the user inputting childcare events via voice through a device such as a smartphone or tablet. First, the device records the voice in real time and converts it into text data using speech recognition. The converted text data is sent to a server, where a generative model summarizes it to a user-specified number of characters. This summarized text data is stored in a storage device and can be accessed later.

[0240] Furthermore, when a user enters a concern related to childcare, the server analyzes that concern and generates appropriate advice and encouraging messages. These messages utilize a generation AI model and are designed to provide users with appropriate and heartwarming content. The generated advice is sent back to the device and presented to the user visually or audibly.

[0241] Furthermore, if a user wishes to have their saved childcare diary bound into a physical book, the server formats this summarized data and sends it to an output device to convert it into printable media. In this way, information stored as electronic data can later be used as a physical album or childcare record.

[0242] For example, if a user inputs an event such as "Today I went down the slide by myself at the park for the first time," the device converts the audio into text, and the server summarizes it and saves it as "First slide success!" Also, if a user inputs a concern such as "Raising children is hard and tiring," the server generates a message of encouragement such as "Thank you for your hard work raising children every day. You're doing a great job," which the device displays.

[0243] In this way, users can easily record childcare information and receive support.

[0244] The following describes the processing flow.

[0245] Step 1:

[0246] The user opens the smartphone app and speaks about events or concerns related to childcare. At this time, they tap the voice input button to start recording.

[0247] Step 2:

[0248] The device records the user's voice and uses a speech recognition engine to convert the audio into text data. This process takes place in real time within the device.

[0249] Step 3:

[0250] The terminal sends the converted text data to the server via an HTTP request. This request includes metadata such as the date and time the audio was recorded and the user ID.

[0251] Step 4:

[0252] The server provides the received text data to an AI model that generates a summary of a specified length. During this process, the most important content of the text is condensed.

[0253] Step 5:

[0254] The server stores the summarized text in a database. This stored data can later be accessed by users or used for bookbinding purposes.

[0255] Step 6:

[0256] If the input is a user's problem or concern, the server analyzes the text data and uses a generative AI model to generate advice and encouraging messages. The generated messages will be tailored to the user's situation.

[0257] Step 7:

[0258] The server sends the generated summary or advice back to the device. The returned data is visually displayed to the user through the app's interface, and can also be presented audibly if voice guidance is available.

[0259] (Example 1)

[0260] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0261] There is a need for a way to easily record daily events and parental concerns related to childcare, making them easily accessible for later reference, and for providing advice and support based on this information. However, parents raising children often have limited time and knowledge, making it difficult to do this using conventional methods. Furthermore, there is the problem of the time and effort required to physically store and bind the recorded information.

[0262] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0263] In this invention, the server includes an input means for a user to input childcare-related activities via voice input through an information terminal, a voice recognition means for converting the voice input into text format, and a summarization means that uses generation technology to compress the text data to a specified number of characters. This makes it possible to quickly and efficiently record childcare events and store them in a format that can be easily referenced later. Furthermore, since advice and support generated using generation technology can be provided in real time, the burden on parents can be reduced and emotional support can be provided.

[0264] An "information terminal" is a device used by a user for voice input, and includes smartphones, tablets, and personal computers.

[0265] "Voice input" refers to the act of users communicating events or questions related to childcare to the system using voice.

[0266] "Speech recognition" is a technology that analyzes input speech data and converts it into text format.

[0267] "Generative technology" refers to techniques that primarily use generative AI models to summarize data and create appropriate messages.

[0268] A "summarization method" is a technique that compresses text data to a specified number of characters and processes it to simplify the information.

[0269] "Memory device" refers to the function of storing summarized text data in a database so that it can be referenced later.

[0270] "Output means" refers to a device that physically or digitally displays or outputs data in order to provide information to users.

[0271] "Bookbinding methods" refer to technologies for formatting digital data into a printable format and providing it as a physical book or album.

[0272] This invention presents a specific embodiment of a system that allows parents raising children to easily record everyday events and concerns related to childcare and receive appropriate advice. The details are described below.

[0273] The term "device" refers to information terminals such as smartphones and tablets, which users use to access the system. Users input information about childcare-related events or topics they wish to discuss using voice input. The device utilizes speech recognition software (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[0274] The server receives the converted text data and utilizes generation technology. This generation technology includes a generative AI model (e.g., GPT-3 or similar models), which enables a summarization method that compresses the input text data to a specified number of characters. It also has a function to generate advice and empathetic messages according to the user's concerns.

[0275] The generated summaries and messages are stored in a database by the server. This storage mechanism allows users to refer to past records at any time. Furthermore, the stored text data can be sent to the terminal as needed and presented to the user visually or audibly.

[0276] Also, when the user wants to save the childcare diary in a physical form, the server formats the digital data and converts it into a printable form such as PDF. As a result, it can be bound as a physical album or diary via an output device.

[0277] As a specific usage example, when the user records an event such as "Today, I slid down the slide alone in the park for the first time", the terminal converts this into text, and the server summarizes it as "First successful slide down the slide!". Also, if the user enters a worry such as "Childcare is very difficult and tiring", the server generates a motivating message such as "Thank you for your daily childcare efforts. You are working very hard", and the terminal presents this message.

[0278] As a specific example of a prompt sentence, there is "The child experienced a new play in the park. I want to record it briefly.", which is used when the generative AI model performs summarization.

[0279] In this way, the system is highly convenient for caregivers during childcare and provides effective recording and support.

[0280] The flow of the specific process in Example 1 will be described using FIG. 11.

[0281] Step 1:

[0282] The user uses the information terminal to voice-input events and worries related to childcare. The terminal captures the user's voice with a microphone. This input data is stored as raw voice data.

[0283] Step 2:

[0284] The terminal uses voice recognition software to convert the captured voice data into text data. Specifically, the waveform data of the voice is converted into text format, and along with this, noise removal and normalization are performed. As a result, the data is output in the form of a text sentence.

[0285] Step 3:

[0286] The terminal sends the converted text data to the server. At this time, the data is encrypted considering security and sent via the network. The sent data is received on the server side and is in a state ready for analysis.

[0287] Step 4:

[0288] The server uses a generation technique to compress the received text data to a specified number of characters. Summarization is performed by inputting a prompt sentence into the generation AI model. This prompt includes specific instructions such as "Please summarize this article briefly." The summary result is obtained as shortened text data.

[0289] Step 5:

[0290] The server saves the summarized text data in the database. In this saving process, the data is organized and written into the database. The saved data can be searched later by a specific date and time or keywords.

[0291] Step 6:

[0292] When the user inputs a child-rearing consultation, the server analyzes the text data and, if necessary, uses the generation AI model to generate advice or encouraging messages. In this process, a customized response according to the user's worries is created. The generated message is output in text or voice format.

[0293] Step 7:

[0294] The server sends the generated message to the terminal. The received message is displayed on the terminal's display or read out using the voice synthesis function. As a result, the user can visually or aurally confirm the message.

[0295] Step 8:

[0296] If a user requests a physical binding of their saved baby journal, the server formats this data and converts it into a printable format. In this process, the text data is converted into a document format such as PDF, sent to an output device, and obtained as a physical album.

[0297] (Application Example 1)

[0298] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0299] In modern parenting, parents lack the time to record daily events and are also required to respond quickly to parenting concerns. However, it is difficult to meet these needs simultaneously using traditional methods. Furthermore, the lack of adequate systems for efficiently recording parenting events and obtaining appropriate support is increasing the mental burden on parents. In addition, there is a need for improved services that comfortably support parenting, such as recommendations for relevant products based on recorded events.

[0300] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0301] In this invention, the server includes information processing means for users to input life events by voice, language processing means for converting the voice input into information data, and information compression means that uses a generation method for summarizing the information data into a predetermined number of symbols. This allows users to easily record events related to childcare and receive efficient and comprehensive support by being recommended related products based on those events.

[0302] "Information processing means" refers to a device that has the function of allowing users to input life events and consultations via voice.

[0303] The "language processing means" refers to the technology that converts voice input into information data, and performs character data conversion through voice recognition.

[0304] The "information compression means" is a function that uses a generation method to summarize the acquired information data into a predetermined number of symbols.

[0305] The "information storage means" is a storage device having a function of accumulating the summarized information data.

[0306] The "media control means" is a system having a function of controlling the summarized information stored in the information storage means and providing it to the user as necessary.

[0307] The "media output means" is a device having a function of converting the summarized information extracted from the information storage means into a recordable format and outputting it.

[0308] The "product presentation means" is a function that plays a role of recommending related products based on the user's life events.

[0309] The "response generation means" is a device having a generation method for generating advice and emotional messages based on the information data and presenting them to the user.

[0310] This invention provides a system for the user to easily record life events and receive support related to child-rearing. The system is mainly composed of a smart device and a server.

[0311] First, the user uses the input device of the smart device to input voice information related to child-rearing. The terminal receives the voice input and converts the voice signal into text data using a voice recognition API. As the voice recognition API, Google Speech-to-Text, etc. are used.

[0312] Next, the server receives the converted text data and uses a generative AI model to summarize the information data into a predetermined number of symbols. This summarized data is stored in a cloud database via an information storage system. OpenAI's GPT model is used as the generative AI model.

[0313] Furthermore, when a user seeks advice regarding childcare, the server uses a generative AI model to generate appropriate advice messages and empathetic messages based on the converted text data, and sends them back to the device. The device then presents these messages to the user visually or audibly.

[0314] In addition, the server can select and recommend relevant childcare products from online stores based on the user's life events. This recommendation function can meet a variety of needs in childcare life.

[0315] For example, if a user inputs the voice message, "My child laughed for the first time today," the AI ​​model will summarize this as "First laugh successful!" and save it. Related products that might be recommended include toys that help elicit smiles. An example of a prompt based on this example would be, "Please record your baby's first smile and recommend related products."

[0316] This invention will improve the efficiency of childcare record-keeping and promote support, allowing users to easily obtain useful information.

[0317] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0318] Step 1:

[0319] Users input childcare-related information by voice via a smart device. The input voice data is captured by the microphone in the smart device. This voice data serves as the starting point for determining what processing will be performed next.

[0320] Step 2:

[0321] The device receives the captured audio data and converts it into text data using a speech recognition API. Google Speech-to-Text is used as the speech recognition API in this case. The audio data is processed as input, and text data is generated as output. This text data semantically preserves the information provided by the user via voice.

[0322] Step 3:

[0323] The server receives text data sent from the terminal. Based on the received text data, it uses a generative AI model (e.g., the OpenAI GPT model) to summarize the information data into a predetermined number of symbols. The goal is to compress the input text data and summarize it into language that is easy for the user to understand.

[0324] Step 4:

[0325] The server saves the summarized text data to an information storage system. A cloud database is used here to securely store the summarized information. The saved information can be accessed later and used like an electronic diary.

[0326] Step 5:

[0327] When a user asks for advice related to childcare, they input the content into the system again via voice. The terminal captures the voice data as described above and converts it into text data using a speech recognition API.

[0328] Step 6:

[0329] The server receives text data related to the consultation and uses a generative AI model to generate advice messages and empathetic messages tailored to the content of the consultation. Here, the generative AI model processes the prompt text and outputs information that is useful to the user.

[0330] Step 7:

[0331] The server sends the generated message back to the terminal, which then presents it to the user visually or audibly. Here, the feedback the user receives greatly influences the user experience.

[0332] Step 8:

[0333] The server complementarily selects and recommends relevant childcare products from a virtual store based on the user's childcare information. This recommendation information aims to provide highly relevant products based on the information entered by the user.

[0334] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0335] This invention relates to a system that records events and concerns related to childcare and provides support combined with emotion analysis. Users input childcare-related events via voice using a device such as a smartphone or tablet. This voice data is recorded on the device and converted into text data by a voice recognition device.

[0336] The terminal then sends the converted text data to the server. The server is equipped with a generative model and an emotion engine, which analyzes the user's emotions based on the content of the text data. The analyzed emotion data is input into the generative model, which generates a summarized text that reflects the emotional information. This results in a summary of the child's growth record that includes emotional nuances.

[0337] Furthermore, for inquiries regarding childcare, advice and empathetic messages are generated that are tailored to the user's emotions. Based on the analysis results of the emotion engine, the generation model responds in a way that is sensitive to the user's feelings. This provides users with more appropriate and personalized support. For example, if a user is feeling anxious, a message such as "I understand your anxiety. Thank you for your daily efforts" is provided.

[0338] The saved summary data retains emotional information even when converted into a format suitable for printing, as instructed by the user. This results in richer records as a parenting diary, allowing users to relive the emotions of the time when looking back. This system enables users to easily record their emotional experiences related to parenting and receive support.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] The user launches the smartphone app and speaks about events or concerns related to childcare. They then press the voice input button to start recording.

[0342] Step 2:

[0343] The device records the user's voice and uses its built-in speech recognition device to convert the voice data into text data in real time.

[0344] Step 3:

[0345] The terminal sends the converted text data to the sentiment engine for sentiment analysis. This analysis yields user sentiment data associated with the input text.

[0346] Step 4:

[0347] The device sends text data and the resulting sentiment data to the server. This data is sent in HTTP request format and includes metadata such as user information and timestamps.

[0348] Step 5:

[0349] The server inputs the received text data and sentiment data into a generative model and generates a summary of the specified length, reflecting the sentiment. The summary will reflect the nuances of the emotions.

[0350] Step 6:

[0351] The server saves the generated sentiment-reflecting summaries to a database. This saved summary data is then used by users for later review.

[0352] Step 7:

[0353] The server analyzes text data entered by the user that is categorized as a consultation or problem, and generates appropriate advice and encouraging messages. Based on the sentiment analysis results, the generative model provides responses that take the user's emotions into consideration.

[0354] Step 8:

[0355] The server sends the generated summary and advice messages to the terminal. The terminal displays the received data in its user interface, providing visual and auditory feedback to the user.

[0356] Step 9:

[0357] Users can instruct the system to bind their saved summary data into a book. The server then formats the summary data into a printable format and sends it to a printing company. During this process, emotional information is also incorporated.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] In recording and seeking advice regarding childcare-related events, there are limited ways to receive support that reflects the user's emotions. Traditional systems struggle to automatically generate emotionally balanced summaries and advice, preventing users from receiving personalized support tailored to their specific situations. This has led to increased childcare burdens and a lack of emotional support.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes an input means for the user to input childcare-related events and consultations via an information device using voice input, a voice recognition means for converting the voice input into text information, and an emotion analysis means for analyzing the text information to identify emotions. This enables the automatic generation of summaries and advice that reflect the user's emotions.

[0363] "Input means" refers to a device or method by which a user inputs voice via an information device.

[0364] "Speech recognition means" refers to a technology or device that converts speech data into text information.

[0365] "Emotional analysis means" refers to a technology or method that analyzes textual information and identifies emotions from its content.

[0366] A "generative model" is an algorithm or computer program that generates information based on given data.

[0367] "Summarization techniques" refer to techniques or methods for shortening information and extracting key points.

[0368] A "recording medium" is a physical or electronic medium used to store data.

[0369] "Output means" refers to a device or method for presenting processed information to a user.

[0370] "Response generation means" refers to a technology or method for generating advice or messages based on input information.

[0371] "Conversion means" refers to a technique or method for changing data format to another format.

[0372] This invention is a system that records events related to childcare, performs emotional analysis, and provides appropriate support to the user. The following describes how this system can be implemented.

[0373] Users input childcare-related events and questions via voice using devices such as smartphones and tablets. These devices are equipped with microphones to convert voice data into digital data. This voice data is then converted into text data on the device using speech recognition software such as Google Speech-to-Text.

[0374] The converted text data is sent to a server via the internet. The server is equipped with an emotion engine that analyzes emotions based on the text information. IBM Watson Natural Language Understanding is a specific example of this process. Once the server receives the results of the emotion analysis, it uses a generative AI model such as OpenAI's GPT to generate summaries and response messages that reflect those emotions.

[0375] In this system, for example, if a user makes a voice input such as, "My child went to kindergarten for the first time today, and I felt lonely," the emotion engine will extract the emotion of "loneliness." Then, a generative AI model will create a summary such as, "As you watch your child grow, you may experience changes in your feelings, but this is a wonderful step."

[0376] Ultimately, the device presents these summaries and messages received from the server to the user. Furthermore, if requested by the user, the saved data can be converted into a printable format, such as PDF. This allows for a more emotionally rich record of childcare experiences, while also providing access to appropriate advice.

[0377] Through this system, users can receive more personalized support in childcare. An example of a prompt message is: "Please input your childcare experiences or questions via voice. The system will analyze your emotions and provide appropriate summaries and messages."

[0378] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0379] Step 1:

[0380] The user uses the device to input childcare-related events and questions via voice. The input voice is recorded by the device's microphone and saved as digital audio data. The input data is the user's voice, and the output is that audio data.

[0381] Step 2:

[0382] The device passes the recorded audio data to speech recognition software, which converts it into text data. At this stage, audio data is input and text data is output as written information. Specifically, the device uses the internet to send the audio data to a cloud-based speech recognition service and receives the text data from there.

[0383] Step 3:

[0384] The terminal sends the converted text data to the server. In this process, text data is input and sent to the server. Specifically, the terminal sends the text data as an HTTP POST request, and communication is securely performed using the SSL / TLS protocol.

[0385] Step 4:

[0386] The server inputs the received text data into the sentiment engine for sentiment analysis. At this stage, the text data is provided as input, and sentiment scores and labels are output. Specifically, the sentiment engine analyzes the content of the text and identifies emotions such as positive, negative, and neutral.

[0387] Step 5:

[0388] The server uses the sentiment analysis results to input information into a generative AI model, which then generates summaries and messages that reflect the emotions. The input here is text data and its sentiment score, while the output is summarized text and response messages. Specifically, the generative AI model receives text and sentiment information and forms a response based on it.

[0389] Step 6:

[0390] The server returns the generated summary and response message to the terminal. In this process, the summary text and message are input and sent as output to the terminal. Specifically, the server generates an HTTP response, the terminal receives it, and displays it to the user.

[0391] Step 7:

[0392] The user operates the terminal to convert the saved summary data into a format suitable for printing, as needed. In this step, the summary data is input, and the output is in a format such as PDF. Specifically, the software on the terminal formats the data, generates a PDF, and saves it to local storage.

[0393] In this way, the system efficiently records events and emotions related to childcare and provides support.

[0394] (Application Example 2)

[0395] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0396] As there is a growing need for real-time and appropriate support for childcare-related events and concerns, there is a need to strengthen childcare-related support at physical stores and provide users with personalized, immediate advice. However, conventional systems have made it difficult to provide emotionally resonant support and to create printed information in a physical store environment.

[0397] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0398] In this invention, the server includes an input means for recognizing information input via voice, a voice recognition means for converting it into text information, and a summarization means using sentiment analysis and a generative model. This makes it possible to provide real-time advice tailored to the user's emotions, as well as compress and compile information, even in a physical store environment.

[0399] "Information input via voice" refers to data and messages that users send to a device using spoken language.

[0400] "Speech recognition means" refers to technologies and devices that analyze speech information and convert it into text information.

[0401] A "generative model" is a type of artificial intelligence technology that generates new data or information based on given data.

[0402] A "summarization tool" refers to a process or device that extracts important information from a large amount of data and summarizes it concisely.

[0403] A "memory device" refers to a device or method for holding and storing data or information.

[0404] "Output means" refers to devices or technologies for displaying or physically outputting processed information to an external source.

[0405] An "interface means" refers to a device or method that serves as a window for information exchange between people and systems.

[0406] A "digital visual device" refers to a wearable device or screen that displays visual information using digital technology.

[0407] "Personalized responses" refer to feedback and messages that are customized based on the user's individual circumstances and emotions.

[0408] "Devices for converting to print media" refer to equipment and methods used to convert digital data into physical printed materials.

[0409] The system realizing this invention works in conjunction with smart glasses or digital visual devices to process the user's childcare-related voice input. When a user asks for advice on childcare while selecting products, the system first collects the voice using the microphone on the smart glasses. The voice is recorded on the device and transmitted to a server via the network.

[0410] The server converts speech data into text using speech recognition software (e.g., Google Speech-to-Text API). This text is then analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer) for sentiment analysis. The analyzed sentiment data is then input into a generative AI model (e.g., OpenAI's GPT) to generate personalized advice and responses based on the user's emotions.

[0411] The generated response is displayed in real time on the smart glasses' screen. This allows users to receive immediate and appropriate support regarding childcare within the store. For example, if a user is unsure about choosing diapers, the smart glasses will display a message such as, "We understand your concerns. Here are some suggestions for improvement..."

[0412] Examples of prompt messages include the following:

[0413] "Users are feeling anxious about choosing diapers. To alleviate their anxiety, please recommend a brand in 10 characters or less."

[0414] This system allows for real-time consultation and information exchange regarding childcare, enabling a more effective shopping experience.

[0415] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0416] Step 1:

[0417] The user inputs childcare-related information by voice through the microphone on their smart glasses. The input voice data is temporarily stored on the device, and then the device transmits this voice data to a server via the network.

[0418] Step 2:

[0419] The server converts the received audio data into text data using speech recognition software. In this process, the input audio data is analyzed, and a corresponding string of characters is generated. The output text data is then used for subsequent processing.

[0420] Step 3:

[0421] The server inputs text data into the emotion engine, which analyzes the user's emotions. The emotion engine detects emotional cues contained in the text data and extracts emotional indicators such as joy, sadness, and anxiety. This analysis result is used as data necessary for the next generation process.

[0422] Step 4:

[0423] The analyzed emotional and textual data are input into a generative AI model. Based on this data, the generative AI model generates personalized advice and empathetic messages that resonate with the user's emotions. These output messages play a crucial role in providing instantaneous support to the user.

[0424] Step 5:

[0425] The server sends the generated advice and messages back to the device. The device displays this information on the smart glasses' screen. Users can immediately receive this feedback visually and use it as a reference for shopping or childcare advice.

[0426] This entire process allows users to seek real-time advice on childcare in physical stores and receive empathetic feedback.

[0427] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0428] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0429] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0430] [Third Embodiment]

[0431] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0432] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0433] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0434] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0435] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0436] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0437] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0438] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0439] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0440] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0441] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0442] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0443] This invention provides technology for implementing a system that allows parents raising children to easily record daily events and receive support for their childcare concerns. The system begins with the user inputting childcare events via voice through a device such as a smartphone or tablet. First, the device records the voice in real time and converts it into text data using speech recognition. The converted text data is sent to a server, where a generative model summarizes it to a user-specified number of characters. This summarized text data is stored in a storage device and can be accessed later.

[0444] Furthermore, when a user enters a concern related to childcare, the server analyzes that concern and generates appropriate advice and encouraging messages. These messages utilize a generation AI model and are designed to provide users with appropriate and heartwarming content. The generated advice is sent back to the device and presented to the user visually or audibly.

[0445] Furthermore, if a user wishes to have their saved childcare diary bound into a physical book, the server formats this summarized data and sends it to an output device to convert it into printable media. In this way, information stored as electronic data can later be used as a physical album or childcare record.

[0446] For example, if a user inputs an event such as "Today I went down the slide by myself at the park for the first time," the device converts the audio into text, and the server summarizes it and saves it as "First slide success!" Also, if a user inputs a concern such as "Raising children is hard and tiring," the server generates a message of encouragement such as "Thank you for your hard work raising children every day. You're doing a great job," which the device displays.

[0447] In this way, users can easily record childcare information and receive support.

[0448] The following describes the processing flow.

[0449] Step 1:

[0450] The user opens the smartphone app and speaks about events or concerns related to childcare. At this time, they tap the voice input button to start recording.

[0451] Step 2:

[0452] The device records the user's voice and uses a speech recognition engine to convert the audio into text data. This process takes place in real time within the device.

[0453] Step 3:

[0454] The terminal sends the converted text data to the server via an HTTP request. This request includes metadata such as the date and time the audio was recorded and the user ID.

[0455] Step 4:

[0456] The server provides the received text data to an AI model that generates a summary of a specified length. During this process, the most important content of the text is condensed.

[0457] Step 5:

[0458] The server stores the summarized text in a database. This stored data can later be accessed by users or used for bookbinding purposes.

[0459] Step 6:

[0460] If the input is a user's problem or concern, the server analyzes the text data and uses a generative AI model to generate advice and encouraging messages. The generated messages will be tailored to the user's situation.

[0461] Step 7:

[0462] The server sends the generated summary or advice back to the device. The returned data is visually displayed to the user through the app's interface, and can also be presented audibly if voice guidance is available.

[0463] (Example 1)

[0464] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0465] There is a need for a way to easily record daily events and parental concerns related to childcare, making them easily accessible for later reference, and for providing advice and support based on this information. However, parents raising children often have limited time and knowledge, making it difficult to do this using conventional methods. Furthermore, there is the problem of the time and effort required to physically store and bind the recorded information.

[0466] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0467] In this invention, the server includes an input means for a user to input childcare-related activities via voice input through an information terminal, a voice recognition means for converting the voice input into text format, and a summarization means that uses generation technology to compress the text data to a specified number of characters. This makes it possible to quickly and efficiently record childcare events and store them in a format that can be easily referenced later. Furthermore, since advice and support generated using generation technology can be provided in real time, the burden on parents can be reduced and emotional support can be provided.

[0468] An "information terminal" is a device used by a user for voice input, and includes smartphones, tablets, and personal computers.

[0469] "Voice input" refers to the act of users communicating events or questions related to childcare to the system using voice.

[0470] "Speech recognition" is a technology that analyzes input speech data and converts it into text format.

[0471] "Generative technology" refers to techniques that primarily use generative AI models to summarize data and create appropriate messages.

[0472] A "summarization method" is a technique that compresses text data to a specified number of characters and processes it to simplify the information.

[0473] "Memory device" refers to the function of storing summarized text data in a database so that it can be referenced later.

[0474] "Output means" refers to a device that physically or digitally displays or outputs data in order to provide information to users.

[0475] "Bookbinding methods" refer to technologies for formatting digital data into a printable format and providing it as a physical book or album.

[0476] This invention presents a specific embodiment of a system that allows parents raising children to easily record everyday events and concerns related to childcare and receive appropriate advice. The details are described below.

[0477] The term "device" refers to information terminals such as smartphones and tablets, which users use to access the system. Users input information about childcare-related events or topics they wish to discuss using voice input. The device utilizes speech recognition software (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[0478] The server receives the converted text data and utilizes generation technology. This generation technology includes a generative AI model (e.g., GPT-3 or similar models), which enables a summarization method that compresses the input text data to a specified number of characters. It also has a function to generate advice and empathetic messages according to the user's concerns.

[0479] The generated summaries and messages are stored in a database by the server. This storage mechanism allows users to refer to past records at any time. Furthermore, the stored text data can be sent to the terminal as needed and presented to the user visually or audibly.

[0480] Furthermore, if a user wishes to save their childcare diary in a physical form, the server will format the digital data and convert it into a printable format such as PDF. This allows it to be bound into a physical album or diary via an output device.

[0481] As a concrete example of its use, if a user records an event such as, "Today I went down the slide by myself at the park for the first time," the device will transcribe it into text, and the server will summarize it as, "First slide success!" Also, if a user inputs a complaint such as, "Raising children is tough and exhausting," the server will generate an encouraging message such as, "You're doing a great job with childcare every day. You're doing a wonderful job," and the device will display this message.

[0482] A concrete example of a prompt sentence is, "My child experienced a new game at the park. I would like to briefly record it." This is used when the generative AI model performs summarization.

[0483] In this way, the system provides convenient and effective record-keeping and support for parents raising children.

[0484] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0485] Step 1:

[0486] Users input events and concerns related to childcare using voice commands via an information terminal. The terminal captures the user's voice using a microphone. This input data is stored as raw audio data.

[0487] Step 2:

[0488] The device uses speech recognition software to convert captured audio data into text data. Specifically, the audio waveform data is converted into text format, and noise reduction and normalization are performed accordingly. As a result, the data is output in the form of text sentences.

[0489] Step 3:

[0490] The terminal sends the converted text data to the server. For security reasons, the data is encrypted and transmitted over the network. The transmitted data is received by the server and ready for analysis.

[0491] Step 4:

[0492] The server uses generation technology to compress the received text data to a specified number of characters. Summarization is performed by inputting a prompt into the generation AI model. This prompt includes specific instructions such as, "Please summarize this text briefly." The summarization result is obtained as shortened text data.

[0493] Step 5:

[0494] The server saves the summarized text data to a database. This saving process involves organizing the data and writing it to the database. The saved data can then be searched later using specific dates, keywords, or other criteria.

[0495] Step 6:

[0496] When a user submits a parenting consultation request, the server analyzes the text data and, if necessary, uses a generative AI model to generate advice and encouraging messages. This process creates a customized response tailored to the user's concerns. The generated messages are output in text or audio format.

[0497] Step 7:

[0498] The server sends the generated message to the terminal. The received message is displayed on the terminal's screen or read aloud using text-to-speech functionality. This allows the user to confirm the message visually or audibly.

[0499] Step 8:

[0500] If a user requests a physical binding of their saved baby journal, the server formats this data and converts it into a printable format. In this process, the text data is converted into a document format such as PDF, sent to an output device, and obtained as a physical album.

[0501] (Application Example 1)

[0502] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0503] In modern parenting, parents lack the time to record daily events and are also required to respond quickly to parenting concerns. However, it is difficult to meet these needs simultaneously using traditional methods. Furthermore, the lack of adequate systems for efficiently recording parenting events and obtaining appropriate support is increasing the mental burden on parents. In addition, there is a need for improved services that comfortably support parenting, such as recommendations for relevant products based on recorded events.

[0504] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0505] In this invention, the server includes information processing means for users to input life events by voice, language processing means for converting the voice input into information data, and information compression means that uses a generation method for summarizing the information data into a predetermined number of symbols. This allows users to easily record events related to childcare and receive efficient and comprehensive support by being recommended related products based on those events.

[0506] "Information processing means" refers to a device that has the function of allowing users to input life events and consultations via voice.

[0507] "Language processing means" refers to technology that converts speech input into information data, specifically by converting it into text data through speech recognition.

[0508] "Information compression means" refers to a function that uses a generation method to summarize acquired information data into a predetermined number of symbols.

[0509] An "information storage device" is a memory device that has the function of accumulating summarized information data.

[0510] A "media control means" is a system that controls summary information stored in an information storage means and provides it to users as needed.

[0511] "Media output means" refers to a device that has the function of converting summary information extracted from information storage means into a recordable format and outputting it.

[0512] A "product recommendation tool" is a function that recommends relevant products based on the user's life events.

[0513] A "response generation means" is a device that has a generation method for generating advice or emotional messages based on information data and presenting them to the user.

[0514] This invention provides a system that allows users to easily record life events and receive childcare-related support. The system mainly consists of smart devices and a server.

[0515] First, users input childcare-related information using voice commands via the input device of their smart device. The device receives the voice input and converts the voice signal into text data using a speech recognition API. Google Speech-to-Text is one of the speech recognition APIs used.

[0516] Next, the server receives the converted text data and uses a generative AI model to summarize the information data into a predetermined number of symbols. This summarized data is stored in a cloud database via an information storage system. OpenAI's GPT model is used as the generative AI model.

[0517] Furthermore, when a user requests advice regarding childcare, the server uses a generative AI model to generate appropriate advice messages and empathetic messages based on the converted text data, and sends them back to the device. The device then presents these messages to the user visually or audibly.

[0518] In addition, the server can select and recommend relevant childcare products from online stores based on the user's life events. This recommendation function can meet a variety of needs in childcare life.

[0519] For example, if a user inputs the voice message, "My child laughed for the first time today," the AI ​​model will summarize this as "First laugh successful!" and save it. Related products that might be recommended include toys that help elicit smiles. An example of a prompt based on this example would be, "Please record your baby's first smile and recommend related products."

[0520] This invention will improve the efficiency of childcare record-keeping and promote support, allowing users to easily obtain useful information.

[0521] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0522] Step 1:

[0523] Users input childcare-related information by voice via a smart device. The input voice data is captured by the microphone in the smart device. This voice data serves as the starting point for determining what processing will be performed next.

[0524] Step 2:

[0525] The device receives the captured audio data and converts it into text data using a speech recognition API. Google Speech-to-Text is used as the speech recognition API in this case. The audio data is processed as input, and text data is generated as output. This text data semantically preserves the information provided by the user via voice.

[0526] Step 3:

[0527] The server receives text data sent from the terminal. Based on the received text data, it uses a generative AI model (e.g., the OpenAI GPT model) to summarize the information data into a predetermined number of symbols. The goal is to compress the input text data to summarize it into language that is easy for the user to understand.

[0528] Step 4:

[0529] The server saves the summarized text data to an information storage system. A cloud database is used here to securely store the summarized information. The saved information can be accessed later and used like an electronic diary.

[0530] Step 5:

[0531] When a user asks for advice related to childcare, they input the content into the system again via voice. The terminal captures the voice data as described above and converts it into text data using a speech recognition API.

[0532] Step 6:

[0533] The server receives text data related to the consultation and uses a generative AI model to generate advice messages and empathetic messages tailored to the content of the consultation. Here, the generative AI model processes the prompt text and outputs information that is useful to the user.

[0534] Step 7:

[0535] The server sends the generated message back to the terminal, which then presents it to the user visually or audibly. Here, the feedback the user receives greatly influences the user experience.

[0536] Step 8:

[0537] The server complementarily selects and recommends relevant childcare products from a virtual store based on the user's childcare information. This recommendation information aims to provide highly relevant products based on the information entered by the user.

[0538] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0539] This invention relates to a system that records events and concerns related to childcare and provides support combined with emotion analysis. Users input childcare-related events via voice using a device such as a smartphone or tablet. This voice data is recorded on the device and converted into text data by a voice recognition device.

[0540] The terminal then sends the converted text data to the server. The server is equipped with a generative model and an emotion engine, which analyzes the user's emotions based on the content of the text data. The analyzed emotion data is input into the generative model, which generates a summarized text that reflects the emotional information. This results in a summary of the child's growth record that includes emotional nuances.

[0541] Furthermore, for inquiries regarding childcare, advice and empathetic messages are generated that are tailored to the user's emotions. Based on the analysis results of the emotion engine, the generation model responds in a way that is sensitive to the user's feelings. This provides users with more appropriate and personalized support. For example, if a user is feeling anxious, a message such as "I understand your anxiety. Thank you for your daily efforts" is provided.

[0542] The saved summary data retains emotional information even when converted into a format suitable for printing, as instructed by the user. This results in richer records as a parenting diary, allowing users to relive the emotions of the time when looking back. This system enables users to easily record their emotional experiences related to parenting and receive support.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The user launches the smartphone app and speaks about events or concerns related to childcare. They then press the voice input button to start recording.

[0546] Step 2:

[0547] The device records the user's voice and uses its built-in speech recognition device to convert the voice data into text data in real time.

[0548] Step 3:

[0549] The terminal sends the converted text data to the sentiment engine for sentiment analysis. This analysis yields user sentiment data associated with the input text.

[0550] Step 4:

[0551] The device sends text data and the resulting sentiment data to the server. This data is sent in HTTP request format and includes metadata such as user information and timestamps.

[0552] Step 5:

[0553] The server inputs the received text data and sentiment data into a generative model and generates a summary of the specified length, reflecting the sentiment. The summary will reflect the nuances of the emotions.

[0554] Step 6:

[0555] The server saves the generated sentiment-reflecting summaries to a database. This saved summary data is then used by users for later review.

[0556] Step 7:

[0557] The server analyzes text data entered by the user that is categorized as a consultation or problem, and generates appropriate advice and encouraging messages. Based on the sentiment analysis results, the generative model provides responses that take the user's emotions into consideration.

[0558] Step 8:

[0559] The server sends the generated summary and advice messages to the terminal. The terminal displays the received data in its user interface, providing visual and auditory feedback to the user.

[0560] Step 9:

[0561] Users can instruct the system to bind their saved summary data into a book. The server then formats the summary data into a printable format and sends it to a printing company. During this process, emotional information is also incorporated.

[0562] (Example 2)

[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0564] In recording and seeking advice regarding childcare-related events, there are limited ways to receive support that reflects the user's emotions. Traditional systems struggle to automatically generate emotionally balanced summaries and advice, preventing users from receiving personalized support tailored to their specific situations. This has led to increased childcare burdens and a lack of emotional support.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0566] In this invention, the server includes an input means for the user to input childcare-related events and consultations via an information device using voice input, a voice recognition means for converting the voice input into text information, and an emotion analysis means for analyzing the text information to identify emotions. This enables the automatic generation of summaries and advice that reflect the user's emotions.

[0567] "Input means" refers to a device or method by which a user inputs voice via an information device.

[0568] "Speech recognition means" refers to a technology or device that converts speech data into text information.

[0569] "Emotional analysis means" refers to a technology or method that analyzes textual information and identifies emotions from its content.

[0570] A "generative model" is an algorithm or computer program that generates information based on given data.

[0571] "Summarization techniques" refer to techniques or methods for shortening information and extracting key points.

[0572] A "recording medium" is a physical or electronic medium used to store data.

[0573] "Output means" refers to a device or method for presenting processed information to a user.

[0574] "Response generation means" refers to a technology or method for generating advice or messages based on input information.

[0575] "Conversion means" refers to a technique or method for changing data format to another format.

[0576] This invention is a system that records events related to childcare, performs emotional analysis, and provides appropriate support to the user. The following describes how this system can be implemented.

[0577] Users input childcare-related events and questions via voice using devices such as smartphones and tablets. These devices are equipped with microphones to convert voice data into digital data. This voice data is then converted into text data on the device using speech recognition software such as Google Speech-to-Text.

[0578] The converted text data is sent to a server via the internet. The server is equipped with an emotion engine that analyzes emotions based on the text information. IBM Watson Natural Language Understanding is a specific example of this process. Once the server receives the results of the emotion analysis, it uses a generative AI model such as OpenAI's GPT to generate summaries and response messages that reflect those emotions.

[0579] In this system, for example, if a user makes a voice input such as, "My child went to kindergarten for the first time today, and I felt lonely," the emotion engine will extract the emotion of "loneliness." Then, a generative AI model will create a summary such as, "As you watch your child grow, you may experience changes in your feelings, but this is a wonderful step."

[0580] Ultimately, the device presents these summaries and messages received from the server to the user. Furthermore, if requested by the user, the saved data can be converted into a printable format, such as PDF. This allows for a more emotionally rich record of childcare experiences, while also providing access to appropriate advice.

[0581] Through this system, users can receive more personalized support in childcare. An example of a prompt message is: "Please input your childcare experiences or questions via voice. The system will analyze your emotions and provide appropriate summaries and messages."

[0582] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0583] Step 1:

[0584] Users input childcare-related events and questions via voice using a device. The input voice is recorded by the device's microphone and saved as digital audio data. The input data is the user's voice, and the output is that audio data.

[0585] Step 2:

[0586] The device passes the recorded audio data to speech recognition software, which converts it into text data. At this stage, audio data is input and text data is output as written information. Specifically, the device uses the internet to send the audio data to a cloud-based speech recognition service and receives the text data from there.

[0587] Step 3:

[0588] The terminal sends the converted text data to the server. In this process, text data is input and sent to the server. Specifically, the terminal sends the text data as an HTTP POST request, and communication is securely performed using the SSL / TLS protocol.

[0589] Step 4:

[0590] The server inputs the received text data into the sentiment engine for sentiment analysis. At this stage, the text data is provided as input, and sentiment scores and labels are output. Specifically, the sentiment engine analyzes the content of the text and identifies emotions such as positive, negative, and neutral.

[0591] Step 5:

[0592] The server uses the sentiment analysis results to input information into a generative AI model, which then generates summaries and messages that reflect the emotions. The input here is text data and its sentiment score, while the output is summarized text and response messages. Specifically, the generative AI model receives text and sentiment information and forms a response based on it.

[0593] Step 6:

[0594] The server returns the generated summary and response message to the terminal. In this process, the summary text and message are input and sent as output to the terminal. Specifically, the server generates an HTTP response, the terminal receives it, and displays it to the user.

[0595] Step 7:

[0596] The user operates the terminal to convert the saved summary data into a format suitable for printing, as needed. In this step, the summary data is input, and the output is in a format such as PDF. Specifically, the software on the terminal formats the data, generates a PDF, and saves it to local storage.

[0597] In this way, the system efficiently records events and emotions related to childcare and provides support.

[0598] (Application Example 2)

[0599] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0600] As there is a growing need for real-time and appropriate support for childcare-related events and concerns, there is a need to strengthen childcare-related support at physical stores and provide users with personalized, immediate advice. However, conventional systems have made it difficult to provide emotionally resonant support and to create printed information in a physical store environment.

[0601] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0602] In this invention, the server includes an input means for recognizing information input via voice, a voice recognition means for converting it into text information, and a summarization means using sentiment analysis and a generative model. This makes it possible to provide real-time advice tailored to the user's emotions, as well as compress and compile information, even in a physical store environment.

[0603] "Information input via voice" refers to data and messages that users send to a device using spoken language.

[0604] "Speech recognition means" refers to technologies and devices that analyze speech information and convert it into text information.

[0605] A "generative model" is a type of artificial intelligence technology that generates new data or information based on given data.

[0606] A "summarization tool" refers to a process or device that extracts important information from a large amount of data and summarizes it concisely.

[0607] A "memory device" refers to a device or method for holding and storing data or information.

[0608] "Output means" refers to devices or technologies for displaying or physically outputting processed information to an external source.

[0609] An "interface means" refers to a device or method that serves as a window for information exchange between people and systems.

[0610] A "digital visual device" refers to a wearable device or screen that displays visual information using digital technology.

[0611] "Personalized responses" refer to feedback and messages that are customized based on the user's individual circumstances and emotions.

[0612] "Devices for converting to print media" refer to equipment and methods used to convert digital data into physical printed materials.

[0613] The system realizing this invention works in conjunction with smart glasses or digital visual devices to process the user's childcare-related voice input. When a user asks for advice on childcare while selecting products, the system first collects the voice using the microphone on the smart glasses. The voice is recorded on the device and transmitted to a server via the network.

[0614] The server converts speech data into text using speech recognition software (e.g., Google Speech-to-Text API). This text is then analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer) for sentiment analysis. The analyzed sentiment data is then input into a generative AI model (e.g., OpenAI's GPT) to generate personalized advice and responses based on the user's emotions.

[0615] The generated response is displayed in real time on the smart glasses' screen. This allows users to receive immediate and appropriate support regarding childcare within the store. For example, if a user is unsure about choosing diapers, the smart glasses will display a message such as, "We understand your concerns. Here are some suggestions for improvement..."

[0616] Examples of prompt messages include the following:

[0617] "Users are feeling anxious about choosing diapers. To alleviate their anxiety, please recommend a brand in 10 characters or less."

[0618] This system allows for real-time consultation and information exchange regarding childcare, enabling a more effective shopping experience.

[0619] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0620] Step 1:

[0621] The user inputs childcare-related information via voice through the microphone on their smart glasses. The input voice data is temporarily stored on the device, and then the device transmits this voice data to a server over the network.

[0622] Step 2:

[0623] The server converts the received audio data into text data using speech recognition software. In this process, the input audio data is analyzed, and a corresponding string of characters is generated. The output text data is then used for subsequent processing.

[0624] Step 3:

[0625] The server inputs text data into the emotion engine, which analyzes the user's emotions. The emotion engine detects emotional cues contained in the text data and extracts emotional indicators such as joy, sadness, and anxiety. This analysis result is used as data necessary for the next generation process.

[0626] Step 4:

[0627] The analyzed emotional and textual data is input into a generative AI model. Based on this data, the generative AI model generates personalized advice and empathetic messages that resonate with the user's emotions. These output messages play a crucial role in providing instantaneous support to the user.

[0628] Step 5:

[0629] The server sends the generated advice and messages back to the device. The device displays this information on the smart glasses' screen. Users can immediately receive this feedback visually and use it as a reference for shopping or childcare consultations.

[0630] This entire process allows users to seek real-time advice on childcare in physical stores and receive empathetic feedback.

[0631] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0632] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0633] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0634] [Fourth Embodiment]

[0635] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0636] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0637] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0638] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0639] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0640] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0641] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0642] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0643] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0644] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0645] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0646] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0647] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0648] This invention provides technology for implementing a system that allows parents raising children to easily record daily events and receive support for their childcare concerns. The system begins with the user inputting childcare events via voice through a device such as a smartphone or tablet. First, the device records the voice in real time and converts it into text data using speech recognition. The converted text data is sent to a server, where a generative model summarizes it to a user-specified number of characters. This summarized text data is stored in a storage device and can be accessed later.

[0649] Furthermore, when a user enters a concern related to childcare, the server analyzes that concern and generates appropriate advice and encouraging messages. These messages utilize a generation AI model and are designed to provide users with appropriate and heartwarming content. The generated advice is sent back to the device and presented to the user visually or audibly.

[0650] Furthermore, if a user wishes to have their saved childcare diary bound into a physical book, the server formats this summarized data and sends it to an output device to convert it into printable media. In this way, information stored as electronic data can later be used as a physical album or childcare record.

[0651] For example, if a user inputs an event such as "Today I went down the slide by myself at the park for the first time," the device converts the audio into text, and the server summarizes it and saves it as "First slide success!" Also, if a user inputs a concern such as "Raising children is hard and tiring," the server generates a message of encouragement such as "Thank you for your hard work raising children every day. You're doing a great job," which the device displays.

[0652] In this way, users can easily record childcare information and receive support.

[0653] The following describes the processing flow.

[0654] Step 1:

[0655] The user opens the smartphone app and speaks about events or concerns related to childcare. At this time, they tap the voice input button to start recording.

[0656] Step 2:

[0657] The device records the user's voice and uses a speech recognition engine to convert the audio into text data. This process takes place in real time within the device.

[0658] Step 3:

[0659] The terminal sends the converted text data to the server via an HTTP request. This request includes metadata such as the date and time the audio was recorded and the user ID.

[0660] Step 4:

[0661] The server provides the received text data to an AI model that generates a summary of a specified length. During this process, the most important content of the text is condensed.

[0662] Step 5:

[0663] The server stores the summarized text in a database. This stored data can later be accessed by users or used for bookbinding purposes.

[0664] Step 6:

[0665] If the input is a user's problem or concern, the server analyzes the text data and uses a generative AI model to generate advice and encouraging messages. The generated messages will be tailored to the user's situation.

[0666] Step 7:

[0667] The server sends the generated summary or advice back to the device. The returned data is visually displayed to the user through the app's interface, and can also be presented audibly if voice guidance is available.

[0668] (Example 1)

[0669] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0670] There is a need for a way to easily record daily events and parental concerns related to childcare, making them easily accessible for later reference, and for providing advice and support based on this information. However, parents raising children often have limited time and knowledge, making it difficult to do this using conventional methods. Furthermore, there is the problem of the time and effort required to physically store and bind the recorded information.

[0671] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0672] In this invention, the server includes an input means for a user to input childcare-related activities via voice input through an information terminal, a voice recognition means for converting the voice input into text format, and a summarization means that uses generation technology to compress the text data to a specified number of characters. This makes it possible to quickly and efficiently record childcare events and store them in a format that can be easily referenced later. Furthermore, since advice and support generated using generation technology can be provided in real time, the burden on parents can be reduced and emotional support can be provided.

[0673] An "information terminal" is a device used by a user for voice input, and includes smartphones, tablets, and personal computers.

[0674] "Voice input" refers to the act of users communicating events or questions related to childcare to the system using voice.

[0675] "Speech recognition" is a technology that analyzes input speech data and converts it into text format.

[0676] "Generative technology" refers to techniques that primarily use generative AI models to summarize data and create appropriate messages.

[0677] A "summarization method" is a technique that compresses text data to a specified number of characters and processes it to simplify the information.

[0678] "Memory device" refers to the function of storing summarized text data in a database so that it can be referenced later.

[0679] "Output means" refers to a device that physically or digitally displays or outputs data in order to provide information to users.

[0680] "Bookbinding methods" refer to technologies for formatting digital data into a printable format and providing it as a physical book or album.

[0681] This invention presents a specific embodiment of a system that allows parents raising children to easily record everyday events and concerns related to childcare and receive appropriate advice. The details are described below.

[0682] The term "device" refers to information terminals such as smartphones and tablets, which users use to access the system. Users input information about childcare-related events or topics they wish to discuss using voice input. The device utilizes speech recognition software (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[0683] The server receives the converted text data and utilizes generation technology. This generation technology includes a generative AI model (e.g., GPT-3 or similar models), which enables a summarization method that compresses the input text data to a specified number of characters. It also has a function to generate advice and empathetic messages according to the user's concerns.

[0684] The generated summaries and messages are stored in a database by the server. This storage mechanism allows users to refer to past records at any time. Furthermore, the stored text data can be sent to the terminal as needed and presented to the user visually or audibly.

[0685] Furthermore, if a user wishes to save their childcare diary in a physical form, the server will format the digital data and convert it into a printable format such as PDF. This allows it to be bound into a physical album or diary via an output device.

[0686] As a concrete example of its use, if a user records an event such as, "Today I went down the slide by myself at the park for the first time," the device will transcribe it into text, and the server will summarize it as, "First slide success!" Also, if a user inputs a complaint such as, "Raising children is tough and exhausting," the server will generate an encouraging message such as, "You're doing a great job with childcare every day. You're doing a wonderful job," and the device will display this message.

[0687] A concrete example of a prompt sentence is, "My child experienced a new game at the park. I would like to briefly record it." This is used when the generative AI model performs summarization.

[0688] In this way, the system provides convenient and effective record-keeping and support for parents raising children.

[0689] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0690] Step 1:

[0691] Users input events and concerns related to childcare using voice commands via an information terminal. The terminal captures the user's voice using a microphone. This input data is stored as raw audio data.

[0692] Step 2:

[0693] The device uses speech recognition software to convert captured audio data into text data. Specifically, the audio waveform data is converted into text format, and noise reduction and normalization are performed accordingly. As a result, the data is output in the form of text sentences.

[0694] Step 3:

[0695] The terminal sends the converted text data to the server. For security reasons, the data is encrypted and transmitted over the network. The transmitted data is received by the server and ready for analysis.

[0696] Step 4:

[0697] The server uses generation technology to compress the received text data to a specified number of characters. Summarization is performed by inputting a prompt into the generation AI model. This prompt includes specific instructions such as, "Please summarize this text briefly." The summarization result is obtained as shortened text data.

[0698] Step 5:

[0699] The server saves the summarized text data to a database. This saving process involves organizing the data and writing it to the database. The saved data can then be searched later using specific dates, keywords, or other criteria.

[0700] Step 6:

[0701] When a user submits a parenting consultation request, the server analyzes the text data and, if necessary, uses a generative AI model to generate advice and encouraging messages. This process creates a customized response tailored to the user's concerns. The generated messages are output in text or audio format.

[0702] Step 7:

[0703] The server sends the generated message to the terminal. The received message is displayed on the terminal's screen or read aloud using text-to-speech functionality. This allows the user to confirm the message visually or audibly.

[0704] Step 8:

[0705] If a user requests a physical binding of their saved baby journal, the server formats this data and converts it into a printable format. In this process, the text data is converted into a document format such as PDF, sent to an output device, and obtained as a physical album.

[0706] (Application Example 1)

[0707] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0708] In modern parenting, parents lack the time to record daily events and are also required to respond quickly to parenting concerns. However, it is difficult to meet these needs simultaneously using traditional methods. Furthermore, the lack of adequate systems for efficiently recording parenting events and obtaining appropriate support is increasing the mental burden on parents. In addition, there is a need for improved services that comfortably support parenting, such as recommendations for relevant products based on recorded events.

[0709] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0710] In this invention, the server includes information processing means for users to input life events by voice, language processing means for converting the voice input into information data, and information compression means that uses a generation method for summarizing the information data into a predetermined number of symbols. This allows users to easily record events related to childcare and receive efficient and comprehensive support by being recommended related products based on those events.

[0711] "Information processing means" refers to a device that has the function of allowing users to input life events and consultations via voice.

[0712] "Language processing means" refers to technology that converts speech input into information data, specifically by converting it into text data through speech recognition.

[0713] "Information compression means" refers to a function that uses a generation method to summarize acquired information data into a predetermined number of symbols.

[0714] An "information storage device" is a memory device that has the function of accumulating summarized information data.

[0715] A "media control means" is a system that controls summary information stored in an information storage means and provides it to users as needed.

[0716] "Media output means" refers to a device that has the function of converting summary information extracted from information storage means into a recordable format and outputting it.

[0717] A "product recommendation tool" is a function that recommends relevant products based on the user's life events.

[0718] A "response generation means" is a device that has a generation method for generating advice or emotional messages based on information data and presenting them to the user.

[0719] This invention provides a system that allows users to easily record life events and receive childcare-related support. The system mainly consists of smart devices and a server.

[0720] First, users input childcare-related information using voice commands via the input device of their smart device. The device receives the voice input and converts the voice signal into text data using a speech recognition API. Google Speech-to-Text is one of the speech recognition APIs used.

[0721] Next, the server receives the converted text data and uses a generative AI model to summarize the information data into a predetermined number of symbols. This summarized data is stored in a cloud database via an information storage system. OpenAI's GPT model is used as the generative AI model.

[0722] Furthermore, when a user requests advice regarding childcare, the server uses a generative AI model to generate appropriate advice messages and empathetic messages based on the converted text data, and sends them back to the device. The device then presents these messages to the user visually or audibly.

[0723] In addition, the server can select and recommend relevant childcare products from online stores based on the user's life events. This recommendation function can meet a variety of needs in childcare life.

[0724] For example, if a user inputs the voice message, "My child laughed for the first time today," the AI ​​model will summarize this as "First laugh successful!" and save it. Related products that might be recommended include toys that help elicit smiles. An example of a prompt based on this example would be, "Please record your baby's first smile and recommend related products."

[0725] This invention will improve the efficiency of childcare record-keeping and promote support, allowing users to easily obtain useful information.

[0726] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0727] Step 1:

[0728] Users input childcare-related information by voice via a smart device. The input voice data is captured by the microphone in the smart device. This voice data serves as the starting point for determining what processing will be performed next.

[0729] Step 2:

[0730] The device receives the captured audio data and converts it into text data using a speech recognition API. Google Speech-to-Text is used as the speech recognition API in this case. The audio data is processed as input, and text data is generated as output. This text data semantically preserves the information provided by the user via voice.

[0731] Step 3:

[0732] The server receives text data sent from the terminal. Based on the received text data, it uses a generative AI model (e.g., the OpenAI GPT model) to summarize the information data into a predetermined number of symbols. The goal is to compress the input text data to summarize it into language that is easy for the user to understand.

[0733] Step 4:

[0734] The server saves the summarized text data to an information storage system. A cloud database is used here to securely store the summarized information. The saved information can be accessed later and used like an electronic diary.

[0735] Step 5:

[0736] When a user asks for advice related to childcare, they input the content into the system again via voice. The terminal captures the voice data as described above and converts it into text data using a speech recognition API.

[0737] Step 6:

[0738] The server receives text data related to the consultation and uses a generative AI model to generate advice messages and empathetic messages tailored to the content of the consultation. Here, the generative AI model processes the prompt text and outputs information that is useful to the user.

[0739] Step 7:

[0740] The server sends the generated message back to the terminal, which then presents it to the user visually or audibly. Here, the feedback the user receives greatly influences the user experience.

[0741] Step 8:

[0742] The server complementarily selects and recommends relevant childcare products from a virtual store based on the user's childcare information. This recommendation information aims to provide highly relevant products based on the information entered by the user.

[0743] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0744] This invention relates to a system that records events and concerns related to childcare and provides support combined with emotion analysis. Users input childcare-related events via voice using a device such as a smartphone or tablet. This voice data is recorded on the device and converted into text data by a voice recognition device.

[0745] The terminal then sends the converted text data to the server. The server is equipped with a generative model and an emotion engine, which analyzes the user's emotions based on the content of the text data. The analyzed emotion data is input into the generative model, which generates a summarized text that reflects the emotional information. This results in a summary of the child's growth record that includes emotional nuances.

[0746] Furthermore, for inquiries regarding childcare, advice and empathetic messages are generated that are tailored to the user's emotions. Based on the analysis results of the emotion engine, the generation model responds in a way that is sensitive to the user's feelings. This provides users with more appropriate and personalized support. For example, if a user is feeling anxious, a message such as "I understand your anxiety. Thank you for your daily efforts" is provided.

[0747] The saved summary data retains emotional information even when converted into a format suitable for printing, as instructed by the user. This results in richer records as a parenting diary, allowing users to relive the emotions of the time when looking back. This system enables users to easily record their emotional experiences related to parenting and receive support.

[0748] The following describes the processing flow.

[0749] Step 1:

[0750] The user launches the smartphone app and speaks about events or concerns related to childcare. They then press the voice input button to start recording.

[0751] Step 2:

[0752] The device records the user's voice and uses its built-in speech recognition device to convert the voice data into text data in real time.

[0753] Step 3:

[0754] The terminal sends the converted text data to the sentiment engine for sentiment analysis. This analysis yields user sentiment data associated with the input text.

[0755] Step 4:

[0756] The device sends text data and the resulting sentiment data to the server. This data is sent in HTTP request format and includes metadata such as user information and timestamps.

[0757] Step 5:

[0758] The server inputs the received text data and sentiment data into a generative model and generates a summary of the specified length, reflecting the sentiment. The summary will reflect the nuances of the emotions.

[0759] Step 6:

[0760] The server saves the generated sentiment-reflecting summaries to a database. This saved summary data is then used by users for later review.

[0761] Step 7:

[0762] The server analyzes text data entered by the user that is categorized as a consultation or problem, and generates appropriate advice and encouraging messages. Based on the sentiment analysis results, the generative model provides responses that take the user's emotions into consideration.

[0763] Step 8:

[0764] The server sends the generated summary and advice messages to the terminal. The terminal displays the received data in its user interface, providing visual and auditory feedback to the user.

[0765] Step 9:

[0766] Users can instruct the system to bind their saved summary data into a book. The server then formats the summary data into a printable format and sends it to a printing company. During this process, emotional information is also incorporated.

[0767] (Example 2)

[0768] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0769] In recording and seeking advice regarding childcare-related events, there are limited ways to receive support that reflects the user's emotions. Traditional systems struggle to automatically generate emotionally balanced summaries and advice, preventing users from receiving personalized support tailored to their specific situations. This has led to increased childcare burdens and a lack of emotional support.

[0770] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0771] In this invention, the server includes an input means for the user to input childcare-related events and consultations via an information device using voice input, a voice recognition means for converting the voice input into text information, and an emotion analysis means for analyzing the text information to identify emotions. This enables the automatic generation of summaries and advice that reflect the user's emotions.

[0772] "Input means" refers to a device or method by which a user inputs voice via an information device.

[0773] "Speech recognition means" refers to a technology or device that converts speech data into text information.

[0774] "Emotional analysis means" refers to a technology or method that analyzes textual information and identifies emotions from its content.

[0775] A "generative model" is an algorithm or computer program that generates information based on given data.

[0776] "Summarization techniques" refer to techniques or methods for shortening information and extracting key points.

[0777] A "recording medium" is a physical or electronic medium used to store data.

[0778] "Output means" refers to a device or method for presenting processed information to a user.

[0779] "Response generation means" refers to a technology or method for generating advice or messages based on input information.

[0780] "Conversion means" refers to a technique or method for changing data format to another format.

[0781] This invention is a system that records events related to childcare, performs emotional analysis, and provides appropriate support to the user. The following describes how this system can be implemented.

[0782] Users input childcare-related events and questions via voice using devices such as smartphones and tablets. These devices are equipped with microphones to convert voice data into digital data. This voice data is then converted into text data on the device using speech recognition software such as Google Speech-to-Text.

[0783] The converted text data is sent to a server via the internet. The server is equipped with an emotion engine that analyzes emotions based on the text information. IBM Watson Natural Language Understanding is a specific example of this process. Once the server receives the results of the emotion analysis, it uses a generative AI model such as OpenAI's GPT to generate summaries and response messages that reflect those emotions.

[0784] In this system, for example, if a user makes a voice input such as, "My child went to kindergarten for the first time today, and I felt lonely," the emotion engine will extract the emotion of "loneliness." Then, a generative AI model will create a summary such as, "As you watch your child grow, you may experience changes in your feelings, but this is a wonderful step."

[0785] Ultimately, the device presents these summaries and messages received from the server to the user. Furthermore, if requested by the user, the saved data can be converted into a printable format, such as PDF. This allows for a more emotionally rich record of childcare experiences, while also providing access to appropriate advice.

[0786] Through this system, users can receive more personalized support in childcare. An example of a prompt message is: "Please input your childcare experiences or questions via voice. The system will analyze your emotions and provide appropriate summaries and messages."

[0787] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0788] Step 1:

[0789] Users input childcare-related events and questions via voice using a device. The input voice is recorded by the device's microphone and saved as digital audio data. The input data is the user's voice, and the output is that audio data.

[0790] Step 2:

[0791] The device passes the recorded audio data to speech recognition software, which converts it into text data. At this stage, audio data is input and text data is output as written information. Specifically, the device uses the internet to send the audio data to a cloud-based speech recognition service and receives the text data from there.

[0792] Step 3:

[0793] The terminal sends the converted text data to the server. In this process, text data is input and sent to the server. Specifically, the terminal sends the text data as an HTTP POST request, and communication is securely performed using the SSL / TLS protocol.

[0794] Step 4:

[0795] The server inputs the received text data into the sentiment engine for sentiment analysis. At this stage, the text data is provided as input, and sentiment scores and labels are output. Specifically, the sentiment engine analyzes the content of the text and identifies emotions such as positive, negative, and neutral.

[0796] Step 5:

[0797] The server uses the sentiment analysis results to input information into a generative AI model, which then generates summaries and messages that reflect the emotions. The input here is text data and its sentiment score, while the output is summarized text and response messages. Specifically, the generative AI model receives text and sentiment information and forms a response based on it.

[0798] Step 6:

[0799] The server returns the generated summary and response message to the terminal. In this process, the summary text and message are input and sent as output to the terminal. Specifically, the server generates an HTTP response, the terminal receives it, and displays it to the user.

[0800] Step 7:

[0801] The user operates the terminal to convert the saved summary data into a format suitable for printing, as needed. In this step, the summary data is input, and the output is in a format such as PDF. Specifically, the software on the terminal formats the data, generates a PDF, and saves it to local storage.

[0802] In this way, the system efficiently records events and emotions related to childcare and provides support.

[0803] (Application Example 2)

[0804] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0805] As there is a growing need for real-time and appropriate support for childcare-related events and concerns, there is a need to strengthen childcare-related support at physical stores and provide users with personalized, immediate advice. However, conventional systems have made it difficult to provide emotionally resonant support and to create printed information in a physical store environment.

[0806] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0807] In this invention, the server includes an input means for recognizing information input via voice, a voice recognition means for converting it into text information, and a summarization means using sentiment analysis and a generative model. This makes it possible to provide real-time advice tailored to the user's emotions, as well as compress and compile information, even in a physical store environment.

[0808] "Information input via voice" refers to data and messages that users send to a device using spoken language.

[0809] "Speech recognition means" refers to technologies and devices that analyze speech information and convert it into text information.

[0810] A "generative model" is a type of artificial intelligence technology that generates new data or information based on given data.

[0811] A "summarization tool" refers to a process or device that extracts important information from a large amount of data and summarizes it concisely.

[0812] A "memory device" refers to a device or method for holding and storing data or information.

[0813] "Output means" refers to devices or technologies for displaying or physically outputting processed information to an external source.

[0814] An "interface means" refers to a device or method that serves as a window for information exchange between people and systems.

[0815] A "digital visual device" refers to a wearable device or screen that displays visual information using digital technology.

[0816] "Personalized responses" refer to feedback and messages that are customized based on the user's individual circumstances and emotions.

[0817] "Devices for converting to print media" refer to equipment and methods used to convert digital data into physical printed materials.

[0818] The system realizing this invention works in conjunction with smart glasses or digital visual devices to process the user's childcare-related voice input. When a user asks for advice on childcare while selecting products, the system first collects the voice using the microphone on the smart glasses. The voice is recorded on the device and transmitted to a server via the network.

[0819] The server converts speech data into text using speech recognition software (e.g., Google Speech-to-Text API). This text is then analyzed by an emotion engine (e.g., IBM Watson Tone Analyzer) for sentiment analysis. The analyzed sentiment data is then input into a generative AI model (e.g., OpenAI's GPT) to generate personalized advice and responses based on the user's emotions.

[0820] The generated response is displayed in real time on the smart glasses' screen. This allows users to receive immediate and appropriate support regarding childcare within the store. For example, if a user is unsure about choosing diapers, the smart glasses will display a message such as, "We understand your concerns. Here are some suggestions for improvement..."

[0821] Examples of prompt messages include the following:

[0822] "Users are feeling anxious about choosing diapers. To alleviate their anxiety, please recommend a brand in 10 characters or less."

[0823] This system allows for real-time consultation and information exchange regarding childcare, enabling a more effective shopping experience.

[0824] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0825] Step 1:

[0826] The user inputs childcare-related information via voice through the microphone on their smart glasses. The input voice data is temporarily stored on the device, and then the device transmits this voice data to a server over the network.

[0827] Step 2:

[0828] The server converts the received audio data into text data using speech recognition software. In this process, the input audio data is analyzed, and a corresponding string of characters is generated. The output text data is then used for subsequent processing.

[0829] Step 3:

[0830] The server inputs text data into the emotion engine, which analyzes the user's emotions. The emotion engine detects emotional cues contained in the text data and extracts emotional indicators such as joy, sadness, and anxiety. This analysis result is used as data necessary for the next generation process.

[0831] Step 4:

[0832] The analyzed emotional and textual data is input into a generative AI model. Based on this data, the generative AI model generates personalized advice and empathetic messages that resonate with the user's emotions. These output messages play a crucial role in providing instantaneous support to the user.

[0833] Step 5:

[0834] The server sends the generated advice and messages back to the device. The device displays this information on the smart glasses' screen. Users can immediately receive this feedback visually and use it as a reference for shopping or childcare consultations.

[0835] This entire process allows users to seek real-time advice on childcare in physical stores and receive empathetic feedback.

[0836] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0837] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0838] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0839] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0840] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0841] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0842] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0843] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0844] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0845] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0846] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0847] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0848] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0849] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0850] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0851] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0852] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0853] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0854] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0855] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0856] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0857] The following is further disclosed regarding the embodiments described above.

[0858] (Claim 1)

[0859] An input device for users to voice-input events related to childcare,

[0860] A speech recognition device that converts the aforementioned voice input into text data,

[0861] A summarization device that uses a generation model to summarize the aforementioned text data into a predetermined number of characters,

[0862] A storage means for storing the summarized text data in a storage device,

[0863] An output device that converts the summary text stored in the aforementioned storage device into a format suitable for bookbinding,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] An input device for users to voice-input questions about childcare via a smart device,

[0867] A speech recognition device that converts the aforementioned voice input into text data,

[0868] A response generation device that generates advice or empathy messages using a generative model based on the aforementioned text data,

[0869] The system according to claim 1, further comprising an output device for presenting the generated message to a user.

[0870] (Claim 3)

[0871] The input device transmits the summary data stored in the storage device to the output device based on the user's instructions.

[0872] The system according to claim 1, wherein the output device further comprises binding means for converting the summary data into a printable medium.

[0873] "Example 1"

[0874] (Claim 1)

[0875] An input means for users to input childcare-related activities via voice input using an information terminal,

[0876] A speech recognition means that converts the aforementioned voice input into text format,

[0877] A summarization means that uses generation techniques to compress the aforementioned text data to a specified number of characters,

[0878] A storage means for storing the compressed text data in a database,

[0879] An output means for converting saved summary data into a printable format and outputting it to a physical medium,

[0880] A system that includes this.

[0881] (Claim 2)

[0882] An input means for users to voice input childcare-related issues via an information terminal,

[0883] A speech recognition means for converting the aforementioned voice input into text,

[0884] A response generation means that generates messages of advice and empathy using generation technology based on the aforementioned text data,

[0885] The system according to claim 1, further comprising output means for providing the generated message to a user.

[0886] (Claim 3)

[0887] The input means transfers the summary data stored in the storage device to the output means in response to the user's request.

[0888] The system according to claim 1, wherein the output means further comprises a binding means for formatting the summary data with a printable paperweight.

[0889] "Application Example 1"

[0890] (Claim 1)

[0891] Information processing means for users to input life events by voice,

[0892] Language processing means for converting the aforementioned voice input into information data,

[0893] Information compression means that uses a generation method for summarizing the aforementioned information data into a predetermined number of symbols,

[0894] A media control means for storing the summarized information data in an information storage means,

[0895] A media output means that converts the summary information stored in the information storage means into a recordable format,

[0896] A product presentation method that recommends relevant products based on the user's life events,

[0897] A system that includes this.

[0898] (Claim 2)

[0899] An information processing means for users to input consultations via voice input using a smart device,

[0900] Language processing means for converting the aforementioned voice input into information data,

[0901] A response generation means that generates advice or emotional messages using a generation method based on the aforementioned information data,

[0902] The system according to claim 1, further comprising information presentation means for visually or audibly presenting the generated message.

[0903] (Claim 3)

[0904] The information processing means transmits the summary information stored in the information storage means to the media output means based on the user's instructions.

[0905] The system according to claim 1, wherein the media output means further comprises a bookbinding means for converting the summary information into a recordable medium.

[0906] "Example 2 of combining an emotion engine"

[0907] (Claim 1)

[0908] An input means for users to input childcare-related events via voice input using an information device,

[0909] A speech recognition means that converts the aforementioned voice input into text information,

[0910] An emotion analysis means for analyzing the aforementioned textual information to identify emotions,

[0911] A summarization means that uses a generative model based on the aforementioned sentiment analysis results to generate a summary that reflects sentiment information,

[0912] A storage means for storing the aforementioned summary on a recording medium,

[0913] Output means for converting a summary stored on the recording medium into a printable format,

[0914] A system that includes this.

[0915] (Claim 2)

[0916] An input method for users to input childcare-related consultations via voice input using an information device,

[0917] A speech recognition means that converts the aforementioned voice input into text information,

[0918] A response generation means that uses a generation model based on the aforementioned textual information and sentiment analysis results to generate advice or empathetic messages,

[0919] The system is equipped with an output means for presenting the generated message to the user.

[0920] The system according to claim 1.

[0921] (Claim 3)

[0922] The input means further comprises a conversion means for transmitting a summary stored on a recording medium to an output means based on user instructions, and for converting the summary into a print medium.

[0923] The system according to claim 1.

[0924] "Application example 2 when combining with an emotional engine"

[0925] (Claim 1)

[0926] An input means for recognizing information input via voice,

[0927] A speech recognition means that converts the aforementioned speech information into text information,

[0928] A summarization means using a generation model for compressing the aforementioned character information to a predetermined number of characters,

[0929] A storage means for storing the compressed character information,

[0930] Output means for converting the stored character information into a format suitable for bookbinding,

[0931] An interface for facilitating consultations regarding childcare in a real-world retail environment,

[0932] A system that includes this.

[0933] (Claim 2)

[0934] Audio information is input in real time through a digital visual device, and speech recognition is performed.

[0935] The system according to claim 1, comprising means for providing a personalized response using a generative model based on analyzed sentiment data.

[0936] (Claim 3)

[0937] The system according to claim 1, further comprising a device for transmitting stored character information to an output means and converting it into a print medium. [Explanation of Symbols]

[0938] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An input device for users to voice-input events related to childcare, A speech recognition device that converts the aforementioned voice input into text data, A summarization device that uses a generation model to summarize the aforementioned text data into a predetermined number of characters, A storage means for storing the summarized text data in a storage device, An output device that converts the summary text stored in the aforementioned storage device into a format suitable for bookbinding, A system that includes this.

2. An input device for users to voice-input questions about childcare via a smart device, A speech recognition device that converts the aforementioned voice input into text data, A response generation device that generates advice or empathy messages using a generative model based on the aforementioned text data, The system according to claim 1, further comprising an output device for presenting the generated message to a user.

3. The input device transmits the summary data stored in the storage device to the output device based on the user's instructions. The system according to claim 1, wherein the output device further comprises binding means for converting the summary data into a printable medium.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A