system

JP2026085724APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Conventional voice dialogue systems face challenges in accurately recognizing user intentions and emotions, generating flexible responses, and effectively utilizing information across different devices, leading to poor user experience, particularly in childcare and customer service settings.

Method used

A system integrating speech recognition, natural language processing, conversation generation, speech synthesis, and data sharing technologies, utilizing a server to analyze voice data, generate appropriate responses, and adjust conversation patterns based on user feedback, with terminals capturing and outputting voice and video data to enhance interaction.

Benefits of technology

Enables natural and effective dialogue by accurately recognizing user intentions and emotions, providing flexible responses, and improving user experience through continuous learning and data sharing across devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085724000001_ABST
    Figure 2026085724000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A speech recognition means that receives voice input and converts the voice into text data, A natural language understanding means that analyzes the text data and determines the user's intent and emotions, A conversation generation means that generates a response based on the said determination, A speech synthesis means that converts the response into speech data, An output means for outputting the audio data, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0003] , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

[0006] "Speech recognition means" refers to a device or software that has the function of receiving speech data and converting it into text data.

[0007] "Natural language understanding means" refers to a process or technology for analyzing text data and determining the user's intent and emotions.

[0008] A "conversation generation means" is a device or software that has the function of generating an appropriate response based on the output of a natural language understanding means.

[0009] "Speech synthesis means" refers to a technology or device that has the function of converting generated text data into speech data and outputting it as speech.

[0010] "Output means" refers to devices such as speakers that physically play back synthesized speech data and provide it to the user.

[0011] "Adjustment means" refers to processes or software that have the function of adjusting and optimizing conversation patterns based on feedback data obtained from users.

[0012] A "data sharing method" is a technology or system that has the function of improving the accuracy of communication with users by enabling multiple terminals to communicate with each other and share learning data. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention provides a system that reduces the burden on parents in childcare and supports interaction with children by coordinating the functions of speech recognition, natural language processing, conversation generation, speech synthesis, and data output. This system mainly consists of a server, terminals, and users.

[0035] The server, acting as the system's central hub, is responsible for analyzing voice data and generating conversations. Upon receiving text data from a terminal, the server uses natural language understanding technology to analyze the child's intentions and emotions, and generates the most appropriate response. The generated response is then converted into voice data using speech synthesis technology and sent back to the terminal.

[0036] The device is installed in the home and serves as a direct interface with the child. Equipped with a microphone and speaker, it detects and captures the child's voice. The captured audio data is converted into text data using speech recognition technology and sent to a server. The server then plays the audio data back through the speaker, providing a response to the child. For example, if a child says, "Read me a picture book," the device recognizes this and can read the story aloud using audio generated from the server.

[0037] Users (parents) can manage system settings and feedback through devices such as smartphones. Parents can use the app to check their child's learning progress and conversation history, and customize specific conversation patterns as needed. For example, by entering the phrase "It's bath time," the system can be set to communicate this to the child at a specific time.

[0038] In this way, this system utilizes voice technology and artificial intelligence to support parents in childcare and promote interactive dialogue with their children.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The device receives audio from the child using its built-in microphone to capture voice input. The received audio is processed in real time, and noise reduction is performed.

[0042] Step 2:

[0043] The device uses speech recognition software to convert the captured audio into text data. The converted text data, along with metadata about the child's intentions, is sent to the server.

[0044] Step 3:

[0045] The server receives text data sent from the terminal and uses natural language processing technology to analyze the intent and emotions behind the child's statements.

[0046] Step 4:

[0047] The server executes a conversation generation engine to generate the optimal response based on the analysis results. The generated response is then passed to the speech synthesis engine.

[0048] Step 5:

[0049] The server uses speech synthesis technology to convert the text responses generated by conversation generation into speech data. The synthesized speech data is then sent to the terminal.

[0050] Step 6:

[0051] The device plays audio data received from the server. The played audio is delivered to the child via a speaker, facilitating natural conversation.

[0052] Step 7:

[0053] Users can review conversations and system behavior through a smartphone app, providing feedback and customization as needed. For example, they can adjust specific phrases to determine when and how the system uses those phrases.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] Conventional voice dialogue systems often suffer from insufficient accuracy in speech recognition and generated responses that do not adequately reflect the user's intentions or emotions, resulting in a poor user experience. Furthermore, the system's response patterns are difficult to adjust flexibly, making it challenging to effectively utilize information data across different devices. Therefore, particularly in childcare support settings, there is a need for technology that enables satisfying dialogue between parents and children.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes speech recognition means, natural language processing means, and response generation means. This makes it possible to analyze the user's intentions and emotions with high accuracy and generate appropriate responses. Furthermore, it enables the sharing of information data between multiple devices, thereby improving the user experience.

[0059] "Speech recognition means" refers to a technology that receives speech as input and converts said speech into text data.

[0060] "Natural language processing means" refers to technologies that analyze converted character data to understand the user's intentions and emotions.

[0061] "Response generation means" refers to a technology that generates an appropriate response based on the judgment results of natural language processing means.

[0062] "Speech synthesis means" refers to a technology that converts a generated response into speech data, making it available for output as speech.

[0063] "Output means" refers to a device or technology that outputs the generated audio data in real time.

[0064] "Adjustment means" refers to a technology that dynamically adjusts response patterns based on information data obtained from the user.

[0065] A "data sharing method" is a technology that communicates data and shares information between multiple terminals or processing devices.

[0066] This system aims to enable natural communication through voice interaction in homes and educational settings. Specifically, it is a system that supports dialogue with children by integrating speech recognition, natural language processing, conversation generation, and speech synthesis technologies.

[0067] The server plays a central role in the system. The server receives text data transmitted from terminals and performs natural language processing using a generative AI model. Specifically, the generative AI model utilizes high-performance natural language understanding models (e.g., GPT and BERT). The server leverages these models to analyze the child's intentions and emotions and generate the optimal response. The generated response is then converted into speech data using speech synthesis technology (e.g., Amazon Polly or Google® Text-to-Speech).

[0068] The device is installed in the home and functions as a direct interface with the child. It has a built-in microphone and speaker, and uses technologies such as Google Speech-to-Text API and IBM Watson® to convert the child's voice into text data. This converted text data is then sent to a server via the internet. The device also plays back the received audio data through the speaker, providing direct responses to the child. For example, if a child asks, "Read this picture book," the device recognizes this request and can read the story aloud using audio generated from the server.

[0069] Users (parents) manage system settings and feedback via mobile devices such as smartphones and tablets. Through the application, users can check their child's learning progress and conversation history, and customize specific conversation patterns. For example, by setting the phrase "It's bath time," the device can be configured to notify the child at a specific time. Also, by entering text such as "Suggest a dinner my child will enjoy," the generative AI model can provide appropriate suggestions.

[0070] In this way, this system utilizes voice technology and artificial intelligence to reduce the burden of childcare on parents and support smooth communication with their children.

[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0072] Step 1:

[0073] The device captures the child's voice using a microphone. The input is the child's voice, which is acquired as audio data. The device then applies noise reduction to this audio data to produce clear sound quality.

[0074] Step 2:

[0075] The device converts captured audio data into text data using speech recognition software. The input is the audio data processed in the previous step, and the output is text data. Specifically, the Google Speech-to-Text API is used to accurately convert the audio "Read the picture book" into text.

[0076] Step 3:

[0077] The terminal sends the converted text data to the server. The input is text data, and the output is the transmission of data to the server. Encryption technology is used to ensure the data is transmitted securely.

[0078] Step 4:

[0079] The server inputs the received text data into a generative AI model and performs natural language processing. The input is text data from the terminal, and the output is the analysis result. The generative AI model understands the child's intentions and emotions, and derives a conclusion such as, "The child wants to hear a story."

[0080] Step 5:

[0081] The server generates a response based on the results of natural language processing. The input is the parsing result, and the output is the generated response text. For example, it generates the beginning of a story, "Once upon a time..."

[0082] Step 6:

[0083] The server converts the generated response text into speech data using a text-to-speech tool. The input is the response text, and the output is speech data. The Google Text-to-Speech API is used to generate fluent speech.

[0084] Step 7:

[0085] The server transmits synthesized audio data to the terminal. The input is audio data, and the output is data transmission to the terminal. Transmission is performed in real time to avoid delays.

[0086] Step 8:

[0087] The device plays the received audio data through its speaker. The input is audio data, and the output is audio played through the speaker. This allows the server to read aloud stories to children.

[0088] (Application Example 1)

[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0090] In modern brick-and-mortar stores, it is a challenge for customers to effectively obtain product information and receive the necessary guidance. In particular, there is a growing need for efficient systems that can respond quickly and accurately to a variety of questions. Furthermore, there is a demand to reduce the workload on store staff and provide consistent service to customers.

[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0092] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language understanding means for analyzing the text data and determining the intentions and emotions of the person speaking, and dialogue generation means for generating a response based on the determination. This makes it possible to provide information in response to questions from customers.

[0093] "Speech recognition means" refers to technology that receives speech input and converts that speech into text data.

[0094] "Natural language understanding methods" are technologies that analyze text data to determine the intentions and emotions of the person speaking.

[0095] A "dialogue generation method" is a technology for generating appropriate responses based on analyzed intentions and emotions.

[0096] "Speech synthesis means" refers to a technology that converts generated responses into speech data.

[0097] "Output means" refers to a system for providing audio data to the user.

[0098] "Information presentation means" refers to the function of providing appropriate information in response to questions from customers.

[0099] "Adjustment means" refers to techniques for optimizing dialogue patterns based on feedback data.

[0100] A "data sharing method" is a system that allows multiple information terminals to share learning data and improve the quality of dialogue.

[0101] To implement this invention, first, the server receives voice input from a customer using speech recognition means and converts it into text data. The server then uses natural language understanding means to analyze the text data and determine the customer's intentions and emotions. Next, using dialogue generation means, the server generates an optimal response based on the determined intentions. This generated response is converted into voice data by speech synthesis means and provided to the customer through output means.

[0102] The terminal is installed in the physical store and is equipped with a microphone and speaker. The terminal captures the voice of the customer using the microphone and sends the audio data to the server. It also plays back the audio data sent from the server through the speaker, providing a response to the customer.

[0103] Users (store staff and operators) can input feedback into the system and optimize dialogue patterns through adjustment mechanisms. Multiple terminals communicate to share learning data, and the accuracy of interactions with customers improves by utilizing data sharing mechanisms.

[0104] For example, if a customer asks, "Where are the cereals on sale?", the system can respond, "The sale cereals are in the food section on the second floor. Shall I show you the way?"

[0105] An example of a prompt for a generative AI model is: "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the best response."

[0106] Thus, the present invention provides a system that allows customers to easily obtain necessary information within a physical store, thereby reducing the burden on store staff and improving service.

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The terminal captures the customer's voice input via the microphone. It receives the voice data as input. This is sent to speech recognition and converted into text data. The speech_recognition library is used for speech recognition. Text data is generated as output.

[0110] Step 2:

[0111] The server receives the generated text data and analyzes it using natural language understanding (NLP) tools. Specifically, it analyzes the text data to determine the intentions and emotions of customers. To achieve this, it utilizes a generative AI model and processes the input text. The output is the analysis result.

[0112] Step 3:

[0113] The server generates the optimal response using a dialogue generation mechanism based on the analysis results. The AI ​​model generates a prompt using the analysis results as input. For example, a prompt such as "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the optimal response." is generated. The response text is obtained as output.

[0114] Step 4:

[0115] The server converts the generated response text into speech data using a speech synthesis system. Specific software, such as the pyttsx3 library, is used for speech synthesis. Speech data is generated as output.

[0116] Step 5:

[0117] The terminal provides the generated audio data to the customer through the speaker. It receives audio data as input and plays it back through the physical speaker. Information is effectively provided by outputting an appropriate audio response to the customer.

[0118] By following these steps, it becomes possible to provide information in real time in response to customers' questions.

[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0120] This invention is a system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, aiming to make interactions with users (especially children) more natural and effective. The system consists of a server, a terminal, and a user.

[0121] The server plays a central role in analyzing the voice data and performing natural language processing to determine the child's intentions and emotions. The server receives the speech-recognized text data and uses an emotion engine to recognize the emotions. Based on this result, the conversation generation engine determines the optimal response and generates the response. The generated response is synthesized into speech and sent to the terminal as voice data. For example, if a child excitedly says, "Let's play quickly!", the server recognizes that emotion and generates a response in a lively tone, "What should we play?"

[0122] The device is responsible for the initial process of capturing voice input and converting it to text. It receives voice from the child via a microphone and performs real-time speech recognition. This text data is sent to a server, and upon receiving a response voice data from the server, it plays it back to the child through a speaker. The device can also capture the user's facial expressions with a camera as supplementary information for emotion recognition.

[0123] The user (parent) has the role of managing interactions with the system through a smartphone application. The parent can review the child's speech and the emotional data recognized by the system, and adjust the conversation response patterns as needed. For example, if the system detects that the child is sounding sad, the parent can customize the system's response by adding phrases to offer words of encouragement.

[0124] In this way, this system utilizes voice and emotion data to enable interactive dialogue with children and reduce the burden of childcare.

[0125] The following describes the processing flow.

[0126] Step 1:

[0127] The device receives the child's voice via a microphone and converts it to text in real time using speech recognition software. This text data also includes attributes such as the timing and volume of the speech.

[0128] Step 2:

[0129] The device sends text data and speech attributes to the server. At the same time, it also sends facial expression data of the child captured by the camera, which improves the accuracy of emotion recognition.

[0130] Step 3:

[0131] The server uses the received text and facial expression data to run an emotion engine and analyze the child's emotions. The results of the emotion analysis are labeled as joy, sadness, anger, etc.

[0132] Step 4:

[0133] The server uses natural language processing technology to understand the user's intent from text data and runs a conversation generation engine that combines this with sentiment analysis results to generate the optimal response.

[0134] Step 5:

[0135] The server uses speech synthesis technology to convert the generated text response into speech and sends the synthesized speech data to the terminal.

[0136] Step 6:

[0137] The device plays audio data received from the server through its speaker and provides a response to the child. The response reflects a tone and tempo based on analyzed emotions.

[0138] Step 7:

[0139] Users can view the system's operation and conversation history through a smartphone app. Parents can provide feedback on response patterns based on recognized emotion data and incorporate it into future interactions. For example, they can change the settings so that the system uses gentle words when a child is feeling anxious.

[0140] (Example 2)

[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0142] Current voice dialogue systems face the challenge of accurately recognizing user emotions and generating natural, appropriate responses based on those emotions. Interacting with children, in particular, requires flexible responses due to their rapidly changing emotional states. Furthermore, continuous learning through user feedback is necessary to improve response accuracy.

[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0144] In this invention, the server includes speech recognition means for acquiring voice input and converting the voice into text information, natural language understanding means for analyzing the text information and determining the user's intentions and emotions, and means for creating a response corresponding to a prompt sentence using a generative model. This enables natural and effective dialogue that takes the user's emotions into consideration, as well as continuous response improvement based on feedback.

[0145] "Speech recognition means" refers to a technical element that receives speech input, analyzes the speech, and converts it into text information.

[0146] "Natural language understanding means" refers to technological elements that provide a process for analyzing and understanding the user's intentions and emotions based on text information.

[0147] A "conversation generation means" is a technological element that generates natural and appropriate responses while taking into account the user's intentions and emotions.

[0148] "Speech synthesis means" refers to a technical element that converts the generated response into speech information and achieves natural-sounding speech output.

[0149] "Output means" refers to a technical element for outputting the response generated as audio data and transmitting it to the user.

[0150] A "generative model" is a technology that uses algorithms learned from large amounts of data to generate responses and content in response to user input.

[0151] A "prompt statement" is text that provides instructions or context to elicit a desired response using a generative model.

[0152] This invention is a dialogue system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, and is particularly aimed at making dialogue with children more natural and effective. The system consists of a server, a terminal, and a parent who is the user.

[0153] The server receives the speech-recognized text information and uses natural language understanding to determine the user's intent and emotions. This analysis uses a natural language processing model (e.g., a large-scale language model). Based on the determination results, a conversation generation system generates a response, and a generation AI model is used to create a natural response corresponding to the prompt sentence. For example, if a child says, "I want to go to the park today," the server receives this and generates a response such as, "That sounds fun! What do you want to do at the park?" This response is converted into speech information by a speech synthesis system and sent to the terminal.

[0154] The device uses a microphone to acquire voice input and converts the speech to text in real time using speech recognition software (e.g., various APIs). The text information is sent to a server, which then receives the generated voice information and outputs it through the speaker. The device is also equipped with a camera, which can identify the user's facial expressions for facial recognition.

[0155] The user (parent) manages the system through a smartphone application. They can see what conversations the user has had through the user interface and adjust the dialogue response patterns and conversation generation methods as needed. For example, the parent can input a prompt such as "Generate a conversation suitable for when the child is interested in sports," and then create an appropriate conversation based on that prompt.

[0156] This system utilizes generative models to enable interactive and emotionally rich dialogue, thereby reducing the burden of childcare.

[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0158] Step 1:

[0159] The device captures the child's voice in real time via a microphone. Speech recognition software is used to convert this audio data into text data. The input is the child's voice, and the output is text data. This conversion uses phonological analysis and pattern matching techniques to replace the audio signal with a string of characters.

[0160] Step 2:

[0161] The terminal sends the converted text data to the server. The server receives this text data and analyzes its intent and sentiment using natural language understanding tools. The input is text data, and the output is the result of the intent and sentiment analysis. This analysis applies natural language processing techniques and sentiment analysis algorithms to identify the user's intent and emotions.

[0162] Step 3:

[0163] The server utilizes a generative AI model to generate natural responses corresponding to prompt sentences based on the analysis results. The input consists of the analysis results of intent and sentiment, and the prompt sentence; the output is a text-based response. This employs predictive generation using a language model to create appropriate and contextually relevant sentences.

[0164] Step 4:

[0165] The server passes the generated text-based response to a speech synthesis system for speech synthesis. The input is the response text, and the output is speech data. This process uses phonological synthesis techniques and speech adjustment algorithms to generate speech with a natural tone.

[0166] Step 5:

[0167] The server sends the generated audio data to the terminal. The terminal receives this audio data and plays it for the child through the speaker. The input is the audio data, and the output is the sound from the speaker. This audio playback involves decoding and amplifying the audio data, ensuring that it is transmitted to the child in clear sound.

[0168] Step 6:

[0169] The user (parent) can view the system's dialogue logs and the child's emotional data through a smartphone application and adjust the conversation response patterns as needed. Input is the dialogue log and emotional data, while output is the updated response pattern or prompt text. This allows parents to adjust system settings and provide better childcare support.

[0170] (Application Example 2)

[0171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0172] Conventional voice dialogue systems simply convert user speech into text and generate responses based on simple rules, which has resulted in a lack of natural conversation that fully understands the user's emotions and intentions. Furthermore, there was a lack of means to utilize video information to more accurately recognize user emotions and reflect them in responses. As a result, improvements in user experience and the provision of personalized services have not been fully realized.

[0173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0174] In this invention, the server includes speech recognition means that receive voice input and convert the voice into text data, natural language understanding means that analyze the text data and determine the user's intentions and emotions, and video analysis means that capture the customer's video and recognize emotions using the video data. This makes it possible to more accurately recognize emotions based on the user's voice and video information and generate natural and effective responses.

[0175] "Voice recognition means" refers to a device or software that receives voice input using a microphone or sensor and converts that voice into text data.

[0176] "Natural language understanding means" refers to technologies that include algorithms and processes for analyzing text data and extracting and determining the user's intentions and emotions.

[0177] A "conversation generation means" is a system or technology for generating a situation-appropriate response based on information extracted by a natural language understanding means.

[0178] "Speech synthesis means" refers to a system or device that converts text information into speech signals in order to represent the generated response as speech data.

[0179] "Output means" refers to a device or interface for allowing the user to hear the generated audio data through a speaker or display.

[0180] "Video analysis means" refers to technologies and devices that analyze customer videos acquired using cameras and sensors and recognize emotions from them.

[0181] A "coordination mechanism" is a system that dynamically optimizes response conversation patterns and the content of services provided based on feedback information and environmental information.

[0182] "Information sharing methods" refer to protocols and functions that improve the accuracy of communication with users by sharing learning information and generated data among multiple devices.

[0183] This invention is a system that enables interactive dialogue with users based on audio and video data. Three main components play crucial roles in implementing this system: the server, the terminal, and the user.

[0184] The server is responsible for the central processing of this system. Upon receiving voice input, the server performs speech recognition and converts the voice into text data. For this process, the Google Speech-to-Text API can be used as the speech recognition software. Subsequently, Google Dialogflow is utilized as a natural language understanding tool to analyze the text data and determine the user's intent and emotions. Based on this determination, the conversation generation engine generates the optimal response and determines the content of the response to provide an emotionally appropriate service. The response is then converted into audio data using Amazon Polly as the speech synthesis software.

[0185] The terminal captures voice input from the user and transmits it to the server in real time. It also uses a camera to capture video of the customer and provides auxiliary data for emotion recognition based on that video. Image analysis software is used for video analysis from the camera to recognize emotions from the acquired video data.

[0186] Users operate the system through smartphones or other interfaces. In particular, parents (users) can manage the system, monitor interactions with their children, and adjust conversation patterns as needed. Feedback from parents is used to adjust response patterns, and the service content is optimized through the aforementioned adjustment methods.

[0187] For example, if the device receives an excited voice message from a child saying, "I want to see the toys!", the server analyzes the voice and generates a response such as, "What kind of toys do you like?", and performs speech synthesis. Finally, this voice data is played from the device's speaker, and the device sends related information to a smartphone, allowing the parent to check detailed product information.

[0188] An example of a prompt to input into the generative AI model is, "What emotion is this child expressing? Please provide the emotion analysis result for 'I want to see the toys!'" This can enrich the user experience and support communication between parents and children.

[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0190] Step 1:

[0191] The device receives user voice input via a microphone. This voice data is converted into text data in real time using the Google Speech-to-Text API. The input is raw voice data, and the output is recognized text data.

[0192] Step 2:

[0193] The server receives text data sent from the terminal. Next, it processes this text data using a natural language understanding process with Google Dialogflow to analyze the user's intent and emotions. In this process, text data is the input, and the user's intent and emotions as the analysis result are output.

[0194] Step 3:

[0195] The server uses a conversation generation engine based on the analysis results to generate an appropriate response. This process takes the analysis results, including the user's intent and emotions, as input and outputs a corresponding text-based response.

[0196] Step 4:

[0197] The server converts the generated text response into audio data using Amazon Polly. The input is the text data of the response, and the output is audio data that can be played back through a speaker.

[0198] Step 5:

[0199] The terminal plays the audio data received from the server through its speaker and communicates the response to the user. In this case, the audio data is the input, and the played audio is the output.

[0200] Step 6:

[0201] The device's camera captures the user's image, and video analysis software is used to analyze that video data. The input for this step is the captured video, and the output is data that recognizes the user's emotions.

[0202] Step 7:

[0203] The user (parent) monitors the system's operation via a smartphone app and adjusts the conversation patterns as needed. Here, feedback data from the system is input, and the adjusted pattern information is output.

[0204] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0205] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0206] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0207] [Second Embodiment]

[0208] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0209] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0210] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0211] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0212] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0214] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0215] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0216] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0217] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0218] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0219] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0220] This invention provides a system that reduces the burden on parents in childcare and supports interaction with children by coordinating the functions of speech recognition, natural language processing, conversation generation, speech synthesis, and data output. This system mainly consists of a server, terminals, and users.

[0221] The server, acting as the system's central hub, is responsible for analyzing voice data and generating conversations. Upon receiving text data from a terminal, the server uses natural language understanding technology to analyze the child's intentions and emotions, and generates the most appropriate response. The generated response is then converted into voice data using speech synthesis technology and sent back to the terminal.

[0222] The device is installed in the home and serves as a direct interface with the child. Equipped with a microphone and speaker, it detects and captures the child's voice. The captured audio data is converted into text data using speech recognition technology and sent to a server. The server then plays the audio data back through the speaker, providing a response to the child. For example, if a child says, "Read me a picture book," the device recognizes this and can read the story aloud using audio generated from the server.

[0223] Users (parents) can manage system settings and feedback through devices such as smartphones. Parents can use the app to check their child's learning progress and conversation history, and customize specific conversation patterns as needed. For example, by entering the phrase "It's bath time," the system can be set to communicate this to the child at a specific time.

[0224] In this way, this system utilizes voice technology and artificial intelligence to support parents in childcare and promote interactive dialogue with their children.

[0225] The following describes the processing flow.

[0226] Step 1:

[0227] The device receives audio from the child using its built-in microphone to capture voice input. The received audio is processed in real time, and noise reduction is performed.

[0228] Step 2:

[0229] The device uses speech recognition software to convert the captured audio into text data. The converted text data, along with metadata about the child's intentions, is sent to the server.

[0230] Step 3:

[0231] The server receives text data sent from the terminal and uses natural language processing technology to analyze the intent and emotions behind the child's statements.

[0232] Step 4:

[0233] The server executes a conversation generation engine to generate the optimal response based on the analysis results. The generated response is then passed to the speech synthesis engine.

[0234] Step 5:

[0235] The server uses speech synthesis technology to convert the text responses generated by conversation generation into speech data. The synthesized speech data is then sent to the terminal.

[0236] Step 6:

[0237] The device plays audio data received from the server. The played audio is delivered to the child via a speaker, facilitating natural conversation.

[0238] Step 7:

[0239] Users can review conversations and system behavior through a smartphone app, providing feedback and customization as needed. For example, they can adjust specific phrases to determine when and how the system uses those phrases.

[0240] (Example 1)

[0241] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0242] Conventional voice dialogue systems often suffer from insufficient accuracy in speech recognition and generated responses that do not adequately reflect the user's intentions or emotions, resulting in a poor user experience. Furthermore, the system's response patterns are difficult to adjust flexibly, making it challenging to effectively utilize information data across different devices. Therefore, particularly in childcare support settings, there is a need for technology that enables satisfying dialogue between parents and children.

[0243] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0244] In this invention, the server includes speech recognition means, natural language processing means, and response generation means. This makes it possible to analyze the user's intentions and emotions with high accuracy and generate appropriate responses. Furthermore, it enables the sharing of information data between multiple devices, thereby improving the user experience.

[0245] "Speech recognition means" refers to a technology that receives speech as input and converts said speech into text data.

[0246] "Natural language processing means" refers to technologies that analyze converted character data to understand the user's intentions and emotions.

[0247] "Response generation means" refers to a technology that generates an appropriate response based on the judgment results of natural language processing means.

[0248] "Speech synthesis means" refers to a technology that converts a generated response into speech data, making it available for output as speech.

[0249] "Output means" refers to a device or technology that outputs the generated audio data in real time.

[0250] "Adjustment means" refers to a technology that dynamically adjusts response patterns based on information data obtained from the user.

[0251] A "data sharing method" is a technology that communicates data and shares information between multiple terminals or processing devices.

[0252] This system aims to enable natural communication through voice interaction in homes and educational settings. Specifically, it is a system that supports dialogue with children by integrating speech recognition, natural language processing, conversation generation, and speech synthesis technologies.

[0253] The server plays a central role in the system. The server receives text data sent from terminals and performs natural language processing using a generative AI model. Specifically, the generative AI model utilizes high-performance natural language understanding models (e.g., GPT and BERT). The server leverages these models to analyze the child's intentions and emotions and generate the optimal response. The generated response is then converted into speech data using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[0254] The device is installed in the home and functions as a direct interface with the child. It has a built-in microphone and speaker, and uses technologies such as Google Speech-to-Text API and IBM Watson to convert the child's voice into text data. This converted text data is then sent to a server via the internet. The device also plays back the received audio data through the speaker, providing direct responses to the child. For example, if a child asks, "Read this picture book," the device recognizes this request and can read the story aloud using audio generated from the server.

[0255] Users (parents) manage system settings and feedback via mobile devices such as smartphones and tablets. Through the application, users can check their child's learning progress and conversation history, and customize specific conversation patterns. For example, by setting the phrase "It's bath time," the device can be configured to notify the child at a specific time. Also, by entering text such as "Suggest a dinner my child will enjoy," the generative AI model can provide appropriate suggestions.

[0256] In this way, this system utilizes voice technology and artificial intelligence to reduce the burden of childcare on parents and support smooth communication with their children.

[0257] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0258] Step 1:

[0259] The device captures the child's voice using a microphone. The input is the child's voice, which is acquired as audio data. The device then applies noise reduction to this audio data to produce clear sound quality.

[0260] Step 2:

[0261] The device converts captured audio data into text data using speech recognition software. The input is the audio data processed in the previous step, and the output is text data. Specifically, the Google Speech-to-Text API is used to accurately convert the audio "Read the picture book" into text.

[0262] Step 3:

[0263] The terminal sends the converted text data to the server. The input is text data, and the output is the transmission of data to the server. Encryption technology is used to ensure the data is transmitted securely.

[0264] Step 4:

[0265] The server inputs the received text data into a generative AI model and performs natural language processing. The input is text data from the terminal, and the output is the analysis result. The generative AI model understands the child's intentions and emotions and derives a conclusion, for example, that "the child wants to hear a story."

[0266] Step 5:

[0267] The server generates a response based on the results of natural language processing. The input is the parsing result, and the output is the generated response text. For example, it generates the beginning of a story, "Once upon a time..."

[0268] Step 6:

[0269] The server converts the generated response text into speech data using a text-to-speech tool. The input is the response text, and the output is speech data. The Google Text-to-Speech API is used to generate fluent speech.

[0270] Step 7:

[0271] The server transmits synthesized audio data to the terminal. The input is audio data, and the output is data transmission to the terminal. Transmission is performed in real time to avoid delays.

[0272] Step 8:

[0273] The device plays the received audio data through its speaker. The input is audio data, and the output is audio played through the speaker. This allows the server to read aloud stories to children.

[0274] (Application Example 1)

[0275] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0276] In modern physical stores, it is a challenge for customers to effectively obtain product information and receive necessary guidance. In particular, there is an increasing need for an efficient system that can respond quickly and accurately to a variety of questions. In addition, it is required to reduce the burden on store staff and provide uniform service to customers.

[0277] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0278] In this invention, the server includes speech recognition means for receiving speech input and converting the speech into text data, natural language understanding means for analyzing the text data and judging the intention and emotion of the interlocutor, and dialogue generation means for generating a response based on the judgment. Thereby, it is possible to provide information corresponding to the questions of customers.

[0279] "Speech recognition means" is a technology for receiving speech input and converting the speech into text data.

[0280] "Natural language understanding means" is a technology for analyzing text data and judging the intention and emotion of the interlocutor.

[0281] "Dialogue generation means" is a technology for generating an appropriate response based on the analyzed intention and emotion.

[0282] "Speech synthesis means" is a technology for converting the generated response into speech data.

[0283] "Output means" is a system for providing speech data to the user.

[0284] "Information presentation means" is a function for providing appropriate information according to the questions of customers.

[0285] The "adjustment means" is a technology for optimizing the dialogue pattern based on feedback data.

[0286] The "data sharing means" is a system in which multiple information terminals share learning data to improve the quality of dialogue.

[0287] To implement this invention, first, the server uses the speech recognition means to receive the voice input of the customer entering the store and converts it into text data. The server then analyzes the text data by utilizing the natural language understanding means to judge the intention and emotion of the customer entering the store. Next, using the dialogue generation means, the server generates an optimal response based on the judged intention. This generated response is converted into voice data by the voice synthesis means and provided to the customer entering the store through the output means.

[0288] The terminal is installed inside the physical store and equipped with a microphone and a speaker. The terminal captures the voice of the customer entering the store with the microphone and transmits the voice data to the server. Also, it plays the voice data transmitted from the server with the speaker and plays a role in presenting the response to the customer entering the store.

[0289] The user (store staff or operator) can input feedback into the system and optimize the dialogue pattern through the adjustment means. Multiple terminals communicate to share learning data, and by utilizing the data sharing means, the accuracy of the dialogue with the customer entering the store is improved.

[0290] As a specific example, when a customer entering the store asks "Where is the serial on sale?", the system can answer "The serial on sale is at the food corner on the second floor. Shall I guide you?"

[0291] An example of the prompt text for the generation AI model is "Received a question from the customer 'Where is the serial on sale?'. Please generate an optimal response."

[0292] Thus, the present invention provides a system that allows customers to easily obtain necessary information within a physical store, thereby reducing the burden on store staff and improving service.

[0293] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0294] Step 1:

[0295] The terminal captures the customer's voice input via the microphone. It receives the voice data as input. This is sent to speech recognition and converted into text data. The speech_recognition library is used for speech recognition. Text data is generated as output.

[0296] Step 2:

[0297] The server receives the generated text data and analyzes it using natural language understanding (NLP) tools. Specifically, it analyzes the text data to determine the intentions and emotions of customers. To achieve this, it utilizes a generative AI model and processes the input text. The output is the analysis result.

[0298] Step 3:

[0299] The server generates the optimal response using a dialogue generation mechanism based on the analysis results. The AI ​​model generates a prompt using the analysis results as input. For example, a prompt such as "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the optimal response." is generated. The response text is obtained as output.

[0300] Step 4:

[0301] The server converts the generated response text into speech data using a speech synthesis system. Specific software, such as the pyttsx3 library, is used for speech synthesis. Speech data is generated as output.

[0302] Step 5:

[0303] The terminal provides the generated voice data to the customer through the speaker. It receives the voice data as input and plays it back using a physical speaker. By outputting an appropriate voice response to the customer, information is effectively provided.

[0304] Through the above steps, real-time information provision for the customer's questions is achieved.

[0305] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.

[0306] The present invention is a system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, and aims to make conversations with users (especially children) more natural and effective. This system is composed of a server, a terminal, and a user. <00009​​​​​​The device is responsible for the initial process of capturing voice input and converting it to text. It receives voice from the child via a microphone and performs real-time speech recognition. This text data is sent to a server, and upon receiving a response voice data from the server, it plays it back to the child through a speaker. The device can also capture the user's facial expressions with a camera as supplementary information for emotion recognition.

[0309] The user (parent) has the role of managing interactions with the system through a smartphone application. The parent can review the child's speech and the emotional data recognized by the system, and adjust the conversation response patterns as needed. For example, if the system detects that the child is sounding sad, the parent can customize the system's response by adding phrases to offer words of encouragement.

[0310] In this way, this system utilizes voice and emotion data to enable interactive dialogue with children and reduce the burden of childcare.

[0311] The following describes the processing flow.

[0312] Step 1:

[0313] The device receives the child's voice via a microphone and converts it to text in real time using speech recognition software. This text data also includes attributes such as the timing and volume of the speech.

[0314] Step 2:

[0315] The device sends text data and speech attributes to the server. At the same time, it also sends facial expression data of the child captured by the camera, which improves the accuracy of emotion recognition.

[0316] Step 3:

[0317] The server uses the received text and facial expression data to run an emotion engine and analyze the child's emotions. The results of the emotion analysis are labeled as joy, sadness, anger, etc.

[0318] Step 4:

[0319] The server uses natural language processing technology to understand the user's intent from text data and runs a conversation generation engine that combines this with sentiment analysis results to generate the optimal response.

[0320] Step 5:

[0321] The server uses speech synthesis technology to convert the generated text response into speech and sends the synthesized speech data to the terminal.

[0322] Step 6:

[0323] The device plays audio data received from the server through its speaker and provides a response to the child. The response reflects a tone and tempo based on analyzed emotions.

[0324] Step 7:

[0325] Users can view the system's operation and conversation history through a smartphone app. Parents can provide feedback on response patterns based on recognized emotion data and incorporate it into future interactions. For example, they can change the settings so that the system uses gentle words when a child is feeling anxious.

[0326] (Example 2)

[0327] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0328] Current voice dialogue systems face the challenge of accurately recognizing user emotions and generating natural, appropriate responses based on those emotions. Interacting with children, in particular, requires flexible responses due to their rapidly changing emotional states. Furthermore, continuous learning through user feedback is necessary to improve response accuracy.

[0329] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0330] In this invention, the server includes speech recognition means for acquiring voice input and converting the voice into text information, natural language understanding means for analyzing the text information and determining the user's intentions and emotions, and means for creating a response corresponding to a prompt sentence using a generative model. This enables natural and effective dialogue that takes the user's emotions into consideration, as well as continuous response improvement based on feedback.

[0331] "Speech recognition means" refers to a technical element that receives speech input, analyzes the speech, and converts it into text information.

[0332] "Natural language understanding means" refers to technological elements that provide a process for analyzing and understanding the user's intentions and emotions based on text information.

[0333] A "conversation generation means" is a technological element that generates natural and appropriate responses while taking into account the user's intentions and emotions.

[0334] "Speech synthesis means" refers to a technical element that converts the generated response into speech information and achieves natural-sounding speech output.

[0335] "Output means" refers to a technical element for outputting the response generated as audio data and transmitting it to the user.

[0336] A "generative model" is a technology that uses algorithms learned from large amounts of data to generate responses and content in response to user input.

[0337] A "prompt statement" is text that provides instructions or context to elicit a desired response using a generative model.

[0338] This invention is a dialogue system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, and is particularly aimed at making dialogue with children more natural and effective. The system consists of a server, a terminal, and a parent who is the user.

[0339] The server receives the speech-recognized text information and uses natural language understanding to determine the user's intent and emotions. This analysis uses a natural language processing model (e.g., a large-scale language model). Based on the determination results, a conversation generation system generates a response, and a generation AI model is used to create a natural response corresponding to the prompt sentence. For example, if a child says, "I want to go to the park today," the server receives this and generates a response such as, "That sounds fun! What do you want to do at the park?" This response is converted into speech information by a speech synthesis system and sent to the terminal.

[0340] The device uses a microphone to acquire voice input and converts the speech to text in real time using speech recognition software (e.g., various APIs). The text information is sent to a server, which then receives the generated voice information and outputs it through the speaker. The device is also equipped with a camera, which can identify the user's facial expressions for facial recognition.

[0341] The user (parent) manages the system through a smartphone application. They can see what conversations the user has had through the user interface and adjust the dialogue response patterns and conversation generation methods as needed. For example, the parent can input a prompt such as "Generate a conversation suitable for when the child is interested in sports," and then create an appropriate conversation based on that prompt.

[0342] This system utilizes generative models to enable interactive and emotionally rich dialogue, thereby reducing the burden of childcare.

[0343] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0344] Step 1:

[0345] The device captures the child's voice in real time via a microphone. Speech recognition software is used to convert this audio data into text data. The input is the child's voice, and the output is text data. This conversion uses phonological analysis and pattern matching techniques to replace the audio signal with a string of characters.

[0346] Step 2:

[0347] The terminal sends the converted text data to the server. The server receives this text data and analyzes its intent and sentiment using natural language understanding tools. The input is text data, and the output is the result of the intent and sentiment analysis. This analysis applies natural language processing techniques and sentiment analysis algorithms to identify the user's intent and emotions.

[0348] Step 3:

[0349] The server utilizes a generative AI model to generate natural responses corresponding to prompt sentences based on the analysis results. The input consists of the analysis results of intent and sentiment, and the prompt sentence; the output is a text-based response. This employs predictive generation using a language model to create appropriate and contextually relevant sentences.

[0350] Step 4:

[0351] The server passes the generated text-based response to a speech synthesis system for speech synthesis. The input is the response text, and the output is speech data. This process uses phonological synthesis techniques and speech adjustment algorithms to generate speech with a natural tone.

[0352] Step 5:

[0353] The server sends the generated audio data to the terminal. The terminal receives this audio data and plays it for the child through the speaker. The input is the audio data, and the output is the sound from the speaker. This audio playback involves decoding and amplifying the audio data, ensuring that it is transmitted to the child in clear sound.

[0354] Step 6:

[0355] The user (parent) can view the system's dialogue logs and the child's emotional data through a smartphone application and adjust the conversation response patterns as needed. Input is the dialogue log and emotional data, while output is the updated response pattern or prompt text. This allows parents to adjust system settings and provide better childcare support.

[0356] (Application Example 2)

[0357] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0358] Conventional voice dialogue systems simply convert user speech into text and generate responses based on simple rules, which has resulted in a lack of natural conversation that fully understands the user's emotions and intentions. Furthermore, there was a lack of means to utilize video information to more accurately recognize user emotions and reflect them in responses. As a result, improvements in user experience and the provision of personalized services have not been fully realized.

[0359] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0360] In this invention, the server includes speech recognition means that receive voice input and convert the voice into text data, natural language understanding means that analyze the text data and determine the user's intentions and emotions, and video analysis means that capture the customer's video and recognize emotions using the video data. This makes it possible to more accurately recognize emotions based on the user's voice and video information and generate natural and effective responses.

[0361] "Voice recognition means" refers to a device or software that receives voice input using a microphone or sensor and converts that voice into text data.

[0362] "Natural language understanding means" refers to technologies that include algorithms and processes for analyzing text data and extracting and determining the user's intentions and emotions.

[0363] A "conversation generation means" is a system or technology for generating a situation-appropriate response based on information extracted by a natural language understanding means.

[0364] "Speech synthesis means" refers to a system or device that converts text information into speech signals in order to represent the generated response as speech data.

[0365] "Output means" refers to a device or interface for allowing the user to hear the generated audio data through a speaker or display.

[0366] "Video analysis means" refers to technologies and devices that analyze customer videos acquired using cameras and sensors and recognize emotions from them.

[0367] A "coordination mechanism" is a system that dynamically optimizes response conversation patterns and the content of services provided based on feedback information and environmental information.

[0368] "Information sharing methods" refer to protocols and functions that improve the accuracy of communication with users by sharing learning information and generated data among multiple devices.

[0369] This invention is a system that enables interactive dialogue with users based on audio and video data. Three main components play crucial roles in implementing this system: the server, the terminal, and the user.

[0370] The server is responsible for the central processing of this system. Upon receiving voice input, the server performs speech recognition and converts the voice into text data. For this process, the Google Speech-to-Text API can be used as the speech recognition software. Subsequently, Google Dialogflow is utilized as a natural language understanding tool to analyze the text data and determine the user's intent and emotions. Based on this determination, the conversation generation engine generates the optimal response and determines the content of the response to provide an emotionally appropriate service. The response is then converted into audio data using Amazon Polly as the speech synthesis software.

[0371] The terminal captures voice input from the user and transmits it to the server in real time. It also uses a camera to capture video of the customer and provides auxiliary data for emotion recognition based on that video. Image analysis software is used for video analysis from the camera to recognize emotions from the acquired video data.

[0372] Users operate the system through smartphones or other interfaces. In particular, parents (users) can manage the system, monitor interactions with their children, and adjust conversation patterns as needed. Feedback from parents is used to adjust response patterns, and the service content is optimized through the aforementioned adjustment methods.

[0373] For example, if the device receives an excited voice message from a child saying, "I want to see the toys!", the server analyzes the voice and generates a response such as, "What kind of toys do you like?", and performs speech synthesis. Finally, this voice data is played from the device's speaker, and the device sends related information to a smartphone, allowing the parent to check detailed product information.

[0374] An example of a prompt to input into the generative AI model is, "What emotion is this child expressing? Please provide the emotion analysis result for 'I want to see the toys!'" This can enrich the user experience and support communication between parents and children.

[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0376] Step 1:

[0377] The device receives user voice input via a microphone. This voice data is converted into text data in real time using the Google Speech-to-Text API. The input is raw voice data, and the output is recognized text data.

[0378] Step 2:

[0379] The server receives text data sent from the terminal. Next, it processes this text data using a natural language understanding process with Google Dialogflow to analyze the user's intent and emotions. In this process, text data is the input, and the user's intent and emotions as the analysis result are output.

[0380] Step 3:

[0381] The server uses a conversation generation engine based on the analysis results to generate an appropriate response. This process takes the analysis results, including the user's intent and emotions, as input and outputs a corresponding text-based response.

[0382] Step 4:

[0383] The server converts the generated text response into audio data using Amazon Polly. The input is the text data of the response, and the output is audio data that can be played back through a speaker.

[0384] Step 5:

[0385] The terminal plays the audio data received from the server through its speaker and communicates the response to the user. In this case, the audio data is the input, and the played audio is the output.

[0386] Step 6:

[0387] The device's camera captures the user's image, and video analysis software is used to analyze that video data. The input for this step is the captured video, and the output is data that recognizes the user's emotions.

[0388] Step 7:

[0389] The user (parent) monitors the system's operation via a smartphone app and adjusts the conversation patterns as needed. Here, feedback data from the system is input, and the adjusted pattern information is output.

[0390] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0391] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0392] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0393] [Third Embodiment]

[0394] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0395] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0396] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0397] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0398] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0399] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0400] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0401] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0402] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0403] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0404] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0405] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0406] This invention provides a system that reduces the burden on parents in childcare and supports interaction with children by coordinating the functions of speech recognition, natural language processing, conversation generation, speech synthesis, and data output. This system mainly consists of a server, terminals, and users.

[0407] The server, acting as the system's central hub, is responsible for analyzing voice data and generating conversations. Upon receiving text data from a terminal, the server uses natural language understanding technology to analyze the child's intentions and emotions, and generates the most appropriate response. The generated response is then converted into voice data using speech synthesis technology and sent back to the terminal.

[0408] The device is installed in the home and serves as a direct interface with the child. Equipped with a microphone and speaker, it detects and captures the child's voice. The captured audio data is converted into text data using speech recognition technology and sent to a server. The server then plays the audio data back through the speaker, providing a response to the child. For example, if a child says, "Read me a picture book," the device recognizes this and can read the story aloud using audio generated from the server.

[0409] Users (parents) can manage system settings and feedback through devices such as smartphones. Parents can use the app to check their child's learning progress and conversation history, and customize specific conversation patterns as needed. For example, by entering the phrase "It's bath time," the system can be set to communicate this to the child at a specific time.

[0410] In this way, this system utilizes voice technology and artificial intelligence to support parents in childcare and promote interactive dialogue with their children.

[0411] The following describes the processing flow.

[0412] Step 1:

[0413] The device receives audio from the child using its built-in microphone to capture voice input. The received audio is processed in real time, and noise reduction is performed.

[0414] Step 2:

[0415] The device uses speech recognition software to convert the captured audio into text data. The converted text data, along with metadata about the child's intentions, is sent to the server.

[0416] Step 3:

[0417] The server receives text data sent from the terminal and uses natural language processing technology to analyze the intent and emotions behind the child's statements.

[0418] Step 4:

[0419] The server executes a conversation generation engine to generate the optimal response based on the analysis results. The generated response is then passed to the speech synthesis engine.

[0420] Step 5:

[0421] The server uses speech synthesis technology to convert the text responses generated by conversation generation into speech data. The synthesized speech data is then sent to the terminal.

[0422] Step 6:

[0423] The device plays audio data received from the server. The played audio is delivered to the child via a speaker, facilitating natural conversation.

[0424] Step 7:

[0425] Users can review conversations and system behavior through a smartphone app, providing feedback and customization as needed. For example, they can adjust specific phrases to determine when and how the system uses those phrases.

[0426] (Example 1)

[0427] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0428] Conventional voice dialogue systems often suffer from insufficient accuracy in speech recognition and generated responses that do not adequately reflect the user's intentions or emotions, resulting in a poor user experience. Furthermore, the system's response patterns are difficult to adjust flexibly, making it challenging to effectively utilize information data across different devices. Therefore, particularly in childcare support settings, there is a need for technology that enables satisfying dialogue between parents and children.

[0429] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0430] In this invention, the server includes speech recognition means, natural language processing means, and response generation means. This makes it possible to analyze the user's intentions and emotions with high accuracy and generate appropriate responses. Furthermore, it enables the sharing of information data between multiple devices, thereby improving the user experience.

[0431] "Speech recognition means" refers to a technology that receives speech as input and converts said speech into text data.

[0432] "Natural language processing means" refers to technologies that analyze converted character data to understand the user's intentions and emotions.

[0433] "Response generation means" refers to a technology that generates an appropriate response based on the judgment results of natural language processing means.

[0434] "Speech synthesis means" refers to a technology that converts a generated response into speech data, making it available for output as speech.

[0435] "Output means" refers to a device or technology that outputs the generated audio data in real time.

[0436] "Adjustment means" refers to a technology that dynamically adjusts response patterns based on information data obtained from the user.

[0437] A "data sharing method" is a technology that communicates data and shares information between multiple terminals or processing devices.

[0438] This system aims to enable natural communication through voice interaction in homes and educational settings. Specifically, it is a system that supports dialogue with children by integrating speech recognition, natural language processing, conversation generation, and speech synthesis technologies.

[0439] The server plays a central role in the system. The server receives text data sent from terminals and performs natural language processing using a generative AI model. Specifically, the generative AI model utilizes high-performance natural language understanding models (e.g., GPT and BERT). The server leverages these models to analyze the child's intentions and emotions and generate the optimal response. The generated response is then converted into speech data using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[0440] The device is installed in the home and functions as a direct interface with the child. It has a built-in microphone and speaker, and uses technologies such as Google Speech-to-Text API and IBM Watson to convert the child's voice into text data. This converted text data is then sent to a server via the internet. The device also plays back the received audio data through the speaker, providing direct responses to the child. For example, if a child asks, "Read this picture book," the device recognizes this request and can read the story aloud using audio generated from the server.

[0441] Users (parents) manage system settings and feedback via mobile devices such as smartphones and tablets. Through the application, users can check their child's learning progress and conversation history, and customize specific conversation patterns. For example, by setting the phrase "It's bath time," the device can be configured to notify the child at a specific time. Also, by entering text such as "Suggest a dinner my child will enjoy," the generative AI model can provide appropriate suggestions.

[0442] In this way, this system utilizes voice technology and artificial intelligence to reduce the burden of childcare on parents and support smooth communication with their children.

[0443] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0444] Step 1:

[0445] The device captures the child's voice using a microphone. The input is the child's voice, which is acquired as audio data. The device then applies noise reduction to this audio data to produce clear sound quality.

[0446] Step 2:

[0447] The device converts captured audio data into text data using speech recognition software. The input is the audio data processed in the previous step, and the output is text data. Specifically, the Google Speech-to-Text API is used to accurately convert the audio "Read the picture book" into text.

[0448] Step 3:

[0449] The terminal sends the converted text data to the server. The input is text data, and the output is the transmission of data to the server. Encryption technology is used to ensure the data is transmitted securely.

[0450] Step 4:

[0451] The server inputs the received text data into a generative AI model and performs natural language processing. The input is text data from the terminal, and the output is the analysis result. The generative AI model understands the child's intentions and emotions and derives a conclusion, for example, that "the child wants to hear a story."

[0452] Step 5:

[0453] The server generates a response based on the results of natural language processing. The input is the parsing result, and the output is the generated response text. For example, it generates the beginning of a story, "Once upon a time..."

[0454] Step 6:

[0455] The server converts the generated response text into speech data using a text-to-speech tool. The input is the response text, and the output is speech data. The Google Text-to-Speech API is used to generate fluent speech.

[0456] Step 7:

[0457] The server transmits synthesized audio data to the terminal. The input is audio data, and the output is data transmission to the terminal. Transmission is performed in real time to avoid delays.

[0458] Step 8:

[0459] The device plays the received audio data through its speaker. The input is audio data, and the output is audio played through the speaker. This allows the server to read aloud stories to children.

[0460] (Application Example 1)

[0461] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0462] In modern brick-and-mortar stores, it is a challenge for customers to effectively obtain product information and receive the necessary guidance. In particular, there is a growing need for efficient systems that can respond quickly and accurately to a variety of questions. Furthermore, there is a demand to reduce the workload on store staff and provide consistent service to customers.

[0463] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0464] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language understanding means for analyzing the text data and determining the intentions and emotions of the person speaking, and dialogue generation means for generating a response based on the determination. This makes it possible to provide information in response to questions from customers.

[0465] "Speech recognition means" refers to technology that receives speech input and converts that speech into text data.

[0466] "Natural language understanding methods" are technologies that analyze text data to determine the intentions and emotions of the person speaking.

[0467] A "dialogue generation method" is a technology for generating appropriate responses based on analyzed intentions and emotions.

[0468] "Speech synthesis means" refers to a technology that converts generated responses into speech data.

[0469] "Output means" refers to a system for providing audio data to the user.

[0470] "Information presentation means" refers to the function of providing appropriate information in response to questions from customers.

[0471] "Adjustment means" refers to techniques for optimizing dialogue patterns based on feedback data.

[0472] A "data sharing method" is a system that allows multiple information terminals to share learning data and improve the quality of dialogue.

[0473] To implement this invention, first, the server receives voice input from a customer using speech recognition means and converts it into text data. The server then uses natural language understanding means to analyze the text data and determine the customer's intentions and emotions. Next, using dialogue generation means, the server generates an optimal response based on the determined intentions. This generated response is converted into voice data by speech synthesis means and provided to the customer through output means.

[0474] The terminal is installed in the physical store and is equipped with a microphone and speaker. The terminal captures the voice of the customer using the microphone and sends the audio data to the server. It also plays back the audio data sent from the server through the speaker, providing a response to the customer.

[0475] Users (store staff and operators) can input feedback into the system and optimize dialogue patterns through adjustment mechanisms. Multiple terminals communicate to share learning data, and the accuracy of interactions with customers improves by utilizing data sharing mechanisms.

[0476] For example, if a customer asks, "Where are the cereals on sale?", the system can respond, "The sale cereals are in the food section on the second floor. Shall I show you the way?"

[0477] An example of a prompt to a generative AI model is: "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the best response."

[0478] Thus, the present invention provides a system that allows customers to easily obtain necessary information within a physical store, thereby reducing the burden on store staff and improving service.

[0479] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0480] Step 1:

[0481] The terminal captures the customer's voice input via the microphone. It receives the voice data as input. This is sent to speech recognition and converted into text data. The speech_recognition library is used for speech recognition. Text data is generated as output.

[0482] Step 2:

[0483] The server receives the generated text data and analyzes it using natural language understanding (NLP) tools. Specifically, it analyzes the text data to determine the intentions and emotions of customers. To achieve this, it utilizes a generative AI model and processes the input text. The output is the analysis result.

[0484] Step 3:

[0485] The server generates the optimal response using a dialogue generation mechanism based on the analysis results. The AI ​​model generates a prompt using the analysis results as input. For example, a prompt such as "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the optimal response." is generated. The response text is obtained as output.

[0486] Step 4:

[0487] The server converts the generated response text into speech data using a speech synthesis system. Specific software, such as the pyttsx3 library, is used for speech synthesis. Speech data is generated as output.

[0488] Step 5:

[0489] The terminal provides the generated audio data to the customer through the speaker. It receives audio data as input and plays it back through the physical speaker. Information is effectively provided by outputting an appropriate audio response to the customer.

[0490] By following these steps, it becomes possible to provide information in real time in response to customers' questions.

[0491] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0492] This invention is a system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, aiming to make interactions with users (especially children) more natural and effective. The system consists of a server, a terminal, and a user.

[0493] The server plays a central role in analyzing the voice data and performing natural language processing to determine the child's intentions and emotions. The server receives the speech-recognized text data and uses an emotion engine to recognize the emotions. Based on this result, the conversation generation engine determines the optimal response and generates the response. The generated response is synthesized into speech and sent to the terminal as voice data. For example, if a child excitedly says, "Let's play quickly!", the server recognizes that emotion and generates a response in a lively tone, "What should we play?"

[0494] The device is responsible for the initial process of capturing voice input and converting it to text. It receives voice from the child via a microphone and performs real-time speech recognition. This text data is sent to a server, and upon receiving a response voice data from the server, it plays it back to the child through a speaker. The device can also capture the user's facial expressions with a camera as supplementary information for emotion recognition.

[0495] The user (parent) has the role of managing interactions with the system through a smartphone application. The parent can review the child's speech and the emotional data recognized by the system, and adjust the conversation response patterns as needed. For example, if the system detects that the child is sounding sad, the parent can customize the system's response by adding phrases to offer words of encouragement.

[0496] In this way, this system utilizes voice and emotion data to enable interactive dialogue with children and reduce the burden of childcare.

[0497] The following describes the processing flow.

[0498] Step 1:

[0499] The device receives the child's voice via a microphone and converts it to text in real time using speech recognition software. This text data also includes attributes such as the timing and volume of the speech.

[0500] Step 2:

[0501] The device sends text data and speech attributes to the server. At the same time, it also sends facial expression data of the child captured by the camera, which improves the accuracy of emotion recognition.

[0502] Step 3:

[0503] The server uses the received text and facial expression data to run an emotion engine and analyze the child's emotions. The results of the emotion analysis are labeled as joy, sadness, anger, etc.

[0504] Step 4:

[0505] The server uses natural language processing technology to understand the user's intent from text data and runs a conversation generation engine that combines this with sentiment analysis results to generate the optimal response.

[0506] Step 5:

[0507] The server uses speech synthesis technology to convert the generated text response into speech and sends the synthesized speech data to the terminal.

[0508] Step 6:

[0509] The device plays audio data received from the server through its speaker and provides a response to the child. The response reflects the tone and tempo based on analyzed emotions.

[0510] Step 7:

[0511] Users can view the system's operation and conversation history through a smartphone app. Parents can provide feedback on response patterns based on recognized emotion data and incorporate it into future interactions. For example, they can change the settings so that the system uses gentle words when a child is feeling anxious.

[0512] (Example 2)

[0513] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0514] Current voice dialogue systems face the challenge of accurately recognizing user emotions and generating natural, appropriate responses based on those emotions. Interacting with children, in particular, requires flexible responses due to their rapidly changing emotional states. Furthermore, continuous learning through user feedback is necessary to improve response accuracy.

[0515] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0516] In this invention, the server includes speech recognition means for acquiring voice input and converting the voice into text information, natural language understanding means for analyzing the text information and determining the user's intentions and emotions, and means for creating a response corresponding to a prompt sentence using a generative model. This enables natural and effective dialogue that takes the user's emotions into consideration, as well as continuous response improvement based on feedback.

[0517] "Speech recognition means" refers to a technical element that receives speech input, analyzes the speech, and converts it into text information.

[0518] "Natural language understanding means" refers to technological elements that provide a process for analyzing and understanding the user's intentions and emotions based on text information.

[0519] A "conversation generation means" is a technological element that generates natural and appropriate responses while taking into account the user's intentions and emotions.

[0520] "Speech synthesis means" refers to a technical element that converts the generated response into speech information and achieves natural-sounding speech output.

[0521] "Output means" refers to a technical element for outputting the response generated as audio data and transmitting it to the user.

[0522] A "generative model" is a technology that uses algorithms learned from large amounts of data to generate responses and content in response to user input.

[0523] A "prompt statement" is text that provides instructions or context to elicit a desired response using a generative model.

[0524] This invention is a dialogue system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, and is particularly aimed at making dialogue with children more natural and effective. The system consists of a server, a terminal, and a parent who is the user.

[0525] The server receives the speech-recognized text information and uses natural language understanding to determine the user's intent and emotions. This analysis uses a natural language processing model (e.g., a large-scale language model). Based on the determination results, a conversation generation system generates a response, and a generation AI model is used to create a natural response corresponding to the prompt sentence. For example, if a child says, "I want to go to the park today," the server receives this and generates a response such as, "That sounds fun! What do you want to do at the park?" This response is converted into speech information by a speech synthesis system and sent to the terminal.

[0526] The device uses a microphone to acquire voice input and converts the speech to text in real time using speech recognition software (e.g., various APIs). The text information is sent to a server, which then receives the generated voice information and outputs it through the speaker. The device is also equipped with a camera, which can identify the user's facial expressions for facial recognition.

[0527] The user (parent) manages the system through a smartphone application. They can see what conversations the user has had through the user interface and adjust the dialogue response patterns and conversation generation methods as needed. For example, the parent can input a prompt such as "Generate a conversation suitable for when the child is interested in sports," and then create an appropriate conversation based on that prompt.

[0528] This system utilizes generative models to enable interactive and emotionally rich dialogue, thereby reducing the burden of childcare.

[0529] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0530] Step 1:

[0531] The device captures the child's voice in real time via a microphone. Speech recognition software is used to convert this audio data into text data. The input is the child's voice, and the output is text data. This conversion uses phonological analysis and pattern matching techniques to replace the audio signal with a string of characters.

[0532] Step 2:

[0533] The terminal sends the converted text data to the server. The server receives this text data and analyzes its intent and sentiment using natural language understanding tools. The input is text data, and the output is the result of the intent and sentiment analysis. This analysis applies natural language processing techniques and sentiment analysis algorithms to identify the user's intent and emotions.

[0534] Step 3:

[0535] The server utilizes a generative AI model to generate natural responses corresponding to prompt sentences based on the analysis results. The input consists of the analysis results of intent and sentiment, and the prompt sentence; the output is a text-based response. This employs predictive generation using a language model to create appropriate and contextually relevant sentences.

[0536] Step 4:

[0537] The server passes the generated text-based response to a speech synthesis system for speech synthesis. The input is the response text, and the output is speech data. This process uses phonological synthesis techniques and speech adjustment algorithms to generate speech with a natural tone.

[0538] Step 5:

[0539] The server sends the generated audio data to the terminal. The terminal receives this audio data and plays it for the child through the speaker. The input is the audio data, and the output is the sound from the speaker. This audio playback involves decoding and amplifying the audio data, ensuring that it is transmitted to the child in clear sound.

[0540] Step 6:

[0541] The user (parent) can view the system's dialogue logs and the child's emotional data through a smartphone application and adjust the conversation response patterns as needed. Input is the dialogue log and emotional data, while output is the updated response pattern or prompt text. This allows parents to adjust system settings and provide better childcare support.

[0542] (Application Example 2)

[0543] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0544] Conventional voice dialogue systems simply convert user speech into text and generate responses based on simple rules, which has resulted in a lack of natural conversation that fully understands the user's emotions and intentions. Furthermore, there was a lack of means to utilize video information to more accurately recognize user emotions and reflect them in responses. As a result, improvements in user experience and the provision of personalized services have not been fully realized.

[0545] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0546] In this invention, the server includes speech recognition means that receive voice input and convert the voice into text data, natural language understanding means that analyze the text data and determine the user's intentions and emotions, and video analysis means that capture the customer's video and recognize emotions using the video data. This makes it possible to more accurately recognize emotions based on the user's voice and video information and generate natural and effective responses.

[0547] "Voice recognition means" refers to a device or software that receives voice input using a microphone or sensor and converts that voice into text data.

[0548] "Natural language understanding means" refers to technologies that include algorithms and processes for analyzing text data and extracting and determining the user's intentions and emotions.

[0549] A "conversation generation means" is a system or technology for generating a situation-appropriate response based on information extracted by a natural language understanding means.

[0550] "Speech synthesis means" refers to a system or device that converts text information into speech signals in order to represent the generated response as speech data.

[0551] "Output means" refers to a device or interface for allowing the user to hear the generated audio data through a speaker or display.

[0552] "Video analysis means" refers to technologies and devices that analyze customer videos acquired using cameras and sensors and recognize emotions from them.

[0553] A "coordination mechanism" is a system that dynamically optimizes response conversation patterns and the content of services provided based on feedback information and environmental information.

[0554] "Information sharing methods" refer to protocols and functions that improve the accuracy of communication with users by sharing learning information and generated data among multiple devices.

[0555] This invention is a system that enables interactive dialogue with users based on audio and video data. Three main components play crucial roles in implementing this system: the server, the terminal, and the user.

[0556] The server is responsible for the central processing of this system. Upon receiving voice input, the server performs speech recognition and converts the voice into text data. For this process, the Google Speech-to-Text API can be used as the speech recognition software. Subsequently, Google Dialogflow is utilized as a natural language understanding tool to analyze the text data and determine the user's intent and emotions. Based on this determination, the conversation generation engine generates the optimal response and determines the content of the response to provide an emotionally appropriate service. The response is then converted into audio data using Amazon Polly as the speech synthesis software.

[0557] The terminal captures voice input from the user and transmits it to the server in real time. It also uses a camera to capture video of the customer and provides auxiliary data for emotion recognition based on that video. Image analysis software is used for video analysis from the camera to recognize emotions from the acquired video data.

[0558] Users operate the system through smartphones or other interfaces. In particular, parents (users) can manage the system, monitor interactions with their children, and adjust conversation patterns as needed. Feedback from parents is used to adjust response patterns, and the service content is optimized through the aforementioned adjustment methods.

[0559] For example, if the device receives an excited voice message from a child saying, "I want to see the toys!", the server analyzes the voice and generates a response such as, "What kind of toys do you like?", and performs speech synthesis. Finally, this voice data is played from the device's speaker, and the device sends related information to a smartphone, allowing the parent to check detailed product information.

[0560] An example of a prompt to input into the generative AI model is, "What emotion is this child expressing? Please provide the emotion analysis result for 'I want to see the toys!'" This can enrich the user experience and support communication between parents and children.

[0561] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0562] Step 1:

[0563] The device receives user voice input via a microphone. This voice data is converted into text data in real time using the Google Speech-to-Text API. The input is raw voice data, and the output is recognized text data.

[0564] Step 2:

[0565] The server receives text data sent from the terminal. Next, it processes this text data using a natural language understanding process with Google Dialogflow to analyze the user's intent and emotions. In this process, text data is the input, and the user's intent and emotions as the analysis result are output.

[0566] Step 3:

[0567] The server uses a conversation generation engine based on the analysis results to generate an appropriate response. This process takes the analysis results, including the user's intent and emotions, as input and outputs a corresponding text-based response.

[0568] Step 4:

[0569] The server converts the generated text response into audio data using Amazon Polly. The input is the text data of the response, and the output is audio data that can be played back through a speaker.

[0570] Step 5:

[0571] The terminal plays the audio data received from the server through its speaker and communicates the response to the user. In this case, the audio data is the input, and the played audio is the output.

[0572] Step 6:

[0573] The device's camera captures the user's image, and video analysis software is used to analyze that video data. The input for this step is the captured video, and the output is data that recognizes the user's emotions.

[0574] Step 7:

[0575] The user (parent) monitors the system's operation via a smartphone app and adjusts the conversation patterns as needed. Here, feedback data from the system is input, and the adjusted pattern information is output.

[0576] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0577] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0578] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0579] [Fourth Embodiment]

[0580] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0581] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0582] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0583] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0584] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0585] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0586] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0587] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0588] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0589] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0590] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0591] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0592] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0593] This invention provides a system that reduces the burden on parents in childcare and supports interaction with children by coordinating the functions of speech recognition, natural language processing, conversation generation, speech synthesis, and data output. This system mainly consists of a server, terminals, and users.

[0594] The server, acting as the system's central hub, is responsible for analyzing voice data and generating conversations. Upon receiving text data from a terminal, the server uses natural language understanding technology to analyze the child's intentions and emotions, and generates the most appropriate response. The generated response is then converted into voice data using speech synthesis technology and sent back to the terminal.

[0595] The device is installed in the home and serves as a direct interface with the child. Equipped with a microphone and speaker, it detects and captures the child's voice. The captured audio data is converted into text data using speech recognition technology and sent to a server. The server then plays the audio data back through the speaker, providing a response to the child. For example, if a child says, "Read me a picture book," the device recognizes this and can read the story aloud using audio generated from the server.

[0596] Users (parents) can manage system settings and feedback through devices such as smartphones. Parents can use the app to check their child's learning progress and conversation history, and customize specific conversation patterns as needed. For example, by entering the phrase "It's bath time," the system can be set to communicate this to the child at a specific time.

[0597] In this way, this system utilizes voice technology and artificial intelligence to support parents in childcare and promote interactive dialogue with their children.

[0598] The following describes the processing flow.

[0599] Step 1:

[0600] The device receives audio from the child using its built-in microphone to capture voice input. The received audio is processed in real time, and noise reduction is performed.

[0601] Step 2:

[0602] The device uses speech recognition software to convert the captured audio into text data. The converted text data, along with metadata about the child's intentions, is sent to the server.

[0603] Step 3:

[0604] The server receives text data sent from the terminal and uses natural language processing technology to analyze the intent and emotions behind the child's statements.

[0605] Step 4:

[0606] The server executes a conversation generation engine to generate the optimal response based on the analysis results. The generated response is then passed to the speech synthesis engine.

[0607] Step 5:

[0608] The server uses speech synthesis technology to convert the text responses generated by conversation generation into speech data. The synthesized speech data is then sent to the terminal.

[0609] Step 6:

[0610] The device plays audio data received from the server. The played audio is delivered to the child via a speaker, facilitating natural conversation.

[0611] Step 7:

[0612] Users can review conversations and system behavior through a smartphone app, providing feedback and customization as needed. For example, they can adjust specific phrases to determine when and how the system uses those phrases.

[0613] (Example 1)

[0614] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0615] Conventional voice dialogue systems often suffer from insufficient accuracy in speech recognition and generated responses that do not adequately reflect the user's intentions or emotions, resulting in a poor user experience. Furthermore, the system's response patterns are difficult to adjust flexibly, making it challenging to effectively utilize information data across different devices. Therefore, particularly in childcare support settings, there is a need for technology that enables satisfying dialogue between parents and children.

[0616] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0617] In this invention, the server includes speech recognition means, natural language processing means, and response generation means. This makes it possible to analyze the user's intentions and emotions with high accuracy and generate appropriate responses. Furthermore, it enables the sharing of information data between multiple devices, thereby improving the user experience.

[0618] "Speech recognition means" refers to a technology that receives speech as input and converts said speech into text data.

[0619] "Natural language processing means" refers to technologies that analyze converted character data to understand the user's intentions and emotions.

[0620] "Response generation means" refers to a technology that generates an appropriate response based on the judgment results of natural language processing means.

[0621] "Speech synthesis means" refers to a technology that converts a generated response into speech data, making it available for output as speech.

[0622] "Output means" refers to a device or technology that outputs the generated audio data in real time.

[0623] "Adjustment means" refers to a technology that dynamically adjusts response patterns based on information data obtained from the user.

[0624] A "data sharing method" is a technology that communicates data and shares information between multiple terminals or processing devices.

[0625] This system aims to enable natural communication through voice interaction in homes and educational settings. Specifically, it is a system that supports dialogue with children by integrating speech recognition, natural language processing, conversation generation, and speech synthesis technologies.

[0626] The server plays a central role in the system. The server receives text data sent from terminals and performs natural language processing using a generative AI model. Specifically, the generative AI model utilizes high-performance natural language understanding models (e.g., GPT and BERT). The server leverages these models to analyze the child's intentions and emotions and generate the optimal response. The generated response is then converted into speech data using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[0627] The device is installed in the home and functions as a direct interface with the child. It has a built-in microphone and speaker, and uses technologies such as Google Speech-to-Text API and IBM Watson to convert the child's voice into text data. This converted text data is then sent to a server via the internet. The device also plays back the received audio data through the speaker, providing direct responses to the child. For example, if a child asks, "Read this picture book," the device recognizes this request and can read the story aloud using audio generated from the server.

[0628] Users (parents) manage system settings and feedback via mobile devices such as smartphones and tablets. Through the application, users can check their child's learning progress and conversation history, and customize specific conversation patterns. For example, by setting the phrase "It's bath time," the device can be configured to notify the child at a specific time. Also, by entering text such as "Suggest a dinner my child will enjoy," the generative AI model can provide appropriate suggestions.

[0629] In this way, this system utilizes voice technology and artificial intelligence to reduce the burden of childcare on parents and support smooth communication with their children.

[0630] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0631] Step 1:

[0632] The device captures the child's voice using a microphone. The input is the child's voice, which is acquired as audio data. The device then applies noise reduction to this audio data to produce clear sound quality.

[0633] Step 2:

[0634] The device converts captured audio data into text data using speech recognition software. The input is the audio data processed in the previous step, and the output is text data. Specifically, the Google Speech-to-Text API is used to accurately convert the audio "Read the picture book" into text.

[0635] Step 3:

[0636] The terminal sends the converted text data to the server. The input is text data, and the output is the transmission of data to the server. Encryption technology is used to ensure the data is transmitted securely.

[0637] Step 4:

[0638] The server inputs the received text data into a generative AI model and performs natural language processing. The input is text data from the terminal, and the output is the analysis result. The generative AI model understands the child's intentions and emotions and derives a conclusion, for example, that "the child wants to hear a story."

[0639] Step 5:

[0640] The server generates a response based on the results of natural language processing. The input is the parsing result, and the output is the generated response text. For example, it generates the beginning of a story, "Once upon a time..."

[0641] Step 6:

[0642] The server converts the generated response text into speech data using a text-to-speech tool. The input is the response text, and the output is speech data. The Google Text-to-Speech API is used to generate fluent speech.

[0643] Step 7:

[0644] The server transmits synthesized audio data to the terminal. The input is audio data, and the output is data transmission to the terminal. Transmission is performed in real time to avoid delays.

[0645] Step 8:

[0646] The device plays the received audio data through its speaker. The input is audio data, and the output is audio played through the speaker. This allows the server to read aloud stories to children.

[0647] (Application Example 1)

[0648] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] In modern brick-and-mortar stores, it is a challenge for customers to effectively obtain product information and receive the necessary guidance. In particular, there is a growing need for efficient systems that can respond quickly and accurately to a variety of questions. Furthermore, there is a demand to reduce the workload on store staff and provide consistent service to customers.

[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0651] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language understanding means for analyzing the text data and determining the intentions and emotions of the person speaking, and dialogue generation means for generating a response based on the determination. This makes it possible to provide information in response to questions from customers.

[0652] "Speech recognition means" refers to technology that receives speech input and converts that speech into text data.

[0653] "Natural language understanding methods" are technologies that analyze text data to determine the intentions and emotions of the person speaking.

[0654] A "dialogue generation method" is a technology for generating appropriate responses based on analyzed intentions and emotions.

[0655] "Speech synthesis means" refers to a technology that converts generated responses into speech data.

[0656] "Output means" refers to a system for providing audio data to the user.

[0657] "Information presentation means" refers to the function of providing appropriate information in response to questions from customers.

[0658] "Adjustment means" refers to techniques for optimizing dialogue patterns based on feedback data.

[0659] A "data sharing method" is a system that allows multiple information terminals to share learning data and improve the quality of dialogue.

[0660] To implement this invention, first, the server receives voice input from a customer using speech recognition means and converts it into text data. The server then uses natural language understanding means to analyze the text data and determine the customer's intentions and emotions. Next, using dialogue generation means, the server generates an optimal response based on the determined intentions. This generated response is converted into voice data by speech synthesis means and provided to the customer through output means.

[0661] The terminal is installed in the physical store and is equipped with a microphone and speaker. The terminal captures the voice of the customer using the microphone and sends the audio data to the server. It also plays back the audio data sent from the server through the speaker, providing a response to the customer.

[0662] Users (store staff and operators) can input feedback into the system and optimize dialogue patterns through adjustment mechanisms. Multiple terminals communicate to share learning data, and the accuracy of interactions with customers improves by utilizing data sharing mechanisms.

[0663] For example, if a customer asks, "Where are the cereals on sale?", the system can respond, "The sale cereals are in the food section on the second floor. Shall I show you the way?"

[0664] An example of a prompt to a generative AI model is: "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the best response."

[0665] Thus, the present invention provides a system that allows customers to easily obtain necessary information within a physical store, thereby reducing the burden on store staff and improving service.

[0666] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0667] Step 1:

[0668] The terminal captures the customer's voice input via the microphone. It receives the voice data as input. This is sent to speech recognition and converted into text data. The speech_recognition library is used for speech recognition. Text data is generated as output.

[0669] Step 2:

[0670] The server receives the generated text data and analyzes it using natural language understanding (NLP) tools. Specifically, it analyzes the text data to determine the intentions and emotions of customers. To achieve this, it utilizes a generative AI model and processes the input text. The output is the analysis result.

[0671] Step 3:

[0672] The server generates the optimal response using a dialogue generation mechanism based on the analysis results. The AI ​​model generates a prompt using the analysis results as input. For example, a prompt such as "A customer asked, 'Where can I find the cereal that's on sale?' Please generate the optimal response." is generated. The response text is obtained as output.

[0673] Step 4:

[0674] The server converts the generated response text into speech data using a speech synthesis system. Specific software, such as the pyttsx3 library, is used for speech synthesis. Speech data is generated as output.

[0675] Step 5:

[0676] The terminal provides the generated audio data to the customer through the speaker. It receives audio data as input and plays it back through the physical speaker. Information is effectively provided by outputting an appropriate audio response to the customer.

[0677] By following these steps, it becomes possible to provide information in real time in response to customers' questions.

[0678] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0679] This invention is a system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, aiming to make interactions with users (especially children) more natural and effective. The system consists of a server, a terminal, and a user.

[0680] The server plays a central role in analyzing the voice data and performing natural language processing to determine the child's intentions and emotions. The server receives the speech-recognized text data and uses an emotion engine to recognize the emotions. Based on this result, the conversation generation engine determines the optimal response and generates the response. The generated response is synthesized into speech and sent to the terminal as voice data. For example, if a child excitedly says, "Let's play quickly!", the server recognizes that emotion and generates a response in a lively tone, "What should we play?"

[0681] The device is responsible for the initial process of capturing voice input and converting it to text. It receives voice from the child via a microphone and performs real-time speech recognition. This text data is sent to a server, and upon receiving a response voice data from the server, it plays it back to the child through a speaker. The device can also capture the user's facial expressions with a camera as supplementary information for emotion recognition.

[0682] The user (parent) has the role of managing interactions with the system through a smartphone application. The parent can review the child's speech and the emotional data recognized by the system, and adjust the conversation response patterns as needed. For example, if the system detects that the child is sounding sad, the parent can customize the system's response by adding phrases to offer words of encouragement.

[0683] In this way, this system utilizes voice and emotion data to enable interactive dialogue with children and reduce the burden of childcare.

[0684] The following describes the processing flow.

[0685] Step 1:

[0686] The device receives the child's voice via a microphone and converts it to text in real time using speech recognition software. This text data also includes attributes such as the timing and volume of the speech.

[0687] Step 2:

[0688] The device sends text data and speech attributes to the server. At the same time, it also sends facial expression data of the child captured by the camera, which improves the accuracy of emotion recognition.

[0689] Step 3:

[0690] The server uses the received text and facial expression data to run an emotion engine and analyze the child's emotions. The results of the emotion analysis are labeled as joy, sadness, anger, etc.

[0691] Step 4:

[0692] The server uses natural language processing technology to understand the user's intent from text data and runs a conversation generation engine that combines this with sentiment analysis results to generate the optimal response.

[0693] Step 5:

[0694] The server uses speech synthesis technology to convert the generated text response into speech and sends the synthesized speech data to the terminal.

[0695] Step 6:

[0696] The device plays audio data received from the server through its speaker and provides a response to the child. The response reflects the tone and tempo based on analyzed emotions.

[0697] Step 7:

[0698] Users can view the system's operation and conversation history through a smartphone app. Parents can provide feedback on response patterns based on recognized emotion data and incorporate it into future interactions. For example, they can change the settings so that the system uses gentle words when a child is feeling anxious.

[0699] (Example 2)

[0700] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0701] Current voice dialogue systems face the challenge of accurately recognizing user emotions and generating natural, appropriate responses based on those emotions. Interacting with children, in particular, requires flexible responses due to their rapidly changing emotional states. Furthermore, continuous learning through user feedback is necessary to improve response accuracy.

[0702] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0703] In this invention, the server includes speech recognition means for acquiring voice input and converting the voice into text information, natural language understanding means for analyzing the text information and determining the user's intentions and emotions, and means for creating a response corresponding to a prompt sentence using a generative model. This enables natural and effective dialogue that takes the user's emotions into consideration, as well as continuous response improvement based on feedback.

[0704] "Speech recognition means" refers to a technical element that receives speech input, analyzes the speech, and converts it into text information.

[0705] "Natural language understanding means" refers to technological elements that provide a process for analyzing and understanding the user's intentions and emotions based on text information.

[0706] A "conversation generation means" is a technological element that generates natural and appropriate responses while taking into account the user's intentions and emotions.

[0707] "Speech synthesis means" refers to a technical element that converts the generated response into speech information and achieves natural-sounding speech output.

[0708] "Output means" refers to a technical element for outputting the response generated as audio data and transmitting it to the user.

[0709] A "generative model" is a technology that uses algorithms learned from large amounts of data to generate responses and content in response to user input.

[0710] A "prompt statement" is text that provides instructions or context to elicit a desired response using a generative model.

[0711] This invention is a dialogue system that combines speech recognition, natural language processing, conversation generation, speech synthesis, and emotion recognition, and is particularly aimed at making dialogue with children more natural and effective. The system consists of a server, a terminal, and a parent who is the user.

[0712] The server receives the speech-recognized text information and uses natural language understanding to determine the user's intent and emotions. This analysis uses a natural language processing model (e.g., a large-scale language model). Based on the determination results, a conversation generation system generates a response, and a generation AI model is used to create a natural response corresponding to the prompt sentence. For example, if a child says, "I want to go to the park today," the server receives this and generates a response such as, "That sounds fun! What do you want to do at the park?" This response is converted into speech information by a speech synthesis system and sent to the terminal.

[0713] The device uses a microphone to acquire voice input and converts the speech to text in real time using speech recognition software (e.g., various APIs). The text information is sent to a server, which then receives the generated voice information and outputs it through the speaker. The device is also equipped with a camera, which can identify the user's facial expressions for facial recognition.

[0714] The user (parent) manages the system through a smartphone application. They can see what conversations the user has had through the user interface and adjust the dialogue response patterns and conversation generation methods as needed. For example, the parent can input a prompt such as "Generate a conversation suitable for when the child is interested in sports," and then create an appropriate conversation based on that prompt.

[0715] This system utilizes generative models to enable interactive and emotionally rich dialogue, thereby reducing the burden of childcare.

[0716] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0717] Step 1:

[0718] The device captures the child's voice in real time via a microphone. Speech recognition software is used to convert this audio data into text data. The input is the child's voice, and the output is text data. This conversion uses phonological analysis and pattern matching techniques to replace the audio signal with a string of characters.

[0719] Step 2:

[0720] The terminal sends the converted text data to the server. The server receives this text data and analyzes its intent and sentiment using natural language understanding tools. The input is text data, and the output is the result of the intent and sentiment analysis. This analysis applies natural language processing techniques and sentiment analysis algorithms to identify the user's intent and emotions.

[0721] Step 3:

[0722] The server utilizes a generative AI model to generate natural responses corresponding to prompt sentences based on the analysis results. The input consists of the analysis results of intent and sentiment, and the prompt sentence; the output is a text-based response. This employs predictive generation using a language model to create appropriate and contextually relevant sentences.

[0723] Step 4:

[0724] The server passes the generated text-based response to a speech synthesis system for speech synthesis. The input is the response text, and the output is speech data. This process uses phonological synthesis techniques and speech adjustment algorithms to generate speech with a natural tone.

[0725] Step 5:

[0726] The server sends the generated audio data to the terminal. The terminal receives this audio data and plays it for the child through the speaker. The input is the audio data, and the output is the sound from the speaker. This audio playback involves decoding and amplifying the audio data, ensuring that it is transmitted to the child in clear sound.

[0727] Step 6:

[0728] The user (parent) can view the system's dialogue logs and the child's emotional data through a smartphone application and adjust the conversation response patterns as needed. Input is the dialogue log and emotional data, while output is the updated response pattern or prompt text. This allows parents to adjust system settings and provide better childcare support.

[0729] (Application Example 2)

[0730] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0731] Conventional voice dialogue systems simply convert user speech into text and generate responses based on simple rules, which has resulted in a lack of natural conversation that fully understands the user's emotions and intentions. Furthermore, there was a lack of means to utilize video information to more accurately recognize user emotions and reflect them in responses. As a result, improvements in user experience and the provision of personalized services have not been fully realized.

[0732] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0733] In this invention, the server includes speech recognition means that receive voice input and convert the voice into text data, natural language understanding means that analyze the text data and determine the user's intentions and emotions, and video analysis means that capture the customer's video and recognize emotions using the video data. This makes it possible to more accurately recognize emotions based on the user's voice and video information and generate natural and effective responses.

[0734] "Voice recognition means" refers to a device or software that receives voice input using a microphone or sensor and converts that voice into text data.

[0735] "Natural language understanding means" refers to technologies that include algorithms and processes for analyzing text data and extracting and determining the user's intentions and emotions.

[0736] A "conversation generation means" is a system or technology for generating a situation-appropriate response based on information extracted by a natural language understanding means.

[0737] "Speech synthesis means" refers to a system or device that converts text information into speech signals in order to represent the generated response as speech data.

[0738] "Output means" refers to a device or interface for allowing the user to hear the generated audio data through a speaker or display.

[0739] "Video analysis means" refers to technologies and devices that analyze customer videos acquired using cameras and sensors and recognize emotions from them.

[0740] A "coordination mechanism" is a system that dynamically optimizes response conversation patterns and the content of services provided based on feedback information and environmental information.

[0741] "Information sharing methods" refer to protocols and functions that improve the accuracy of communication with users by sharing learning information and generated data among multiple devices.

[0742] This invention is a system that enables interactive dialogue with users based on audio and video data. Three main components play crucial roles in implementing this system: the server, the terminal, and the user.

[0743] The server is responsible for the central processing of this system. Upon receiving voice input, the server performs speech recognition and converts the voice into text data. For this process, the Google Speech-to-Text API can be used as the speech recognition software. Subsequently, Google Dialogflow is utilized as a natural language understanding tool to analyze the text data and determine the user's intent and emotions. Based on this determination, the conversation generation engine generates the optimal response and determines the content of the response to provide an emotionally appropriate service. The response is then converted into audio data using Amazon Polly as the speech synthesis software.

[0744] The terminal captures voice input from the user and transmits it to the server in real time. It also uses a camera to capture video of the customer and provides auxiliary data for emotion recognition based on that video. Image analysis software is used for video analysis from the camera to recognize emotions from the acquired video data.

[0745] Users operate the system through smartphones or other interfaces. In particular, parents (users) can manage the system, monitor interactions with their children, and adjust conversation patterns as needed. Feedback from parents is used to adjust response patterns, and the service content is optimized through the aforementioned adjustment methods.

[0746] For example, if the device receives an excited voice message from a child saying, "I want to see the toys!", the server analyzes the voice and generates a response such as, "What kind of toys do you like?", and performs speech synthesis. Finally, this voice data is played from the device's speaker, and the device sends related information to a smartphone, allowing the parent to check detailed product information.

[0747] An example of a prompt to input into the generative AI model is, "What emotion is this child expressing? Please provide the emotion analysis result for 'I want to see the toys!'" This can enrich the user experience and support communication between parents and children.

[0748] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0749] Step 1:

[0750] The device receives user voice input via a microphone. This voice data is converted into text data in real time using the Google Speech-to-Text API. The input is raw voice data, and the output is recognized text data.

[0751] Step 2:

[0752] The server receives text data sent from the terminal. Next, it processes this text data using a natural language understanding process with Google Dialogflow to analyze the user's intent and emotions. In this process, text data is the input, and the user's intent and emotions as the analysis result are output.

[0753] Step 3:

[0754] The server uses a conversation generation engine based on the analysis results to generate an appropriate response. This process takes the analysis results, including the user's intent and emotions, as input and outputs a corresponding text-based response.

[0755] Step 4:

[0756] The server converts the generated text response into audio data using Amazon Polly. The input is the text data of the response, and the output is audio data that can be played back through a speaker.

[0757] Step 5:

[0758] The terminal plays the audio data received from the server through its speaker and communicates the response to the user. In this case, the audio data is the input, and the played audio is the output.

[0759] Step 6:

[0760] The device's camera captures the user's image, and video analysis software is used to analyze that video data. The input for this step is the captured video, and the output is data that recognizes the user's emotions.

[0761] Step 7:

[0762] The user (parent) monitors the system's operation via a smartphone app and adjusts the conversation patterns as needed. Here, feedback data from the system is input, and the adjusted pattern information is output.

[0763] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0764] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0765] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0766] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0767] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0768] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0769] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0770] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0771] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0772] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0773] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0774] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0775] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0776] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0777] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0778] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0779] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0780] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0781] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0782] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0783] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0784] The following is further disclosed regarding the embodiments described above.

[0785] (Claim 1)

[0786] A speech recognition means that receives voice input and converts the voice into text data,

[0787] A natural language understanding means that analyzes the text data and determines the user's intent and emotions,

[0788] A conversation generation means that generates a response based on the said determination,

[0789] A speech synthesis means that converts the response into speech data,

[0790] An output means for outputting the audio data,

[0791] A system that includes this.

[0792] (Claim 2)

[0793] The system according to claim 1, wherein the conversation generation means further includes an adjustment means for adjusting the conversation pattern based on feedback data obtained from the user.

[0794] (Claim 3)

[0795] The system according to claim 1, further comprising data sharing means for multiple terminals to communicate and share learning data to improve the accuracy of communication with the user.

[0796] "Example 1"

[0797] (Claim 1)

[0798] A speech recognition means that detects sound and converts the sound into text data,

[0799] A natural language processing means that analyzes the character data and determines the user's intent and emotions,

[0800] A response generation means that generates a response based on the said determination,

[0801] A speech synthesis means that converts the response into speech data,

[0802] An output means that communicates the audio data to multiple devices and outputs the generated responses,

[0803] A system that includes this.

[0804] (Claim 2)

[0805] The system according to claim 1, wherein the response generation means further includes an adjustment means for adjusting the response pattern based on information data obtained from the user.

[0806] (Claim 3)

[0807] The system according to claim 1, further comprising data sharing means for multiple processing units to communicate data and share information data to improve the accuracy of interaction with the user.

[0808] "Application Example 1"

[0809] (Claim 1)

[0810] A speech recognition means that receives voice input and converts the voice into text data,

[0811] A natural language understanding means that analyzes the text data and determines the intentions and emotions of the interlocutor,

[0812] A dialogue generation means that generates a response based on the judgment,

[0813] A speech synthesis means that converts the response into speech data,

[0814] An output means for outputting the audio data,

[0815] A means of providing information in response to questions from customers,

[0816] A system that includes this.

[0817] (Claim 2)

[0818] The system according to claim 1, wherein the dialogue generation means further includes an adjustment means for adjusting the dialogue pattern based on feedback data obtained from a user.

[0819] (Claim 3)

[0820] The system according to claim 1, further comprising a data sharing means for improving the quality of interaction with users by enabling multiple information terminals to communicate and share learning data.

[0821] "Example 2 of combining an emotion engine"

[0822] (Claim 1)

[0823] A speech recognition means that acquires voice input and converts the voice into text information,

[0824] A natural language understanding means that analyzes the text information and determines the user's intent and emotions,

[0825] A conversation generation means that generates a response based on the determination,

[0826] A speech synthesis means that converts the response into speech information,

[0827] An output means for outputting the audio information,

[0828] A means of creating a response corresponding to a prompt using a generative model,

[0829] A system that includes this.

[0830] (Claim 2)

[0831] The system according to claim 1, wherein the conversation generation means further includes an adjustment means for adjusting the conversation format based on feedback information collected from the user.

[0832] (Claim 3)

[0833] The system according to claim 1, further comprising information sharing means for multiple terminals to communicate and share learning information to improve the accuracy of communication with the user.

[0834] "Application example 2 when combining with an emotional engine"

[0835] (Claim 1)

[0836] A speech recognition means that receives voice input and converts the voice into text data,

[0837] A natural language understanding means that analyzes the text data and determines the user's intent and emotions,

[0838] A conversation generation means that generates a response based on the judgment and provides a service that responds to emotions,

[0839] A speech synthesis means that converts the response into speech data,

[0840] An output means for outputting the audio data,

[0841] A video analysis means that captures images of customers and uses the video data to recognize their emotions,

[0842] A system that includes this.

[0843] (Claim 2)

[0844] The system according to claim 1, wherein the conversation generation means further includes an adjustment means that adjusts the conversation pattern and optimizes the content of the services to be provided based on feedback information obtained from the user.

[0845] (Claim 3)

[0846] The system according to claim 1, further comprising information sharing means for multiple terminals to communicate and share learning information to improve the accuracy of interactions with the user. [Explanation of symbols]

[0847] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A speech recognition means that receives voice input and converts the voice into text data, A natural language understanding means that analyzes the text data and determines the user's intent and emotions, A conversation generation means that generates a response based on the said determination, A speech synthesis means that converts the response into speech data, An output means for outputting the audio data, A system that includes this.

2. The system according to claim 1, wherein the conversation generation means further includes an adjustment means for adjusting the conversation pattern based on feedback data obtained from the user.

3. The system according to claim 1, further comprising data sharing means for multiple terminals to communicate and share learning data to improve the accuracy of communication with the user.