system

The system addresses the limitations of conventional AI systems by integrating voice input and natural language processing in an earphone-type device, facilitating natural interactions and efficient information retrieval.

JP2026068466APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional AI systems require users to operate specific interfaces and lack efficient data processing for natural conversations, limiting user convenience and information acquisition efficiency.

Method used

A system equipped with voice input, speech recognition, natural language processing, and text-to-speech technologies implemented in an earphone-type device for natural interactions, enabling seamless communication with AI systems.

Benefits of technology

Enables users to interact naturally with AI systems, improving information acquisition efficiency and user experience by allowing hands-free operation and personalized responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068466000001_ABST
    Figure 2026068466000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Audio input means for receiving audio signals, A speech recognition means that converts a received audio signal into text data, A communication means for transmitting the converted text data to a data processing device, A data processing device includes an analysis means that analyzes text data and generates corresponding response data, A text-to-speech conversion means that converts the generated response data into audio data, A means of providing audio data to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, due to the progress of artificial intelligence technology, the utilization of AI in daily life has been emphasized. However, in conventional AI systems, users need to operate a specific interface, and the convenience is limited. Also, in technologies using voice input, efficient data processing for realizing natural conversations is required. The purpose of the present invention is to enable users to interact with an AI system more naturally and to improve the efficiency of information acquisition and the user experience.

Means for Solving the Problems

[0005] This invention provides a system equipped with a voice input means that receives an audio signal and converts it into text using speech recognition technology. The converted text data is transmitted to a data processing device via a communication means and analyzed using natural language processing technology. The response generated by the analysis is converted into speech using text-to-speech technology and finally provided to the user by a voice output means. In particular, this system is designed to be implemented in an earphone-type device so that users can have natural conversations on a daily basis. Furthermore, the data processing device can cooperate with external information sources and can accurately respond to user requests by acquiring additional information as needed.

[0006] "Voice input means" refers to a device that has the function of receiving voice signals from a user.

[0007] "Speech recognition means" refers to a technology or system that analyzes received speech signals and converts them into text data.

[0008] "Communication means" refers to an interface or protocol for sending and receiving data to and from other devices or systems.

[0009] "Analysis means" refers to a device or software that processes received text data, understands the user's intent, and generates appropriate response data.

[0010] "Text-to-speech conversion means" refers to a technology or system that converts data in text format into data in audio format.

[0011] "Audio output means" refers to a device or function for providing generated audio data to the user in a format that can be heard.

[0012] An "earphone-type device" is a device worn in the ear that has the function of inputting and outputting audio.

[0013] A "data processing device" is a computer or system that analyzes received data, performs processing according to user requests, and generates results.

[0014] An "external information source" is a database or service located outside the system that is accessed to provide additional information. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The present invention provides a voice platform system that enables natural voice interaction with AI using an earphone-type device that users can use on a daily basis. This system receives voice, converts it into text data using speech recognition technology, generates an appropriate response using a generative AI, and then converts that response back into voice for the user.

[0037] Program processing

[0038] 1. The user puts on an earphone-type device and says, "Tell me today's news."

[0039] 2. The device receives the user's voice and generates the text "Tell me today's news" via speech recognition.

[0040] 3. The terminal sends this text to the server. The server analyzes this text using an analysis tool.

[0041] 4. Based on the analysis results, the server accesses external information sources to obtain news information. For example, it might use a news API to retrieve the latest news articles.

[0042] 5. Based on the information it has retrieved, the server generates a response that says, "Today's top news is about the recovery of the domestic economy."

[0043] 6. The server converts this response into audio data using a text-to-speech conversion method.

[0044] 7. The device receives audio data sent from the server and plays it back to the user through the earphones. The user hears the audio message, "Today's top news is about the recovery of the domestic economy."

[0045] This system allows users to obtain necessary information using only their voice, without using their hands, lowering the barrier to using AI in daily life. Furthermore, by linking with external information sources, it is possible to provide a wide range of information and respond to diverse user requests.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user voice-inputs "Tell me today's news" into the earphone-type device.

[0049] Step 2:

[0050] The device receives the user's voice signal through its built-in microphone.

[0051] Step 3:

[0052] The device activates its voice recognition system and converts the received voice signal into text data that reads, "Tell me today's news."

[0053] Step 4:

[0054] The terminal sends the converted text data to the server using a communication method.

[0055] Step 5:

[0056] The server analyzes the received text data using parsing tools to understand the user's request.

[0057] Step 6:

[0058] Based on the analysis results, the server sends a request to an external news API to retrieve the latest news information.

[0059] Step 7:

[0060] The server retrieves the latest news from the news API and generates response data that says, "Today's top news is about the recovery of the domestic economy."

[0061] Step 8:

[0062] The server converts the response data into audio data using a text-to-speech conversion method.

[0063] Step 9:

[0064] The server sends the generated audio data to the terminal.

[0065] Step 10:

[0066] The device plays the received audio data through the earphones and tells the user, "Today's top news is about the recovery of the domestic economy."

[0067] (Example 1)

[0068] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0069] Conventional voice dialogue systems offer the convenience of users obtaining information without using their hands, but they have limited capabilities in generating natural conversations and providing smooth access to external information sources. Therefore, a challenge is their inability to adequately respond to diverse user needs. In particular, there is a need for advanced conversation generation using generative artificial intelligence and real-time information retrieval that responds to user intent.

[0070] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0071] In this invention, the server includes a voice input means for receiving voice signals, a voice recognition means for converting the received voice signals into text data, and an analysis means for analyzing the text data and generating a response using generative artificial intelligence. This allows the user to obtain high-quality responses based on the latest external information through smooth and natural dialogue without using their hands.

[0072] "Voice input means" refers to a device or technology for receiving a voice signal and passing it on to the next processing stage.

[0073] "Speech recognition means" refers to a technology that analyzes received audio signals and converts them into corresponding text data.

[0074] "Communication means" refers to means for transmitting processed data to another device or system.

[0075] An "information processing device" is a computer or system used to analyze received data and generate the necessary response.

[0076] "Analysis means" refers to a method or apparatus that uses generative artificial intelligence to analyze received text data and generate appropriate response data.

[0077] "Generative artificial intelligence" is a technology for generating natural language responses based on received data, and it is an artificial intelligence that uses models to generate and understand text.

[0078] A "text-to-speech conversion method" is a technology that converts generated text data into an audio signal and outputs it to the user.

[0079] "Audio output means" refers to a device or technology for transmitting converted audio data to the user.

[0080] "External information acquisition means" refers to the means of accessing external information sources and obtaining necessary information when processing data.

[0081] A "holding device" is a device that can be used by a user by attaching it to a part of their body.

[0082] A "prompt sentence" is an input sentence used to give instructions or questions to a generative artificial intelligence.

[0083] The voice dialogue system of the present invention enables daily interaction with AI using a wearable, retainable device. This system mainly consists of voice input means, voice recognition means, communication means, information processing device, analysis means, generative artificial intelligence, text-to-speech conversion means, voice output means, and external information acquisition means.

[0084] The user wears a hold-type device and gives instructions in natural language, such as "Tell me today's news." This voice is received by a microphone built into the device and captured by the device's voice input mechanism. The received voice is then converted into text data using speech recognition technology. This process utilizes a common speech recognition API; for example, a cloud-based speech recognition service may be used.

[0085] The terminal sends this converted text data to the server via secure communication. The server receives the data through an information processing device and uses analysis tools to understand the user's intent. Here, generative artificial intelligence is utilized to create a prompt sentence and input it into the AI ​​model. A possible prompt sentence would be, "The user wants the latest news. Please look up the news."

[0086] The server accesses external information sources, such as news APIs, using external information acquisition methods to obtain the necessary information. The acquired data is analyzed by generative artificial intelligence and output as a natural language response. This response will have specific content, such as, "Today's top news is about the recovery of the domestic economy."

[0087] Subsequently, the server uses text-to-speech conversion to convert the generated text response into audio data. A commercial text-to-speech service is used for speech synthesis. This audio data is transmitted to the terminal and played back to the user via a holding device. This allows the user to obtain the latest information in audio format without using their hands.

[0088] This system provides users with a natural and intuitive interface, and helps AI provide useful information for everyday life.

[0089] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0090] Step 1:

[0091] The user wears a holding device and says, "Tell me today's news." The voice signal, as input, is received by the microphone inside the device. Here, the microphone converts the voice into a digital signal. The output is the digitized voice data.

[0092] Step 2:

[0093] The device uses speech recognition to convert digitized speech data into text data. Specifically, speech recognition software analyzes the speech waveform and identifies phonemes. This process outputs the text data "Tell me today's news."

[0094] Step 3:

[0095] The terminal sends the converted text data to the server. Here, the data is packetized using a communication method and securely transmitted over the internet. The input is text data, and the output is text data received by the server.

[0096] Step 4:

[0097] The server analyzes the received text data in the information processing device using an analysis mechanism. A prompt sentence is created using a generative AI model. Based on this prompt sentence, the server understands the user's intent and generates the instruction "The user wants news." The input is the received text, and the output is the analyzed result and the prompt sentence.

[0098] Step 5:

[0099] The server accesses a news API using an external information retrieval method to obtain the latest news data. Specifically, it sends a query to the API and receives news data in JSON format. The input is a news request, and the output is the retrieved news information.

[0100] Step 6:

[0101] The server uses generative artificial intelligence to generate a response to the user based on the acquired news information. In this generation process, the AI ​​model understands the context and creates a response sentence such as, "Today's top news is about the recovery of the domestic economy." The input is the news information and the prompt sentence, and the output is the response text.

[0102] Step 7:

[0103] The server uses text-to-speech conversion to convert the generated response text into speech data. Speech synthesis technology analyzes the text and generates a corresponding speech waveform. The input is the response text, and the output is the speech data.

[0104] Step 8:

[0105] The terminal receives audio data transmitted from the server and provides it to the user through a holding device. The earphones convert the digital audio into an analog signal and play it back to the user. The input is audio data, and the output is the audio the user hears. This entire process allows the user to obtain necessary information by voice without using their hands.

[0106] (Application Example 1)

[0107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0108] In modern society, there is a growing demand for systems that allow users to easily complete food orders using only their voice, without directly operating a smartphone or computer. However, existing applications using voice recognition technology struggle to smoothly handle complex ordering processes, and improvements in usability are needed.

[0109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0110] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a communication means for transmitting the converted text data to a data processing device. This allows users to easily order food without using their hands by using an earphone-type device.

[0111] "Voice input means" refers to a device or apparatus for receiving voice signals.

[0112] "Speech recognition means" refers to a technology or process that converts received speech signals into corresponding text data.

[0113] "Communication means" refers to means of transmitting voice signals or data to other devices or networks.

[0114] "Analysis means" refers to a means for analyzing text data and generating corresponding response data.

[0115] "Text-to-speech conversion means" refers to a technology or device for converting generated text data into speech data.

[0116] "Audio output means" refers to a device or means for providing generated audio data to a user.

[0117] A "food ordering support system" is a function that assists the food ordering process based on voice commands and generates instructions corresponding to each step of the order.

[0118] A system for carrying out the present invention enables a user to send voice instructions using an earphone-type device and order food through an external food delivery service. This system includes voice input means, voice recognition means, communication means, analysis means, text-to-speech conversion means, voice output means, and food ordering support means.

[0119] The user first puts on an earphone-type device and speaks a voice command, such as "I want to order a pizza," into the device. This voice is received by the terminal via a voice input device. The terminal uses speech recognition software, such as Google® Speech-to-Text API, to convert this voice signal into text data. This text data is then sent to a data processing device on a smartphone or in the cloud.

[0120] The data processing device analyzes received text data using generative AI models such as the OpenAI® GPT model. This analysis generates the information and questions necessary to proceed with the ordering process. For example, it generates appropriate responses in the form of prompts, such as prompting the user to respond with "What kind of pizza would you like?"

[0121] The generated text response is then converted into audio data using text-to-speech software such as Amazon Polly. Finally, this audio data is sent through the terminal to an earphone-type device and provided to the user as audio.

[0122] As a concrete example of the system's operation, if a user says "I want to order sushi," the system will generate a question such as "What kind of sushi would you like?" and wait for further instructions from the user. An example of a prompt message would be, "The user said 'I want to order sushi.' Please generate the next appropriate question."

[0123] This system allows users to complete food orders using only their voice, providing a convenient way for users to easily and quickly utilize food delivery services.

[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0125] Step 1:

[0126] The user wears an earphone-type device and verbally instructs, "I want to order a pizza." The input includes the user's voice. The earphone receives this voice signal and transmits it to the device.

[0127] Step 2:

[0128] The device converts the received audio signal into text data using speech recognition. In this step, the Google Speech-to-Text API analyzes the audio data and outputs it in text format. The output is the text data "I want to order a pizza."

[0129] Step 3:

[0130] The terminal sends text data to the server via a communication method. The input is the converted text data, and the output is the text data received by the server.

[0131] Step 4:

[0132] The server analyzes the received text data using a generation AI model. The input is text data, and it uses prompts to generate response data such as "What kind of pizza would you like?". This process determines the next step in the ordering process. The output is the generated response data.

[0133] Step 5:

[0134] The server converts the generated response data into audio data using a text-to-speech conversion method. In this step, Amazon Polly converts the text data into an audio file. The output is audio data.

[0135] Step 6:

[0136] The device receives audio data sent from the server and plays it back to the user as audio through an earphone-type device. The user can hear the question, "What kind of pizza would you like?" through the earphone.

[0137] As a result, users can proceed with their orders using only their voice. This series of steps completes the process from voice input to final voice output.

[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0139] The voice platform system incorporating the emotion engine of the present invention makes user interaction more personal and effective. This system aims to generate more appropriate responses by receiving the user's voice and analyzing the emotions from the voice using the emotion engine. Specific embodiments are described below.

[0140] In this system, the user speaks "I've been feeling stressed lately" through an earphone-type device. The device receives this audio and uses speech recognition to convert it into text, "I've been feeling stressed lately." This text data and audio data are analyzed by an emotion engine to detect that the user is experiencing stress.

[0141] The device then sends text and emotion data to the server. The server analyzes the received data and uses it as a basis for generating responses that take the user's emotions into account. For example, if the user is feeling stressed, the server might consider relaxation-related content helpful and obtain information on relaxation methods from external sources.

[0142] In addition, the server uses text-to-speech technology to generate audio data, which responds with the message, "To alleviate recent stress, we recommend practicing deep breathing." This audio data is transmitted to the earphones via the terminal for the user to listen to.

[0143] In this way, by providing information that takes into account the user's emotional state, users can more easily obtain effective information that is tailored to their situation. This system aims to improve the user experience using AI by understanding and responding to the user's emotions in daily life.

[0144] The following describes the processing flow.

[0145] Step 1:

[0146] The user voice-inputs "I've been feeling stressed lately" into an earphone-type device.

[0147] Step 2:

[0148] The device uses a microphone to receive audio signals.

[0149] Step 3:

[0150] The device processes the audio signal using speech recognition technology and converts it into text data that says, "I've been under a lot of stress lately."

[0151] Step 4:

[0152] The device inputs the converted text data and the original audio data into the emotion engine.

[0153] Step 5:

[0154] The emotion engine analyzes the user's emotions from voice data and detects when they are feeling stressed.

[0155] Step 6:

[0156] The device sends text data and sentiment analysis data to the server.

[0157] Step 7:

[0158] The server analyzes the received data and determines how to generate a response based on the user's emotional state.

[0159] Step 8:

[0160] The server queries external sources to retrieve recommended content related to relaxation.

[0161] Step 9:

[0162] Based on the information it retrieves, the server generates a text response that reads, "To alleviate recent stress, we recommend practicing deep breathing."

[0163] Step 10:

[0164] The server converts the text response into audio data using a text-to-speech conversion method.

[0165] Step 11:

[0166] The server sends the generated audio data to the terminal.

[0167] Step 12:

[0168] The device plays the received audio data through the earphones and tells the user in voice, "To relieve recent stress, we recommend practicing deep breathing."

[0169] (Example 2)

[0170] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0171] Modern users expect voice assistants to understand their emotional state and respond appropriately during interactions. However, existing voice platforms lack the ability to adequately analyze user emotions and generate appropriate responses based on them. As a result, personalized experiences that users desire are less likely to be provided, potentially leading to lower user satisfaction.

[0172] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0173] In this invention, the server includes an audio input means for receiving audio signals, an emotion analysis means for analyzing audio signals and emotion data, and a means for analyzing and generating responses using a generative AI model. This enables personalized responses that correspond to the user's emotional state.

[0174] A "voice input means" is a device that has the function of receiving a user's speech and converting it into an audio signal that can be processed within the system.

[0175] A "speech recognition means" is a device that has the function of analyzing a received speech signal and converting it into corresponding text data.

[0176] An "emotion analysis device" is a device that analyzes a user's emotional state based on voice data and text data and extracts emotional data.

[0177] A "communication device" is a device that has the function of transmitting text data and emotional data to a data processing device.

[0178] A "generative AI model" is an artificial intelligence model that executes algorithms to generate responses that are appropriate to the user's intentions and emotions based on input data.

[0179] "Analysis means" refers to a device that has the function of generating response data based on the text data and sentiment data received in a data processing device.

[0180] A "text-to-speech conversion device" is a device that has the function of converting generated text data into speech data.

[0181] "Audio output means" refers to a device that has the function of providing generated audio data in a format that can be listened to by the user.

[0182] This invention is a voice platform system that analyzes the user's emotions based on voice input and provides individually personalized responses. The following describes embodiments of the system in detail.

[0183] The user speaks naturally to the system through an earphone-type device. This earphone-type device functions as a voice input means and receives the voice signal as digital data. The terminal then converts the received voice data into text data using a voice recognition means, such as commercially available voice recognition software.

[0184] The converted text and audio data are input by the terminal into an emotion analysis device. Natural language processing technology is used for this analysis to extract the emotions contained in the audio data. Existing emotion analysis software can be used as the emotion analysis device.

[0185] The terminal transmits the obtained text and sentiment data to the server using a communication method. The server uses a processing unit to run a generative AI model based on the received data and generates a response that is appropriate to the user's context and emotions.

[0186] Furthermore, the server can access external information sources and retrieve additional information tailored to the user's emotional state as needed. For example, if it detects stress, it can retrieve information from online resources on relaxation techniques.

[0187] The conversion from generated text to speech is performed using the server's text-to-speech conversion capabilities. This process sends the audio data back to the terminal in a usable format, and the user ultimately receives the response through their earphones.

[0188] For example, if a user says, "I want to do something fun this weekend, but I'm not sure what to do," the system can analyze this statement and provide information about hobbies and local events. This response will be as personalized as possible, taking into account the user's emotional state.

[0189] An example of a prompt is, "What suggestions would be best when a user is looking for activities to enjoy on their day off?" The system is expected to use a generative AI model to provide appropriate information suggestions in response to this prompt.

[0190] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0191] Step 1:

[0192] The user speaks into an earphone-type device. This voice input is received as an audio signal by the microphone built into the earphone. In this step, the input is the user's raw voice, and the output is a digital audio signal.

[0193] Step 2:

[0194] The device converts the received audio signal into text data using your speech recognition software. In this process, the digital audio signal (input) is analyzed, and a language model generates text data (output) corresponding to words and phrases.

[0195] Step 3:

[0196] The device extracts the user's emotions from the audio data using emotion analysis methods, simultaneously with the analyzed text data. At this stage, the input consists of text data and audio signals, which are analyzed and output in the form of emotion data. Emotion analysis includes tone analysis of the voice and extraction of emotional words from the text.

[0197] Step 4:

[0198] The terminal sends the converted text data and sentiment data to the server via a communication method. The input for this step is the text data and sentiment data, and the output is the transmitted data packets that reach the server.

[0199] Step 5:

[0200] The server uses the received data to run a generative AI model that generates a response appropriate to the user's input. The input consists of text data and sentiment data, and by analyzing these, it generates response data (output) that includes content. In particular, the generative AI model is trained to produce appropriate responses to prompt sentences.

[0201] Step 6:

[0202] The server uses a text-to-speech conversion mechanism to convert the generated response data into speech data. The input to this process is response data in text format, and the output is data in speech format.

[0203] Step 7:

[0204] The generated audio data is transmitted to the earphones via the terminal and ultimately provided to the user. The input in this step is data in audio format, and the output is the audio the user hears. Through this audio, the user can receive the suggested information.

[0205] (Application Example 2)

[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0207] In the modern real world, consumers demand personalized services and product recommendations tailored to their individual emotional states. However, conventional technologies struggle to accurately grasp consumer emotions in real time and propose appropriate products and services based on that information. Furthermore, standardized approaches often fail to provide adequate interaction with consumers, resulting in insufficient satisfaction.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0209] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a response generation means for performing sentiment analysis and proposing products and services based on the consumer's emotions. This enables real-time understanding of the consumer's emotional state and effective product recommendations based on that understanding.

[0210] An "audio signal" is the physical representation of sound transmitted through vibrations in the air.

[0211] "Voice input means" refers to a device or method for acquiring a voice signal and inputting it into a system.

[0212] "Speech recognition means" refers to the technology and process for converting received speech signals into text data.

[0213] "Communication means" refers to a means of sending and receiving data between different devices.

[0214] A "data processing device" is a system that analyzes received data and generates or processes information.

[0215] "Analysis means" refers to a method or technique for analyzing input data and generating necessary information.

[0216] "Text-to-speech conversion means" refers to a device or software that has the function of converting text data into speech output.

[0217] A "sound output device" is a device that delivers generated sound data to the user as sound.

[0218] A "user" is a person or entity that utilizes the system's voice input / output functions.

[0219] A "response generation means" is a means of generating appropriate product suggestions and services based on user sentiment data.

[0220] An "external information source" is a source that provides data or information that exists outside the system.

[0221] In the system for carrying out this invention, the user first provides information using a voice input device. The voice input device digitizes the received voice signal and converts it into text data using speech recognition software (e.g., Google Speech-to-Text). The converted text data is transmitted to a server via a communication means.

[0222] The server first receives data and then uses sentiment analysis APIs (such as IBM Watson® or Microsoft® Azure® Cognitive Services) to analyze the user's emotional state. This sentiment analysis makes it possible to determine what the user wants and what their emotional state is.

[0223] Based on the analysis results, the server uses a response generation mechanism to generate product and service suggestions that correspond to the user's emotions. These suggestions are then converted back into audio data using a text-to-speech tool (such as Google Cloud Text-to-Speech) and provided to the user. The user can listen to these suggestions through the audio output mechanism.

[0224] A concrete example of this system would be a department store's in-store guidance robot that analyzes the user's stress level and suggests purchasing appropriate aromatherapy products for relaxation. An example of a prompt message might be, "Generate a message suggesting relaxation products for when the user is feeling stressed."

[0225] This invention aims to enhance real-world interactions via smartphones and robots, thereby improving consumer satisfaction.

[0226] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0227] Step 1:

[0228] The user provides information using a voice input device. The voice input device converts the analog voice signal into a digital signal and outputs it as data for input to voice recognition software.

[0229] Step 2:

[0230] The terminal passes the converted digital audio signal to speech recognition software, which then converts it into text data. In this process, the system analyzes the audio signal and outputs text that is recognized as words.

[0231] Step 3:

[0232] The terminal sends the generated text data to the server via a communication method. In this step, the data is wrapped to ensure that the text data reaches the server securely over the network.

[0233] Step 4:

[0234] The server passes the received text data to a sentiment analysis API to identify the user's emotional state. The API analyzes the text, extracts emotional information, and outputs a detailed report on the emotions the user is experiencing.

[0235] Step 5:

[0236] The server passes the analyzed sentiment information to the response generation system, which then generates appropriate product and service suggestions. This response generation process creates a prompt statement based on the sentiment data, which is then output.

[0237] Step 6:

[0238] The server passes the generated suggestions to a text-to-speech tool, which converts them into audio data. In this process, the suggested text is output as natural-sounding speech.

[0239] Step 7:

[0240] The server delivers the generated audio data to the user through an audio output device. The audio output device emits the audio in a format that makes it easy for the user to hear the suggestions.

[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0244] [Second Embodiment]

[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0257] The present invention provides a voice platform system that enables natural voice interaction with AI using an earphone-type device that users can use on a daily basis. This system receives voice, converts it into text data using speech recognition technology, generates an appropriate response using a generative AI, and then converts that response back into voice for the user.

[0258] Program processing

[0259] 1. The user puts on an earphone-type device and says, "Tell me today's news."

[0260] 2. The device receives the user's voice and generates the text "Tell me today's news" via speech recognition.

[0261] 3. The terminal sends this text to the server. The server analyzes this text using an analysis tool.

[0262] 4. Based on the analysis results, the server accesses external information sources to obtain news information. For example, it might use a news API to retrieve the latest news articles.

[0263] 5. Based on the information it has retrieved, the server generates a response that says, "Today's top news is about the recovery of the domestic economy."

[0264] 6. The server converts this response into audio data using a text-to-speech conversion method.

[0265] 7. The device receives audio data sent from the server and plays it back to the user through the earphones. The user hears the audio message, "Today's top news is about the recovery of the domestic economy."

[0266] This system allows users to obtain necessary information using only their voice, without using their hands, lowering the barrier to using AI in daily life. Furthermore, by linking with external information sources, it is possible to provide a wide range of information and respond to diverse user requests.

[0267] The following describes the processing flow.

[0268] Step 1:

[0269] The user voice-inputs "Tell me today's news" into the earphone-type device.

[0270] Step 2:

[0271] The terminal receives the user's voice signal through the built-in microphone.

[0272] Step 3:

[0273] The terminal activates the voice recognition means and converts the received voice signal into text data of "Tell me today's news".

[0274] Step 4:

[0275] The terminal transmits the converted text data to the server using the communication means.

[0276] Step 5:

[0277] The server analyzes the received text data using the analysis means to understand the user's request.

[0278] Step 6:

[0279] Based on the analysis result, the server sends a request to an external news API to obtain the latest news information.

[0280] Step 7:

[0281] The server takes in the latest news obtained from the news API and generates response data of "Today's top news is about the recovery of the domestic economy".

[0282] Step 8:

[0283] The server converts the response data into voice data using the text-to-speech conversion means.

[0284] Step 9:

[0285] The server transmits the generated voice data to the terminal.

[0286] Step 10:

[0287] The device plays the received audio data through the earphones and tells the user, "Today's top news is about the recovery of the domestic economy."

[0288] (Example 1)

[0289] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0290] Conventional voice dialogue systems offer the convenience of users obtaining information without using their hands, but they have limited capabilities in generating natural conversations and providing smooth access to external information sources. Therefore, a challenge is their inability to adequately respond to diverse user needs. In particular, there is a need for advanced conversation generation using generative artificial intelligence and real-time information retrieval that responds to user intent.

[0291] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0292] In this invention, the server includes a voice input means for receiving voice signals, a voice recognition means for converting the received voice signals into text data, and an analysis means for analyzing the text data and generating a response using generative artificial intelligence. This allows the user to obtain high-quality responses based on the latest external information through smooth and natural dialogue without using their hands.

[0293] "Voice input means" refers to a device or technology for receiving a voice signal and passing it on to the next processing stage.

[0294] "Speech recognition means" refers to a technology that analyzes received audio signals and converts them into corresponding text data.

[0295] "Communication means" refers to means for transmitting processed data to another device or system.

[0296] An "information processing device" is a computer or system used to analyze received data and generate the necessary response.

[0297] "Analysis means" refers to a method or apparatus that uses generative artificial intelligence to analyze received text data and generate appropriate response data.

[0298] "Generative artificial intelligence" is a technology for generating natural language responses based on received data, and it is an artificial intelligence that uses models to generate and understand text.

[0299] A "text-to-speech conversion method" is a technology that converts generated text data into an audio signal and outputs it to the user.

[0300] "Audio output means" refers to a device or technology for transmitting converted audio data to the user.

[0301] "External information acquisition means" refers to the means of accessing external information sources and obtaining necessary information when processing data.

[0302] A "holding device" is a device that can be used by a user by attaching it to a part of their body.

[0303] A "prompt sentence" is an input sentence used to give instructions or questions to a generative artificial intelligence.

[0304] The voice dialogue system of the present invention enables daily interaction with AI using a wearable, retainable device. This system mainly consists of voice input means, voice recognition means, communication means, information processing device, analysis means, generative artificial intelligence, text-to-speech conversion means, voice output means, and external information acquisition means.

[0305] The user wears a holding device-type device and gives an instruction in natural language such as "Tell me today's news". The voice is received by a microphone built into the device and captured by the voice input means in the device. The received voice is converted into text data using voice recognition means. In this process, a commonly used voice recognition API as voice recognition technology is utilized. For example, a cloud-based voice recognition service may be used.

[0306] The terminal transmits this converted text data to the server through secure communication. The server receives the data through an information processing device and understands the user's intention using analysis means. Here, generative artificial intelligence is utilized to create a prompt sentence and input it into the AI model. As this prompt sentence, instructions such as "The user hopes for the latest news. Please search for news." can be considered.

[0307] The server accesses an external information source such as a news API using external information acquisition means and obtains the necessary information. The acquired data is analyzed by generative artificial intelligence and output as a response in natural language. This response sentence will have specific content such as "Today's top news is about the recovery of the domestic economy".

[0308] After that, the server uses text-to-speech conversion means to convert the generated text response into voice data. For text-to-speech synthesis, a commercial text-to-speech conversion service is utilized. This voice data is transmitted to the terminal and played back to the user through the holding device-type device. Thereby, the user can obtain the latest information in voice without using their hands.

[0309] This system provides a natural and intuitive interface to the user and helps AI provide useful information in daily life.

[0310] The flow of the specific process in Example 1 will be described using FIG. 11.

[0311] Step 1:

[0312] The user wears a holding device and says, "Tell me today's news." The voice signal, as input, is received by the microphone inside the device. Here, the microphone converts the voice into a digital signal. The output is the digitized voice data.

[0313] Step 2:

[0314] The device uses speech recognition to convert digitized speech data into text data. Specifically, speech recognition software analyzes the speech waveform and identifies phonemes. This process outputs the text data "Tell me today's news."

[0315] Step 3:

[0316] The terminal sends the converted text data to the server. Here, the data is packetized using a communication method and securely transmitted over the internet. The input is text data, and the output is text data received by the server.

[0317] Step 4:

[0318] The server analyzes the received text data in the information processing device using an analysis mechanism. A prompt sentence is created using a generative AI model. Based on this prompt sentence, the server understands the user's intent and generates the instruction "The user wants news." The input is the received text, and the output is the analyzed result and the prompt sentence.

[0319] Step 5:

[0320] The server accesses a news API using an external information retrieval method to obtain the latest news data. Specifically, it sends a query to the API and receives news data in JSON format. The input is a news request, and the output is the retrieved news information.

[0321] Step 6:

[0322] The server uses generative artificial intelligence to generate a response to the user based on the acquired news information. In this generation process, the AI ​​model understands the context and creates a response sentence such as, "Today's top news is about the recovery of the domestic economy." The input is the news information and the prompt sentence, and the output is the response text.

[0323] Step 7:

[0324] The server uses text-to-speech conversion to convert the generated response text into speech data. Speech synthesis technology analyzes the text and generates a corresponding speech waveform. The input is the response text, and the output is the speech data.

[0325] Step 8:

[0326] The terminal receives audio data transmitted from the server and provides it to the user through a holding device. The earphones convert the digital audio into an analog signal and play it back to the user. The input is audio data, and the output is the audio the user hears. This entire process allows the user to obtain necessary information by voice without using their hands.

[0327] (Application Example 1)

[0328] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0329] In modern society, there is a growing demand for systems that allow users to easily complete food orders using only their voice, without directly operating a smartphone or computer. However, existing applications using voice recognition technology struggle to smoothly handle complex ordering processes, and improvements in usability are needed.

[0330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0331] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a communication means for transmitting the converted text data to a data processing device. This allows users to easily order food without using their hands by using an earphone-type device.

[0332] "Voice input means" refers to a device or apparatus for receiving voice signals.

[0333] "Speech recognition means" refers to a technology or process that converts received speech signals into corresponding text data.

[0334] "Communication means" refers to means of transmitting voice signals or data to other devices or networks.

[0335] "Analysis means" refers to a means for analyzing text data and generating corresponding response data.

[0336] "Text-to-speech conversion means" refers to a technology or device for converting generated text data into speech data.

[0337] "Audio output means" refers to a device or means for providing generated audio data to a user.

[0338] A "food ordering support system" is a function that assists the food ordering process based on voice commands and generates instructions corresponding to each step of the order.

[0339] A system for carrying out the present invention enables a user to send voice instructions using an earphone-type device and order food through an external food delivery service. This system includes voice input means, voice recognition means, communication means, analysis means, text-to-speech conversion means, voice output means, and food ordering support means.

[0340] The user first puts on an earphone-type device and speaks a voice command, such as "I want to order a pizza," into the device. This voice is received by the terminal via a voice input device. The terminal uses speech recognition software, such as the Google Speech-to-Text API, to convert this voice signal into text data. This text data is then sent to a data processing device on a smartphone or in the cloud.

[0341] The data processing unit analyzes received text data using generative AI models such as the OpenAI GPT model. This analysis generates the information and questions necessary to proceed with the ordering process. For example, it generates appropriate responses in the form of prompts, such as prompting the user to respond with "What kind of pizza would you like?"

[0342] The generated text response is then converted into audio data using text-to-speech software such as Amazon Polly. Finally, this audio data is sent through the terminal to an earphone-type device and provided to the user as audio.

[0343] As a concrete example of the system's operation, if a user says "I want to order sushi," the system will generate a question such as "What kind of sushi would you like?" and wait for further instructions from the user. An example of a prompt message would be, "The user said 'I want to order sushi.' Please generate the next appropriate question."

[0344] This system allows users to complete food orders using only their voice, providing a convenient way for users to easily and quickly utilize food delivery services.

[0345] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0346] Step 1:

[0347] The user wears an earphone-type device and verbally instructs, "I want to order a pizza." The input includes the user's voice. The earphone receives this voice signal and transmits it to the device.

[0348] Step 2:

[0349] The device converts the received audio signal into text data using speech recognition. In this step, the Google Speech-to-Text API analyzes the audio data and outputs it in text format. The output is the text data "I want to order a pizza."

[0350] Step 3:

[0351] The terminal sends text data to the server via a communication method. The input is the converted text data, and the output is the text data received by the server.

[0352] Step 4:

[0353] The server analyzes the received text data using a generation AI model. The input is text data, and it uses prompts to generate response data such as "What kind of pizza would you like?". This process determines the next step in the ordering process. The output is the generated response data.

[0354] Step 5:

[0355] The server converts the generated response data into audio data using a text-to-speech conversion method. In this step, Amazon Polly converts the text data into an audio file. The output is audio data.

[0356] Step 6:

[0357] The device receives audio data sent from the server and plays it back to the user as audio through an earphone-type device. The user can hear the question, "What kind of pizza would you like?" through the earphone.

[0358] As a result, users can proceed with their orders using only their voice. This series of steps completes the process from voice input to final voice output.

[0359] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0360] The voice platform system incorporating the emotion engine of the present invention makes user interaction more personal and effective. This system aims to generate more appropriate responses by receiving the user's voice and analyzing the emotions from the voice using the emotion engine. Specific embodiments are described below.

[0361] In this system, the user speaks "I've been feeling stressed lately" through an earphone-type device. The device receives this audio and uses speech recognition to convert it into text, "I've been feeling stressed lately." This text data and audio data are analyzed by an emotion engine to detect that the user is experiencing stress.

[0362] The device then sends text and emotion data to the server. The server analyzes the received data and uses it as a basis for generating responses that take the user's emotions into account. For example, if the user is feeling stressed, the server might consider relaxation-related content helpful and obtain information on relaxation methods from external sources.

[0363] In addition, the server uses text-to-speech technology to generate audio data, which responds with the message, "To alleviate recent stress, we recommend practicing deep breathing." This audio data is transmitted to the earphones via the terminal for the user to listen to.

[0364] In this way, by providing information that takes into account the user's emotional state, users can more easily obtain effective information that is tailored to their situation. This system aims to improve the user experience using AI by understanding and responding to the user's emotions in daily life.

[0365] The following describes the processing flow.

[0366] Step 1:

[0367] The user voice-inputs "I've been feeling stressed lately" into an earphone-type device.

[0368] Step 2:

[0369] The device uses a microphone to receive audio signals.

[0370] Step 3:

[0371] The device processes the audio signal using speech recognition technology and converts it into text data that says, "I've been under a lot of stress lately."

[0372] Step 4:

[0373] The device inputs the converted text data and the original audio data into the emotion engine.

[0374] Step 5:

[0375] The emotion engine analyzes the user's emotions from voice data and detects when they are feeling stressed.

[0376] Step 6:

[0377] The device sends text data and sentiment analysis data to the server.

[0378] Step 7:

[0379] The server analyzes the received data and determines how to generate a response based on the user's emotional state.

[0380] Step 8:

[0381] The server queries external sources to retrieve recommended content related to relaxation.

[0382] Step 9:

[0383] Based on the information it retrieves, the server generates a text response that reads, "To alleviate recent stress, we recommend practicing deep breathing."

[0384] Step 10:

[0385] The server converts the text response into audio data using a text-to-speech conversion method.

[0386] Step 11:

[0387] The server sends the generated audio data to the terminal.

[0388] Step 12:

[0389] The device plays the received audio data through the earphones and tells the user in voice, "To relieve recent stress, we recommend practicing deep breathing."

[0390] (Example 2)

[0391] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0392] Modern users expect voice assistants to understand their emotional state and respond appropriately during interactions. However, existing voice platforms lack the ability to adequately analyze user emotions and generate appropriate responses based on them. As a result, personalized experiences that users desire are less likely to be provided, potentially leading to lower user satisfaction.

[0393] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0394] In this invention, the server includes an audio input means for receiving audio signals, an emotion analysis means for analyzing audio signals and emotion data, and a means for analyzing and generating responses using a generative AI model. This enables personalized responses that correspond to the user's emotional state.

[0395] A "voice input means" is a device that has the function of receiving a user's speech and converting it into an audio signal that can be processed within the system.

[0396] A "speech recognition means" is a device that has the function of analyzing a received speech signal and converting it into corresponding text data.

[0397] An "emotion analysis device" is a device that analyzes a user's emotional state based on voice data and text data and extracts emotional data.

[0398] A "communication device" is a device that has the function of transmitting text data and emotional data to a data processing device.

[0399] A "generative AI model" is an artificial intelligence model that executes algorithms to generate responses that are appropriate to the user's intentions and emotions based on input data.

[0400] "Analysis means" refers to a device that has the function of generating response data based on the text data and sentiment data received in a data processing device.

[0401] A "text-to-speech conversion device" is a device that has the function of converting generated text data into speech data.

[0402] "Audio output means" refers to a device that has the function of providing generated audio data in a format that can be listened to by the user.

[0403] This invention is a voice platform system that analyzes the user's emotions based on voice input and provides individually personalized responses. The following describes embodiments of the system in detail.

[0404] The user speaks naturally to the system through an earphone-type device. This earphone-type device functions as a voice input means and receives the voice signal as digital data. The terminal then converts the received voice data into text data using a voice recognition means, such as commercially available voice recognition software.

[0405] The converted text and audio data are input by the terminal into an emotion analysis device. Natural language processing technology is used for this analysis to extract the emotions contained in the audio data. Existing emotion analysis software can be used as the emotion analysis device.

[0406] The terminal transmits the obtained text and sentiment data to the server using a communication method. The server uses a processing unit to run a generative AI model based on the received data and generates a response that is appropriate to the user's context and emotions.

[0407] Furthermore, the server can access external information sources and retrieve additional information tailored to the user's emotional state as needed. For example, if it detects stress, it can retrieve information from online resources on relaxation techniques.

[0408] The conversion from generated text to speech is performed using the server's text-to-speech conversion capabilities. This process sends the audio data back to the terminal in a usable format, and the user ultimately receives the response through their earphones.

[0409] For example, if a user says, "I want to do something fun this weekend, but I'm not sure what to do," the system can analyze this statement and provide information about hobbies and local events. This response will be as personalized as possible, taking into account the user's emotional state.

[0410] An example of a prompt is, "What suggestions would be best when a user is looking for activities to enjoy on their day off?" The system is expected to use a generative AI model to provide appropriate information suggestions in response to this prompt.

[0411] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0412] Step 1:

[0413] The user speaks into an earphone-type device. This voice input is received as an audio signal by the microphone built into the earphone. In this step, the input is the user's raw voice, and the output is a digital audio signal.

[0414] Step 2:

[0415] The device converts the received audio signal into text data using your speech recognition software. In this process, the digital audio signal (input) is analyzed, and a language model generates text data (output) corresponding to words and phrases.

[0416] Step 3:

[0417] The device extracts the user's emotions from the audio data using emotion analysis methods, simultaneously with the analyzed text data. At this stage, the input consists of text data and audio signals, which are analyzed and output in the form of emotion data. Emotion analysis includes tone analysis of the voice and extraction of emotional words from the text.

[0418] Step 4:

[0419] The terminal sends the converted text data and sentiment data to the server via a communication method. The input for this step is the text data and sentiment data, and the output is the transmitted data packets that reach the server.

[0420] Step 5:

[0421] The server uses the received data to run a generative AI model that generates a response appropriate to the user's input. The input consists of text data and sentiment data, and by analyzing these, it generates response data (output) that includes content. In particular, the generative AI model is trained to produce appropriate responses to prompt sentences.

[0422] Step 6:

[0423] The server uses a text-to-speech conversion mechanism to convert the generated response data into speech data. The input to this process is response data in text format, and the output is data in speech format.

[0424] Step 7:

[0425] The generated audio data is transmitted to the earphones via the terminal and ultimately provided to the user. The input in this step is data in audio format, and the output is the audio the user hears. Through this audio, the user can receive the suggested information.

[0426] (Application Example 2)

[0427] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0428] In the modern real world, consumers demand personalized services and product recommendations tailored to their individual emotional states. However, conventional technologies struggle to accurately grasp consumer emotions in real time and propose appropriate products and services based on that information. Furthermore, standardized approaches often fail to provide adequate interaction with consumers, resulting in insufficient satisfaction.

[0429] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0430] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a response generation means for performing sentiment analysis and proposing products and services based on the consumer's emotions. This enables real-time understanding of the consumer's emotional state and effective product recommendations based on that understanding.

[0431] An "audio signal" is the physical representation of sound transmitted through vibrations in the air.

[0432] "Voice input means" refers to a device or method for acquiring a voice signal and inputting it into a system.

[0433] "Speech recognition means" refers to the technology and process for converting received speech signals into text data.

[0434] "Communication means" refers to a means of sending and receiving data between different devices.

[0435] A "data processing device" is a system that analyzes received data and generates or processes information.

[0436] "Analysis means" refers to a method or technique for analyzing input data and generating necessary information.

[0437] "Text-to-speech conversion means" refers to a device or software that has the function of converting text data into speech output.

[0438] A "sound output device" is a device that delivers generated sound data to the user as sound.

[0439] A "user" is a person or entity that utilizes the system's voice input / output functions.

[0440] A "response generation means" is a means of generating appropriate product suggestions and services based on user sentiment data.

[0441] An "external information source" is a source that provides data or information that exists outside the system.

[0442] In the system for carrying out this invention, the user first provides information using a voice input device. The voice input device digitizes the received voice signal and converts it into text data using speech recognition software (e.g., Google Speech-to-Text). The converted text data is transmitted to a server via a communication means.

[0443] The server first receives the data and then uses a sentiment analysis API (such as IBM Watson or Microsoft Azure Cognitive Services) to analyze the user's emotional state. This sentiment analysis makes it possible to determine what the user wants and what their emotional state is.

[0444] Based on the analysis results, the server uses a response generation mechanism to generate product and service suggestions that correspond to the user's emotions. These suggestions are then converted back into audio data using a text-to-speech tool (such as Google Cloud Text-to-Speech) and provided to the user. The user can listen to these suggestions through the audio output mechanism.

[0445] A concrete example of this system would be a department store's in-store guidance robot that analyzes the user's stress level and suggests purchasing appropriate aromatherapy products for relaxation. An example of a prompt message might be, "Generate a message suggesting relaxation products for when the user is feeling stressed."

[0446] This invention aims to enhance real-world interactions via smartphones and robots, thereby improving consumer satisfaction.

[0447] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0448] Step 1:

[0449] The user provides information using a voice input device. The voice input device converts the analog voice signal into a digital signal and outputs it as data for input to voice recognition software.

[0450] Step 2:

[0451] The terminal passes the converted digital audio signal to speech recognition software, which then converts it into text data. In this process, the system analyzes the audio signal and outputs text that is recognized as words.

[0452] Step 3:

[0453] The terminal sends the generated text data to the server via a communication method. In this step, the data is wrapped to ensure that the text data reaches the server securely over the network.

[0454] Step 4:

[0455] The server passes the received text data to a sentiment analysis API to identify the user's emotional state. The API analyzes the text, extracts emotional information, and outputs a detailed report on the emotions the user is experiencing.

[0456] Step 5:

[0457] The server passes the analyzed sentiment information to the response generation system, which then generates appropriate product and service suggestions. This response generation process creates a prompt statement based on the sentiment data, which is then output.

[0458] Step 6:

[0459] The server passes the generated suggestions to a text-to-speech tool, which converts them into audio data. In this process, the suggested text is output as natural-sounding speech.

[0460] Step 7:

[0461] The server delivers the generated audio data to the user through an audio output device. The audio output device emits the audio in a format that makes it easy for the user to hear the suggestions.

[0462] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0463] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0464] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0465] [Third Embodiment]

[0466] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0467] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0468] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0469] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0470] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0471] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0472] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0473] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0474] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0475] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0476] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0477] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0478] The present invention provides a voice platform system that enables natural voice interaction with AI using an earphone-type device that users can use on a daily basis. This system receives voice, converts it into text data using speech recognition technology, generates an appropriate response using a generative AI, and then converts that response back into voice for the user.

[0479] Program processing

[0480] 1. The user puts on an earphone-type device and says, "Tell me today's news."

[0481] 2. The device receives the user's voice and generates the text "Tell me today's news" via speech recognition.

[0482] 3. The terminal sends this text to the server. The server analyzes this text using an analysis tool.

[0483] 4. Based on the analysis results, the server accesses external information sources to obtain news information. For example, it might use a news API to retrieve the latest news articles.

[0484] 5. Based on the information it has retrieved, the server generates a response that says, "Today's top news is about the recovery of the domestic economy."

[0485] 6. The server converts this response into audio data using a text-to-speech conversion method.

[0486] 7. The device receives audio data sent from the server and plays it back to the user through the earphones. The user hears the audio message, "Today's top news is about the recovery of the domestic economy."

[0487] This system allows users to obtain necessary information using only their voice, without using their hands, lowering the barrier to using AI in daily life. Furthermore, by linking with external information sources, it is possible to provide a wide range of information and respond to diverse user requests.

[0488] The following describes the processing flow.

[0489] Step 1:

[0490] The user voice-inputs "Tell me today's news" into the earphone-type device.

[0491] Step 2:

[0492] The device receives the user's voice signal through its built-in microphone.

[0493] Step 3:

[0494] The device activates its voice recognition system and converts the received voice signal into text data that reads, "Tell me today's news."

[0495] Step 4:

[0496] The terminal sends the converted text data to the server using a communication method.

[0497] Step 5:

[0498] The server analyzes the received text data using parsing tools to understand the user's request.

[0499] Step 6:

[0500] Based on the analysis results, the server sends a request to an external news API to retrieve the latest news information.

[0501] Step 7:

[0502] The server retrieves the latest news from the news API and generates response data that says, "Today's top news is about the recovery of the domestic economy."

[0503] Step 8:

[0504] The server converts the response data into audio data using a text-to-speech conversion method.

[0505] Step 9:

[0506] The server sends the generated audio data to the terminal.

[0507] Step 10:

[0508] The device plays the received audio data through the earphones and tells the user, "Today's top news is about the recovery of the domestic economy."

[0509] (Example 1)

[0510] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0511] Conventional voice dialogue systems offer the convenience of users obtaining information without using their hands, but they have limited capabilities in generating natural conversations and providing smooth access to external information sources. Therefore, a challenge is their inability to adequately respond to diverse user needs. In particular, there is a need for advanced conversation generation using generative artificial intelligence and real-time information retrieval that responds to user intent.

[0512] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0513] In this invention, the server includes a voice input means for receiving voice signals, a voice recognition means for converting the received voice signals into text data, and an analysis means for analyzing the text data and generating a response using generative artificial intelligence. This allows the user to obtain high-quality responses based on the latest external information through smooth and natural dialogue without using their hands.

[0514] "Voice input means" refers to a device or technology for receiving a voice signal and passing it on to the next processing stage.

[0515] "Speech recognition means" refers to a technology that analyzes received audio signals and converts them into corresponding text data.

[0516] "Communication means" refers to means for transmitting processed data to another device or system.

[0517] An "information processing device" is a computer or system used to analyze received data and generate the necessary response.

[0518] "Analysis means" refers to a method or apparatus that uses generative artificial intelligence to analyze received text data and generate appropriate response data.

[0519] "Generative artificial intelligence" is a technology for generating natural language responses based on received data, and it is an artificial intelligence that uses models to generate and understand text.

[0520] A "text-to-speech conversion method" is a technology that converts generated text data into an audio signal and outputs it to the user.

[0521] "Audio output means" refers to a device or technology for transmitting converted audio data to the user.

[0522] "External information acquisition means" refers to the means of accessing external information sources and obtaining necessary information when processing data.

[0523] A "holding device" is a device that can be used by a user by attaching it to a part of their body.

[0524] A "prompt sentence" is an input sentence used to give instructions or questions to a generative artificial intelligence.

[0525] The voice dialogue system of the present invention enables daily interaction with AI using a wearable, retainable device. This system mainly consists of voice input means, voice recognition means, communication means, information processing device, analysis means, generative artificial intelligence, text-to-speech conversion means, voice output means, and external information acquisition means.

[0526] The user wears a hold-type device and gives instructions in natural language, such as "Tell me today's news." This voice is received by a microphone built into the device and captured by the device's voice input mechanism. The received voice is then converted into text data using speech recognition technology. This process utilizes a common speech recognition API; for example, a cloud-based speech recognition service may be used.

[0527] The terminal sends this converted text data to the server via secure communication. The server receives the data through an information processing device and uses analysis tools to understand the user's intent. Here, generative artificial intelligence is utilized to create a prompt sentence and input it into the AI ​​model. A possible prompt sentence would be, "The user wants the latest news. Please look up the news."

[0528] The server accesses external information sources, such as news APIs, using external information acquisition methods to obtain the necessary information. The acquired data is analyzed by generative artificial intelligence and output as a natural language response. This response will have specific content, such as, "Today's top news is about the recovery of the domestic economy."

[0529] Subsequently, the server uses text-to-speech conversion to convert the generated text response into audio data. A commercial text-to-speech service is used for speech synthesis. This audio data is transmitted to the terminal and played back to the user via a holding device. This allows the user to obtain the latest information in audio format without using their hands.

[0530] This system provides users with a natural and intuitive interface, and helps AI provide useful information for everyday life.

[0531] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0532] Step 1:

[0533] The user wears a holding device and says, "Tell me today's news." The voice signal, as input, is received by the microphone inside the device. Here, the microphone converts the voice into a digital signal. The output is the digitized voice data.

[0534] Step 2:

[0535] The device uses speech recognition to convert digitized speech data into text data. Specifically, speech recognition software analyzes the speech waveform and identifies phonemes. This process outputs the text data "Tell me today's news."

[0536] Step 3:

[0537] The terminal sends the converted text data to the server. Here, the data is packetized using a communication method and securely transmitted over the internet. The input is text data, and the output is text data received by the server.

[0538] Step 4:

[0539] The server analyzes the received text data in the information processing device using an analysis mechanism. A prompt sentence is created using a generative AI model. Based on this prompt sentence, the server understands the user's intent and generates the instruction "The user wants news." The input is the received text, and the output is the analyzed result and the prompt sentence.

[0540] Step 5:

[0541] The server accesses a news API using an external information retrieval method to obtain the latest news data. Specifically, it sends a query to the API and receives news data in JSON format. The input is a news request, and the output is the retrieved news information.

[0542] Step 6:

[0543] The server uses generative artificial intelligence to generate a response to the user based on the acquired news information. In this generation process, the AI ​​model understands the context and creates a response sentence such as, "Today's top news is about the recovery of the domestic economy." The input is the news information and the prompt sentence, and the output is the response text.

[0544] Step 7:

[0545] The server uses text-to-speech conversion to convert the generated response text into speech data. Speech synthesis technology analyzes the text and generates a corresponding speech waveform. The input is the response text, and the output is the speech data.

[0546] Step 8:

[0547] The terminal receives audio data transmitted from the server and provides it to the user through a holding device. The earphones convert the digital audio into an analog signal and play it back to the user. The input is audio data, and the output is the audio the user hears. This entire process allows the user to obtain necessary information by voice without using their hands.

[0548] (Application Example 1)

[0549] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0550] In modern society, there is a growing demand for systems that allow users to easily complete food orders using only their voice, without directly operating a smartphone or computer. However, existing applications using voice recognition technology struggle to smoothly handle complex ordering processes, and improvements in usability are needed.

[0551] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0552] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a communication means for transmitting the converted text data to a data processing device. This allows users to easily order food without using their hands by using an earphone-type device.

[0553] "Voice input means" refers to a device or apparatus for receiving voice signals.

[0554] "Speech recognition means" refers to a technology or process that converts received speech signals into corresponding text data.

[0555] "Communication means" refers to means of transmitting voice signals or data to other devices or networks.

[0556] "Analysis means" refers to a means for analyzing text data and generating corresponding response data.

[0557] "Text-to-speech conversion means" refers to a technology or device for converting generated text data into speech data.

[0558] "Audio output means" refers to a device or means for providing generated audio data to a user.

[0559] A "food ordering support system" is a function that assists the food ordering process based on voice commands and generates instructions corresponding to each step of the order.

[0560] A system for carrying out the present invention enables a user to send voice instructions using an earphone-type device and order food through an external food delivery service. This system includes voice input means, voice recognition means, communication means, analysis means, text-to-speech conversion means, voice output means, and food ordering support means.

[0561] The user first puts on an earphone-type device and speaks a voice command, such as "I want to order a pizza," into the device. This voice is received by the terminal via a voice input device. The terminal uses speech recognition software, such as the Google Speech-to-Text API, to convert this voice signal into text data. This text data is then sent to a data processing device on a smartphone or in the cloud.

[0562] The data processing unit analyzes received text data using generative AI models such as the OpenAI GPT model. This analysis generates the information and questions necessary to proceed with the ordering process. For example, it generates appropriate responses in the form of prompts, such as prompting the user to respond with "What kind of pizza would you like?"

[0563] The generated text response is then converted into audio data using text-to-speech software such as Amazon Polly. Finally, this audio data is sent through the terminal to an earphone-type device and provided to the user as audio.

[0564] As a concrete example of the system's operation, if a user says "I want to order sushi," the system will generate a question such as "What kind of sushi would you like?" and wait for further instructions from the user. An example of a prompt message would be, "The user said 'I want to order sushi.' Please generate the next appropriate question."

[0565] This system allows users to complete food orders using only their voice, providing a convenient way for users to easily and quickly utilize food delivery services.

[0566] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0567] Step 1:

[0568] The user wears an earphone-type device and verbally instructs, "I want to order a pizza." The input includes the user's voice. The earphone receives this voice signal and transmits it to the device.

[0569] Step 2:

[0570] The device converts the received audio signal into text data using speech recognition. In this step, the Google Speech-to-Text API analyzes the audio data and outputs it in text format. The output is the text data "I want to order a pizza."

[0571] Step 3:

[0572] The terminal sends text data to the server via a communication method. The input is the converted text data, and the output is the text data received by the server.

[0573] Step 4:

[0574] The server analyzes the received text data using a generation AI model. The input is text data, and it uses prompts to generate response data such as "What kind of pizza would you like?". This process determines the next step in the ordering process. The output is the generated response data.

[0575] Step 5:

[0576] The server converts the generated response data into audio data using a text-to-speech conversion method. In this step, Amazon Polly converts the text data into an audio file. The output is audio data.

[0577] Step 6:

[0578] The device receives audio data sent from the server and plays it back to the user as audio through an earphone-type device. The user can hear the question, "What kind of pizza would you like?" through the earphone.

[0579] As a result, users can proceed with their orders using only their voice. This series of steps completes the process from voice input to final voice output.

[0580] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0581] The voice platform system incorporating the emotion engine of the present invention makes user interaction more personal and effective. This system aims to generate more appropriate responses by receiving the user's voice and analyzing the emotions from the voice using the emotion engine. Specific embodiments are described below.

[0582] In this system, the user speaks "I've been feeling stressed lately" through an earphone-type device. The device receives this audio and uses speech recognition to convert it into text, "I've been feeling stressed lately." This text data and audio data are analyzed by an emotion engine to detect that the user is experiencing stress.

[0583] The device then sends text and emotion data to the server. The server analyzes the received data and uses it as a basis for generating responses that take the user's emotions into account. For example, if the user is feeling stressed, the server might consider relaxation-related content helpful and obtain information on relaxation methods from external sources.

[0584] In addition, the server uses text-to-speech technology to generate audio data, which responds with the message, "To alleviate recent stress, we recommend practicing deep breathing." This audio data is transmitted to the earphones via the terminal for the user to listen to.

[0585] In this way, by providing information that takes into account the user's emotional state, users can more easily obtain effective information that is tailored to their situation. This system aims to improve the user experience using AI by understanding and responding to the user's emotions in daily life.

[0586] The following describes the processing flow.

[0587] Step 1:

[0588] The user voice-inputs "I've been feeling stressed lately" into an earphone-type device.

[0589] Step 2:

[0590] The device uses a microphone to receive audio signals.

[0591] Step 3:

[0592] The device processes the audio signal using speech recognition technology and converts it into text data that says, "I've been under a lot of stress lately."

[0593] Step 4:

[0594] The device inputs the converted text data and the original audio data into the emotion engine.

[0595] Step 5:

[0596] The emotion engine analyzes the user's emotions from voice data and detects when they are feeling stressed.

[0597] Step 6:

[0598] The device sends text data and sentiment analysis data to the server.

[0599] Step 7:

[0600] The server analyzes the received data and determines how to generate a response based on the user's emotional state.

[0601] Step 8:

[0602] The server queries external sources to retrieve recommended content related to relaxation.

[0603] Step 9:

[0604] Based on the information it retrieves, the server generates a text response that reads, "To alleviate recent stress, we recommend practicing deep breathing."

[0605] Step 10:

[0606] The server converts the text response into audio data using a text-to-speech conversion method.

[0607] Step 11:

[0608] The server sends the generated audio data to the terminal.

[0609] Step 12:

[0610] The device plays the received audio data through the earphones and tells the user in voice, "To relieve recent stress, we recommend practicing deep breathing."

[0611] (Example 2)

[0612] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0613] Modern users expect voice assistants to understand their emotional state and respond appropriately during interactions. However, existing voice platforms lack the ability to adequately analyze user emotions and generate appropriate responses based on them. As a result, personalized experiences that users desire are less likely to be provided, potentially leading to lower user satisfaction.

[0614] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0615] In this invention, the server includes an audio input means for receiving audio signals, an emotion analysis means for analyzing audio signals and emotion data, and a means for analyzing and generating responses using a generative AI model. This enables personalized responses that correspond to the user's emotional state.

[0616] A "voice input means" is a device that has the function of receiving a user's speech and converting it into an audio signal that can be processed within the system.

[0617] A "speech recognition means" is a device that has the function of analyzing a received speech signal and converting it into corresponding text data.

[0618] An "emotion analysis device" is a device that analyzes a user's emotional state based on voice data and text data and extracts emotional data.

[0619] A "communication device" is a device that has the function of transmitting text data and emotional data to a data processing device.

[0620] A "generative AI model" is an artificial intelligence model that executes algorithms to generate responses that are appropriate to the user's intentions and emotions based on input data.

[0621] "Analysis means" refers to a device that has the function of generating response data based on the text data and sentiment data received in a data processing device.

[0622] A "text-to-speech conversion device" is a device that has the function of converting generated text data into speech data.

[0623] "Audio output means" refers to a device that has the function of providing generated audio data in a format that can be listened to by the user.

[0624] This invention is a voice platform system that analyzes the user's emotions based on voice input and provides individually personalized responses. The following describes embodiments of the system in detail.

[0625] The user speaks naturally to the system through an earphone-type device. This earphone-type device functions as a voice input means and receives the voice signal as digital data. The terminal then converts the received voice data into text data using a voice recognition means, such as commercially available voice recognition software.

[0626] The converted text and audio data are input by the terminal into an emotion analysis device. Natural language processing technology is used for this analysis to extract the emotions contained in the audio data. Existing emotion analysis software can be used as the emotion analysis device.

[0627] The terminal transmits the obtained text and sentiment data to the server using a communication method. The server uses a processing unit to run a generative AI model based on the received data and generates a response that is appropriate to the user's context and emotions.

[0628] Furthermore, the server can access external information sources and retrieve additional information tailored to the user's emotional state as needed. For example, if it detects stress, it can retrieve information from online resources on relaxation techniques.

[0629] The conversion from generated text to speech is performed using the server's text-to-speech conversion capabilities. This process sends the audio data back to the terminal in a usable format, and the user ultimately receives the response through their earphones.

[0630] For example, if a user says, "I want to do something fun this weekend, but I'm not sure what to do," the system can analyze this statement and provide information about hobbies and local events. This response will be as personalized as possible, taking into account the user's emotional state.

[0631] An example of a prompt is, "What suggestions would be best when a user is looking for activities to enjoy on their day off?" The system is expected to use a generative AI model to provide appropriate information suggestions in response to this prompt.

[0632] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0633] Step 1:

[0634] The user speaks into an earphone-type device. This voice input is received as an audio signal by the microphone built into the earphone. In this step, the input is the user's raw voice, and the output is a digital audio signal.

[0635] Step 2:

[0636] The device converts the received audio signal into text data using your speech recognition software. In this process, the digital audio signal (input) is analyzed, and a language model generates text data (output) corresponding to words and phrases.

[0637] Step 3:

[0638] The device extracts the user's emotions from the audio data using emotion analysis methods, simultaneously with the analyzed text data. At this stage, the input consists of text data and audio signals, which are analyzed and output in the form of emotion data. Emotion analysis includes tone analysis of the voice and extraction of emotional words from the text.

[0639] Step 4:

[0640] The terminal sends the converted text data and sentiment data to the server via a communication method. The input for this step is the text data and sentiment data, and the output is the transmitted data packets that reach the server.

[0641] Step 5:

[0642] The server uses the received data to run a generative AI model that generates a response appropriate to the user's input. The input consists of text data and sentiment data, and by analyzing these, it generates response data (output) that includes content. In particular, the generative AI model is trained to produce appropriate responses to prompt sentences.

[0643] Step 6:

[0644] The server uses a text-to-speech conversion mechanism to convert the generated response data into speech data. The input to this process is response data in text format, and the output is data in speech format.

[0645] Step 7:

[0646] The generated audio data is transmitted to the earphones via the terminal and ultimately provided to the user. The input in this step is data in audio format, and the output is the audio the user hears. Through this audio, the user can receive the suggested information.

[0647] (Application Example 2)

[0648] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0649] In the modern real world, consumers demand personalized services and product recommendations tailored to their individual emotional states. However, conventional technologies struggle to accurately grasp consumer emotions in real time and propose appropriate products and services based on that information. Furthermore, standardized approaches often fail to provide adequate interaction with consumers, resulting in insufficient satisfaction.

[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0651] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a response generation means for performing sentiment analysis and proposing products and services based on the consumer's emotions. This enables real-time understanding of the consumer's emotional state and effective product recommendations based on that understanding.

[0652] An "audio signal" is the physical representation of sound transmitted through vibrations in the air.

[0653] "Voice input means" refers to a device or method for acquiring a voice signal and inputting it into a system.

[0654] "Speech recognition means" refers to the technology and process for converting received speech signals into text data.

[0655] "Communication means" refers to a means of sending and receiving data between different devices.

[0656] A "data processing device" is a system that analyzes received data and generates or processes information.

[0657] "Analysis means" refers to a method or technique for analyzing input data and generating necessary information.

[0658] "Text-to-speech conversion means" refers to a device or software that has the function of converting text data into speech output.

[0659] A "sound output device" is a device that delivers generated sound data to the user as sound.

[0660] A "user" is a person or entity that utilizes the system's voice input / output functions.

[0661] A "response generation means" is a means of generating appropriate product suggestions and services based on user sentiment data.

[0662] An "external information source" is a source that provides data or information that exists outside the system.

[0663] In the system for carrying out this invention, the user first provides information using a voice input device. The voice input device digitizes the received voice signal and converts it into text data using speech recognition software (e.g., Google Speech-to-Text). The converted text data is transmitted to a server via a communication means.

[0664] The server first receives the data and then uses a sentiment analysis API (such as IBM Watson or Microsoft Azure Cognitive Services) to analyze the user's emotional state. This sentiment analysis makes it possible to determine what the user wants and what their emotional state is.

[0665] Based on the analysis results, the server uses a response generation mechanism to generate product and service suggestions that correspond to the user's emotions. These suggestions are then converted back into audio data using a text-to-speech tool (such as Google Cloud Text-to-Speech) and provided to the user. The user can listen to these suggestions through the audio output mechanism.

[0666] A concrete example of this system would be a department store's in-store guidance robot that analyzes the user's stress level and suggests purchasing appropriate aromatherapy products for relaxation. An example of a prompt message might be, "Generate a message suggesting relaxation products for when the user is feeling stressed."

[0667] This invention aims to enhance real-world interactions via smartphones and robots, thereby improving consumer satisfaction.

[0668] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0669] Step 1:

[0670] The user provides information using a voice input device. The voice input device converts the analog voice signal into a digital signal and outputs it as data for input to voice recognition software.

[0671] Step 2:

[0672] The terminal passes the converted digital audio signal to speech recognition software, which then converts it into text data. In this process, the system analyzes the audio signal and outputs text that is recognized as words.

[0673] Step 3:

[0674] The terminal sends the generated text data to the server via a communication method. In this step, the data is wrapped to ensure that the text data reaches the server securely over the network.

[0675] Step 4:

[0676] The server passes the received text data to a sentiment analysis API to identify the user's emotional state. The API analyzes the text, extracts emotional information, and outputs a detailed report on the emotions the user is experiencing.

[0677] Step 5:

[0678] The server passes the analyzed sentiment information to the response generation system, which then generates appropriate product and service suggestions. This response generation process creates a prompt statement based on the sentiment data, which is then output.

[0679] Step 6:

[0680] The server passes the generated suggestions to a text-to-speech tool, which converts them into audio data. In this process, the suggested text is output as natural-sounding speech.

[0681] Step 7:

[0682] The server delivers the generated audio data to the user through an audio output device. The audio output device emits the audio in a format that makes it easy for the user to hear the suggestions.

[0683] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0684] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0685] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0686] [Fourth Embodiment]

[0687] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0688] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0689] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0690] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0691] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0692] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0693] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0694] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0695] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0696] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0697] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0698] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0699] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] The present invention provides a voice platform system that enables natural voice interaction with AI using an earphone-type device that users can use on a daily basis. This system receives voice, converts it into text data using speech recognition technology, generates an appropriate response using a generative AI, and then converts that response back into voice for the user.

[0701] Program processing

[0702] 1. The user puts on an earphone-type device and says, "Tell me today's news."

[0703] 2. The device receives the user's voice and generates the text "Tell me today's news" via speech recognition.

[0704] 3. The terminal sends this text to the server. The server analyzes this text using an analysis tool.

[0705] 4. Based on the analysis results, the server accesses external information sources to obtain news information. For example, it might use a news API to retrieve the latest news articles.

[0706] 5. Based on the information it has retrieved, the server generates a response that says, "Today's top news is about the recovery of the domestic economy."

[0707] 6. The server converts this response into audio data using a text-to-speech conversion method.

[0708] 7. The device receives audio data sent from the server and plays it back to the user through the earphones. The user hears the audio message, "Today's top news is about the recovery of the domestic economy."

[0709] This system allows users to obtain necessary information using only their voice, without using their hands, lowering the barrier to using AI in daily life. Furthermore, by linking with external information sources, it is possible to provide a wide range of information and respond to diverse user requests.

[0710] The following describes the processing flow.

[0711] Step 1:

[0712] The user voice-inputs "Tell me today's news" into the earphone-type device.

[0713] Step 2:

[0714] The device receives the user's voice signal through its built-in microphone.

[0715] Step 3:

[0716] The device activates its voice recognition system and converts the received voice signal into text data that reads, "Tell me today's news."

[0717] Step 4:

[0718] The terminal sends the converted text data to the server using a communication method.

[0719] Step 5:

[0720] The server analyzes the received text data using parsing tools to understand the user's request.

[0721] Step 6:

[0722] Based on the analysis results, the server sends a request to an external news API to retrieve the latest news information.

[0723] Step 7:

[0724] The server retrieves the latest news from the news API and generates response data that says, "Today's top news is about the recovery of the domestic economy."

[0725] Step 8:

[0726] The server converts the response data into audio data using a text-to-speech conversion method.

[0727] Step 9:

[0728] The server sends the generated audio data to the terminal.

[0729] Step 10:

[0730] The device plays the received audio data through the earphones and tells the user, "Today's top news is about the recovery of the domestic economy."

[0731] (Example 1)

[0732] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0733] Conventional voice dialogue systems offer the convenience of users obtaining information without using their hands, but they have limited capabilities in generating natural conversations and providing smooth access to external information sources. Therefore, a challenge is their inability to adequately respond to diverse user needs. In particular, there is a need for advanced conversation generation using generative artificial intelligence and real-time information retrieval that responds to user intent.

[0734] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0735] In this invention, the server includes a voice input means for receiving voice signals, a voice recognition means for converting the received voice signals into text data, and an analysis means for analyzing the text data and generating a response using generative artificial intelligence. This allows the user to obtain high-quality responses based on the latest external information through smooth and natural dialogue without using their hands.

[0736] "Voice input means" refers to a device or technology for receiving a voice signal and passing it on to the next processing stage.

[0737] "Speech recognition means" refers to a technology that analyzes received audio signals and converts them into corresponding text data.

[0738] "Communication means" refers to means for transmitting processed data to another device or system.

[0739] An "information processing device" is a computer or system used to analyze received data and generate the necessary response.

[0740] "Analysis means" refers to a method or apparatus that uses generative artificial intelligence to analyze received text data and generate appropriate response data.

[0741] "Generative artificial intelligence" is a technology for generating natural language responses based on received data, and it is an artificial intelligence that uses models to generate and understand text.

[0742] A "text-to-speech conversion method" is a technology that converts generated text data into an audio signal and outputs it to the user.

[0743] "Audio output means" refers to a device or technology for transmitting converted audio data to the user.

[0744] "External information acquisition means" refers to the means of accessing external information sources and obtaining necessary information when processing data.

[0745] A "holding device" is a device that can be used by a user by attaching it to a part of their body.

[0746] A "prompt sentence" is an input sentence used to give instructions or questions to a generative artificial intelligence.

[0747] The voice dialogue system of the present invention enables daily interaction with AI using a wearable, retainable device. This system mainly consists of voice input means, voice recognition means, communication means, information processing device, analysis means, generative artificial intelligence, text-to-speech conversion means, voice output means, and external information acquisition means.

[0748] The user wears a hold-type device and gives instructions in natural language, such as "Tell me today's news." This voice is received by a microphone built into the device and captured by the device's voice input mechanism. The received voice is then converted into text data using speech recognition technology. This process utilizes a common speech recognition API; for example, a cloud-based speech recognition service may be used.

[0749] The terminal sends this converted text data to the server via secure communication. The server receives the data through an information processing device and uses analysis tools to understand the user's intent. Here, generative artificial intelligence is utilized to create a prompt sentence and input it into the AI ​​model. A possible prompt sentence would be, "The user wants the latest news. Please look up the news."

[0750] The server accesses external information sources, such as news APIs, using external information acquisition methods to obtain the necessary information. The acquired data is analyzed by generative artificial intelligence and output as a natural language response. This response will have specific content, such as, "Today's top news is about the recovery of the domestic economy."

[0751] Subsequently, the server uses text-to-speech conversion to convert the generated text response into audio data. A commercial text-to-speech service is used for speech synthesis. This audio data is transmitted to the terminal and played back to the user via a holding device. This allows the user to obtain the latest information in audio format without using their hands.

[0752] This system provides users with a natural and intuitive interface, and helps AI provide useful information for everyday life.

[0753] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0754] Step 1:

[0755] The user wears a holding device and says, "Tell me today's news." The voice signal, as input, is received by the microphone inside the device. Here, the microphone converts the voice into a digital signal. The output is the digitized voice data.

[0756] Step 2:

[0757] The device uses speech recognition to convert digitized speech data into text data. Specifically, speech recognition software analyzes the speech waveform and identifies phonemes. This process outputs the text data "Tell me today's news."

[0758] Step 3:

[0759] The terminal sends the converted text data to the server. Here, the data is packetized using a communication method and securely transmitted over the internet. The input is text data, and the output is text data received by the server.

[0760] Step 4:

[0761] The server analyzes the received text data in the information processing device using an analysis mechanism. A prompt sentence is created using a generative AI model. Based on this prompt sentence, the server understands the user's intent and generates the instruction "The user wants news." The input is the received text, and the output is the analyzed result and the prompt sentence.

[0762] Step 5:

[0763] The server accesses a news API using an external information retrieval method to obtain the latest news data. Specifically, it sends a query to the API and receives news data in JSON format. The input is a news request, and the output is the retrieved news information.

[0764] Step 6:

[0765] The server uses generative artificial intelligence to generate a response to the user based on the acquired news information. In this generation process, the AI ​​model understands the context and creates a response sentence such as, "Today's top news is about the recovery of the domestic economy." The input is the news information and the prompt sentence, and the output is the response text.

[0766] Step 7:

[0767] The server uses text-to-speech conversion to convert the generated response text into speech data. Speech synthesis technology analyzes the text and generates a corresponding speech waveform. The input is the response text, and the output is the speech data.

[0768] Step 8:

[0769] The terminal receives audio data transmitted from the server and provides it to the user through a holding device. The earphones convert the digital audio into an analog signal and play it back to the user. The input is audio data, and the output is the audio the user hears. This entire process allows the user to obtain necessary information by voice without using their hands.

[0770] (Application Example 1)

[0771] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0772] In modern society, there is a growing demand for systems that allow users to easily complete food orders using only their voice, without directly operating a smartphone or computer. However, existing applications using voice recognition technology struggle to smoothly handle complex ordering processes, and improvements in usability are needed.

[0773] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0774] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a communication means for transmitting the converted text data to a data processing device. This allows users to easily order food without using their hands by using an earphone-type device.

[0775] "Voice input means" refers to a device or apparatus for receiving voice signals.

[0776] "Speech recognition means" refers to a technology or process that converts received speech signals into corresponding text data.

[0777] "Communication means" refers to means of transmitting voice signals or data to other devices or networks.

[0778] "Analysis means" refers to a means for analyzing text data and generating corresponding response data.

[0779] "Text-to-speech conversion means" refers to a technology or device for converting generated text data into speech data.

[0780] "Audio output means" refers to a device or means for providing generated audio data to a user.

[0781] A "food ordering support system" is a function that assists the food ordering process based on voice commands and generates instructions corresponding to each step of the order.

[0782] A system for carrying out the present invention enables a user to send voice instructions using an earphone-type device and order food through an external food delivery service. This system includes voice input means, voice recognition means, communication means, analysis means, text-to-speech conversion means, voice output means, and food ordering support means.

[0783] The user first puts on an earphone-type device and speaks a voice command, such as "I want to order a pizza," into the device. This voice is received by the terminal via a voice input device. The terminal uses speech recognition software, such as the Google Speech-to-Text API, to convert this voice signal into text data. This text data is then sent to a data processing device on a smartphone or in the cloud.

[0784] The data processing unit analyzes received text data using generative AI models such as the OpenAI GPT model. This analysis generates the information and questions necessary to proceed with the ordering process. For example, it generates appropriate responses in the form of prompts, such as prompting the user to respond with "What kind of pizza would you like?"

[0785] The generated text response is then converted into audio data using text-to-speech software such as Amazon Polly. Finally, this audio data is sent through the terminal to an earphone-type device and provided to the user as audio.

[0786] As a concrete example of the system's operation, if a user says "I want to order sushi," the system will generate a question such as "What kind of sushi would you like?" and wait for further instructions from the user. An example of a prompt message would be, "The user said 'I want to order sushi.' Please generate the next appropriate question."

[0787] This system allows users to complete food orders using only their voice, providing a convenient way for users to easily and quickly utilize food delivery services.

[0788] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0789] Step 1:

[0790] The user wears an earphone-type device and verbally instructs, "I want to order a pizza." The input includes the user's voice. The earphone receives this voice signal and transmits it to the device.

[0791] Step 2:

[0792] The device converts the received audio signal into text data using speech recognition. In this step, the Google Speech-to-Text API analyzes the audio data and outputs it in text format. The output is the text data "I want to order a pizza."

[0793] Step 3:

[0794] The terminal sends text data to the server via a communication method. The input is the converted text data, and the output is the text data received by the server.

[0795] Step 4:

[0796] The server analyzes the received text data using a generation AI model. The input is text data, and it uses prompts to generate response data such as "What kind of pizza would you like?". This process determines the next step in the ordering process. The output is the generated response data.

[0797] Step 5:

[0798] The server converts the generated response data into audio data using a text-to-speech conversion method. In this step, Amazon Polly converts the text data into an audio file. The output is audio data.

[0799] Step 6:

[0800] The device receives audio data sent from the server and plays it back to the user as audio through an earphone-type device. The user can hear the question, "What kind of pizza would you like?" through the earphone.

[0801] As a result, users can proceed with their orders using only their voice. This series of steps completes the process from voice input to final voice output.

[0802] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0803] The voice platform system incorporating the emotion engine of the present invention makes user interaction more personal and effective. This system aims to generate more appropriate responses by receiving the user's voice and analyzing the emotions from the voice using the emotion engine. Specific embodiments are described below.

[0804] In this system, the user speaks "I've been feeling stressed lately" through an earphone-type device. The device receives this audio and uses speech recognition to convert it into text, "I've been feeling stressed lately." This text data and audio data are analyzed by an emotion engine to detect that the user is experiencing stress.

[0805] The device then sends text and emotion data to the server. The server analyzes the received data and uses it as a basis for generating responses that take the user's emotions into account. For example, if the user is feeling stressed, the server might consider relaxation-related content helpful and obtain information on relaxation methods from external sources.

[0806] In addition, the server uses text-to-speech technology to generate audio data, which responds with the message, "To alleviate recent stress, we recommend practicing deep breathing." This audio data is transmitted to the earphones via the terminal for the user to listen to.

[0807] In this way, by providing information that takes into account the user's emotional state, users can more easily obtain effective information that is tailored to their situation. This system aims to improve the user experience using AI by understanding and responding to the user's emotions in daily life.

[0808] The following describes the processing flow.

[0809] Step 1:

[0810] The user voice-inputs "I've been feeling stressed lately" into an earphone-type device.

[0811] Step 2:

[0812] The device uses a microphone to receive audio signals.

[0813] Step 3:

[0814] The device processes the audio signal using speech recognition technology and converts it into text data that says, "I've been under a lot of stress lately."

[0815] Step 4:

[0816] The device inputs the converted text data and the original audio data into the emotion engine.

[0817] Step 5:

[0818] The emotion engine analyzes the user's emotions from voice data and detects when they are feeling stressed.

[0819] Step 6:

[0820] The device sends text data and sentiment analysis data to the server.

[0821] Step 7:

[0822] The server analyzes the received data and determines how to generate a response based on the user's emotional state.

[0823] Step 8:

[0824] The server queries external sources to retrieve recommended content related to relaxation.

[0825] Step 9:

[0826] Based on the information it retrieves, the server generates a text response that reads, "To alleviate recent stress, we recommend practicing deep breathing."

[0827] Step 10:

[0828] The server converts the text response into audio data using a text-to-speech conversion method.

[0829] Step 11:

[0830] The server sends the generated audio data to the terminal.

[0831] Step 12:

[0832] The device plays the received audio data through the earphones and tells the user in voice, "To relieve recent stress, we recommend practicing deep breathing."

[0833] (Example 2)

[0834] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0835] Modern users expect voice assistants to understand their emotional state and respond appropriately during interactions. However, existing voice platforms lack the ability to adequately analyze user emotions and generate appropriate responses based on them. As a result, personalized experiences that users desire are less likely to be provided, potentially leading to lower user satisfaction.

[0836] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0837] In this invention, the server includes an audio input means for receiving audio signals, an emotion analysis means for analyzing audio signals and emotion data, and a means for analyzing and generating responses using a generative AI model. This enables personalized responses that correspond to the user's emotional state.

[0838] A "voice input means" is a device that has the function of receiving a user's speech and converting it into an audio signal that can be processed within the system.

[0839] A "speech recognition means" is a device that has the function of analyzing a received speech signal and converting it into corresponding text data.

[0840] An "emotion analysis device" is a device that analyzes a user's emotional state based on voice data and text data and extracts emotional data.

[0841] A "communication device" is a device that has the function of transmitting text data and emotional data to a data processing device.

[0842] A "generative AI model" is an artificial intelligence model that executes algorithms to generate responses that are appropriate to the user's intentions and emotions based on input data.

[0843] "Analysis means" refers to a device that has the function of generating response data based on the text data and sentiment data received in a data processing device.

[0844] A "text-to-speech conversion device" is a device that has the function of converting generated text data into speech data.

[0845] "Audio output means" refers to a device that has the function of providing generated audio data in a format that can be listened to by the user.

[0846] This invention is a voice platform system that analyzes the user's emotions based on voice input and provides individually personalized responses. The following describes embodiments of the system in detail.

[0847] The user speaks naturally to the system through an earphone-type device. This earphone-type device functions as a voice input means and receives the voice signal as digital data. The terminal then converts the received voice data into text data using a voice recognition means, such as commercially available voice recognition software.

[0848] The converted text and audio data are input by the terminal into an emotion analysis device. Natural language processing technology is used for this analysis to extract the emotions contained in the audio data. Existing emotion analysis software can be used as the emotion analysis device.

[0849] The terminal transmits the obtained text and sentiment data to the server using a communication method. The server uses a processing unit to run a generative AI model based on the received data and generates a response that is appropriate to the user's context and emotions.

[0850] Furthermore, the server can access external information sources and retrieve additional information tailored to the user's emotional state as needed. For example, if it detects stress, it can retrieve information from online resources on relaxation techniques.

[0851] The conversion from generated text to speech is performed using the server's text-to-speech conversion capabilities. This process sends the audio data back to the terminal in a usable format, and the user ultimately receives the response through their earphones.

[0852] For example, if a user says, "I want to do something fun this weekend, but I'm not sure what to do," the system can analyze this statement and provide information about hobbies and local events. This response will be as personalized as possible, taking into account the user's emotional state.

[0853] An example of a prompt is, "What suggestions would be best when a user is looking for activities to enjoy on their day off?" The system is expected to use a generative AI model to provide appropriate information suggestions in response to this prompt.

[0854] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0855] Step 1:

[0856] The user speaks into an earphone-type device. This voice input is received as an audio signal by the microphone built into the earphone. In this step, the input is the user's raw voice, and the output is a digital audio signal.

[0857] Step 2:

[0858] The device converts the received audio signal into text data using your speech recognition software. In this process, the digital audio signal (input) is analyzed, and a language model generates text data (output) corresponding to words and phrases.

[0859] Step 3:

[0860] The device extracts the user's emotions from the audio data using emotion analysis methods, simultaneously with the analyzed text data. At this stage, the input consists of text data and audio signals, which are analyzed and output in the form of emotion data. Emotion analysis includes tone analysis of the voice and extraction of emotional words from the text.

[0861] Step 4:

[0862] The terminal sends the converted text data and sentiment data to the server via a communication method. The input for this step is the text data and sentiment data, and the output is the transmitted data packets that reach the server.

[0863] Step 5:

[0864] The server uses the received data to run a generative AI model that generates a response appropriate to the user's input. The input consists of text data and sentiment data, and by analyzing these, it generates response data (output) that includes content. In particular, the generative AI model is trained to produce appropriate responses to prompt sentences.

[0865] Step 6:

[0866] The server uses a text-to-speech conversion mechanism to convert the generated response data into speech data. The input to this process is response data in text format, and the output is data in speech format.

[0867] Step 7:

[0868] The generated audio data is transmitted to the earphones via the terminal and ultimately provided to the user. The input in this step is data in audio format, and the output is the audio the user hears. Through this audio, the user can receive the suggested information.

[0869] (Application Example 2)

[0870] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0871] In the modern real world, consumers demand personalized services and product recommendations tailored to their individual emotional states. However, conventional technologies struggle to accurately grasp consumer emotions in real time and propose appropriate products and services based on that information. Furthermore, standardized approaches often fail to provide adequate interaction with consumers, resulting in insufficient satisfaction.

[0872] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0873] In this invention, the server includes an audio input means for receiving audio signals, an audio recognition means for converting the received audio signals into text data, and a response generation means for performing sentiment analysis and proposing products and services based on the consumer's emotions. This enables real-time understanding of the consumer's emotional state and effective product recommendations based on that understanding.

[0874] An "audio signal" is the physical representation of sound transmitted through vibrations in the air.

[0875] "Voice input means" refers to a device or method for acquiring a voice signal and inputting it into a system.

[0876] "Speech recognition means" refers to the technology and process for converting received speech signals into text data.

[0877] "Communication means" refers to a means of sending and receiving data between different devices.

[0878] A "data processing device" is a system that analyzes received data and generates or processes information.

[0879] "Analysis means" refers to a method or technique for analyzing input data and generating necessary information.

[0880] "Text-to-speech conversion means" refers to a device or software that has the function of converting text data into speech output.

[0881] A "sound output device" is a device that delivers generated sound data to the user as sound.

[0882] A "user" is a person or entity that utilizes the system's voice input / output functions.

[0883] A "response generation means" is a means of generating appropriate product suggestions and services based on user sentiment data.

[0884] An "external information source" is a source that provides data or information that exists outside the system.

[0885] In the system for carrying out this invention, the user first provides information using a voice input device. The voice input device digitizes the received voice signal and converts it into text data using speech recognition software (e.g., Google Speech-to-Text). The converted text data is transmitted to a server via a communication means.

[0886] The server first receives the data and then uses a sentiment analysis API (such as IBM Watson or Microsoft Azure Cognitive Services) to analyze the user's emotional state. This sentiment analysis makes it possible to determine what the user wants and what their emotional state is.

[0887] Based on the analysis results, the server uses a response generation mechanism to generate product and service suggestions that correspond to the user's emotions. These suggestions are then converted back into audio data using a text-to-speech tool (such as Google Cloud Text-to-Speech) and provided to the user. The user can listen to these suggestions through the audio output mechanism.

[0888] A concrete example of this system would be a department store's in-store guidance robot that analyzes the user's stress level and suggests purchasing appropriate aromatherapy products for relaxation. An example of a prompt message might be, "Generate a message suggesting relaxation products for when the user is feeling stressed."

[0889] This invention aims to enhance real-world interactions via smartphones and robots, thereby improving consumer satisfaction.

[0890] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0891] Step 1:

[0892] The user provides information using a voice input device. The voice input device converts the analog voice signal into a digital signal and outputs it as data for input to voice recognition software.

[0893] Step 2:

[0894] The terminal passes the converted digital audio signal to speech recognition software, which then converts it into text data. In this process, the system analyzes the audio signal and outputs text that is recognized as words.

[0895] Step 3:

[0896] The terminal sends the generated text data to the server via a communication method. In this step, the data is wrapped to ensure that the text data reaches the server securely over the network.

[0897] Step 4:

[0898] The server passes the received text data to a sentiment analysis API to identify the user's emotional state. The API analyzes the text, extracts emotional information, and outputs a detailed report on the emotions the user is experiencing.

[0899] Step 5:

[0900] The server passes the analyzed sentiment information to the response generation system, which then generates appropriate product and service suggestions. This response generation process creates a prompt statement based on the sentiment data, which is then output.

[0901] Step 6:

[0902] The server passes the generated suggestions to a text-to-speech tool, which converts them into audio data. In this process, the suggested text is output as natural-sounding speech.

[0903] Step 7:

[0904] The server delivers the generated audio data to the user through an audio output device. The audio output device emits the audio in a format that makes it easy for the user to hear the suggestions.

[0905] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0906] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0907] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0908] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0909] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0910] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0911] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0912] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0913] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0914] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0915] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0916] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0917] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0918] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0919] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0920] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0921] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0922] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0923] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0924] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0925] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0926] The following is further disclosed regarding the embodiments described above.

[0927] (Claim 1)

[0928] Audio input means for receiving audio signals,

[0929] A speech recognition means that converts a received audio signal into text data,

[0930] A communication means for transmitting the converted text data to a data processing device,

[0931] A data processing device includes an analysis means that analyzes text data and generates corresponding response data,

[0932] A text-to-speech conversion means that converts the generated response data into audio data,

[0933] A means of providing audio data to the user,

[0934] A system that includes this.

[0935] (Claim 2)

[0936] The system according to claim 1, wherein the voice input means is incorporated into an earphone-type device.

[0937] (Claim 3)

[0938] The system according to claim 1, wherein the data processing device includes means for accessing an external information source and obtaining additional information.

[0939] "Example 1"

[0940] (Claim 1)

[0941] Audio input means for receiving audio signals,

[0942] A speech recognition means that converts a received audio signal into text data,

[0943] A communication means for transmitting the converted text data to an information processing device,

[0944] An analysis means for analyzing text data in an information processing device and generating corresponding response data using generative artificial intelligence,

[0945] A text-to-speech conversion means that converts the generated response data into speech data,

[0946] A voice output means that provides the generated voice data to the user,

[0947] An information processing device accesses an external information source and obtains the necessary information;

[0948] A system that includes this.

[0949] (Claim 2)

[0950] The system according to claim 1, wherein the voice input means is incorporated into a holding device.

[0951] (Claim 3)

[0952] The system according to claim 1, wherein the analysis means includes means for creating prompt sentences to be input to a generative artificial intelligence.

[0953] "Application Example 1"

[0954] (Claim 1)

[0955] Audio input means for receiving audio signals,

[0956] A speech recognition means that converts a received audio signal into text data,

[0957] A communication means for transmitting the converted text data to a data processing device,

[0958] A data processing device includes an analysis means that analyzes text data and generates corresponding response data,

[0959] A text-to-speech conversion means that converts the generated response data into audio data,

[0960] A means of providing audio data to the user,

[0961] A food ordering support means that initiates a food order based on voice commands and generates instructions corresponding to each step of the order,

[0962] A system that includes this.

[0963] (Claim 2)

[0964] The system according to claim 1, wherein the voice input means is incorporated into an earphone-type device.

[0965] (Claim 3)

[0966] The system according to claim 1, wherein the data processing device includes means for accessing an external information source and obtaining additional information.

[0967] "Example 2 of combining an emotion engine"

[0968] (Claim 1)

[0969] Audio input means for receiving audio signals,

[0970] A speech recognition means that converts a received audio signal into text data,

[0971] A sentiment analysis means for analyzing sentiment data obtained through converted text data and audio signals,

[0972] A communication means for transmitting text data and sentiment data to a data processing device,

[0973] An analysis means for analyzing text data and sentiment data using a generation AI model in a data processing device, and generating corresponding response data,

[0974] A text-to-speech conversion means that converts the generated response data into audio data,

[0975] A voice output means that provides the generated voice data to the user,

[0976] A system that includes this.

[0977] (Claim 2)

[0978] The system according to claim 1, wherein the voice input means is incorporated into an earphone-type device.

[0979] (Claim 3)

[0980] The system according to claim 1, wherein the data processing device includes means for accessing an external information source and obtaining additional information based on the user's emotional state.

[0981] "Application example 2 when combining with an emotional engine"

[0982] (Claim 1)

[0983] Audio input means for receiving audio signals,

[0984] A speech recognition means that converts a received audio signal into text data,

[0985] A communication means for transmitting the converted text data to a data processing device,

[0986] An analysis means in a data processing device that analyzes text data, performs sentiment analysis, and generates corresponding response data,

[0987] A text-to-speech conversion means that converts the generated response data into audio data,

[0988] A means of providing audio data to the user,

[0989] A response generation means that proposes products and services based on the user's emotions,

[0990] A system that includes this.

[0991] (Claim 2)

[0992] The system according to claim 1, wherein the voice input means is incorporated into the voice transmission device.

[0993] (Claim 3)

[0994] The system according to claim 1, wherein the data processing device has means for accessing an external information source to obtain additional information and for generating sentiment-based suggestion information. [Explanation of Symbols]

[0995] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Audio input means for receiving audio signals, A speech recognition means that converts a received audio signal into text data, A communication means for transmitting the converted text data to a data processing device, A data processing device includes an analysis means that analyzes text data and generates corresponding response data, A text-to-speech conversion means that converts the generated response data into audio data, A means of providing audio data to the user, A system that includes this.

2. The system according to claim 1, wherein the voice input means is incorporated into an earphone-type device.

3. The system according to claim 1, wherein the data processing device includes means for accessing an external information source and obtaining additional information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A