system

The system addresses the limitations of conventional voice assistants by enabling accurate and comprehensive information delivery to diverse user groups through speech recognition, natural language processing, and generative AI, ensuring quick and appropriate responses.

JP2026063807APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Conventional voice assistant systems struggle to provide accurate and comprehensive information to diverse user groups, including the elderly and children, due to limitations in speech recognition, natural language processing, and the inability to access the latest information and specialized knowledge.

Method used

A system that includes means for receiving user voice, converting it into digital data, transmitting it to a server for speech recognition and natural language processing, searching internal and external databases for information, generating answers using generative artificial intelligence, and converting the answers into voice data for response, ensuring accurate and appropriate information delivery.

Benefits of technology

The system provides quick and accurate information to a wide range of users, including seniors and children, by accessing both internal and external databases and utilizing generative AI to generate appropriate responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063807000001_ABST
    Figure 2026063807000001_ABST
Patent Text Reader

Abstract

We provide a system that can deliver accurate and appropriate information to diverse user groups, including seniors and children. [Solution] A system comprising: means for receiving user voice and converting it into digital data; means for transmitting the digital data to a server; means for converting the digital data into text data using a speech recognition engine on the server; means for passing the text data to a natural language processing engine to analyze its intent and content; means for searching for necessary information from an internal database and an external database based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; and means for transmitting the voice data to a terminal and responding to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional voice assistant systems have the problem that they cannot adequately respond to a variety of user groups such as the elderly and children, and particularly lack the ability to answer specialized information and extensive knowledge. In addition, there are also limitations in the accuracy of speech recognition and the ability of natural language processing, and there are many situations where appropriate answers cannot be provided for specific questions. Furthermore, since information is obtained only within the scope of the internal database, there is also a problem that the latest information and specialized knowledge cannot be accessed.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides a system comprising means for receiving user voice and converting it into digital data, means for transmitting the digital data to a server, means for converting the digital data into text data using a speech recognition engine, means for passing the text data to a natural language processing engine to analyze intent and content, means for searching for necessary information from an internal database and an external database based on the analysis results, means for generating answers to user questions using generative artificial intelligence, means including a speech synthesis engine that converts the generated answers into voice data, and means for transmitting the voice data to a terminal and responding to the user. This makes it possible to provide accurate and appropriate information to diverse user groups such as seniors and children, and if information does not exist in the internal database, it can obtain information from an external database to handle the latest information and specialized knowledge.

[0006] "User" refers to an individual or non-individual who uses the system for voice input and information retrieval.

[0007] A "terminal" refers to a device that captures the user's voice, converts it into digital data, and sends it to a server.

[0008] A "server" refers to a centralized computing resource that performs processing such as speech recognition, natural language processing, information retrieval, response generation using generative artificial intelligence, and speech synthesis.

[0009] A "speech recognition engine" refers to a software module that converts speech data into text data.

[0010] A "natural language processing engine" refers to a software module used to analyze the intent and content of text data.

[0011] An "internal database" refers to a database managed within a system that stores information that allows users to search for answers to their questions.

[0012] An "external database" refers to an external source of information that is consulted when the information is not found in the internal database.

[0013] "Generative artificial intelligence" refers to artificial intelligence technology that generates appropriate responses in natural language based on user input.

[0014] A "speech synthesis engine" refers to a software module that converts text data into speech data.

[0015] "Audio data" refers to digitized audio information.

[0016] "Text data" refers to character information converted by a speech recognition engine.

[0017] "Answer generation" refers to the process of creating appropriate responses to user questions using generative artificial intelligence.

[0018] "Analysis" refers to the process of understanding the intent and content of text data using a natural language processing engine.

[0019] "Information retrieval" refers to the process of obtaining necessary information from internal and external databases.

[0020] "Response" refers to the voice-based answer provided to the user. [Brief explanation of the drawing]

[0021] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0022] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0025] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0026] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0027] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0029] [First Embodiment]

[0030] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0031] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0034] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0037] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0041] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0042] The system according to the present invention is a voice assistant system that caters to a diverse range of users and provides appropriate answers to questions about information in a wide range of fields. This system encompasses a series of processes that analyze voice input and generate appropriate answers, and the details of its operation are as follows.

[0043] Receiving voice input

[0044] The user speaks a question to the voice assistant. The device captures this audio through the microphone and converts the analog audio into digital data. The converted audio data is then sent from the device to the server.

[0045] Speech recognition and language analysis

[0046] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. The converted text data is then passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[0047] Information retrieval and generation

[0048] Based on the analysis results, the server first searches its internal database for the necessary information. If the information is not found in the internal database, it retrieves it from an external database. This process may utilize various external APIs.

[0049] After obtaining the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. This generated text answer is then passed to a speech synthesis engine and converted into speech data.

[0050] Speech synthesis and response return

[0051] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[0052] Specific example

[0053] Example 1: Weather question

[0054] When a user asks, "What's the weather like in Tokyo today?",

[0055] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0056] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0057] 3. Analyze that the server is requesting weather information using an NLP engine.

[0058] 4. The server queries the weather information API for the weather in Tokyo.

[0059] 5. The server retrieves the information "Today in Tokyo it is sunny and the maximum temperature is 25 degrees Celsius," and a generative AI generates an appropriate response.

[0060] 6. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0061] 7. The device plays the answer and says, "Today in Tokyo it's sunny and the high temperature is 25 degrees Celsius."

[0062] Example 2: History Questions

[0063] When a user asks, "When was Alexander the Great born?",

[0064] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0065] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[0066] 3. Analyze that the server is requesting historical information using the NLP engine.

[0067] 4. The server searches its internal database and confirms the relevant information.

[0068] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and generates an answer using a generative AI.

[0069] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0070] 8. The device plays the answer and says, "Alexander the Great was born in 356 BC."

[0071] A system configured in this way can provide users with answers from a wide range of sources quickly and accurately, and can accommodate diverse user groups such as seniors and children.

[0072] The following describes the processing flow.

[0073] Step 1:

[0074] The user speaks a question to the voice assistant. An example of such a question is, "What's the weather like in Tokyo today?"

[0075] Step 2:

[0076] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. This digital audio data is then encoded into a format suitable for speech recognition.

[0077] Step 3:

[0078] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[0079] Step 4:

[0080] The server receives the audio data. The received audio data is then passed to the speech recognition engine.

[0081] Step 5:

[0082] The server uses a speech recognition engine to convert the audio data into text data. For example, text data such as "What's the weather like in Tokyo today?" is generated.

[0083] Step 6:

[0084] The server passes the generated text data to a natural language processing (NLP) engine, which analyzes the intent and content of the question. The NLP engine then identifies the category of the question (e.g., weather information, historical information, etc.).

[0085] Step 7:

[0086] Based on the analysis results, the server searches its internal database to check if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[0087] Step 8:

[0088] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[0089] Step 9:

[0090] Based on the information acquired by the server, generative artificial intelligence (e.g., a large-scale language model) is used to generate natural-sounding answers to the user's questions. For example, an answer such as "Today in Tokyo it is sunny and the highest temperature is 25 degrees Celsius" might be generated.

[0091] Step 10:

[0092] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[0093] Step 11:

[0094] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[0095] Step 12:

[0096] The device receives the audio data and plays it back. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[0097] This series of steps allows users to receive accurate and rapid voice answers to their questions.

[0098] (Example 1)

[0099] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Conventional systems using speech recognition and natural language processing have limitations in their ability to access internal and external databases when searching for specific information, making it difficult to provide flexible and comprehensive information. Furthermore, the lack of sufficient technology to accurately analyze user intent and quickly retrieve corresponding information results in a difficulty in improving the user experience.

[0101] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0102] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to a network; means for converting the digital data into text data using a speech recognition engine on a computer on the network; means for passing the text data to a natural language processing engine to analyze its intent and content; means for retrieving necessary information from internal and external storage devices based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; and means for transmitting the voice data to a terminal and responding to the user. This makes it possible to provide users with information from a wide range of sources quickly and accurately, and to provide appropriate answers to user questions in voice.

[0103] A "user" is an individual or organization that uses the system to input questions by voice.

[0104] A "network" is a means of communication for sending and receiving digital data between various computers.

[0105] A "speech recognition engine" is a software or hardware system that analyzes input speech data and converts it into text data.

[0106] "Text data" refers to digital data in sentence format converted by a speech recognition engine.

[0107] A "natural language processing engine" is a software or hardware system that analyzes text data and understands the user's intent and content from it.

[0108] An "internal storage device" is a device installed within a system for storing and retrieving data.

[0109] "External storage devices" are devices or services that exist outside the system and are used to store and retrieve data.

[0110] "Generative artificial intelligence" is a system of algorithms or software that generates appropriate answers based on questions from users.

[0111] A "speech synthesis engine" is a software or hardware system that converts generated text data into speech data.

[0112] A "terminal" is a device used by a user to input voice, receive audio data, and play it back.

[0113] The system according to the present invention is a voice assistant system that analyzes the user's voice input, generates an appropriate response, and provides a voice reply. Specific embodiments of the present invention are described below.

[0114] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the audio using its built-in microphone and converts this analog audio data into digital data. The converted audio data is then sent to the server in binary format.

[0115] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0116] Next, this text data is passed to a natural language processing (NLP) engine (e.g., a typical NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information.

[0117] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve the current weather information for Tokyo.

[0118] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI model, e.g., a general-purpose generative AI model) to generate an appropriate answer to the user's question. For example, it might generate the sentence, "Today in Tokyo it is sunny, and the high temperature is 25 degrees Celsius." The prompt used would be: "Today's weather in Tokyo is sunny, and the high temperature is 25 degrees Celsius."

[0119] The generated text response is passed to a speech synthesis engine (e.g., general speech synthesis software) and converted into audio data. The speech synthesis engine converts the generated text data into an audio format (e.g., MP3) and returns that audio data to the server.

[0120] The server sends the returned audio data to the terminal. The terminal plays the received audio data through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the highest temperature is 25 degrees Celsius."

[0121] This system allows users to quickly and accurately obtain information from a wide range of sources simply by asking questions by voice. Furthermore, the system can accommodate diverse user groups, including seniors and children.

[0122] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0123] Step 1: Receiving voice input

[0124] The user speaks a question into the microphone. The device uses its built-in microphone to capture the audio and converts this analog audio data into digital data. Specifically, the device's audio capture function is activated, the audio signal is converted to a digital format according to the sampling rate, and stored in a buffer in binary format.

[0125] Input: User's analog voice

[0126] Output: Digital audio data (binary format)

[0127] Step 2: Sending the audio data

[0128] The terminal divides the digital audio data stored in the buffer into packets of a fixed size and sends them to the server via secure and high-speed network communication (e.g., HTTPS). HTTP POST requests are used to send the audio data to the server-side API endpoint.

[0129] Input: Digital audio data (binary format)

[0130] Output: Request to send audio data to the server

[0131] Step 3: Speech recognition and text conversion

[0132] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. Specifically, it performs audio waveform analysis and uses a phonological and lexical model to convert the audio into corresponding text. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0133] Input: Audio data (binary format)

[0134] Output: Text data

[0135] Step 4: Text analysis and intent identification

[0136] The server passes the converted text data to a natural language processing (NLP) engine (e.g., general NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information. The analysis results are stored in JSON format.

[0137] Input: Text data

[0138] Output: Analysis result (JSON format)

[0139] Step 5: Information Search

[0140] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve current weather information for Tokyo. Here, it uses an HTTP request to query the external database and receives a response in JSON format.

[0141] Input: Analysis result (JSON format)

[0142] Output: Acquired information (JSON format)

[0143] Step 6: Generating the answer

[0144] Based on the weather information it receives, the server uses generative artificial intelligence (generative AI model, e.g., a general generative AI model) to generate appropriate answers to the user's questions. Specifically, it inputs the prompt "What's the weather like in Tokyo today?" into the generative AI model and generates the answer "It's sunny in Tokyo today, and the maximum temperature is 25 degrees Celsius."

[0145] Input: Acquired information (in JSON format) and prompt message

[0146] Output: Generated response (text data)

[0147] Step 7: Speech Synthesis

[0148] The server passes the generated text to a speech synthesis engine (e.g., common speech synthesis software) and converts it into audio data. The speech synthesis engine then converts the generated text data into an audio format (e.g., MP3). This conversion uses a speech model and involves acoustic processing to produce natural-sounding speech.

[0149] Input: Generated response (text data)

[0150] Output: Audio data (MP3 format)

[0151] Step 8: Return of audio data and response playback

[0152] The server sends the returned audio data to the terminal as an HTTP response. The terminal stores the received audio data in a buffer and plays it back through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the maximum temperature is 25 degrees Celsius."

[0153] Input: Audio data (MP3 format)

[0154] Output: Voice response to the user

[0155] (Application Example 1)

[0156] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0157] In physical stores, there is a need for a system that allows customers to receive quick and appropriate answers to their questions regarding product information, store layout, and inventory. Traditional methods require customers to ask store staff directly or search for store signs themselves, which is inconvenient and time-consuming. In particular, when customers ask a variety of questions, responding to them takes time, making it difficult to provide satisfactory service.

[0158] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0159] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to the server; means for the server to convert the digital data into text data using a speech recognition engine; means for passing the text data to a natural language processing engine to analyze its intent and content; means for searching for necessary information from an internal database and an external database based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; means for transmitting the voice data to a terminal and responding to the user; and means for providing quick and appropriate voice answers to questions from customers in physical stores regarding product information, store guidance, inventory checks, etc. This enables users to quickly obtain necessary information through a voice assistant no matter where they are in the store.

[0160] A "user" is someone who uses a voice assistant system to ask questions or give instructions.

[0161] "Digital data" refers to data obtained by converting analog audio signals into a digital format.

[0162] A "server" is a computer system used for speech recognition, natural language processing, database retrieval, generative artificial intelligence, and speech synthesis.

[0163] A "speech recognition engine" is a program or function that analyzes speech data and converts it into text data.

[0164] "Text data" refers to character information generated by a speech recognition engine.

[0165] A "natural language processing engine" is a program or function that analyzes text data to understand the user's intent and content.

[0166] An "internal database" is a database stored within the server where the information to be searched is stored.

[0167] An "external database" is a database that exists outside the server and can be accessed via an API.

[0168] "Generative artificial intelligence" is artificial intelligence that generates appropriate answers based on analyzed information.

[0169] A "speech synthesis engine" is a program or function that converts text data into speech data.

[0170] A "device" is a device used by a user to operate a voice assistant system, and includes smartphones and smart glasses.

[0171] A "physical store" refers to a sales facility or service provider that actually exists in a physical space.

[0172] "Product information" refers to detailed product descriptions, prices, and stock availability information provided within a physical store.

[0173] "Prompt and appropriate responses" refer to providing accurate information immediately in response to user questions.

[0174] Using the above definitions, we will clarify each function and element of the voice assistant system.

[0175] The system according to this invention is a voice assistant system that provides quick and appropriate answers to questions from users in physical stores. Specific embodiments thereof will be described below.

[0176] System Configuration

[0177] 1. User voice input

[0178] The user inputs their questions by voice into a device such as a smartphone or smart glasses. The voice input is captured through the device's built-in microphone.

[0179] 2. Digital conversion and transmission of audio

[0180] The terminal converts the captured analog audio into digital data. The converted digital data is then sent to the server via the network.

[0181] 3. Speech Recognition and Natural Language Processing

[0182] The server uses a speech recognition engine (e.g., Google® Speech Recognition API) to convert the received digital audio data into text data. Next, it passes this text data to a natural language processing engine (e.g., an NLP engine) to analyze the intent and content of the user's question.

[0183] 4. Information Retrieval and Generation

[0184] The server searches for necessary information from internal and external databases based on the analysis results. It uses external APIs as needed to retrieve information from external databases. This search result is then used with generative artificial intelligence (generative AI model) to generate appropriate answers.

[0185] 5. Speech synthesis and response

[0186] The generated response is converted into audio data using a speech synthesis engine (e.g., gTTS). The converted audio data is then sent back to the terminal via the network, where it plays the audio data and responds to the user.

[0187] Specific usage examples

[0188] Example 1: Product Information

[0189] The user asks, "Where can I find this product?" inside the store.

[0190] The voice assistant system captures this question and sends it to the server.

[0191] The server converts the speech into text using a speech recognition engine, and then analyzes it as "product information" using a natural language processing engine.

[0192] The server retrieves the product's location information from an internal database and uses generative artificial intelligence to generate a response such as, "This product is located next to the cash register on the third floor."

[0193] The generated response is converted by a speech synthesis engine and sent to the terminal.

[0194] The device plays a message to the user saying, "This item is located next to the cash register on the 3rd floor."

[0195] Usage example 2: Inventory check

[0196] The user asks, "Do you have this item in stock?"

[0197] The voice assistant system captures the user's question and sends it to the server.

[0198] The server uses a speech recognition engine to convert the speech into text data, and then a natural language processing engine analyzes it to read "check inventory".

[0199] The server uses an external API to retrieve the latest inventory information and generates responses such as "This product is still in stock" using generative artificial intelligence.

[0200] The generated response is converted into audio data by a speech synthesis engine and sent to the terminal.

[0201] The device plays the product and informs the user that "this item is still in stock."

[0202] Examples of prompt statements

[0203] Example of a prompt message for product information:

[0204] "The user is asking for the location information of this product within the store. Please provide the specific location of the product."

[0205] Example prompt message for checking inventory:

[0206] "A user is asking about the availability of a specific product. Please provide the stock status for this product."

[0207] Through the above explanation, we have provided specific examples of voice assistant systems and clarified methods for improving user convenience in physical stores.

[0208] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0209] Step 1:

[0210] The user inputs a question by voice. The device's microphone captures the voice and converts the analog audio signal into digital data. This digital data is transmitted to the server via the network. The input is the user's voice, and the output is digital audio data.

[0211] Step 2:

[0212] The server receives digital data and passes it to a speech recognition engine (e.g., Google Speech Recognition API), which converts the audio into text data. This conversion process analyzes the frequency information of the audio data and generates the corresponding string. The input is digital audio data, and the output is text data.

[0213] Step 3:

[0214] The server passes text data to a natural language processing (NLP) engine, which analyzes the intent and content of the user's question. The NLP engine grammatically and semantically analyzes the text data to identify which information the user's question relates to. The input is text data, and the output is the analysis result.

[0215] Step 4:

[0216] The server searches for the necessary information from internal and external databases based on the analysis results. The internal database contains product information and in-store guidance, while the external database contains the latest information accessible via API. It generates database queries and retrieves information based on them. The input is the analysis results, and the output is the necessary information.

[0217] Step 5:

[0218] Based on the information acquired by the server, a generative artificial intelligence (generative AI model) is used to generate appropriate answers to the user's questions. In this process, the generative AI model automatically generates answers based on the prompt text. The input consists of the required information and the prompt text, and the output is the generated text answer.

[0219] Step 6:

[0220] The generated text response is passed to a speech synthesis engine (e.g., gTTS) and converted into audio data. The speech synthesis engine then converts the text data into an audio file. The input is the generated text response, and the output is the audio data.

[0221] Step 7:

[0222] The server sends audio data to the terminal. The terminal plays the received audio data and provides the answer to the user verbally. The input is audio data, and the output is the verbal answer to the user.

[0223] The above outlines the specific processing steps in the invention's system. This series of processes enables users to quickly obtain necessary information via a voice assistant, regardless of their location within the store.

[0224] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0225] The system according to the present invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. The details of the operation of the present invention are as follows.

[0226] Receiving voice input

[0227] The user speaks a question to the voice assistant. An example question might be, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. This digital audio data is then sent from the device to the server.

[0228] Speech recognition and language analysis

[0229] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. For example, the text data "What's the weather like in Tokyo today?" is generated. This text data is passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[0230] Addition of emotion recognition

[0231] In the process described above, the server uses an emotion engine to recognize emotional information from the user's voice. For example, if the user is angry, the emotion engine recognizes that anger and adds it to the analysis results.

[0232] Information retrieval and generation

[0233] Based on the analysis results, the server first searches its internal database for the necessary information. If the necessary information is not found in the internal database, it retrieves it from an external database. Various external APIs may be used in this process.

[0234] After acquiring the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone. This generated text response is then passed to a speech synthesis engine and converted into audio data.

[0235] Speech synthesis and response return

[0236] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[0237] Specific example

[0238] Example 1: Weather question

[0239] When a user asks, "What's the weather like in Tokyo today?",

[0240] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0241] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0242] 3. Analyze that the server is requesting weather information using an NLP engine.

[0243] 4. The server uses an emotion engine to recognize the user's emotions (e.g., relaxed).

[0244] 5. The server queries the weather information API for the weather in Tokyo.

[0245] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[0246] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0247] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0248] Example 2: History Questions

[0249] When a user asks, "When was Alexander the Great born?",

[0250] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0251] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[0252] 3. Analyze that the server is requesting historical information using the NLP engine.

[0253] 4. The server uses an emotion engine to recognize the user's emotions (e.g., excited).

[0254] 5. The server searches its internal database and confirms the relevant information.

[0255] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and uses a generative AI to generate an excited response.

[0256] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0257] 8. The device plays the answer back, excitedly stating, "Alexander the Great was born in 356 BC."

[0258] A system configured in this way can quickly and accurately provide answers from a wide range of sources, including appropriate responses tailored to the user's emotions. This allows users to experience more satisfying interactions.

[0259] The following describes the processing flow.

[0260] Step 1:

[0261] The user speaks a question to the voice assistant, such as, "What's the weather like in Tokyo today?"

[0262] Step 2:

[0263] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. The converted audio data is sent to a server for processing.

[0264] Step 3:

[0265] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[0266] Step 4:

[0267] The server receives the transmitted audio data. The received audio data is then passed to the speech recognition engine.

[0268] Step 5:

[0269] The server uses a speech recognition engine to convert the audio data into text data. For example, the text "What's the weather like in Tokyo today?" is generated.

[0270] Step 6:

[0271] The server passes the generated text data to a natural language processing (NLP) engine to analyze the intent and content of the question. Based on the analysis results, it identifies that the question is related to weather information.

[0272] Step 7:

[0273] The server uses an emotion engine to analyze the user's emotional state from the voice data. For example, it can determine whether the user is speaking in a relaxed state or is angry.

[0274] Step 8:

[0275] Based on the analysis results, the server first searches its internal database to see if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[0276] Step 9:

[0277] The server queries an external database (e.g., weather information API) to obtain the necessary information. For example, using an external API to obtain "the current weather in Tokyo".

[0278] Step 10:

[0279] Based on the information obtained by the server, a generative artificial intelligence is used to generate an answer to the user's question. At this time, reflecting the recognized sentiment information, the tone and content of the answer sentence are adjusted. For example, an answer such as "Today in Tokyo, it is sunny and the highest temperature is 25 degrees" is generated.

[0280] Step 11:

[0281] The server passes the generated text answer to a text-to-speech engine and converts it into audio data. The text-to-speech engine converts the text into speech.

[0282] Step 12:

[0283] The server transmits the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[0284] Step 13:

[0285] The terminal receives the audio data and plays the received audio data. As a result, the user can receive an audio answer such as "Today in Tokyo, it is sunny and the highest temperature is 25 degrees". This answer is provided in an appropriate tone according to the user's sentiment.

[0286] (Example 2)

[0287] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".

[0288] Conventional voice assistant systems have been required to provide quick and accurate responses to user utterances, but they have faced challenges in natural communication with users, particularly due to their inability to respond with consideration for emotions. Furthermore, when the information corresponding to a user's question is not found in the internal database, the means of retrieving information from external databases are limited, making it difficult to provide an appropriate answer.

[0289] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0290] In this invention, the server includes means for converting digital data into text data using a speech recognition engine, means for analyzing intent and content using a natural language processing engine, and means for analyzing the user's emotions. This enables natural communication that takes emotions into consideration. Furthermore, by providing means for retrieving information from an external database when the information does not exist in the internal database, it becomes possible to provide quick and appropriate answers to a wide range of inquiries.

[0291] A "speech recognition engine" is a software or hardware technology used to convert speech data into text data.

[0292] A "natural language processing engine" is a software or hardware technology used to analyze user intent and content from text data.

[0293] "Sentiment analysis" is a technology that extracts emotional information from a user's voice and text data to recognize the user's state of mind.

[0294] An "internal database" is a database that stores information managed within a system.

[0295] An "external database" is a database that provides information located outside of the system.

[0296] "Generative artificial intelligence" is an artificial intelligence technology used to generate answers to user questions.

[0297] A "speech synthesis engine" is a software or hardware technology used to convert text data into speech data.

[0298] A "terminal" is a device used by a user to access a voice assistant system.

[0299] "Emotional information" refers to data that indicates the emotional state expressed by the user.

[0300] SSL / TLS is a protocol for encrypting and securely conducting data communications.

[0301] The system according to this invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. Details of the system for carrying out this invention are described below.

[0302] Receiving voice input

[0303] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. An A / D converter is used for this conversion process.

[0304] Sending audio data

[0305] After the digital audio data is generated, the device sends this data to the server via the internet. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security.

[0306] Speech Recognition and Language Analysis

[0307] The server passes the received voice data to a voice recognition engine (common voice recognition software) and converts it into text data. For example, commonly used voice recognition technologies are used in the voice recognition engine. Next, this text data is passed to a natural language processing (NLP) engine (common natural language processing software) to analyze the user's intention and the content of the question. Based on this analysis result, it is determined what the necessary information is.

[0308] Addition of Emotion Recognition

[0309] The server passes the voice data or text data to an emotion analysis engine (common emotion analysis software) to recognize the user's emotion information. For example, the emotion analysis engine identifies emotions such as the user being relaxed, angry, excited, etc. This emotion information is added to the analysis result.

[0310] Search and Generation of Information

[0311] The server first searches the internal database (common database management system) based on the content of the user's question to see if the necessary information exists. If the information does not exist in the internal database, it obtains the information from an external database (general-purpose API or external data source). For example, in the case of weather information, it obtains data from an external weather information API.

[0312] After the necessary information is obtained, the server uses a generative artificial intelligence (common generative AI technology) to generate an appropriate answer to the user's question. At this time, the response is adjusted based on the recognized emotion information. For example, if the user is angry, the answer is generated in a calm and gentle tone.

[0313] Voice Synthesis and Return of Response

[0314] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into audio data. This converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the response in voice.

[0315] Specific example

[0316] When a user asks "What's the weather like in Tokyo today?", the following process takes place.

[0317] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0318] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0319] 3. Analyze that the server is requesting weather information using an NLP engine.

[0320] 4. The server uses an emotion analysis engine to recognize the user's emotions (e.g., relaxed).

[0321] 5. The server queries the weather information API for Tokyo's weather.

[0322] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[0323] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0324] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0325] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0326] Step 1: Receiving voice input

[0327] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone. It receives the voice signal from the user as input and converts the analog voice signal into digital data through an A / D converter. The output here is digital voice data. Specifically, when the user says "What's the weather like in Tokyo today?", the microphone picks up the voice and converts it into digital data.

[0328] Step 2: Sending the audio data

[0329] After the digital audio data is generated, the terminal sends this data to the server via the internet. The terminal receives the digital audio data as input and sends the encrypted digital audio data to the server as output. Specifically, the terminal sends the digital audio data to the server using an encryption protocol such as SSL / TLS.

[0330] Step 3: Speech Recognition and Language Analysis

[0331] The server receives audio data and passes it to a speech recognition engine (general speech recognition software) to convert it into text data. It receives encrypted digital audio data as input and generates text data as output. Specifically, the server uses the speech recognition engine to convert "What's the weather like in Tokyo today?" into text data. Then, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. The input is text data, and the output is the analysis result. Specifically, the server passes the text data to the NLP engine, which analyzes that the user is requesting weather information.

[0332] Step 4: Adding emotion recognition

[0333] The server also passes the voice or text data to an emotion analysis engine (general emotion analysis software) to recognize the user's emotions. It receives voice or text data as input and generates emotion information as output. Specifically, the server uses the emotion analysis engine to identify emotions such as relaxed, angry, or excited based on voice tone and word choice. This emotion information is then added to the subsequent analysis results.

[0334] Step 5: Information retrieval and generation

[0335] Based on the analysis results and sentiment analysis results, the server first searches its internal database (a general database management system) for the necessary information. It receives the analysis results and sentiment information as input and retrieves the information present in the database as output. Specifically, the server searches its internal database for weather information queries, and if not found, it uses a weather information API (a general API) to retrieve information from an external database. Next, it uses generative artificial intelligence (a general generative AI technology) to generate an appropriate answer to the user's question. In this process, it adjusts the response based on the sentiment information. The input is the search results and sentiment information, and the output is the generated answer. Specifically, the server generates a relaxed-sounding answer such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[0336] Step 6: Speech synthesis and response return

[0337] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into speech data. The generated text response is received as input, and speech data is generated as output. Specifically, the server passes the text response to the speech synthesis engine and generates speech data. This converted speech data is sent from the server to the terminal. The terminal plays this speech data and provides the user with the answer to the question in speech. Speech data is received as input, and speech playback is performed as output. Specifically, the terminal plays the received speech data and says in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0338] (Application Example 2)

[0339] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0340] Conventional voice assistant systems can provide appropriate responses to user voice commands, but they often come across as mechanical and cold because they cannot take user emotions into consideration. Furthermore, in factory work instructions, disregarding the emotional state of workers makes it difficult to create an efficient and comfortable work environment.

[0341] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing emotional information from the user's voice and adjusting the response, and means for adjusting the speed and volume of the response to provide an appropriate response according to the user's emotions. This enables natural and human-like interaction that responds to the user's emotions, and in particular, it is possible to provide an appropriate response that takes into account the emotional state of the worker, even in work instructions within a factory.

[0342] "User voice" refers to the voice spoken by a human being to ask questions or give instructions to a system.

[0343] "Means of converting to digital data" refers to equipment or software that converts analog audio into digital data.

[0344] "Means of sending to a server" refers to equipment or software used to transmit digital data to a server via a network.

[0345] A "speech recognition engine" is software that converts received audio data into text data.

[0346] A "natural language processing engine" is software used to analyze the intent and content of text data.

[0347] An "internal database" is a database containing information data stored within a system.

[0348] An "external database" is a database used to retrieve necessary data from information sources outside the system.

[0349] "Generative artificial intelligence" is a type of artificial intelligence that generates appropriate answers to user questions.

[0350] A "speech synthesis engine" is software that converts generated text responses into speech data.

[0351] "Means of sending to a terminal and responding to the user" refers to a device or software that plays audio data sent from a server and provides a response to the user.

[0352] "Means for recognizing emotional information" refers to software or hardware that analyzes emotions from a user's voice.

[0353] "Means of adjusting responses" refers to software that modifies the content and method of responses based on recognized emotional information.

[0354] "Means for adjusting response speed and volume" refers to software or hardware that dynamically sets the speed and volume of voice responses according to the user's emotional state.

[0355] The system of this invention is designed to receive voice commands from a user and provide appropriate responses based on their emotions, using voice assistant technology. The system begins by converting the user's voice into digital data and sending it to a server. The server consists of a speech recognition engine, a natural language processing engine, an emotion recognition engine, an internal database, an external database, a generative artificial intelligence system, and a speech synthesis engine.

[0356] First, the user inputs audio via a microphone. The terminal captures this audio as an analog signal and converts it into a digital signal. This digital data is then transmitted to the server via the network.

[0357] On the server, the speech recognition engine first converts digital data into text data. This text data is then passed to the natural language processing engine, which analyzes the user's intent and the content of the question. Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information does not exist in the internal database, it can access the external database to retrieve the information.

[0358] Furthermore, an emotion recognition engine is used to analyze emotional information from the user's voice. This emotional information is added to the analysis results and used by the generative artificial intelligence to generate appropriate answers to the user's questions. For example, if the user is angry, the answer will be generated in a calm tone, and if they are relaxed, the answer will be generated in a natural tone.

[0359] The generated text response is converted into audio data by a speech synthesis engine and sent back to the terminal. The terminal plays this audio data, providing a response to the user. The response speed and volume are adjusted based on the user's emotions, which are analyzed by an emotion recognition engine.

[0360] For example, if a user instructs, "Please take this part to Unit 5," the system will recognize this voice, determine the emotion, and respond appropriately. Once an emotion is recognized, the system will adjust the speed and volume of the response accordingly. For instance, if the user is excited, it will respond quickly and loudly; if the user is relaxed, it will provide a response at a normal speed and standard volume.

[0361] Examples of prompt statements include the following:

[0362] User voice input: "Please take this part to Unit 5."

[0363] Output: "We will transport the parts. Where would you like us to transport them?"

[0364] In this way, the system provides natural and human-like interactions that respond to the user's emotions, and can provide appropriate responses that take into account the emotional state of the worker, especially when giving work instructions in a factory.

[0365] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0366] Step 1:

[0367] The user provides voice input via a microphone. For example, they might say, "Please take this part to unit number 5." This analog voice signal is captured by the terminal's microphone.

[0368] Input: User's analog audio signal

[0369] Output: Captured analog audio signal

[0370] Operation: The user gives instructions to the system via the microphone.

[0371] Step 2:

[0372] The terminal converts the captured analog audio signal into digital data. This digital data must be in a format that the speech recognition engine can process.

[0373] Input: Captured analog audio signal

[0374] Output: Digital audio data

[0375] Operation: The device's audio conversion module converts analog audio signals into digital data.

[0376] Step 3:

[0377] The terminal sends the converted digital audio data to the server via the network. The server receives this data.

[0378] Input: Digital audio data

[0379] Output: Digital audio data transferred to the server

[0380] Operation: The device sends digital audio data to the server.

[0381] Step 4:

[0382] The server uses a speech recognition engine to convert digital speech data into text data. For example, it might generate text data such as, "Please take this part to Unit 5."

[0383] Input: Digital audio data

[0384] Output: Text data

[0385] Operation: The server's speech recognition engine converts digital speech data into text data.

[0386] Step 5:

[0387] The server passes this text data to a natural language processing engine, which analyzes the user's intent and the content of the question. For example, it might analyze that the user is giving instructions to transport parts.

[0388] Input: Text data

[0389] Output: Analyzed intent and content

[0390] Operation: The server's natural language processing engine analyzes the text data to identify the user's intent.

[0391] Step 6:

[0392] The server uses an emotion engine to recognize emotional information from the user's voice. For example, it can determine whether the user is excited or calm.

[0393] Input: Text data

[0394] Output: User sentiment information

[0395] Operation: The server's emotion engine analyzes and extracts emotional information from the audio data.

[0396] Step 7:

[0397] Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information is not found in the internal database, it retrieves it from the external database.

[0398] Input: Analyzed intent and content, user sentiment information

[0399] Output: Required information

[0400] Operation: The server searches for the necessary information from internal and external databases.

[0401] Step 8:

[0402] Generative artificial intelligence is used to generate appropriate answers to user questions. During this process, the content and tone of the responses are adjusted based on emotional information.

[0403] Input: Required information, user sentiment information

[0404] Output: Generated text answer

[0405] Operation: The server's generational artificial intelligence generates the appropriate answer.

[0406] Step 9:

[0407] The generated text response is passed to a speech synthesis engine and converted into speech data. For example, speech data such as "I will transport the parts. Where do you want me to transport them?" is generated.

[0408] Input: Generated text answer

[0409] Output: Audio data

[0410] Operation: The server's speech synthesis engine converts text data into speech data.

[0411] Step 10:

[0412] The server sends audio data to the terminal, which then plays this audio data to respond to the user. The response speed and volume are adjusted according to the user's emotions.

[0413] Input: Audio data

[0414] Output: Audio response played to the user

[0415] Operation: The device plays audio data and provides a response to the user.

[0416] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0417] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0418] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0419] [Second Embodiment]

[0420] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0421] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0422] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0423] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0424] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0425] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0426] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0427] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0428] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0429] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0430] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0431] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0432] The system according to the present invention is a voice assistant system that caters to a diverse range of users and provides appropriate answers to questions about information in a wide range of fields. This system encompasses a series of processes that analyze voice input and generate appropriate answers, and the details of its operation are as follows.

[0433] Receiving voice input

[0434] The user speaks a question to the voice assistant. The device captures this audio through the microphone and converts the analog audio into digital data. The converted audio data is then sent from the device to the server.

[0435] Speech recognition and language analysis

[0436] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. The converted text data is then passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[0437] Information retrieval and generation

[0438] Based on the analysis results, the server first searches its internal database for the necessary information. If the information is not found in the internal database, it retrieves it from an external database. This process may utilize various external APIs.

[0439] After obtaining the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. This generated text answer is then passed to a speech synthesis engine and converted into speech data.

[0440] Speech synthesis and response return

[0441] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[0442] Specific example

[0443] Example 1: Weather question

[0444] When a user asks, "What's the weather like in Tokyo today?",

[0445] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0446] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0447] 3. Analyze that the server is requesting weather information using an NLP engine.

[0448] 4. The server queries the weather information API for the weather in Tokyo.

[0449] 5. The server retrieves the information "Today in Tokyo it is sunny and the maximum temperature is 25 degrees Celsius," and a generative AI generates an appropriate response.

[0450] 6. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0451] 7. The device plays the answer and says, "Today in Tokyo it's sunny and the high temperature is 25 degrees Celsius."

[0452] Example 2: History Questions

[0453] When a user asks, "When was Alexander the Great born?",

[0454] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0455] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[0456] 3. Analyze that the server is requesting historical information using the NLP engine.

[0457] 4. The server searches its internal database and confirms the relevant information.

[0458] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and generates an answer using a generative AI.

[0459] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0460] 8. The device plays the answer and says, "Alexander the Great was born in 356 BC."

[0461] A system configured in this way can provide users with answers from a wide range of sources quickly and accurately, and can accommodate diverse user groups such as seniors and children.

[0462] The following describes the processing flow.

[0463] Step 1:

[0464] The user speaks a question to the voice assistant. An example of such a question is, "What's the weather like in Tokyo today?"

[0465] Step 2:

[0466] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. This digital audio data is then encoded into a format suitable for speech recognition.

[0467] Step 3:

[0468] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[0469] Step 4:

[0470] The server receives the audio data. The received audio data is then passed to the speech recognition engine.

[0471] Step 5:

[0472] The server uses a speech recognition engine to convert the audio data into text data. For example, text data such as "What's the weather like in Tokyo today?" is generated.

[0473] Step 6:

[0474] The server passes the generated text data to a natural language processing (NLP) engine, which analyzes the intent and content of the question. The NLP engine then identifies the category of the question (e.g., weather information, historical information, etc.).

[0475] Step 7:

[0476] Based on the analysis results, the server searches its internal database to check if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[0477] Step 8:

[0478] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[0479] Step 9:

[0480] Based on the information acquired by the server, generative artificial intelligence (e.g., a large-scale language model) is used to generate natural-sounding answers to the user's questions. For example, an answer such as "Today in Tokyo it is sunny and the highest temperature is 25 degrees Celsius" might be generated.

[0481] Step 10:

[0482] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[0483] Step 11:

[0484] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[0485] Step 12:

[0486] The device receives the audio data and plays it back. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[0487] This series of steps allows users to receive accurate and rapid voice answers to their questions.

[0488] (Example 1)

[0489] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0490] Conventional systems using speech recognition and natural language processing have limitations in their ability to access internal and external databases when searching for specific information, making it difficult to provide flexible and comprehensive information. Furthermore, the lack of sufficient technology to accurately analyze user intent and quickly retrieve corresponding information results in a difficulty in improving the user experience.

[0491] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0492] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to a network; means for converting the digital data into text data using a speech recognition engine on a computer on the network; means for passing the text data to a natural language processing engine to analyze its intent and content; means for retrieving necessary information from internal and external storage devices based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; and means for transmitting the voice data to a terminal and responding to the user. This makes it possible to provide users with information from a wide range of sources quickly and accurately, and to provide appropriate answers to user questions in voice.

[0493] A "user" is an individual or organization that uses the system to input questions by voice.

[0494] A "network" is a means of communication for sending and receiving digital data between various computers.

[0495] A "speech recognition engine" is a software or hardware system that analyzes input speech data and converts it into text data.

[0496] "Text data" refers to digital data in sentence format converted by a speech recognition engine.

[0497] A "natural language processing engine" is a software or hardware system that analyzes text data and understands the user's intent and content from it.

[0498] An "internal storage device" is a device installed within a system for storing and retrieving data.

[0499] "External storage devices" are devices or services that exist outside the system and are used to store and retrieve data.

[0500] "Generative artificial intelligence" is a system of algorithms or software that generates appropriate answers based on questions from users.

[0501] A "speech synthesis engine" is a software or hardware system that converts generated text data into speech data.

[0502] A "terminal" is a device used by a user to input voice, receive audio data, and play it back.

[0503] The system according to the present invention is a voice assistant system that analyzes the user's voice input, generates an appropriate response, and provides a voice reply. Specific embodiments of the present invention are described below.

[0504] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the audio using its built-in microphone and converts this analog audio data into digital data. The converted audio data is then sent to the server in binary format.

[0505] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0506] Next, this text data is passed to a natural language processing (NLP) engine (e.g., a typical NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information.

[0507] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve the current weather information for Tokyo.

[0508] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI model, e.g., a general-purpose generative AI model) to generate an appropriate answer to the user's question. For example, it might generate the sentence, "Today in Tokyo it is sunny, and the high temperature is 25 degrees Celsius." The prompt used would be: "Today's weather in Tokyo is sunny, and the high temperature is 25 degrees Celsius."

[0509] The generated text response is passed to a speech synthesis engine (e.g., general speech synthesis software) and converted into audio data. The speech synthesis engine converts the generated text data into an audio format (e.g., MP3) and returns that audio data to the server.

[0510] The server sends the returned audio data to the terminal. The terminal plays the received audio data through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the highest temperature is 25 degrees Celsius."

[0511] This system allows users to quickly and accurately obtain information from a wide range of sources simply by asking questions by voice. Furthermore, the system can accommodate diverse user groups, including seniors and children.

[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0513] Step 1: Receiving voice input

[0514] The user speaks a question into the microphone. The device uses its built-in microphone to capture the audio and converts this analog audio data into digital data. Specifically, the device's audio capture function is activated, the audio signal is converted to a digital format according to the sampling rate, and stored in a buffer in binary format.

[0515] Input: User's analog voice

[0516] Output: Digital audio data (binary format)

[0517] Step 2: Sending the audio data

[0518] The terminal divides the digital audio data stored in the buffer into packets of a fixed size and sends them to the server via secure and high-speed network communication (e.g., HTTPS). HTTP POST requests are used to send the audio data to the server-side API endpoint.

[0519] Input: Digital audio data (binary format)

[0520] Output: Request to send audio data to the server

[0521] Step 3: Speech recognition and text conversion

[0522] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. Specifically, it performs audio waveform analysis and uses a phonological and lexical model to convert the audio into corresponding text. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0523] Input: Audio data (binary format)

[0524] Output: Text data

[0525] Step 4: Text analysis and intent identification

[0526] The server passes the converted text data to a natural language processing (NLP) engine (e.g., general NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information. The analysis results are stored in JSON format.

[0527] Input: Text data

[0528] Output: Analysis result (JSON format)

[0529] Step 5: Information Search

[0530] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve current weather information for Tokyo. Here, it uses an HTTP request to query the external database and receives a response in JSON format.

[0531] Input: Analysis result (JSON format)

[0532] Output: Acquired information (JSON format)

[0533] Step 6: Generating the answer

[0534] Based on the weather information it receives, the server uses generative artificial intelligence (generative AI model, e.g., a general generative AI model) to generate appropriate answers to the user's questions. Specifically, it inputs the prompt "What's the weather like in Tokyo today?" into the generative AI model and generates the answer "It's sunny in Tokyo today, and the maximum temperature is 25 degrees Celsius."

[0535] Input: Acquired information (in JSON format) and prompt message

[0536] Output: Generated response (text data)

[0537] Step 7: Speech Synthesis

[0538] The server passes the generated text to a speech synthesis engine (e.g., common speech synthesis software) and converts it into audio data. The speech synthesis engine then converts the generated text data into an audio format (e.g., MP3). This conversion uses a speech model and involves acoustic processing to produce natural-sounding speech.

[0539] Input: Generated response (text data)

[0540] Output: Audio data (MP3 format)

[0541] Step 8: Return of audio data and response playback

[0542] The server sends the returned audio data to the terminal as an HTTP response. The terminal stores the received audio data in a buffer and plays it back through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the maximum temperature is 25 degrees Celsius."

[0543] Input: Audio data (MP3 format)

[0544] Output: Voice response to the user

[0545] (Application Example 1)

[0546] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0547] In physical stores, there is a need for a system that allows customers to receive quick and appropriate answers to their questions regarding product information, store layout, and inventory. Traditional methods require customers to ask store staff directly or search for store signs themselves, which is inconvenient and time-consuming. In particular, when customers ask a variety of questions, responding to them takes time, making it difficult to provide satisfactory service.

[0548] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0549] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to the server; means for the server to convert the digital data into text data using a speech recognition engine; means for passing the text data to a natural language processing engine to analyze its intent and content; means for searching for necessary information from an internal database and an external database based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; means for transmitting the voice data to a terminal and responding to the user; and means for providing quick and appropriate voice answers to questions from customers in physical stores regarding product information, store guidance, inventory checks, etc. This enables users to quickly obtain necessary information through a voice assistant no matter where they are in the store.

[0550] A "user" is someone who uses a voice assistant system to ask questions or give instructions.

[0551] "Digital data" refers to data obtained by converting analog audio signals into a digital format.

[0552] A "server" is a computer system used for speech recognition, natural language processing, database retrieval, generative artificial intelligence, and speech synthesis.

[0553] A "speech recognition engine" is a program or function that analyzes speech data and converts it into text data.

[0554] "Text data" refers to character information generated by a speech recognition engine.

[0555] A "natural language processing engine" is a program or function that analyzes text data to understand the user's intent and content.

[0556] An "internal database" is a database stored within the server where the information to be searched is stored.

[0557] An "external database" is a database that exists outside the server and can be accessed via an API.

[0558] "Generative artificial intelligence" is artificial intelligence that generates appropriate answers based on analyzed information.

[0559] A "speech synthesis engine" is a program or function that converts text data into speech data.

[0560] A "device" is a device used by a user to operate a voice assistant system, and includes smartphones and smart glasses.

[0561] A "physical store" refers to a sales facility or service provider that actually exists in a physical space.

[0562] "Product information" refers to detailed product descriptions, prices, and stock availability information provided within a physical store.

[0563] "Prompt and appropriate responses" refer to providing accurate information immediately in response to user questions.

[0564] Using the above definitions, we will clarify each function and element of the voice assistant system.

[0565] The system according to this invention is a voice assistant system that provides quick and appropriate answers to questions from users in physical stores. Specific embodiments thereof will be described below.

[0566] System Configuration

[0567] 1. User voice input

[0568] The user inputs their questions by voice into a device such as a smartphone or smart glasses. The voice input is captured through the device's built-in microphone.

[0569] 2. Digital conversion and transmission of audio

[0570] The terminal converts the captured analog audio into digital data. The converted digital data is then sent to the server via the network.

[0571] 3. Speech Recognition and Natural Language Processing

[0572] The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert the received digital audio data into text data. Next, it passes this text data to a natural language processing engine (e.g., an NLP engine) to analyze the intent and content of the user's question.

[0573] 4. Information Retrieval and Generation

[0574] The server searches for necessary information from internal and external databases based on the analysis results. It uses external APIs as needed to retrieve information from external databases. This search result is then used with generative artificial intelligence (generative AI model) to generate appropriate answers.

[0575] 5. Speech synthesis and response

[0576] The generated response is converted into audio data using a speech synthesis engine (e.g., gTTS). The converted audio data is then sent back to the terminal via the network, where it plays the audio data and responds to the user.

[0577] Specific usage examples

[0578] Example 1: Product Information

[0579] The user asks, "Where can I find this product?" inside the store.

[0580] The voice assistant system captures this question and sends it to the server.

[0581] The server converts the speech into text using a speech recognition engine, and then analyzes it as "product information" using a natural language processing engine.

[0582] The server retrieves the product's location information from an internal database and uses generative artificial intelligence to generate a response such as, "This product is located next to the cash register on the third floor."

[0583] The generated response is converted by a speech synthesis engine and sent to the terminal.

[0584] The device plays a message to the user saying, "This item is located next to the cash register on the 3rd floor."

[0585] Usage example 2: Inventory check

[0586] The user asks, "Do you have this item in stock?"

[0587] The voice assistant system captures the user's question and sends it to the server.

[0588] The server uses a speech recognition engine to convert the speech into text data, and then a natural language processing engine analyzes it to read "check inventory".

[0589] The server uses an external API to retrieve the latest inventory information and generates responses such as "This product is still in stock" using generative artificial intelligence.

[0590] The generated response is converted into audio data by a speech synthesis engine and sent to the terminal.

[0591] The device plays the product and informs the user that "this item is still in stock."

[0592] Examples of prompt statements

[0593] Example of a prompt message for product information:

[0594] "The user is asking for the location information of this product within the store. Please provide the specific location of the product."

[0595] Example prompt message for checking inventory:

[0596] "A user is asking about the availability of a specific product. Please provide the stock status for this product."

[0597] Through the above explanation, we have provided specific examples of voice assistant systems and clarified methods for improving user convenience in physical stores.

[0598] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0599] Step 1:

[0600] The user inputs a question by voice. The device's microphone captures the voice and converts the analog audio signal into digital data. This digital data is transmitted to the server via the network. The input is the user's voice, and the output is digital audio data.

[0601] Step 2:

[0602] The server receives digital data and passes it to a speech recognition engine (e.g., Google Speech Recognition API), which converts the audio into text data. This conversion process analyzes the frequency information of the audio data and generates the corresponding string. The input is digital audio data, and the output is text data.

[0603] Step 3:

[0604] The server passes text data to a natural language processing (NLP) engine, which analyzes the intent and content of the user's question. The NLP engine grammatically and semantically analyzes the text data to identify which information the user's question relates to. The input is text data, and the output is the analysis result.

[0605] Step 4:

[0606] The server searches for the necessary information from internal and external databases based on the analysis results. The internal database contains product information and in-store guidance, while the external database contains the latest information accessible via API. It generates database queries and retrieves information based on them. The input is the analysis results, and the output is the necessary information.

[0607] Step 5:

[0608] Based on the information acquired by the server, a generative artificial intelligence (generative AI model) is used to generate appropriate answers to the user's questions. In this process, the generative AI model automatically generates answers based on the prompt text. The input consists of the required information and the prompt text, and the output is the generated text answer.

[0609] Step 6:

[0610] The generated text response is passed to a speech synthesis engine (e.g., gTTS) and converted into audio data. The speech synthesis engine then converts the text data into an audio file. The input is the generated text response, and the output is the audio data.

[0611] Step 7:

[0612] The server sends audio data to the terminal. The terminal plays the received audio data and provides the answer to the user verbally. The input is audio data, and the output is the verbal answer to the user.

[0613] The above outlines the specific processing steps in the invention's system. This series of processes enables users to quickly obtain necessary information via a voice assistant, regardless of their location within the store.

[0614] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0615] The system according to the present invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. The details of the operation of the present invention are as follows.

[0616] Receiving voice input

[0617] The user speaks a question to the voice assistant. An example question might be, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. This digital audio data is then sent from the device to the server.

[0618] Speech recognition and language analysis

[0619] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. For example, the text data "What's the weather like in Tokyo today?" is generated. This text data is passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[0620] Addition of emotion recognition

[0621] In the process described above, the server uses an emotion engine to recognize emotional information from the user's voice. For example, if the user is angry, the emotion engine recognizes that anger and adds it to the analysis results.

[0622] Information retrieval and generation

[0623] Based on the analysis results, the server first searches its internal database for the necessary information. If the necessary information is not found in the internal database, it retrieves it from an external database. Various external APIs may be used in this process.

[0624] After acquiring the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone. This generated text response is then passed to a speech synthesis engine and converted into audio data.

[0625] Speech synthesis and response return

[0626] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[0627] Specific example

[0628] Example 1: Weather question

[0629] When a user asks, "What's the weather like in Tokyo today?",

[0630] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0631] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0632] 3. Analyze that the server is requesting weather information using an NLP engine.

[0633] 4. The server uses an emotion engine to recognize the user's emotions (e.g., relaxed).

[0634] 5. The server queries the weather information API for the weather in Tokyo.

[0635] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[0636] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0637] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0638] Example 2: History Questions

[0639] When a user asks, "When was Alexander the Great born?",

[0640] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0641] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[0642] 3. Analyze that the server is requesting historical information using the NLP engine.

[0643] 4. The server uses an emotion engine to recognize the user's emotions (e.g., excited).

[0644] 5. The server searches its internal database and confirms the relevant information.

[0645] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and uses a generative AI to generate an excited response.

[0646] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0647] 8. The device plays the answer back, excitedly stating, "Alexander the Great was born in 356 BC."

[0648] A system configured in this way can quickly and accurately provide answers from a wide range of sources, including appropriate responses tailored to the user's emotions. This allows users to experience more satisfying interactions.

[0649] The following describes the processing flow.

[0650] Step 1:

[0651] The user speaks a question to the voice assistant, such as, "What's the weather like in Tokyo today?"

[0652] Step 2:

[0653] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. The converted audio data is sent to a server for processing.

[0654] Step 3:

[0655] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[0656] Step 4:

[0657] The server receives the transmitted audio data. The received audio data is then passed to the speech recognition engine.

[0658] Step 5:

[0659] The server uses a speech recognition engine to convert the audio data into text data. For example, the text "What's the weather like in Tokyo today?" is generated.

[0660] Step 6:

[0661] The server passes the generated text data to a natural language processing (NLP) engine to analyze the intent and content of the question. Based on the analysis results, it identifies that the question is related to weather information.

[0662] Step 7:

[0663] The server uses an emotion engine to analyze the user's emotional state from the voice data. For example, it can determine whether the user is speaking in a relaxed state or is angry.

[0664] Step 8:

[0665] Based on the analysis results, the server first searches its internal database to see if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[0666] Step 9:

[0667] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[0668] Step 10:

[0669] Based on the information acquired by the server, generative artificial intelligence is used to generate answers to the user's questions. In this process, the tone and content of the answer are adjusted to reflect recognized emotional information. For example, an answer such as "Today in Tokyo it is sunny, and the highest temperature is 25 degrees Celsius" might be generated.

[0670] Step 11:

[0671] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[0672] Step 12:

[0673] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[0674] Step 13:

[0675] The device receives and plays the received audio data. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius." This response is delivered in an appropriate tone depending on the user's mood.

[0676] (Example 2)

[0677] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0678] Conventional voice assistant systems have been required to provide quick and accurate responses to user utterances, but they have faced challenges in natural communication with users, particularly due to their inability to respond with consideration for emotions. Furthermore, when the information corresponding to a user's question is not found in the internal database, the means of retrieving information from external databases are limited, making it difficult to provide an appropriate answer.

[0679] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0680] In this invention, the server includes means for converting digital data into text data using a speech recognition engine, means for analyzing intent and content using a natural language processing engine, and means for analyzing the user's emotions. This enables natural communication that takes emotions into consideration. Furthermore, by providing means for retrieving information from an external database when the information does not exist in the internal database, it becomes possible to provide quick and appropriate answers to a wide range of inquiries.

[0681] A "speech recognition engine" is a software or hardware technology used to convert speech data into text data.

[0682] A "natural language processing engine" is a software or hardware technology used to analyze user intent and content from text data.

[0683] "Sentiment analysis" is a technology that extracts emotional information from a user's voice and text data to recognize the user's state of mind.

[0684] An "internal database" is a database that stores information managed within a system.

[0685] An "external database" is a database that provides information located outside of the system.

[0686] "Generative artificial intelligence" is an artificial intelligence technology used to generate answers to user questions.

[0687] A "speech synthesis engine" is a software or hardware technology used to convert text data into speech data.

[0688] A "terminal" is a device used by a user to access a voice assistant system.

[0689] "Emotional information" refers to data that indicates the emotional state expressed by the user.

[0690] SSL / TLS is a protocol for encrypting and securely conducting data communications.

[0691] The system according to this invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. Details of the system for carrying out this invention are described below.

[0692] Receiving voice input

[0693] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. An A / D converter is used for this conversion process.

[0694] Sending audio data

[0695] After the digital audio data is generated, the device sends this data to the server via the internet. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security.

[0696] Speech recognition and language analysis

[0697] The server passes the received audio data to a speech recognition engine (general speech recognition software) and converts it into text data. For example, commonly used speech recognition technologies are used for the speech recognition engine. Next, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. Based on the results of this analysis, the server identifies what information is needed.

[0698] Addition of emotion recognition

[0699] The server passes voice or text data to an emotion analysis engine (common emotion analysis software) to recognize the user's emotional information. For example, the emotion analysis engine identifies emotions such as the user being relaxed, angry, or excited. This emotional information is added to the analysis results.

[0700] Information retrieval and generation

[0701] Based on the user's query, the server first searches its internal database (a general-purpose database management system) to see if the necessary information exists. If the information is not found in the internal database, it retrieves it from an external database (a general-purpose API or external data source). For example, in the case of weather information, it retrieves data from an external weather information API.

[0702] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI technology) to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone.

[0703] Speech synthesis and response return

[0704] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into audio data. This converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the response in voice.

[0705] Specific example

[0706] When a user asks "What's the weather like in Tokyo today?", the following process takes place.

[0707] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0708] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0709] 3. Analyze that the server is requesting weather information using an NLP engine.

[0710] 4. The server uses an emotion analysis engine to recognize the user's emotions (e.g., relaxed).

[0711] 5. The server queries the weather information API for Tokyo's weather.

[0712] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[0713] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0714] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0715] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0716] Step 1: Receiving voice input

[0717] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone. It receives the voice signal from the user as input and converts the analog voice signal into digital data through an A / D converter. The output here is digital voice data. Specifically, when the user says "What's the weather like in Tokyo today?", the microphone picks up the voice and converts it into digital data.

[0718] Step 2: Sending the audio data

[0719] After the digital audio data is generated, the terminal sends this data to the server via the internet. The terminal receives the digital audio data as input and sends the encrypted digital audio data to the server as output. Specifically, the terminal sends the digital audio data to the server using an encryption protocol such as SSL / TLS.

[0720] Step 3: Speech Recognition and Language Analysis

[0721] The server receives audio data and passes it to a speech recognition engine (general speech recognition software) to convert it into text data. It receives encrypted digital audio data as input and generates text data as output. Specifically, the server uses the speech recognition engine to convert "What's the weather like in Tokyo today?" into text data. Then, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. The input is text data, and the output is the analysis result. Specifically, the server passes the text data to the NLP engine, which analyzes that the user is requesting weather information.

[0722] Step 4: Adding emotion recognition

[0723] The server also passes the voice or text data to an emotion analysis engine (general emotion analysis software) to recognize the user's emotions. It receives voice or text data as input and generates emotion information as output. Specifically, the server uses the emotion analysis engine to identify emotions such as relaxed, angry, or excited based on voice tone and word choice. This emotion information is then added to the subsequent analysis results.

[0724] Step 5: Information retrieval and generation

[0725] Based on the analysis results and sentiment analysis results, the server first searches its internal database (a general database management system) for the necessary information. It receives the analysis results and sentiment information as input and retrieves the information present in the database as output. Specifically, the server searches its internal database for weather information queries, and if not found, it uses a weather information API (a general API) to retrieve information from an external database. Next, it uses generative artificial intelligence (a general generative AI technology) to generate an appropriate answer to the user's question. In this process, it adjusts the response based on the sentiment information. The input is the search results and sentiment information, and the output is the generated answer. Specifically, the server generates a relaxed-sounding answer such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[0726] Step 6: Speech synthesis and response return

[0727] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into speech data. The generated text response is received as input, and speech data is generated as output. Specifically, the server passes the text response to the speech synthesis engine and generates speech data. This converted speech data is sent from the server to the terminal. The terminal plays this speech data and provides the user with the answer to the question in speech. Speech data is received as input, and speech playback is performed as output. Specifically, the terminal plays the received speech data and says in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[0728] (Application Example 2)

[0729] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0730] Conventional voice assistant systems can provide appropriate responses to user voice commands, but they often come across as mechanical and cold because they cannot take user emotions into consideration. Furthermore, in factory work instructions, disregarding the emotional state of workers makes it difficult to create an efficient and comfortable work environment.

[0731] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing emotional information from the user's voice and adjusting the response, and means for adjusting the speed and volume of the response to provide an appropriate response according to the user's emotions. This enables natural and human-like interaction that responds to the user's emotions, and in particular, it is possible to provide an appropriate response that takes into account the emotional state of the worker, even in work instructions within a factory.

[0732] "User voice" refers to the voice spoken by a human being to ask questions or give instructions to a system.

[0733] "Means of converting to digital data" refers to equipment or software that converts analog audio into digital data.

[0734] "Means of sending to a server" refers to equipment or software used to transmit digital data to a server via a network.

[0735] A "speech recognition engine" is software that converts received audio data into text data.

[0736] A "natural language processing engine" is software used to analyze the intent and content of text data.

[0737] An "internal database" is a database containing information data stored within a system.

[0738] An "external database" is a database used to retrieve necessary data from information sources outside the system.

[0739] "Generative artificial intelligence" is a type of artificial intelligence that generates appropriate answers to user questions.

[0740] A "speech synthesis engine" is software that converts generated text responses into speech data.

[0741] "Means of sending to a terminal and responding to the user" refers to a device or software that plays audio data sent from a server and provides a response to the user.

[0742] "Means for recognizing emotional information" refers to software or hardware that analyzes emotions from a user's voice.

[0743] "Means of adjusting responses" refers to software that modifies the content and method of responses based on recognized emotional information.

[0744] "Means for adjusting response speed and volume" refers to software or hardware that dynamically sets the speed and volume of voice responses according to the user's emotional state.

[0745] The system of this invention is designed to receive voice commands from a user and provide appropriate responses based on their emotions, using voice assistant technology. The system begins by converting the user's voice into digital data and sending it to a server. The server consists of a speech recognition engine, a natural language processing engine, an emotion recognition engine, an internal database, an external database, a generative artificial intelligence system, and a speech synthesis engine.

[0746] First, the user inputs audio via a microphone. The terminal captures this audio as an analog signal and converts it into a digital signal. This digital data is then transmitted to the server via the network.

[0747] On the server, the speech recognition engine first converts digital data into text data. This text data is then passed to the natural language processing engine, which analyzes the user's intent and the content of the question. Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information does not exist in the internal database, it can access the external database to retrieve the information.

[0748] Furthermore, an emotion recognition engine is used to analyze emotional information from the user's voice. This emotional information is added to the analysis results and used by the generative artificial intelligence to generate appropriate answers to the user's questions. For example, if the user is angry, the answer will be generated in a calm tone, and if they are relaxed, the answer will be generated in a natural tone.

[0749] The generated text response is converted into audio data by a speech synthesis engine and sent back to the terminal. The terminal plays this audio data, providing a response to the user. The response speed and volume are adjusted based on the user's emotions, which are analyzed by an emotion recognition engine.

[0750] For example, if a user instructs, "Please take this part to Unit 5," the system will recognize this voice, determine the emotion, and respond appropriately. Once an emotion is recognized, the system will adjust the speed and volume of the response accordingly. For instance, if the user is excited, it will respond quickly and loudly; if the user is relaxed, it will provide a response at a normal speed and standard volume.

[0751] Examples of prompt statements include the following:

[0752] User voice input: "Please take this part to Unit 5."

[0753] Output: "We will transport the parts. Where would you like us to transport them?"

[0754] In this way, the system provides natural and human-like interactions that respond to the user's emotions, and can provide appropriate responses that take into account the emotional state of the worker, especially when giving work instructions in a factory.

[0755] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0756] Step 1:

[0757] The user provides voice input via a microphone. For example, they might say, "Please take this part to unit number 5." This analog voice signal is captured by the terminal's microphone.

[0758] Input: User's analog audio signal

[0759] Output: Captured analog audio signal

[0760] Operation: The user gives instructions to the system via the microphone.

[0761] Step 2:

[0762] The terminal converts the captured analog audio signal into digital data. This digital data must be in a format that the speech recognition engine can process.

[0763] Input: Captured analog audio signal

[0764] Output: Digital audio data

[0765] Operation: The device's audio conversion module converts analog audio signals into digital data.

[0766] Step 3:

[0767] The terminal sends the converted digital audio data to the server via the network. The server receives this data.

[0768] Input: Digital audio data

[0769] Output: Digital audio data transferred to the server

[0770] Operation: The device sends digital audio data to the server.

[0771] Step 4:

[0772] The server uses a speech recognition engine to convert digital speech data into text data. For example, it might generate text data such as, "Please take this part to Unit 5."

[0773] Input: Digital audio data

[0774] Output: Text data

[0775] Operation: The server's speech recognition engine converts digital speech data into text data.

[0776] Step 5:

[0777] The server passes this text data to a natural language processing engine, which analyzes the user's intent and the content of the question. For example, it might analyze that the user is giving instructions to transport parts.

[0778] Input: Text data

[0779] Output: Analyzed intent and content

[0780] Operation: The server's natural language processing engine analyzes the text data to identify the user's intent.

[0781] Step 6:

[0782] The server uses an emotion engine to recognize emotional information from the user's voice. For example, it can determine whether the user is excited or calm.

[0783] Input: Text data

[0784] Output: User sentiment information

[0785] Operation: The server's emotion engine analyzes and extracts emotional information from the audio data.

[0786] Step 7:

[0787] Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information is not found in the internal database, it retrieves it from the external database.

[0788] Input: Analyzed intent and content, user sentiment information

[0789] Output: Required information

[0790] Operation: The server searches for the necessary information from internal and external databases.

[0791] Step 8:

[0792] Generative artificial intelligence is used to generate appropriate answers to user questions. During this process, the content and tone of the responses are adjusted based on emotional information.

[0793] Input: Required information, user sentiment information

[0794] Output: Generated text answer

[0795] Operation: The server's generational artificial intelligence generates the appropriate answer.

[0796] Step 9:

[0797] The generated text response is passed to a speech synthesis engine and converted into speech data. For example, speech data such as "I will transport the parts. Where do you want me to transport them?" is generated.

[0798] Input: Generated text answer

[0799] Output: Audio data

[0800] Operation: The server's speech synthesis engine converts text data into speech data.

[0801] Step 10:

[0802] The server sends audio data to the terminal, which then plays this audio data to respond to the user. The response speed and volume are adjusted according to the user's emotions.

[0803] Input: Audio data

[0804] Output: Audio response played to the user

[0805] Operation: The device plays audio data and provides a response to the user.

[0806] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0807] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0809] [Third Embodiment]

[0810] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0811] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0812] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0813] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0814] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0815] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0816] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0817] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0818] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0819] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0820] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0821] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0822] The system according to the present invention is a voice assistant system that caters to a diverse range of users and provides appropriate answers to questions about information in a wide range of fields. This system encompasses a series of processes that analyze voice input and generate appropriate answers, and the details of its operation are as follows.

[0823] Receiving voice input

[0824] The user speaks a question to the voice assistant. The device captures this audio through the microphone and converts the analog audio into digital data. The converted audio data is then sent from the device to the server.

[0825] Speech recognition and language analysis

[0826] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. The converted text data is then passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[0827] Information retrieval and generation

[0828] Based on the analysis results, the server first searches its internal database for the necessary information. If the information is not found in the internal database, it retrieves it from an external database. This process may utilize various external APIs.

[0829] After obtaining the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. This generated text answer is then passed to a speech synthesis engine and converted into speech data.

[0830] Speech synthesis and response return

[0831] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[0832] Specific example

[0833] Example 1: Weather question

[0834] When a user asks, "What's the weather like in Tokyo today?",

[0835] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0836] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[0837] 3. Analyze that the server is requesting weather information using an NLP engine.

[0838] 4. The server queries the weather information API for the weather in Tokyo.

[0839] 5. The server retrieves the information "Today in Tokyo it is sunny and the maximum temperature is 25 degrees Celsius," and a generative AI generates an appropriate response.

[0840] 6. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0841] 7. The device plays the answer and says, "Today in Tokyo it's sunny and the high temperature is 25 degrees Celsius."

[0842] Example 2: History Questions

[0843] When a user asks, "When was Alexander the Great born?",

[0844] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[0845] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[0846] 3. Analyze that the server is requesting historical information using the NLP engine.

[0847] 4. The server searches its internal database and confirms the relevant information.

[0848] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and generates an answer using a generative AI.

[0849] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[0850] 8. The device plays the answer and says, "Alexander the Great was born in 356 BC."

[0851] A system configured in this way can provide users with answers from a wide range of sources quickly and accurately, and can accommodate diverse user groups such as seniors and children.

[0852] The following describes the processing flow.

[0853] Step 1:

[0854] The user speaks a question to the voice assistant. An example of such a question is, "What's the weather like in Tokyo today?"

[0855] Step 2:

[0856] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. This digital audio data is then encoded into a format suitable for speech recognition.

[0857] Step 3:

[0858] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[0859] Step 4:

[0860] The server receives the audio data. The received audio data is then passed to the speech recognition engine.

[0861] Step 5:

[0862] The server uses a speech recognition engine to convert the audio data into text data. For example, text data such as "What's the weather like in Tokyo today?" is generated.

[0863] Step 6:

[0864] The server passes the generated text data to a natural language processing (NLP) engine, which analyzes the intent and content of the question. The NLP engine then identifies the category of the question (e.g., weather information, historical information, etc.).

[0865] Step 7:

[0866] Based on the analysis results, the server searches its internal database to check if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[0867] Step 8:

[0868] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[0869] Step 9:

[0870] Based on the information acquired by the server, generative artificial intelligence (e.g., a large-scale language model) is used to generate natural-sounding answers to the user's questions. For example, an answer such as "Today in Tokyo it is sunny and the highest temperature is 25 degrees Celsius" might be generated.

[0871] Step 10:

[0872] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[0873] Step 11:

[0874] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[0875] Step 12:

[0876] The device receives the audio data and plays it back. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[0877] This series of steps allows users to receive accurate and rapid voice answers to their questions.

[0878] (Example 1)

[0879] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0880] Conventional systems using speech recognition and natural language processing have limitations in their ability to access internal and external databases when searching for specific information, making it difficult to provide flexible and comprehensive information. Furthermore, the lack of sufficient technology to accurately analyze user intent and quickly retrieve corresponding information results in a difficulty in improving the user experience.

[0881] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0882] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to a network; means for converting the digital data into text data using a speech recognition engine on a computer on the network; means for passing the text data to a natural language processing engine to analyze its intent and content; means for retrieving necessary information from internal and external storage devices based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; and means for transmitting the voice data to a terminal and responding to the user. This makes it possible to provide users with information from a wide range of sources quickly and accurately, and to provide appropriate answers to user questions in voice.

[0883] A "user" is an individual or organization that uses the system to input questions by voice.

[0884] A "network" is a means of communication for sending and receiving digital data between various computers.

[0885] A "speech recognition engine" is a software or hardware system that analyzes input speech data and converts it into text data.

[0886] "Text data" refers to digital data in sentence format converted by a speech recognition engine.

[0887] A "natural language processing engine" is a software or hardware system that analyzes text data and understands the user's intent and content from it.

[0888] An "internal storage device" is a device installed within a system for storing and retrieving data.

[0889] "External storage devices" are devices or services that exist outside the system and are used to store and retrieve data.

[0890] "Generative artificial intelligence" is a system of algorithms or software that generates appropriate answers based on questions from users.

[0891] A "speech synthesis engine" is a software or hardware system that converts generated text data into speech data.

[0892] A "terminal" is a device used by a user to input voice, receive audio data, and play it back.

[0893] The system according to the present invention is a voice assistant system that analyzes the user's voice input, generates an appropriate response, and provides a voice reply. Specific embodiments of the present invention are described below.

[0894] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the audio using its built-in microphone and converts this analog audio data into digital data. The converted audio data is then sent to the server in binary format.

[0895] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0896] Next, this text data is passed to a natural language processing (NLP) engine (e.g., a typical NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information.

[0897] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve the current weather information for Tokyo.

[0898] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI model, e.g., a general-purpose generative AI model) to generate an appropriate answer to the user's question. For example, it might generate the sentence, "Today in Tokyo it is sunny, and the high temperature is 25 degrees Celsius." The prompt used would be: "Today's weather in Tokyo is sunny, and the high temperature is 25 degrees Celsius."

[0899] The generated text response is passed to a speech synthesis engine (e.g., general speech synthesis software) and converted into audio data. The speech synthesis engine converts the generated text data into an audio format (e.g., MP3) and returns that audio data to the server.

[0900] The server sends the returned audio data to the terminal. The terminal plays the received audio data through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the highest temperature is 25 degrees Celsius."

[0901] This system allows users to quickly and accurately obtain information from a wide range of sources simply by asking questions by voice. Furthermore, the system can accommodate diverse user groups, including seniors and children.

[0902] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0903] Step 1: Receiving voice input

[0904] The user speaks a question into the microphone. The device uses its built-in microphone to capture the audio and converts this analog audio data into digital data. Specifically, the device's audio capture function is activated, the audio signal is converted to a digital format according to the sampling rate, and stored in a buffer in binary format.

[0905] Input: User's analog voice

[0906] Output: Digital audio data (binary format)

[0907] Step 2: Sending the audio data

[0908] The terminal divides the digital audio data stored in the buffer into packets of a fixed size and sends them to the server via secure and high-speed network communication (e.g., HTTPS). HTTP POST requests are used to send the audio data to the server-side API endpoint.

[0909] Input: Digital audio data (binary format)

[0910] Output: Request to send audio data to the server

[0911] Step 3: Speech recognition and text conversion

[0912] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. Specifically, it performs audio waveform analysis and uses a phonological and lexical model to convert the audio into corresponding text. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[0913] Input: Audio data (binary format)

[0914] Output: Text data

[0915] Step 4: Text analysis and intent identification

[0916] The server passes the converted text data to a natural language processing (NLP) engine (e.g., general NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information. The analysis results are stored in JSON format.

[0917] Input: Text data

[0918] Output: Analysis result (JSON format)

[0919] Step 5: Information Search

[0920] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve current weather information for Tokyo. Here, it uses an HTTP request to query the external database and receives a response in JSON format.

[0921] Input: Analysis result (JSON format)

[0922] Output: Acquired information (JSON format)

[0923] Step 6: Generating the answer

[0924] Based on the weather information it receives, the server uses generative artificial intelligence (generative AI model, e.g., a general generative AI model) to generate appropriate answers to the user's questions. Specifically, it inputs the prompt "What's the weather like in Tokyo today?" into the generative AI model and generates the answer "It's sunny in Tokyo today, and the maximum temperature is 25 degrees Celsius."

[0925] Input: Acquired information (in JSON format) and prompt message

[0926] Output: Generated response (text data)

[0927] Step 7: Speech Synthesis

[0928] The server passes the generated text to a speech synthesis engine (e.g., common speech synthesis software) and converts it into audio data. The speech synthesis engine then converts the generated text data into an audio format (e.g., MP3). This conversion uses a speech model and involves acoustic processing to produce natural-sounding speech.

[0929] Input: Generated response (text data)

[0930] Output: Audio data (MP3 format)

[0931] Step 8: Return of audio data and response playback

[0932] The server sends the returned audio data to the terminal as an HTTP response. The terminal stores the received audio data in a buffer and plays it back through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the maximum temperature is 25 degrees Celsius."

[0933] Input: Audio data (MP3 format)

[0934] Output: Voice response to the user

[0935] (Application Example 1)

[0936] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0937] In physical stores, there is a need for a system that allows customers to receive quick and appropriate answers to their questions regarding product information, store layout, and inventory. Traditional methods require customers to ask store staff directly or search for store signs themselves, which is inconvenient and time-consuming. In particular, when customers ask a variety of questions, responding to them takes time, making it difficult to provide satisfactory service.

[0938] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0939] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to the server; means for the server to convert the digital data into text data using a speech recognition engine; means for passing the text data to a natural language processing engine to analyze its intent and content; means for searching for necessary information from an internal database and an external database based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; means for transmitting the voice data to a terminal and responding to the user; and means for providing quick and appropriate voice answers to questions from customers in physical stores regarding product information, store guidance, inventory checks, etc. This enables users to quickly obtain necessary information through a voice assistant no matter where they are in the store.

[0940] A "user" is someone who uses a voice assistant system to ask questions or give instructions.

[0941] "Digital data" refers to data obtained by converting analog audio signals into a digital format.

[0942] A "server" is a computer system used for speech recognition, natural language processing, database retrieval, generative artificial intelligence, and speech synthesis.

[0943] A "speech recognition engine" is a program or function that analyzes speech data and converts it into text data.

[0944] "Text data" refers to character information generated by a speech recognition engine.

[0945] A "natural language processing engine" is a program or function that analyzes text data to understand the user's intent and content.

[0946] An "internal database" is a database stored within the server where the information to be searched is stored.

[0947] An "external database" is a database that exists outside the server and can be accessed via an API.

[0948] "Generative artificial intelligence" is artificial intelligence that generates appropriate answers based on analyzed information.

[0949] A "speech synthesis engine" is a program or function that converts text data into speech data.

[0950] A "device" is a device used by a user to operate a voice assistant system, and includes smartphones and smart glasses.

[0951] A "physical store" refers to a sales facility or service provider that actually exists in a physical space.

[0952] "Product information" refers to detailed product descriptions, prices, and stock availability information provided within a physical store.

[0953] "Prompt and appropriate responses" refer to providing accurate information immediately in response to user questions.

[0954] Using the above definitions, we will clarify each function and element of the voice assistant system.

[0955] The system according to this invention is a voice assistant system that provides quick and appropriate answers to questions from users in physical stores. Specific embodiments thereof will be described below.

[0956] System Configuration

[0957] 1. User voice input

[0958] The user inputs their questions by voice into a device such as a smartphone or smart glasses. The voice input is captured through the device's built-in microphone.

[0959] 2. Digital conversion and transmission of audio

[0960] The terminal converts the captured analog audio into digital data. The converted digital data is then sent to the server via the network.

[0961] 3. Speech Recognition and Natural Language Processing

[0962] The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert the received digital audio data into text data. Next, it passes this text data to a natural language processing engine (e.g., an NLP engine) to analyze the intent and content of the user's question.

[0963] 4. Information Retrieval and Generation

[0964] The server searches for necessary information from internal and external databases based on the analysis results. It uses external APIs as needed to retrieve information from external databases. This search result is then used with generative artificial intelligence (generative AI model) to generate appropriate answers.

[0965] 5. Speech synthesis and response

[0966] The generated response is converted into audio data using a speech synthesis engine (e.g., gTTS). The converted audio data is then sent back to the terminal via the network, where it plays the audio data and responds to the user.

[0967] Specific usage examples

[0968] Example 1: Product Information

[0969] The user asks, "Where can I find this product?" inside the store.

[0970] The voice assistant system captures this question and sends it to the server.

[0971] The server converts the speech into text using a speech recognition engine, and then analyzes it as "product information" using a natural language processing engine.

[0972] The server retrieves the product's location information from an internal database and uses generative artificial intelligence to generate a response such as, "This product is located next to the cash register on the third floor."

[0973] The generated response is converted by a speech synthesis engine and sent to the terminal.

[0974] The device plays a message to the user saying, "This item is located next to the cash register on the 3rd floor."

[0975] Usage example 2: Inventory check

[0976] The user asks, "Do you have this item in stock?"

[0977] The voice assistant system captures the user's question and sends it to the server.

[0978] The server uses a speech recognition engine to convert the speech into text data, and then a natural language processing engine analyzes it to read "check inventory".

[0979] The server uses an external API to retrieve the latest inventory information and generates responses such as "This product is still in stock" using generative artificial intelligence.

[0980] The generated response is converted into audio data by a speech synthesis engine and sent to the terminal.

[0981] The device plays the product and informs the user that "this item is still in stock."

[0982] Examples of prompt statements

[0983] Example of a prompt message for product information:

[0984] "The user is asking for the location information of this product within the store. Please provide the specific location of the product."

[0985] Example prompt message for checking inventory:

[0986] "A user is asking about the availability of a specific product. Please provide the stock status for this product."

[0987] Through the above explanation, we have provided specific examples of voice assistant systems and clarified methods for improving user convenience in physical stores.

[0988] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0989] Step 1:

[0990] The user inputs a question by voice. The device's microphone captures the voice and converts the analog audio signal into digital data. This digital data is transmitted to the server via the network. The input is the user's voice, and the output is digital audio data.

[0991] Step 2:

[0992] The server receives digital data and passes it to a speech recognition engine (e.g., Google Speech Recognition API), which converts the audio into text data. This conversion process analyzes the frequency information of the audio data and generates the corresponding string. The input is digital audio data, and the output is text data.

[0993] Step 3:

[0994] The server passes text data to a natural language processing (NLP) engine, which analyzes the intent and content of the user's question. The NLP engine grammatically and semantically analyzes the text data to identify which information the user's question relates to. The input is text data, and the output is the analysis result.

[0995] Step 4:

[0996] The server searches for the necessary information from internal and external databases based on the analysis results. The internal database contains product information and in-store guidance, while the external database contains the latest information accessible via API. It generates database queries and retrieves information based on them. The input is the analysis results, and the output is the necessary information.

[0997] Step 5:

[0998] Based on the information acquired by the server, a generative artificial intelligence (generative AI model) is used to generate appropriate answers to the user's questions. In this process, the generative AI model automatically generates answers based on the prompt text. The input consists of the required information and the prompt text, and the output is the generated text answer.

[0999] Step 6:

[1000] The generated text response is passed to a speech synthesis engine (e.g., gTTS) and converted into audio data. The speech synthesis engine then converts the text data into an audio file. The input is the generated text response, and the output is the audio data.

[1001] Step 7:

[1002] The server sends audio data to the terminal. The terminal plays the received audio data and provides the answer to the user verbally. The input is audio data, and the output is the verbal answer to the user.

[1003] The above outlines the specific processing steps in the invention's system. This series of processes enables users to quickly obtain necessary information via a voice assistant, regardless of their location within the store.

[1004] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1005] The system according to the present invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. The details of the operation of the present invention are as follows.

[1006] Receiving voice input

[1007] The user speaks a question to the voice assistant. An example question might be, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. This digital audio data is then sent from the device to the server.

[1008] Speech recognition and language analysis

[1009] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. For example, the text data "What's the weather like in Tokyo today?" is generated. This text data is passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[1010] Addition of emotion recognition

[1011] In the process described above, the server uses an emotion engine to recognize emotional information from the user's voice. For example, if the user is angry, the emotion engine recognizes that anger and adds it to the analysis results.

[1012] Information retrieval and generation

[1013] Based on the analysis results, the server first searches its internal database for the necessary information. If the necessary information is not found in the internal database, it retrieves it from an external database. Various external APIs may be used in this process.

[1014] After acquiring the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone. This generated text response is then passed to a speech synthesis engine and converted into audio data.

[1015] Speech synthesis and response return

[1016] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[1017] Specific example

[1018] Example 1: Weather question

[1019] When a user asks, "What's the weather like in Tokyo today?",

[1020] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1021] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[1022] 3. Analyze that the server is requesting weather information using an NLP engine.

[1023] 4. The server uses an emotion engine to recognize the user's emotions (e.g., relaxed).

[1024] 5. The server queries the weather information API for the weather in Tokyo.

[1025] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[1026] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1027] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1028] Example 2: History Questions

[1029] When a user asks, "When was Alexander the Great born?",

[1030] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1031] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[1032] 3. Analyze that the server is requesting historical information using the NLP engine.

[1033] 4. The server uses an emotion engine to recognize the user's emotions (e.g., excited).

[1034] 5. The server searches its internal database and confirms the relevant information.

[1035] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and uses a generative AI to generate an excited response.

[1036] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1037] 8. The device plays the answer back, excitedly stating, "Alexander the Great was born in 356 BC."

[1038] A system configured in this way can quickly and accurately provide answers from a wide range of sources, including appropriate responses tailored to the user's emotions. This allows users to experience more satisfying interactions.

[1039] The following describes the processing flow.

[1040] Step 1:

[1041] The user speaks a question to the voice assistant, such as, "What's the weather like in Tokyo today?"

[1042] Step 2:

[1043] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. The converted audio data is sent to a server for processing.

[1044] Step 3:

[1045] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[1046] Step 4:

[1047] The server receives the transmitted audio data. The received audio data is then passed to the speech recognition engine.

[1048] Step 5:

[1049] The server uses a speech recognition engine to convert the audio data into text data. For example, the text "What's the weather like in Tokyo today?" is generated.

[1050] Step 6:

[1051] The server passes the generated text data to a natural language processing (NLP) engine to analyze the intent and content of the question. Based on the analysis results, it identifies that the question is related to weather information.

[1052] Step 7:

[1053] The server uses an emotion engine to analyze the user's emotional state from the voice data. For example, it can determine whether the user is speaking in a relaxed state or is angry.

[1054] Step 8:

[1055] Based on the analysis results, the server first searches its internal database to see if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[1056] Step 9:

[1057] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[1058] Step 10:

[1059] Based on the information acquired by the server, generative artificial intelligence is used to generate answers to the user's questions. In this process, the tone and content of the answer are adjusted to reflect recognized emotional information. For example, an answer such as "Today in Tokyo it is sunny, and the highest temperature is 25 degrees Celsius" might be generated.

[1060] Step 11:

[1061] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[1062] Step 12:

[1063] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[1064] Step 13:

[1065] The device receives and plays the received audio data. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius." This response is delivered in an appropriate tone depending on the user's mood.

[1066] (Example 2)

[1067] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1068] Conventional voice assistant systems have been required to provide quick and accurate responses to user utterances, but they have faced challenges in natural communication with users, particularly due to their inability to respond with consideration for emotions. Furthermore, when the information corresponding to a user's question is not found in the internal database, the means of retrieving information from external databases are limited, making it difficult to provide an appropriate answer.

[1069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1070] In this invention, the server includes means for converting digital data into text data using a speech recognition engine, means for analyzing intent and content using a natural language processing engine, and means for analyzing the user's emotions. This enables natural communication that takes emotions into consideration. Furthermore, by providing means for retrieving information from an external database when the information does not exist in the internal database, it becomes possible to provide quick and appropriate answers to a wide range of inquiries.

[1071] A "speech recognition engine" is a software or hardware technology used to convert speech data into text data.

[1072] A "natural language processing engine" is a software or hardware technology used to analyze user intent and content from text data.

[1073] "Sentiment analysis" is a technology that extracts emotional information from a user's voice and text data to recognize the user's state of mind.

[1074] An "internal database" is a database that stores information managed within a system.

[1075] An "external database" is a database that provides information located outside of the system.

[1076] "Generative artificial intelligence" is an artificial intelligence technology used to generate answers to user questions.

[1077] A "speech synthesis engine" is a software or hardware technology used to convert text data into speech data.

[1078] A "terminal" is a device used by a user to access a voice assistant system.

[1079] "Emotional information" refers to data that indicates the emotional state expressed by the user.

[1080] SSL / TLS is a protocol for encrypting and securely conducting data communications.

[1081] The system according to this invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. Details of the system for carrying out this invention are described below.

[1082] Receiving voice input

[1083] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. An A / D converter is used for this conversion process.

[1084] Sending audio data

[1085] After the digital audio data is generated, the device sends this data to the server via the internet. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security.

[1086] Speech recognition and language analysis

[1087] The server passes the received audio data to a speech recognition engine (general speech recognition software) and converts it into text data. For example, commonly used speech recognition technologies are used for the speech recognition engine. Next, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. Based on the results of this analysis, the server identifies what information is needed.

[1088] Addition of emotion recognition

[1089] The server passes voice or text data to an emotion analysis engine (common emotion analysis software) to recognize the user's emotional information. For example, the emotion analysis engine identifies emotions such as the user being relaxed, angry, or excited. This emotional information is added to the analysis results.

[1090] Information retrieval and generation

[1091] Based on the user's query, the server first searches its internal database (a general-purpose database management system) to see if the necessary information exists. If the information is not found in the internal database, it retrieves it from an external database (a general-purpose API or external data source). For example, in the case of weather information, it retrieves data from an external weather information API.

[1092] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI technology) to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone.

[1093] Speech synthesis and response return

[1094] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into audio data. This converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the response in voice.

[1095] Specific example

[1096] When a user asks "What's the weather like in Tokyo today?", the following process takes place.

[1097] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1098] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[1099] 3. Analyze that the server is requesting weather information using an NLP engine.

[1100] 4. The server uses an emotion analysis engine to recognize the user's emotions (e.g., relaxed).

[1101] 5. The server queries the weather information API for Tokyo's weather.

[1102] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[1103] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1104] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1105] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1106] Step 1: Receiving voice input

[1107] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone. It receives the voice signal from the user as input and converts the analog voice signal into digital data through an A / D converter. The output here is digital voice data. Specifically, when the user says "What's the weather like in Tokyo today?", the microphone picks up the voice and converts it into digital data.

[1108] Step 2: Sending the audio data

[1109] After the digital audio data is generated, the terminal sends this data to the server via the internet. The terminal receives the digital audio data as input and sends the encrypted digital audio data to the server as output. Specifically, the terminal sends the digital audio data to the server using an encryption protocol such as SSL / TLS.

[1110] Step 3: Speech Recognition and Language Analysis

[1111] The server receives audio data and passes it to a speech recognition engine (general speech recognition software) to convert it into text data. It receives encrypted digital audio data as input and generates text data as output. Specifically, the server uses the speech recognition engine to convert "What's the weather like in Tokyo today?" into text data. Then, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. The input is text data, and the output is the analysis result. Specifically, the server passes the text data to the NLP engine, which analyzes that the user is requesting weather information.

[1112] Step 4: Adding emotion recognition

[1113] The server also passes the voice or text data to an emotion analysis engine (general emotion analysis software) to recognize the user's emotions. It receives voice or text data as input and generates emotion information as output. Specifically, the server uses the emotion analysis engine to identify emotions such as relaxed, angry, or excited based on voice tone and word choice. This emotion information is then added to the subsequent analysis results.

[1114] Step 5: Information retrieval and generation

[1115] Based on the analysis results and sentiment analysis results, the server first searches its internal database (a general database management system) for the necessary information. It receives the analysis results and sentiment information as input and retrieves the information present in the database as output. Specifically, the server searches its internal database for weather information queries, and if not found, it uses a weather information API (a general API) to retrieve information from an external database. Next, it uses generative artificial intelligence (a general generative AI technology) to generate an appropriate answer to the user's question. In this process, it adjusts the response based on the sentiment information. The input is the search results and sentiment information, and the output is the generated answer. Specifically, the server generates a relaxed-sounding answer such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[1116] Step 6: Speech synthesis and response return

[1117] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into speech data. The generated text response is received as input, and speech data is generated as output. Specifically, the server passes the text response to the speech synthesis engine and generates speech data. This converted speech data is sent from the server to the terminal. The terminal plays this speech data and provides the user with the answer to the question in speech. Speech data is received as input, and speech playback is performed as output. Specifically, the terminal plays the received speech data and says in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1118] (Application Example 2)

[1119] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1120] Conventional voice assistant systems can provide appropriate responses to user voice commands, but they often come across as mechanical and cold because they cannot take user emotions into consideration. Furthermore, in factory work instructions, disregarding the emotional state of workers makes it difficult to create an efficient and comfortable work environment.

[1121] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing emotional information from the user's voice and adjusting the response, and means for adjusting the speed and volume of the response to provide an appropriate response according to the user's emotions. This enables natural and human-like interaction that responds to the user's emotions, and in particular, it is possible to provide an appropriate response that takes into account the emotional state of the worker, even in work instructions within a factory.

[1122] "User voice" refers to the voice spoken by a human being to ask questions or give instructions to a system.

[1123] "Means of converting to digital data" refers to equipment or software that converts analog audio into digital data.

[1124] "Means of sending to a server" refers to equipment or software used to transmit digital data to a server via a network.

[1125] A "speech recognition engine" is software that converts received audio data into text data.

[1126] A "natural language processing engine" is software used to analyze the intent and content of text data.

[1127] An "internal database" is a database containing information data stored within a system.

[1128] An "external database" is a database used to retrieve necessary data from information sources outside the system.

[1129] "Generative artificial intelligence" is a type of artificial intelligence that generates appropriate answers to user questions.

[1130] A "speech synthesis engine" is software that converts generated text responses into speech data.

[1131] "Means of sending to a terminal and responding to the user" refers to a device or software that plays audio data sent from a server and provides a response to the user.

[1132] "Means for recognizing emotional information" refers to software or hardware that analyzes emotions from a user's voice.

[1133] "Means of adjusting responses" refers to software that modifies the content and method of responses based on recognized emotional information.

[1134] "Means for adjusting response speed and volume" refers to software or hardware that dynamically sets the speed and volume of voice responses according to the user's emotional state.

[1135] The system of this invention is designed to receive voice commands from a user and provide appropriate responses based on their emotions, using voice assistant technology. The system begins by converting the user's voice into digital data and sending it to a server. The server consists of a speech recognition engine, a natural language processing engine, an emotion recognition engine, an internal database, an external database, a generative artificial intelligence system, and a speech synthesis engine.

[1136] First, the user inputs audio via a microphone. The terminal captures this audio as an analog signal and converts it into a digital signal. This digital data is then transmitted to the server via the network.

[1137] On the server, the speech recognition engine first converts digital data into text data. This text data is then passed to the natural language processing engine, which analyzes the user's intent and the content of the question. Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information does not exist in the internal database, it can access the external database to retrieve the information.

[1138] Furthermore, an emotion recognition engine is used to analyze emotional information from the user's voice. This emotional information is added to the analysis results and used by the generative artificial intelligence to generate appropriate answers to the user's questions. For example, if the user is angry, the answer will be generated in a calm tone, and if they are relaxed, the answer will be generated in a natural tone.

[1139] The generated text response is converted into audio data by a speech synthesis engine and sent back to the terminal. The terminal plays this audio data, providing a response to the user. The response speed and volume are adjusted based on the user's emotions, which are analyzed by an emotion recognition engine.

[1140] For example, if a user instructs, "Please take this part to Unit 5," the system will recognize this voice, determine the emotion, and respond appropriately. Once an emotion is recognized, the system will adjust the speed and volume of the response accordingly. For instance, if the user is excited, it will respond quickly and loudly; if the user is relaxed, it will provide a response at a normal speed and standard volume.

[1141] Examples of prompt statements include the following:

[1142] User voice input: "Please take this part to Unit 5."

[1143] Output: "We will transport the parts. Where would you like us to transport them?"

[1144] In this way, the system provides natural and human-like interactions that respond to the user's emotions, and can provide appropriate responses that take into account the emotional state of the worker, especially when giving work instructions in a factory.

[1145] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1146] Step 1:

[1147] The user provides voice input via a microphone. For example, they might say, "Please take this part to unit number 5." This analog voice signal is captured by the terminal's microphone.

[1148] Input: User's analog audio signal

[1149] Output: Captured analog audio signal

[1150] Operation: The user gives instructions to the system via the microphone.

[1151] Step 2:

[1152] The terminal converts the captured analog audio signal into digital data. This digital data must be in a format that the speech recognition engine can process.

[1153] Input: Captured analog audio signal

[1154] Output: Digital audio data

[1155] Operation: The device's audio conversion module converts analog audio signals into digital data.

[1156] Step 3:

[1157] The terminal sends the converted digital audio data to the server via the network. The server receives this data.

[1158] Input: Digital audio data

[1159] Output: Digital audio data transferred to the server

[1160] Operation: The device sends digital audio data to the server.

[1161] Step 4:

[1162] The server uses a speech recognition engine to convert digital speech data into text data. For example, it might generate text data such as, "Please take this part to Unit 5."

[1163] Input: Digital audio data

[1164] Output: Text data

[1165] Operation: The server's speech recognition engine converts digital speech data into text data.

[1166] Step 5:

[1167] The server passes this text data to a natural language processing engine, which analyzes the user's intent and the content of the question. For example, it might analyze that the user is giving instructions to transport parts.

[1168] Input: Text data

[1169] Output: Analyzed intent and content

[1170] Operation: The server's natural language processing engine analyzes the text data to identify the user's intent.

[1171] Step 6:

[1172] The server uses an emotion engine to recognize emotional information from the user's voice. For example, it can determine whether the user is excited or calm.

[1173] Input: Text data

[1174] Output: User sentiment information

[1175] Operation: The server's emotion engine analyzes and extracts emotional information from the audio data.

[1176] Step 7:

[1177] Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information is not found in the internal database, it retrieves it from the external database.

[1178] Input: Analyzed intent and content, user sentiment information

[1179] Output: Required information

[1180] Operation: The server searches for the necessary information from internal and external databases.

[1181] Step 8:

[1182] Generative artificial intelligence is used to generate appropriate answers to user questions. During this process, the content and tone of the responses are adjusted based on emotional information.

[1183] Input: Required information, user sentiment information

[1184] Output: Generated text answer

[1185] Operation: The server's generational artificial intelligence generates the appropriate answer.

[1186] Step 9:

[1187] The generated text response is passed to a speech synthesis engine and converted into speech data. For example, speech data such as "I will transport the parts. Where do you want me to transport them?" is generated.

[1188] Input: Generated text answer

[1189] Output: Audio data

[1190] Operation: The server's speech synthesis engine converts text data into speech data.

[1191] Step 10:

[1192] The server sends audio data to the terminal, which then plays this audio data to respond to the user. The response speed and volume are adjusted according to the user's emotions.

[1193] Input: Audio data

[1194] Output: Audio response played to the user

[1195] Operation: The device plays audio data and provides a response to the user.

[1196] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1197] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1198] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1199] [Fourth Embodiment]

[1200] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1201] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1202] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1203] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1204] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1205] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1206] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1207] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1208] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1209] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1210] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1211] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1212] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1213] The system according to the present invention is a voice assistant system that caters to a diverse range of users and provides appropriate answers to questions about information in a wide range of fields. This system encompasses a series of processes that analyze voice input and generate appropriate answers, and the details of its operation are as follows.

[1214] Receiving voice input

[1215] The user speaks a question to the voice assistant. The device captures this audio through the microphone and converts the analog audio into digital data. The converted audio data is then sent from the device to the server.

[1216] Speech recognition and language analysis

[1217] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. The converted text data is then passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[1218] Information retrieval and generation

[1219] Based on the analysis results, the server first searches its internal database for the necessary information. If the information is not found in the internal database, it retrieves it from an external database. This process may utilize various external APIs.

[1220] After obtaining the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. This generated text answer is then passed to a speech synthesis engine and converted into speech data.

[1221] Speech synthesis and response return

[1222] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[1223] Specific example

[1224] Example 1: Weather question

[1225] When a user asks, "What's the weather like in Tokyo today?",

[1226] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1227] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[1228] 3. Analyze that the server is requesting weather information using an NLP engine.

[1229] 4. The server queries the weather information API for the weather in Tokyo.

[1230] 5. The server retrieves the information "Today in Tokyo it is sunny and the maximum temperature is 25 degrees Celsius," and a generative AI generates an appropriate response.

[1231] 6. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1232] 7. The device plays the answer and says, "Today in Tokyo it's sunny and the high temperature is 25 degrees Celsius."

[1233] Example 2: History Questions

[1234] When a user asks, "When was Alexander the Great born?",

[1235] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1236] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[1237] 3. Analyze that the server is requesting historical information using the NLP engine.

[1238] 4. The server searches its internal database and confirms the relevant information.

[1239] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and generates an answer using a generative AI.

[1240] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1241] 8. The device plays the answer and says, "Alexander the Great was born in 356 BC."

[1242] A system configured in this way can provide users with answers from a wide range of sources quickly and accurately, and can accommodate diverse user groups such as seniors and children.

[1243] The following describes the processing flow.

[1244] Step 1:

[1245] The user speaks a question to the voice assistant. An example of such a question is, "What's the weather like in Tokyo today?"

[1246] Step 2:

[1247] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. This digital audio data is then encoded into a format suitable for speech recognition.

[1248] Step 3:

[1249] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[1250] Step 4:

[1251] The server receives the audio data. The received audio data is then passed to the speech recognition engine.

[1252] Step 5:

[1253] The server uses a speech recognition engine to convert the audio data into text data. For example, text data such as "What's the weather like in Tokyo today?" is generated.

[1254] Step 6:

[1255] The server passes the generated text data to a natural language processing (NLP) engine, which analyzes the intent and content of the question. The NLP engine then identifies the category of the question (e.g., weather information, historical information, etc.).

[1256] Step 7:

[1257] Based on the analysis results, the server searches its internal database to check if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[1258] Step 8:

[1259] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[1260] Step 9:

[1261] Based on the information acquired by the server, generative artificial intelligence (e.g., a large-scale language model) is used to generate natural-sounding answers to the user's questions. For example, an answer such as "Today in Tokyo it is sunny and the highest temperature is 25 degrees Celsius" might be generated.

[1262] Step 10:

[1263] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[1264] Step 11:

[1265] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[1266] Step 12:

[1267] The device receives the audio data and plays it back. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[1268] This series of steps allows users to receive accurate and rapid voice answers to their questions.

[1269] (Example 1)

[1270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1271] Conventional systems using speech recognition and natural language processing have limitations in their ability to access internal and external databases when searching for specific information, making it difficult to provide flexible and comprehensive information. Furthermore, the lack of sufficient technology to accurately analyze user intent and quickly retrieve corresponding information results in a difficulty in improving the user experience.

[1272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1273] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to a network; means for converting the digital data into text data using a speech recognition engine on a computer on the network; means for passing the text data to a natural language processing engine to analyze its intent and content; means for retrieving necessary information from internal and external storage devices based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; and means for transmitting the voice data to a terminal and responding to the user. This makes it possible to provide users with information from a wide range of sources quickly and accurately, and to provide appropriate answers to user questions in voice.

[1274] A "user" is an individual or organization that uses the system to input questions by voice.

[1275] A "network" is a means of communication for sending and receiving digital data between various computers.

[1276] A "speech recognition engine" is a software or hardware system that analyzes input speech data and converts it into text data.

[1277] "Text data" refers to digital data in sentence format converted by a speech recognition engine.

[1278] A "natural language processing engine" is a software or hardware system that analyzes text data and understands the user's intent and content from it.

[1279] An "internal storage device" is a device installed within a system for storing and retrieving data.

[1280] "External storage devices" are devices or services that exist outside the system and are used to store and retrieve data.

[1281] "Generative artificial intelligence" is a system of algorithms or software that generates appropriate answers based on questions from users.

[1282] A "speech synthesis engine" is a software or hardware system that converts generated text data into speech data.

[1283] A "terminal" is a device used by a user to input voice, receive audio data, and play it back.

[1284] The system according to the present invention is a voice assistant system that analyzes the user's voice input, generates an appropriate response, and provides a voice reply. Specific embodiments of the present invention are described below.

[1285] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the audio using its built-in microphone and converts this analog audio data into digital data. The converted audio data is then sent to the server in binary format.

[1286] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[1287] Next, this text data is passed to a natural language processing (NLP) engine (e.g., a typical NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information.

[1288] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve the current weather information for Tokyo.

[1289] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI model, e.g., a general-purpose generative AI model) to generate an appropriate answer to the user's question. For example, it might generate the sentence, "Today in Tokyo it is sunny, and the high temperature is 25 degrees Celsius." The prompt used would be: "Today's weather in Tokyo is sunny, and the high temperature is 25 degrees Celsius."

[1290] The generated text response is passed to a speech synthesis engine (e.g., general speech synthesis software) and converted into audio data. The speech synthesis engine converts the generated text data into an audio format (e.g., MP3) and returns that audio data to the server.

[1291] The server sends the returned audio data to the terminal. The terminal plays the received audio data through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the highest temperature is 25 degrees Celsius."

[1292] This system allows users to quickly and accurately obtain information from a wide range of sources simply by asking questions by voice. Furthermore, the system can accommodate diverse user groups, including seniors and children.

[1293] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1294] Step 1: Receiving voice input

[1295] The user speaks a question into the microphone. The device uses its built-in microphone to capture the audio and converts this analog audio data into digital data. Specifically, the device's audio capture function is activated, the audio signal is converted to a digital format according to the sampling rate, and stored in a buffer in binary format.

[1296] Input: User's analog voice

[1297] Output: Digital audio data (binary format)

[1298] Step 2: Sending the audio data

[1299] The terminal divides the digital audio data stored in the buffer into packets of a fixed size and sends them to the server via secure and high-speed network communication (e.g., HTTPS). HTTP POST requests are used to send the audio data to the server-side API endpoint.

[1300] Input: Digital audio data (binary format)

[1301] Output: Request to send audio data to the server

[1302] Step 3: Speech recognition and text conversion

[1303] The server passes the received audio data to a speech recognition engine (e.g., general speech recognition software). The speech recognition engine converts this audio data into text data. Specifically, it performs audio waveform analysis and uses a phonological and lexical model to convert the audio into corresponding text. For example, it converts the audio data "What's the weather like in Tokyo today?" directly into text data.

[1304] Input: Audio data (binary format)

[1305] Output: Text data

[1306] Step 4: Text analysis and intent identification

[1307] The server passes the converted text data to a natural language processing (NLP) engine (e.g., general NLP software). The NLP engine analyzes the text to understand the user's intent and the content of the question. Specifically, it extracts the keywords "Tokyo" and "weather" and interprets that the user is requesting weather information. The analysis results are stored in JSON format.

[1308] Input: Text data

[1309] Output: Analysis result (JSON format)

[1310] Step 5: Information Search

[1311] Based on the analysis results, the server first searches its internal storage (e.g., a typical database system) for the necessary information. If the information is not found in the internal storage, it retrieves it from external storage (e.g., a typical external database or API). Specifically, it accesses the API of a weather information service to retrieve current weather information for Tokyo. Here, it uses an HTTP request to query the external database and receives a response in JSON format.

[1312] Input: Analysis result (JSON format)

[1313] Output: Acquired information (JSON format)

[1314] Step 6: Generating the answer

[1315] Based on the weather information it receives, the server uses generative artificial intelligence (generative AI model, e.g., a general generative AI model) to generate appropriate answers to the user's questions. Specifically, it inputs the prompt "What's the weather like in Tokyo today?" into the generative AI model and generates the answer "It's sunny in Tokyo today, and the maximum temperature is 25 degrees Celsius."

[1316] Input: Acquired information (in JSON format) and prompt message

[1317] Output: Generated response (text data)

[1318] Step 7: Speech Synthesis

[1319] The server passes the generated text to a speech synthesis engine (e.g., common speech synthesis software) and converts it into audio data. The speech synthesis engine then converts the generated text data into an audio format (e.g., MP3). This conversion uses a speech model and involves acoustic processing to produce natural-sounding speech.

[1320] Input: Generated response (text data)

[1321] Output: Audio data (MP3 format)

[1322] Step 8: Return of audio data and response playback

[1323] The server sends the returned audio data to the terminal as an HTTP response. The terminal stores the received audio data in a buffer and plays it back through its built-in speaker. Specifically, it provides the user with an audio response such as, "Today in Tokyo it's sunny, and the maximum temperature is 25 degrees Celsius."

[1324] Input: Audio data (MP3 format)

[1325] Output: Voice response to the user

[1326] (Application Example 1)

[1327] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1328] In physical stores, there is a need for a system that allows customers to receive quick and appropriate answers to their questions regarding product information, store layout, and inventory. Traditional methods require customers to ask store staff directly or search for store signs themselves, which is inconvenient and time-consuming. In particular, when customers ask a variety of questions, responding to them takes time, making it difficult to provide satisfactory service.

[1329] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1330] In this invention, the server includes means for receiving user voice and converting it into digital data; means for transmitting the digital data to the server; means for the server to convert the digital data into text data using a speech recognition engine; means for passing the text data to a natural language processing engine to analyze its intent and content; means for searching for necessary information from an internal database and an external database based on the analysis results; means for generating answers to user questions using generative artificial intelligence; means including a speech synthesis engine that converts the generated answers into voice data; means for transmitting the voice data to a terminal and responding to the user; and means for providing quick and appropriate voice answers to questions from customers in physical stores regarding product information, store guidance, inventory checks, etc. This enables users to quickly obtain necessary information through a voice assistant no matter where they are in the store.

[1331] A "user" is someone who uses a voice assistant system to ask questions or give instructions.

[1332] "Digital data" refers to data obtained by converting analog audio signals into a digital format.

[1333] A "server" is a computer system used for speech recognition, natural language processing, database retrieval, generative artificial intelligence, and speech synthesis.

[1334] A "speech recognition engine" is a program or function that analyzes speech data and converts it into text data.

[1335] "Text data" refers to character information generated by a speech recognition engine.

[1336] A "natural language processing engine" is a program or function that analyzes text data to understand the user's intent and content.

[1337] An "internal database" is a database stored within the server where the information to be searched is stored.

[1338] An "external database" is a database that exists outside the server and can be accessed via an API.

[1339] "Generative artificial intelligence" is artificial intelligence that generates appropriate answers based on analyzed information.

[1340] A "speech synthesis engine" is a program or function that converts text data into speech data.

[1341] A "device" is a device used by a user to operate a voice assistant system, and includes smartphones and smart glasses.

[1342] A "physical store" refers to a sales facility or service provider that actually exists in a physical space.

[1343] "Product information" refers to detailed product descriptions, prices, and stock availability information provided within a physical store.

[1344] "Prompt and appropriate responses" refer to providing accurate information immediately in response to user questions.

[1345] Using the above definitions, we will clarify each function and element of the voice assistant system.

[1346] The system according to this invention is a voice assistant system that provides quick and appropriate answers to questions from users in physical stores. Specific embodiments thereof will be described below.

[1347] System Configuration

[1348] 1. User voice input

[1349] The user inputs their questions by voice into a device such as a smartphone or smart glasses. The voice input is captured through the device's built-in microphone.

[1350] 2. Digital conversion and transmission of audio

[1351] The terminal converts the captured analog audio into digital data. The converted digital data is then sent to the server via the network.

[1352] 3. Speech Recognition and Natural Language Processing

[1353] The server uses a speech recognition engine (e.g., Google Speech Recognition API) to convert the received digital audio data into text data. Next, it passes this text data to a natural language processing engine (e.g., an NLP engine) to analyze the intent and content of the user's question.

[1354] 4. Information Retrieval and Generation

[1355] The server searches for necessary information from internal and external databases based on the analysis results. It uses external APIs as needed to retrieve information from external databases. This search result is then used with generative artificial intelligence (generative AI model) to generate appropriate answers.

[1356] 5. Speech synthesis and response

[1357] The generated response is converted into audio data using a speech synthesis engine (e.g., gTTS). The converted audio data is then sent back to the terminal via the network, where it plays the audio data and responds to the user.

[1358] Specific usage examples

[1359] Example 1: Product Information

[1360] The user asks, "Where can I find this product?" inside the store.

[1361] The voice assistant system captures this question and sends it to the server.

[1362] The server converts the speech into text using a speech recognition engine, and then analyzes it as "product information" using a natural language processing engine.

[1363] The server retrieves the product's location information from an internal database and uses generative artificial intelligence to generate a response such as, "This product is located next to the cash register on the third floor."

[1364] The generated response is converted by a speech synthesis engine and sent to the terminal.

[1365] The device plays a message to the user saying, "This item is located next to the cash register on the 3rd floor."

[1366] Usage example 2: Inventory check

[1367] The user asks, "Do you have this item in stock?"

[1368] The voice assistant system captures the user's question and sends it to the server.

[1369] The server uses a speech recognition engine to convert the speech into text data, and then a natural language processing engine analyzes it to read "check inventory".

[1370] The server uses an external API to retrieve the latest inventory information and generates responses such as "This product is still in stock" using generative artificial intelligence.

[1371] The generated response is converted into audio data by a speech synthesis engine and sent to the terminal.

[1372] The device plays the product and informs the user that "this item is still in stock."

[1373] Examples of prompt statements

[1374] Example of a prompt message for product information:

[1375] "The user is asking for the location information of this product within the store. Please provide the specific location of the product."

[1376] Example prompt message for checking inventory:

[1377] "A user is asking about the availability of a specific product. Please provide the stock status for this product."

[1378] Through the above explanation, we have provided specific examples of voice assistant systems and clarified methods for improving user convenience in physical stores.

[1379] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1380] Step 1:

[1381] The user inputs a question by voice. The device's microphone captures the voice and converts the analog audio signal into digital data. This digital data is transmitted to the server via the network. The input is the user's voice, and the output is digital audio data.

[1382] Step 2:

[1383] The server receives digital data and passes it to a speech recognition engine (e.g., Google Speech Recognition API), which converts the audio into text data. This conversion process analyzes the frequency information of the audio data and generates the corresponding string. The input is digital audio data, and the output is text data.

[1384] Step 3:

[1385] The server passes text data to a natural language processing (NLP) engine, which analyzes the intent and content of the user's question. The NLP engine grammatically and semantically analyzes the text data to identify which information the user's question relates to. The input is text data, and the output is the analysis result.

[1386] Step 4:

[1387] The server searches for the necessary information from internal and external databases based on the analysis results. The internal database contains product information and in-store guidance, while the external database contains the latest information accessible via API. It generates database queries and retrieves information based on them. The input is the analysis results, and the output is the necessary information.

[1388] Step 5:

[1389] Based on the information acquired by the server, a generative artificial intelligence (generative AI model) is used to generate appropriate answers to the user's questions. In this process, the generative AI model automatically generates answers based on the prompt text. The input consists of the required information and the prompt text, and the output is the generated text answer.

[1390] Step 6:

[1391] The generated text response is passed to a speech synthesis engine (e.g., gTTS) and converted into audio data. The speech synthesis engine then converts the text data into an audio file. The input is the generated text response, and the output is the audio data.

[1392] Step 7:

[1393] The server sends audio data to the terminal. The terminal plays the received audio data and provides the answer to the user verbally. The input is audio data, and the output is the verbal answer to the user.

[1394] The above outlines the specific processing steps in the invention's system. This series of processes enables users to quickly obtain necessary information via a voice assistant, regardless of their location within the store.

[1395] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1396] The system according to the present invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. The details of the operation of the present invention are as follows.

[1397] Receiving voice input

[1398] The user speaks a question to the voice assistant. An example question might be, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. This digital audio data is then sent from the device to the server.

[1399] Speech recognition and language analysis

[1400] The server receives the transmitted audio data. Next, it uses a speech recognition engine to convert this audio data into text data. For example, the text data "What's the weather like in Tokyo today?" is generated. This text data is passed to a natural language processing (NLP) engine, which analyzes the user's intent and the content of the question. Based on this analysis, the server identifies what information is needed to answer the user's question.

[1401] Addition of emotion recognition

[1402] In the process described above, the server uses an emotion engine to recognize emotional information from the user's voice. For example, if the user is angry, the emotion engine recognizes that anger and adds it to the analysis results.

[1403] Information retrieval and generation

[1404] Based on the analysis results, the server first searches its internal database for the necessary information. If the necessary information is not found in the internal database, it retrieves it from an external database. Various external APIs may be used in this process.

[1405] After acquiring the necessary information, the server uses generative artificial intelligence to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone. This generated text response is then passed to a speech synthesis engine and converted into audio data.

[1406] Speech synthesis and response return

[1407] The converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the answer to the question in audio format. This allows the user to receive the answer to their question in audio form.

[1408] Specific example

[1409] Example 1: Weather question

[1410] When a user asks, "What's the weather like in Tokyo today?",

[1411] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1412] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[1413] 3. Analyze that the server is requesting weather information using an NLP engine.

[1414] 4. The server uses an emotion engine to recognize the user's emotions (e.g., relaxed).

[1415] 5. The server queries the weather information API for the weather in Tokyo.

[1416] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[1417] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1418] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1419] Example 2: History Questions

[1420] When a user asks, "When was Alexander the Great born?",

[1421] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1422] 2. The server uses a speech recognition engine to convert "When was Alexander the Great born?" into text.

[1423] 3. Analyze that the server is requesting historical information using the NLP engine.

[1424] 4. The server uses an emotion engine to recognize the user's emotions (e.g., excited).

[1425] 5. The server searches its internal database and confirms the relevant information.

[1426] 6. The server retrieves the information "Alexander the Great was born in 356 BC" from its internal database and uses a generative AI to generate an excited response.

[1427] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1428] 8. The device plays the answer back, excitedly stating, "Alexander the Great was born in 356 BC."

[1429] A system configured in this way can quickly and accurately provide answers from a wide range of sources, including appropriate responses tailored to the user's emotions. This allows users to experience more satisfying interactions.

[1430] The following describes the processing flow.

[1431] Step 1:

[1432] The user speaks a question to the voice assistant, such as, "What's the weather like in Tokyo today?"

[1433] Step 2:

[1434] The device captures the user's voice through the microphone. The captured analog audio is converted into digital data. The converted audio data is sent to a server for processing.

[1435] Step 3:

[1436] The device sends digital audio data to the server. The data is transmitted using a secure communication protocol (e.g., HTTPS).

[1437] Step 4:

[1438] The server receives the transmitted audio data. The received audio data is then passed to the speech recognition engine.

[1439] Step 5:

[1440] The server uses a speech recognition engine to convert the audio data into text data. For example, the text "What's the weather like in Tokyo today?" is generated.

[1441] Step 6:

[1442] The server passes the generated text data to a natural language processing (NLP) engine to analyze the intent and content of the question. Based on the analysis results, it identifies that the question is related to weather information.

[1443] Step 7:

[1444] The server uses an emotion engine to analyze the user's emotional state from the voice data. For example, it can determine whether the user is speaking in a relaxed state or is angry.

[1445] Step 8:

[1446] Based on the analysis results, the server first searches its internal database to see if the necessary information exists. If the necessary information is not found in the internal database, it proceeds to the next step.

[1447] Step 9:

[1448] The server queries an external database (e.g., a weather information API) to retrieve the necessary information. For example, it might use an external API to retrieve "the current weather in Tokyo."

[1449] Step 10:

[1450] Based on the information acquired by the server, generative artificial intelligence is used to generate answers to the user's questions. In this process, the tone and content of the answer are adjusted to reflect recognized emotional information. For example, an answer such as "Today in Tokyo it is sunny, and the highest temperature is 25 degrees Celsius" might be generated.

[1451] Step 11:

[1452] The server passes the generated text response to the speech synthesis engine, which converts it into audio data. The speech synthesis engine then converts the text into speech.

[1453] Step 12:

[1454] The server sends the generated audio data to the terminal. The audio data is transmitted using a secure protocol.

[1455] Step 13:

[1456] The device receives and plays the received audio data. This allows the user to receive an audio response such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius." This response is delivered in an appropriate tone depending on the user's mood.

[1457] (Example 2)

[1458] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1459] Conventional voice assistant systems have been required to provide quick and accurate responses to user utterances, but they have faced challenges in natural communication with users, particularly due to their inability to respond with consideration for emotions. Furthermore, when the information corresponding to a user's question is not found in the internal database, the means of retrieving information from external databases are limited, making it difficult to provide an appropriate answer.

[1460] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1461] In this invention, the server includes means for converting digital data into text data using a speech recognition engine, means for analyzing intent and content using a natural language processing engine, and means for analyzing the user's emotions. This enables natural communication that takes emotions into consideration. Furthermore, by providing means for retrieving information from an external database when the information does not exist in the internal database, it becomes possible to provide quick and appropriate answers to a wide range of inquiries.

[1462] A "speech recognition engine" is a software or hardware technology used to convert speech data into text data.

[1463] A "natural language processing engine" is a software or hardware technology used to analyze user intent and content from text data.

[1464] "Sentiment analysis" is a technology that extracts emotional information from a user's voice and text data to recognize the user's state of mind.

[1465] An "internal database" is a database that stores information managed within a system.

[1466] An "external database" is a database that provides information located outside of the system.

[1467] "Generative artificial intelligence" is an artificial intelligence technology used to generate answers to user questions.

[1468] A "speech synthesis engine" is a software or hardware technology used to convert text data into speech data.

[1469] A "terminal" is a device used by a user to access a voice assistant system.

[1470] "Emotional information" refers to data that indicates the emotional state expressed by the user.

[1471] SSL / TLS is a protocol for encrypting and securely conducting data communications.

[1472] The system according to this invention is a voice assistant system designed to accommodate diverse user groups and provide appropriate answers to questions about information across a wide range of fields. Furthermore, by incorporating a function that recognizes the user's emotions and adjusts the response accordingly, it achieves a more natural and human-like interaction. Details of the system for carrying out this invention are described below.

[1473] Receiving voice input

[1474] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone and converts the analog audio into digital data. An A / D converter is used for this conversion process.

[1475] Sending audio data

[1476] After the digital audio data is generated, the device sends this data to the server via the internet. During this process, the data is encrypted using encryption protocols such as SSL / TLS to ensure security.

[1477] Speech recognition and language analysis

[1478] The server passes the received audio data to a speech recognition engine (general speech recognition software) and converts it into text data. For example, commonly used speech recognition technologies are used for the speech recognition engine. Next, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. Based on the results of this analysis, the server identifies what information is needed.

[1479] Addition of emotion recognition

[1480] The server passes voice or text data to an emotion analysis engine (common emotion analysis software) to recognize the user's emotional information. For example, the emotion analysis engine identifies emotions such as the user being relaxed, angry, or excited. This emotional information is added to the analysis results.

[1481] Information retrieval and generation

[1482] Based on the user's query, the server first searches its internal database (a general-purpose database management system) to see if the necessary information exists. If the information is not found in the internal database, it retrieves it from an external database (a general-purpose API or external data source). For example, in the case of weather information, it retrieves data from an external weather information API.

[1483] After obtaining the necessary information, the server uses generative artificial intelligence (generative AI technology) to generate appropriate answers to the user's questions. During this process, the response is adjusted based on recognized emotional information. For example, if the user is angry, the response will be generated in a calm and gentle tone.

[1484] Speech synthesis and response return

[1485] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into audio data. This converted audio data is sent from the server to the terminal. The terminal plays this audio data and provides the user with the response in voice.

[1486] Specific example

[1487] When a user asks "What's the weather like in Tokyo today?", the following process takes place.

[1488] 1. The device captures the audio, converts it into digital data, and sends it to the server.

[1489] 2. The server uses a speech recognition engine to convert "What's the weather like in Tokyo today?" into text.

[1490] 3. Analyze that the server is requesting weather information using an NLP engine.

[1491] 4. The server uses an emotion analysis engine to recognize the user's emotions (e.g., relaxed).

[1492] 5. The server queries the weather information API for Tokyo's weather.

[1493] 6. The server retrieves the information "Today in Tokyo it's sunny and the highest temperature is 25 degrees Celsius," and a generative AI generates a response in a relaxed tone.

[1494] 7. The server converts the generated text response into speech data using a speech synthesis engine and sends it to the terminal.

[1495] 8. The device plays the answer back, saying in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1496] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1497] Step 1: Receiving voice input

[1498] The user speaks a question to the voice assistant. For example, "What's the weather like in Tokyo today?" The device captures the user's voice through the microphone. It receives the voice signal from the user as input and converts the analog voice signal into digital data through an A / D converter. The output here is digital voice data. Specifically, when the user says "What's the weather like in Tokyo today?", the microphone picks up the voice and converts it into digital data.

[1499] Step 2: Sending the audio data

[1500] After the digital audio data is generated, the terminal sends this data to the server via the internet. The terminal receives the digital audio data as input and sends the encrypted digital audio data to the server as output. Specifically, the terminal sends the digital audio data to the server using an encryption protocol such as SSL / TLS.

[1501] Step 3: Speech Recognition and Language Analysis

[1502] The server receives audio data and passes it to a speech recognition engine (general speech recognition software) to convert it into text data. It receives encrypted digital audio data as input and generates text data as output. Specifically, the server uses the speech recognition engine to convert "What's the weather like in Tokyo today?" into text data. Then, this text data is passed to a natural language processing (NLP) engine (general natural language processing software) to analyze the user's intent and the content of the question. The input is text data, and the output is the analysis result. Specifically, the server passes the text data to the NLP engine, which analyzes that the user is requesting weather information.

[1503] Step 4: Adding emotion recognition

[1504] The server also passes the voice or text data to an emotion analysis engine (general emotion analysis software) to recognize the user's emotions. It receives voice or text data as input and generates emotion information as output. Specifically, the server uses the emotion analysis engine to identify emotions such as relaxed, angry, or excited based on voice tone and word choice. This emotion information is then added to the subsequent analysis results.

[1505] Step 5: Information retrieval and generation

[1506] Based on the analysis results and sentiment analysis results, the server first searches its internal database (a general database management system) for the necessary information. It receives the analysis results and sentiment information as input and retrieves the information present in the database as output. Specifically, the server searches its internal database for weather information queries, and if not found, it uses a weather information API (a general API) to retrieve information from an external database. Next, it uses generative artificial intelligence (a general generative AI technology) to generate an appropriate answer to the user's question. In this process, it adjusts the response based on the sentiment information. The input is the search results and sentiment information, and the output is the generated answer. Specifically, the server generates a relaxed-sounding answer such as, "Today in Tokyo it's sunny, and the high temperature is 25 degrees Celsius."

[1507] Step 6: Speech synthesis and response return

[1508] The generated text response is passed to a speech synthesis engine (common speech synthesis software) and converted into speech data. The generated text response is received as input, and speech data is generated as output. Specifically, the server passes the text response to the speech synthesis engine and generates speech data. This converted speech data is sent from the server to the terminal. The terminal plays this speech data and provides the user with the answer to the question in speech. Speech data is received as input, and speech playback is performed as output. Specifically, the terminal plays the received speech data and says in a relaxed tone, "It's sunny in Tokyo today, and the high temperature is 25 degrees Celsius."

[1509] (Application Example 2)

[1510] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1511] Conventional voice assistant systems can provide appropriate responses to user voice commands, but they often come across as mechanical and cold because they cannot take user emotions into consideration. Furthermore, in factory work instructions, disregarding the emotional state of workers makes it difficult to create an efficient and comfortable work environment.

[1512] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing emotional information from the user's voice and adjusting the response, and means for adjusting the speed and volume of the response to provide an appropriate response according to the user's emotions. This enables natural and human-like interaction that responds to the user's emotions, and in particular, it is possible to provide an appropriate response that takes into account the emotional state of the worker, even in work instructions within a factory.

[1513] "User voice" refers to the voice spoken by a human being to ask questions or give instructions to a system.

[1514] "Means of converting to digital data" refers to equipment or software that converts analog audio into digital data.

[1515] "Means of sending to a server" refers to equipment or software used to transmit digital data to a server via a network.

[1516] A "speech recognition engine" is software that converts received audio data into text data.

[1517] A "natural language processing engine" is software used to analyze the intent and content of text data.

[1518] An "internal database" is a database containing information data stored within a system.

[1519] An "external database" is a database used to retrieve necessary data from information sources outside the system.

[1520] "Generative artificial intelligence" is a type of artificial intelligence that generates appropriate answers to user questions.

[1521] A "speech synthesis engine" is software that converts generated text responses into speech data.

[1522] "Means of sending to a terminal and responding to the user" refers to a device or software that plays audio data sent from a server and provides a response to the user.

[1523] "Means for recognizing emotional information" refers to software or hardware that analyzes emotions from a user's voice.

[1524] "Means of adjusting responses" refers to software that modifies the content and method of responses based on recognized emotional information.

[1525] "Means for adjusting response speed and volume" refers to software or hardware that dynamically sets the speed and volume of voice responses according to the user's emotional state.

[1526] The system of this invention is designed to receive voice commands from a user and provide appropriate responses based on their emotions, using voice assistant technology. The system begins by converting the user's voice into digital data and sending it to a server. The server consists of a speech recognition engine, a natural language processing engine, an emotion recognition engine, an internal database, an external database, a generative artificial intelligence system, and a speech synthesis engine.

[1527] First, the user inputs audio via a microphone. The terminal captures this audio as an analog signal and converts it into a digital signal. This digital data is then transmitted to the server via the network.

[1528] On the server, the speech recognition engine first converts digital data into text data. This text data is then passed to the natural language processing engine, which analyzes the user's intent and the content of the question. Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information does not exist in the internal database, it can access the external database to retrieve the information.

[1529] Furthermore, an emotion recognition engine is used to analyze emotional information from the user's voice. This emotional information is added to the analysis results and used by the generative artificial intelligence to generate appropriate answers to the user's questions. For example, if the user is angry, the answer will be generated in a calm tone, and if they are relaxed, the answer will be generated in a natural tone.

[1530] The generated text response is converted into audio data by a speech synthesis engine and sent back to the terminal. The terminal plays this audio data, providing a response to the user. The response speed and volume are adjusted based on the user's emotions, which are analyzed by an emotion recognition engine.

[1531] For example, if a user instructs, "Please take this part to Unit 5," the system will recognize this voice, determine the emotion, and respond appropriately. Once an emotion is recognized, the system will adjust the speed and volume of the response accordingly. For instance, if the user is excited, it will respond quickly and loudly; if the user is relaxed, it will provide a response at a normal speed and standard volume.

[1532] Examples of prompt statements include the following:

[1533] User voice input: "Please take this part to Unit 5."

[1534] Output: "We will transport the parts. Where would you like us to transport them?"

[1535] In this way, the system provides natural and human-like interactions that respond to the user's emotions, and can provide appropriate responses that take into account the emotional state of the worker, especially when giving work instructions in a factory.

[1536] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1537] Step 1:

[1538] The user provides voice input via a microphone. For example, they might say, "Please take this part to unit number 5." This analog voice signal is captured by the terminal's microphone.

[1539] Input: User's analog audio signal

[1540] Output: Captured analog audio signal

[1541] Operation: The user gives instructions to the system via the microphone.

[1542] Step 2:

[1543] The terminal converts the captured analog audio signal into digital data. This digital data must be in a format that the speech recognition engine can process.

[1544] Input: Captured analog audio signal

[1545] Output: Digital audio data

[1546] Operation: The device's audio conversion module converts analog audio signals into digital data.

[1547] Step 3:

[1548] The terminal sends the converted digital audio data to the server via the network. The server receives this data.

[1549] Input: Digital audio data

[1550] Output: Digital audio data transferred to the server

[1551] Operation: The device sends digital audio data to the server.

[1552] Step 4:

[1553] The server uses a speech recognition engine to convert digital speech data into text data. For example, it might generate text data such as, "Please take this part to Unit 5."

[1554] Input: Digital audio data

[1555] Output: Text data

[1556] Operation: The server's speech recognition engine converts digital speech data into text data.

[1557] Step 5:

[1558] The server passes this text data to a natural language processing engine, which analyzes the user's intent and the content of the question. For example, it might analyze that the user is giving instructions to transport parts.

[1559] Input: Text data

[1560] Output: Analyzed intent and content

[1561] Operation: The server's natural language processing engine analyzes the text data to identify the user's intent.

[1562] Step 6:

[1563] The server uses an emotion engine to recognize emotional information from the user's voice. For example, it can determine whether the user is excited or calm.

[1564] Input: Text data

[1565] Output: User sentiment information

[1566] Operation: The server's emotion engine analyzes and extracts emotional information from the audio data.

[1567] Step 7:

[1568] Based on the analysis results, the server searches for the necessary information from internal and external databases. If the information is not found in the internal database, it retrieves it from the external database.

[1569] Input: Analyzed intent and content, user sentiment information

[1570] Output: Required information

[1571] Operation: The server searches for the necessary information from internal and external databases.

[1572] Step 8:

[1573] Generative artificial intelligence is used to generate appropriate answers to user questions. During this process, the content and tone of the responses are adjusted based on emotional information.

[1574] Input: Required information, user sentiment information

[1575] Output: Generated text answer

[1576] Operation: The server's generational artificial intelligence generates the appropriate answer.

[1577] Step 9:

[1578] The generated text response is passed to a speech synthesis engine and converted into speech data. For example, speech data such as "I will transport the parts. Where do you want me to transport them?" is generated.

[1579] Input: Generated text answer

[1580] Output: Audio data

[1581] Operation: The server's speech synthesis engine converts text data into speech data.

[1582] Step 10:

[1583] The server sends audio data to the terminal, which then plays this audio data to respond to the user. The response speed and volume are adjusted according to the user's emotions.

[1584] Input: Audio data

[1585] Output: Audio response played to the user

[1586] Operation: The device plays audio data and provides a response to the user.

[1587] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1588] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1589] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1590] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1591] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1592] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1593] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1594] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1595] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1596] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1597] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1598] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1599] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1600] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1601] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1602] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1603] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1604] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1605] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1606] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1607] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1608] The following is further disclosed regarding the embodiments described above.

[1609] (Claim 1)

[1610] A means of receiving the user's voice and converting it into digital data,

[1611] A means of transmitting the above digital data to a server,

[1612] The above server provides a means for converting digital data into text data using a speech recognition engine,

[1613] The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided.

[1614] Based on the above analysis results, a means for retrieving necessary information from internal and external databases,

[1615] A means for generating answers to user questions using generative artificial intelligence,

[1616] A means including a speech synthesis engine that converts the above-generated response into audio data,

[1617] A means of sending the above audio data to the terminal and responding to the user,

[1618] A system that includes this.

[1619] (Claim 2)

[1620] The system according to claim 1, further comprising means for obtaining information from an external database when the information does not exist in the internal database.

[1621] (Claim 3)

[1622] The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content.

[1623] "Example 1"

[1624] (Claim 1)

[1625] A means of receiving the user's voice and converting it into digital data,

[1626] A means of transmitting the above digital data to a network,

[1627] A means of converting digital data into text data using a speech recognition engine on a computer on the above network,

[1628] The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided.

[1629] Based on the above analysis results, a means for retrieving necessary information from internal and external storage devices,

[1630] A means for generating answers to user questions using generative artificial intelligence,

[1631] A means including a speech synthesis engine that converts the above-generated response into audio data,

[1632] A means of sending the above audio data to the terminal and responding to the user,

[1633] A system that includes this.

[1634] (Claim 2)

[1635] The system according to claim 1, further comprising means for obtaining information from an external storage device when no information exists in the internal storage device.

[1636] (Claim 3)

[1637] The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content.

[1638] "Application Example 1"

[1639] (Claim 1)

[1640] A means of receiving the user's voice and converting it into digital data,

[1641] A means of transmitting the above digital data to a server,

[1642] The above server provides a means for converting digital data into text data using a speech recognition engine,

[1643] The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided.

[1644] Based on the above analysis results, a means for retrieving necessary information from internal and external databases,

[1645] A means for generating answers to user questions using generative artificial intelligence,

[1646] A means including a speech synthesis engine that converts the above-generated response into audio data,

[1647] A means of sending the above audio data to the terminal and responding to the user,

[1648] A means of providing customers at physical stores with quick and appropriate voice answers to their questions regarding product information, store guidance, and inventory checks,

[1649] A system that includes this.

[1650] (Claim 2)

[1651] The system according to claim 1, further comprising means by which a customer of a physical store provides location information of the product within the store.

[1652] (Claim 3)

[1653] The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content.

[1654] "Example 2 of combining an emotion engine"

[1655] (Claim 1)

[1656] A means of receiving the user's voice and converting it into digital data,

[1657] A means of transmitting the above digital data to a server,

[1658] The above server provides a means for converting digital data into text data using a speech recognition engine,

[1659] The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided.

[1660] Based on the above text data, a means of analyzing user emotions,

[1661] Based on the above analysis results and sentiment analysis results, a means for retrieving necessary information from internal and external databases,

[1662] A means for generating answers to user questions using generative artificial intelligence,

[1663] A means including a speech synthesis engine that converts the above-generated response into audio data,

[1664] A means of sending the above audio data to the terminal and responding to the user,

[1665] A system that includes this.

[1666] (Claim 2)

[1667] The system according to claim 1, further comprising means for obtaining information from an external database when the information does not exist in the internal database.

[1668] (Claim 3)

[1669] The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content.

[1670] "Application example 2 when combining with an emotional engine"

[1671] (Claim 1)

[1672] A means of receiving the user's voice and converting it into digital data,

[1673] A means of transmitting the above digital data to a server,

[1674] The above server provides a means for converting digital data into text data using a speech recognition engine,

[1675] The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided.

[1676] Based on the above analysis results, a means for retrieving necessary information from internal and external databases,

[1677] A means for generating answers to user questions using generative artificial intelligence,

[1678] A means including a speech synthesis engine that converts the above-generated response into audio data,

[1679] A means of sending the above audio data to the terminal and responding to the user,

[1680] A means of recognizing emotional information from the user's voice and adjusting the response,

[1681] A means of providing an appropriate response according to the user's emotions by adjusting the response speed and volume,

[1682] A system that includes this.

[1683] (Claim 2)

[1684] The system according to claim 1, further comprising means for obtaining information from an external database when the information does not exist in the internal database.

[1685] (Claim 3)

[1686] The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content. [Explanation of symbols]

[1687] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving the user's voice and converting it into digital data, A means of transmitting the above digital data to a server, The above server provides a means for converting digital data into text data using a speech recognition engine, The above text data is passed to a natural language processing engine, and a means of analyzing its intent and content is provided. Based on the above analysis results, a means for retrieving necessary information from internal and external databases, A means for generating answers to user questions using generative artificial intelligence, A means including a speech synthesis engine that converts the above-generated response into audio data, A means of sending the above audio data to the terminal and responding to the user, A system that includes this.

2. The system according to claim 1, further comprising means for obtaining information from an external database when the information does not exist in the internal database.

3. The system according to claim 1, further comprising means for performing analysis based on specific attributes when the above-mentioned natural language processing engine analyzes the user's intent and content.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A