system
The system provides real-time audio responses to spectators' questions by converting audio to text, analyzing intent, and retrieving information from a database, addressing understanding complexities in sports and entertainment events.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Spectators face difficulties in understanding sports and entertainment events due to complex rules and language barriers, with existing systems lacking real-time information provision and response to their queries.
A system that converts audio data to text, analyzes user intent, retrieves relevant information from a database, and generates audio responses in real-time to address spectators' questions, utilizing speech recognition, natural language processing, and text-to-speech conversion.
Enables spectators to receive instant answers to their questions, enhancing their understanding and enjoyment of sports and entertainment events.
Smart Images

Figure 2026047977000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern sports watching and entertainment events, it is difficult for spectators to instantly understand the rules, game situations, or the background of the program or content. In addition, since a high level of language ability is required to understand foreign language content in real time, many spectators cannot fully enjoy it. Moreover, when listening to the explanation of a commentator and the explanation is insufficient, there is no means to solve the doubts of the spectators on the spot. There is a need for technical means to solve such problems.
Means for Solving the Problems
[0005] This invention provides a system for acquiring audio data and transmitting it to a server. The server converts the audio data into text, analyzes the text to understand the user's intent, extracts relevant information from a database based on the analysis results, and generates an answer from the extracted information. The generated answer is converted into audio data and transmitted to the user's terminal. The user's terminal provides the user with appropriate information by playing the audio data. This system allows spectators to resolve their questions in real time and enjoy sports events and entertainment events more deeply.
[0006] "Audio data" refers to data in which audio is recorded and stored in digital format.
[0007] "Means" refers to specific methods or devices used to achieve a particular objective.
[0008] A "server" refers to a computer system that provides resources and services over a network.
[0009] "Text" refers to string information generated by analyzing audio data.
[0010] "Analysis" refers to the process of deciphering text data and understanding its meaning and intent.
[0011] "User" refers to an individual or group that uses the system.
[0012] A "database" refers to a collection of data that systematically organizes and stores specific information, allowing it to be searched and used as needed.
[0013] "Speech conversion" refers to the process of converting text data back into speech data.
[0014] "Playback" refers to the process of outputting audio data so that the user can listen to it.
Brief Description of the Drawings
[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Modes for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0037] User voice input
[0038] When a user has a question while watching a game, they can naturally ask it into their device, such as, "What are the rules of this play?" The device captures the user's voice through its microphone.
[0039] Sending audio data
[0040] The terminal sends the captured audio data to the server. During transmission, the audio data is converted into a byte stream and packaged as part of the HTTP request.
[0041] Speech recognition and analysis
[0042] The server passes the received audio data to a speech recognition module, which converts the audio data into text data. A highly accurate speech recognition engine is used for this conversion. The converted text data is then analyzed by a natural language processing engine to understand the intent of the question.
[0043] Database referencing and response generation
[0044] Based on the analysis results, the server queries a real-time database to retrieve relevant information. For example, if a user requests information about the rules of play, it extracts information about the relevant rules from the database. Based on this information, it generates an appropriate response. The generated response is in text format, which is then converted into audio data.
[0045] Sending and playing audio data
[0046] The server returns the generated audio data to the terminal, which receives the data and plays it back to the user. In this way, the user can get answers in real time, such as "This violates the offside rule."
[0047] Specific example
[0048] For example, let's say a user is watching a soccer match. If they have a question about a particular play and ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text, "What is the rule for this play?", and parses the text. Based on the parse results, the server searches its database and retrieves information about the rule, for example, "offside". It generates this as a text response, converts it back into audio data, and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0049] In this way, users can get real-time answers to their questions the moment they arise, allowing them to enjoy the viewing experience more deeply. Furthermore, this system can be applied to explanations of classical performing arts and translations of foreign films. By providing expert information in response to users' arbitrary questions, it enables them to enjoy various entertainment content from multiple perspectives.
[0050] The following describes the processing flow.
[0051] Step 1: Start voice input
[0052] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0053] Step 2: Acquire audio data
[0054] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0055] Step 3: Preprocessing of audio data
[0056] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0057] Step 4: Sending audio data
[0058] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0059] Step 5: Speech Recognition Processing
[0060] Server: Converts the speech data received by the speech recognition module into text. For example, it uses Google® Cloud Speech-to-Text API or Microsoft® Azure® Cognitive Services.
[0061] Step 6: Natural Language Processing (NLP)
[0062] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0063] Step 7: Database query
[0064] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[0065] Step 8: Generating the answer
[0066] Server: Generates appropriate responses for the user based on information retrieved from the database. For example, it might create a response such as, "This violates the offside rule."
[0067] Step 9: Text to Speech
[0068] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[0069] Step 10: Sending audio data
[0070] Server: Sends the generated audio data to the terminal as an HTTP response.
[0071] Step 11: Receiving and Playing Audio
[0072] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This violates the offside rule."
[0073] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply.
[0074] (Example 1)
[0075] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0076] There is a need to provide real-time information on sports and entertainment events and quickly resolve user questions. However, conventional systems have low accuracy in speech recognition and natural language processing, making it difficult to accurately obtain the information users are looking for. In addition, when real-time is required, the overall processing speed becomes slow. As a result, users cannot enjoy the event and their satisfaction level is low.
[0077] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0078] In this invention, the server includes means for passing voice data to a speech recognition engine, means for converting the voice data into text data, means for passing the text data to a natural language processing engine for analysis, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the generated answer into voice data, and means for transmitting the voice data to a terminal. This enables the entire process from voice input to answering to be carried out quickly and accurately, allowing users to obtain event information in real time and enjoy a highly satisfying viewing experience.
[0079] "Audio data" refers to a digital representation of sound, including user voice input.
[0080] A "byte stream" is a sequence of consecutive bytes, and is a format for transferring data over a communication path.
[0081] A "speech recognition engine" is software or hardware used to convert speech data into text data.
[0082] "Text data" refers to sentence-format data extracted from audio data by a speech recognition engine.
[0083] A "natural language processing engine" is software or hardware that analyzes text data and understands its meaning and intent.
[0084] "Analysis results" refer to information about the meaning and intent of text data analyzed by a natural language processing engine.
[0085] A "database" is a system for storing and managing structured data, and is used by servers to extract relevant information.
[0086] "Answer" refers to the content of the response to a user's question, generated based on analysis results and information extracted from the database.
[0087] A "speech synthesis engine" is software or hardware used to convert text data into speech data.
[0088] A "terminal" is an electronic device used by a user to input voice data and play back audio data.
[0089] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0090] Users can ask questions to their devices while watching sporting or entertainment events. For example, if they have a question about a particular play during a soccer match, they can naturally ask, "What are the rules for this play?" The devices have built-in microphones that capture the user's voice in real time.
[0091] The captured audio data is converted into a byte stream on the device. The converted audio data is then sent to the server as an HTTP POST request. The device utilizes an internet connection during this process.
[0092] On the server side, the received audio data is passed to a high-precision speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio into text data. This speech recognition engine can quickly and accurately convert the audio data into text. The converted text data is then sent to a natural language processing engine (e.g., a generative AI model) for analysis. This natural language processing engine understands the intent of the user's question and performs the necessary analysis to generate an appropriate response.
[0093] Based on the analysis results, the server queries a database (e.g., a relational database) to retrieve relevant information. For example, if a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. The server then generates an appropriate answer based on the retrieved information and uses a speech synthesis engine (e.g., a speech synthesis API) to convert that answer into speech data.
[0094] The generated audio data is sent from the server to the terminal, and the terminal provides the answer to the user by playing back the received audio data. In this way, the user can obtain answers to questions in real time.
[0095] As a concrete example, let's say a user is watching a soccer match. If a question arises about a particular play and they ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text data, "What is the rule for this play?", and analyzes the text. Based on the analysis, the server retrieves rule information about "offside" from its database and generates it as the answer. Then, it uses a speech synthesis engine to convert it into audio data and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0096] The following are some examples of prompt statements.
[0097] User input: "What are the rules of this game?"
[0098] Prompts for generative AI models:
[0099] "The user is watching a soccer match and is asking about the rules regarding a specific play. Please explain the correct rule based on the following question: What are the rules for this play?"
[0100] Thus, the system of the present invention allows users to resolve their questions efficiently in real time and enjoy a richer viewing experience.
[0101] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0102] Step 1:
[0103] The user provides voice input. When a question arises during a game, the user can naturally ask it into the device, such as, "What are the rules of this play?" The device has a built-in microphone that captures the user's voice in real time.
[0104] Input: User's voice
[0105] Output: Captured audio data
[0106] Step 2:
[0107] The terminal converts the audio data into a byte stream. The terminal saves the captured audio data as digital audio data and converts it back into a byte stream.
[0108] Input: Captured audio data
[0109] Output: Audio data in byte stream format
[0110] Step 3:
[0111] The terminal sends audio data to the server. The terminal sends the converted byte stream of audio data to the server as an HTTP POST request. The terminal uses an internet connection during this process.
[0112] Input: Byte stream audio data
[0113] Output: HTTP request to the server
[0114] Step 4:
[0115] The server receives the audio data and passes it to the speech recognition engine. The server passes the received audio data to the speech recognition engine (general-purpose speech recognition API) and starts processing.
[0116] Input: Audio data in an HTTP request
[0117] Output: Audio data passed to the speech recognition engine
[0118] Step 5:
[0119] The server converts the audio data into text data. The speech recognition engine converts the audio data into text data. For example, the text data "What are the rules of this play?" is generated.
[0120] Input: Audio data passed to the speech recognition engine.
[0121] Output: Generated text data
[0122] Step 6:
[0123] The server passes the text data to a natural language processing engine for analysis. The server then passes the converted text data to the natural language processing engine (generative AI model) to analyze the intent of the question.
[0124] Input: Text data
[0125] Output: Data analyzing the intent behind the questions
[0126] Step 7:
[0127] The server queries the database based on the analysis results. The server queries the database (relational database) based on the analysis results and retrieves the relevant rules and information.
[0128] Input: Data analyzing the intent of the question
[0129] Output: Related information extracted from the database
[0130] Step 8:
[0131] The server generates the response and converts it into audio data. Based on the acquired information, the server generates a text-based response, and then converts that text into audio data using a speech synthesis engine (speech synthesis API).
[0132] Input: Related information extracted from the database
[0133] Output: Speech data generated by the speech synthesis engine
[0134] Step 9:
[0135] The server sends audio data to the terminal. The server then sends the generated audio data back to the terminal. During this process, the server utilizes an internet connection.
[0136] Input: Generated audio data
[0137] Output: Sending audio data to the terminal
[0138] Step 10:
[0139] The device receives and plays audio data. The device receives audio data sent from the server and plays it for the user. The user can receive real-time responses from the device, such as "This violates the offside rule."
[0140] Input: Audio data sent to the terminal
[0141] Output: Audio response played back to the user
[0142] (Application Example 1)
[0143] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0144] Within the factory, problems related to robot operation and troubleshooting frequently occur, and employees are unable to respond quickly, posing a challenge. Furthermore, the inability of employees to quickly obtain the information necessary for robot operation and problem solving poses a risk of decreased production efficiency. Traditional methods often involve manual information retrieval, which is time-consuming in finding appropriate solutions.
[0145] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0146] In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data to text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the answer to audio data, means for transmitting audio data to a user terminal, means for playing the audio data on the user terminal, means for capturing audio input to smart glasses, means for playing the audio data on the smart glasses, and means for querying a database for real-time answers to questions regarding the operation and troubleshooting of robots in the factory. This enables factory employees to obtain information regarding the operation and troubleshooting of robots in real time, facilitating rapid problem solving and improved production efficiency.
[0147] "Means for acquiring audio data" refers to a device or software that receives a user's speech as an audio signal and records that audio signal as digital data.
[0148] "Means for transmitting audio data to a server" refers to a device or software that transmits acquired audio data to a remote server via a network.
[0149] "Means for converting audio data to text" refers to a device or software that analyzes an audio signal and outputs its contents in text format.
[0150] "Means for analyzing text and understanding user intent" refers to a device or software that analyzes text data using natural language processing technology, understands its content, and identifies the intent behind the user's questions or instructions.
[0151] "Means for extracting relevant information from a database based on analysis results" refers to a device or software that searches for and extracts appropriate information from a database based on the analyzed user intent.
[0152] "Means for generating answers based on extracted information" refers to a device or software that automatically generates appropriate answers to user questions based on information extracted from a database.
[0153] "Means for converting responses into audio data" refers to a device or software that converts generated text-formatted responses into audio signals.
[0154] "Means for transmitting audio data to a user's terminal" refers to a device or software that transmits generated audio data to a user's terminal via a network.
[0155] "Means for playing audio data on a user terminal" refers to a device or software that plays received audio data and provides the user with an audio response.
[0156] "Means for capturing voice input on smart glasses" refers to a device or software that acquires the user's voice using a microphone installed in smart glasses.
[0157] "Means of playing audio data with smart glasses" refers to a device or software that plays audio data using speakers or bone conduction devices installed in smart glasses.
[0158] "Means of querying a database to provide real-time answers to questions regarding the operation and troubleshooting of robots in a factory" refers to a device or software that searches a database containing information on robot operation and errors in real time and obtains answers to questions.
[0159] This invention relates to a system for rapidly operating and troubleshooting robots in a factory. The system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0160] Hardware and software configuration
[0161] hardware
[0162] Smart glasses: These are devices that have a built-in microphone and speaker, capture the user's voice input, and play back the audio data.
[0163] Server: A computer device equipped with a high-performance processor and large-capacity memory, used for data processing and database retrieval.
[0164] software
[0165] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert user voice input into text data.
[0166] Natural language processing engine: Uses OpenAI® GPT-4® to analyze text data and understand user intent.
[0167] Database: A cloud-based data management system (e.g., MySQL®) is used to store information related to robot operation and troubleshooting.
[0168] Communication technology: The HTTP protocol is used to send and receive voice data and analysis results between the server and the smart glasses.
[0169] System operation
[0170] 1. Voice input capture
[0171] The user voice-inputs "Why did the robot stop working?" into the smart glasses. The smart glasses' microphone captures this voice.
[0172] 2. Sending audio data
[0173] The smart glasses convert the captured audio data into a byte stream and send it to the server as an HTTP request.
[0174] 3. Speech Recognition and Analysis
[0175] The server receives the audio data and converts it to text using Google Cloud Speech-to-Text. Then, OpenAI GPT-4 is used to analyze the text data and understand the user's intent.
[0176] 4. Database referencing and response generation
[0177] The server queries a cloud-based database based on the analysis results and extracts relevant information. Based on the information obtained, it generates an appropriate response.
[0178] 5. Converting and sending responses as audio data.
[0179] The server generates text-based responses, which are then converted into audio data and sent to the smart glasses.
[0180] 6. Playback of audio data
[0181] The smart glasses play back the received audio data and provide the user with a response such as, "There may be a sensor malfunction due to error code 123."
[0182] Examples of specific cases and prompt statements
[0183] For example, if a robot suddenly stops working in a factory, an employee might ask the smart glasses, "Why did the robot stop working?" This voice input is sent to a server, which performs speech recognition and text analysis, then queries a database to generate an answer. This answer is then sent as voice data to the smart glasses, allowing the employee to receive the cause of the problem and its solution in real time via voice.
[0184] Examples of prompt statements are as follows:
[0185] Employee question: Why did the robot stop working?
[0186] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0187] Step 1:
[0188] The user inputs a question by voice into the smart glasses. Specifically, the user says, "Why did the robot stop working?" and the smart glasses' microphone captures the voice. At this stage, the input is in the form of voice data.
[0189] Step 2:
[0190] The smart glasses convert the captured audio data into a byte stream. This processes the audio data into a format that can be transmitted over the network. The converted data is then sent to the server as an HTTP request. The output is audio data in byte stream format.
[0191] Step 3:
[0192] The server receives audio data and passes it to the Google Cloud Speech-to-Text speech recognition engine, which converts it into text data. The input is audio data in byte stream format, and the output is text data. Specifically, the speech recognition engine analyzes the audio signal and converts it from speech to text.
[0193] Step 4:
[0194] The server passes text data to OpenAI GPT-4, which performs natural language processing to analyze the user's intent. The input is text data converted from speech, and the output is text data including the analysis results. Specifically, GPT-4 analyzes the text and performs the process of understanding the user's question.
[0195] Step 5:
[0196] The server queries the database based on the analysis results to retrieve relevant information. The input is text data reflecting the analysis results, and the output is a database record containing the relevant information. Specifically, the server generates and executes an SQL query to retrieve the relevant information.
[0197] Step 6:
[0198] The server generates appropriate answers to user questions based on information retrieved from the database. The input is information retrieved from the database, and the output is answer text data. Specifically, the server combines the retrieved information to generate appropriate answer sentences.
[0199] Step 7:
[0200] The system converts the server-generated response text into audio data. The input is text data, and the output is audio data. Specifically, it uses a text-to-speech conversion engine to create a natural-sounding response.
[0201] Step 8:
[0202] The server converts the audio data and sends it to the smart glasses. The input is audio data, and the output is the audio data that reaches the smart glasses. Specifically, the server performs the process of sending the audio data as an HTTP response.
[0203] Step 9:
[0204] The smart glasses play back the received audio data to the user. The input is the received audio data, and the output is the audio information the user hears. Specifically, the smart glasses use their speakers or bone conduction device to play back the voice response.
[0205] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0206] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0207] User voice input
[0208] When a user has a question while watching a game, they can ask, "What are the rules for this play?" The device captures the user's voice via the microphone.
[0209] Sending audio data
[0210] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request.
[0211] Speech recognition and analysis
[0212] The server receives the audio data and converts it to text using a speech recognition module. This text is then analyzed by a natural language processing (NLP) engine to understand the user's intent.
[0213] Database referencing and response generation
[0214] The server queries a real-time database based on the analysis results to retrieve relevant information. For example, if a query is made about the rules for a specific play, it extracts the relevant rule information from the database. Based on that information, it generates an appropriate response.
[0215] Use of the emotion engine
[0216] The emotion engine recognizes emotions from the user's voice. This emotional information is taken into consideration during the response generation process. For example, if the user is excited, a concise and clear response will be provided. On the other hand, if the user is confused, a response with added detail will be provided.
[0217] Text to speech conversion
[0218] The generated responses are converted into speech data by a speech synthesis module. A Text-to-Speech (TTS) engine is used.
[0219] Sending and playing audio data
[0220] The server sends the generated audio data back to the terminal. The terminal receives the audio data and plays it back through its built-in speaker, providing the user with an appropriate response.
[0221] Specific example
[0222] If a user asks "What is the rule for this play?" while watching a soccer match, the device captures the audio and sends it to a server. The server converts the audio data into text and analyzes it with a natural language processing engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information, such as information about the "offside" rule.
[0223] In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear response. Something like, "This is the offside rule," is converted into audio data, sent to the device, and played back. This allows the user to receive a quick and appropriate response, enhancing their viewing experience.
[0224] The introduction of an emotion engine enables more personalized responses that reflect the user's emotional state, further improving system usability and user satisfaction.
[0225] The following describes the processing flow.
[0226] Step 1: Start voice input
[0227] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0228] Step 2: Acquire audio data
[0229] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0230] Step 3: Preprocessing of audio data
[0231] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0232] Step 4: Sending audio data
[0233] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0234] Step 5: Speech Recognition Processing
[0235] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[0236] Step 6: Natural Language Processing (NLP)
[0237] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0238] Step 7: Emotion Analysis
[0239] Server: Passes text and audio data to the emotion engine to recognize the user's emotions. For example, the emotion engine analyzes whether the user is excited or confused.
[0240] Step 8: Database query
[0241] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[0242] Step 9: Emotion-Based Response Adjustment
[0243] Server: Based on the results of the emotion engine, it generates responses that correspond to the user's emotional state. For example, if the user is excited, the response will be concise and clear.
[0244] Step 10: Generating the answer
[0245] Server: Generates appropriate responses to the user, including information retrieved from the database. For example, it might create a response such as, "This is the offside rule."
[0246] Step 11: Text to Speech
[0247] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[0248] Step 12: Sending audio data
[0249] Server: Sends the generated audio data to the terminal as an HTTP response.
[0250] Step 13: Receiving and Playing Audio
[0251] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This is an offside rule."
[0252] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply. Furthermore, the introduction of an emotion engine provides more personalized answers that reflect the user's emotional state, further improving the system's usability and user satisfaction.
[0253] (Example 2)
[0254] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0255] Conventional question-answering systems struggle to provide timely and appropriate answers to user questions, especially during sporting events and entertainment events where immediate explanations of rules and other information are required. Furthermore, they are unable to generate answers that take into account the user's emotional state, which hinders the improvement of the user experience.
[0256] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for recognizing the user's emotions from the voice data and reflecting that emotional information in the response generation process, and means for converting the response into voice data. This makes it possible to provide appropriate and emotionally personalized responses to user questions in real time.
[0257] "Audio data" refers to data that represents the voice information spoken by a user in digital format.
[0258] A "byte stream" is a data format that represents and transmits digital data in the form of consecutive bytes.
[0259] A "server" is a computer system used for processing, storing, and responding to requests from clients.
[0260] "Text" refers to data obtained by converting audio data into written information.
[0261] "Analyzing text" means analyzing text, which is character information, using natural language processing techniques to understand its meaning and intent.
[0262] "User intent" refers to the specific information or actions that the user is seeking, based on what they have said.
[0263] "Related information" refers to information retrieved from the database that is relevant to the user's intent.
[0264] A "database" is a system for efficiently storing, managing, and searching large amounts of data.
[0265] An "answer" is a piece of informational text generated based on a user's question.
[0266] "Emotional recognition" means analyzing voice and other data to identify the user's emotional state (e.g., excitement, confusion).
[0267] "Emotional information" refers to data about a user's emotional state obtained through the emotion recognition process.
[0268] A "Text-to-Speech (TTS) engine" is a technology used to convert text data into speech data.
[0269] "Playing audio data" refers to the act of playing converted audio data through sound equipment such as speakers.
[0270] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0271] This system includes the following components:
[0272] 1. Device (with microphone and speaker)
[0273] 2. Server
[0274] 3. Speech recognition module
[0275] 4. Natural Language Processing (NLP) Engine
[0276] 5. Real-time database
[0277] 6. Emotional Engine
[0278] 7. Text-to-Speech (TTS) Engine
[0279] User voice input
[0280] When a user has a question while watching a game, they might ask, "What are the rules for this play?" The device captures the user's voice via its microphone. The device uses the microphone as an audio input device to acquire the audio signal as digital data.
[0281] Sending audio data
[0282] The terminal converts the captured audio data into a byte stream and sends it to the server as part of an HTTP request. The terminal uses, for example, the requests library in Python to send the audio data to the server.
[0283] Speech Recognition and Analysis
[0284] The server receives the audio data and converts it to text using a speech recognition module (e.g., Google Speech-to-Text API). This text is analyzed using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API) to understand the user's intent.
[0285] Database Reference and Answer Generation
[0286] The server queries a real-time database (e.g., Firebase or MySQL) based on the analysis results to obtain relevant information. When the user asks "What are the rules of this play?", information about the rules of the corresponding play is extracted from the database. Then, a grammatically correct answer is generated based on the information obtained using a template engine (e.g., Jinja2).
[0287] Use of Sentiment Engine
[0288] The server uses a sentiment engine (e.g., Microsoft Azure's Text Analytics) to recognize the sentiment from the user's audio data. This recognized sentiment is reflected in the answer generation process. For example, if the user is excited, the answer is adjusted to be concise and clear. On the other hand, if the user is confused, detailed explanations are added.
[0289] Text-to-Speech Conversion
[0290] The server passes the generated text response to a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert it into audio data. This audio data is saved as an audio file and used for subsequent processing.
[0291] Sending and playing audio data
[0292] The server sends the generated audio data back to the terminal, which receives the data and plays it back through its built-in speaker. This allows the user to receive quick and appropriate answers via voice.
[0293] Specific example
[0294] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the device captures the audio and sends it to the server. The server converts the audio data into text and analyzes it with an NLP engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information and extracts information about the "offside" rule. In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear answer. Something like "This is the offside rule" is converted into audio data, sent to the device, and played back. This allows the user to get a quick and appropriate answer in audio, improving the viewing experience.
[0295] Example of a prompt
[0296] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[0297] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0298] Step 1:
[0299] The user issues a question. When the user has a doubt during watching a game, the user issues a question such as "What is the rule of this play?". The input is the user's voice, and the output is the digital voice data captured by the microphone.
[0300] Step 2:
[0301] The terminal captures the voice. The microphone of the terminal captures the user's voice in real time. The input is the user's voice, and the output is the digital voice data on the terminal. As a specific operation, the microphone device converts the voice signal into digital data.
[0302] <http: / / www.example.com / Step 3:
[0303] The terminal converts the voice data into a byte stream and sends it to the server. The terminal encodes the voice data into a byte stream and sends it to the server as part of an HTTP request. The input is the captured digital voice data, and the output is the byte stream sent to the server. As a specific operation, the requests library of Python is used.
[0304] Step 4: <http: / / www.example.com /
[0305] The server receives the voice data. The server receives the HTTP request and obtains the voice data included in the body. The input is the voice data in byte stream format, and the output is the digital voice data on the server.
[0306] Step 5:
[0307] The server converts the voice data into text. The server uses a voice recognition module (for example, Speech-to-Text API) to convert the voice data into text. The input is the digital voice data, and the output is the text data. As a specific operation, an API call is made to convert the voice into text information.
[0308] Step 6:
[0309] The server analyzes the text and understands the user's intent. The server analyzes the text data using a natural language processing engine (e.g., Cloud Natural Language API) to identify the user's intent. The input is text data, and the output is information about the user's intent. For example, it analyzes the question "What are the rules of this game?" and understands the intent "I want to know the rules of the game."
[0310] Step 7:
[0311] The server references a database based on the analysis results. The server queries a real-time database (e.g., Firebase) based on the analysis results and retrieves relevant information. The input is user intent information, and the output is information from the relevant database. For example, it might retrieve rule information related to a specific play from the database.
[0312] Step 8:
[0313] The server generates an answer based on the information it extracts. The server uses a template engine (e.g., Jinja2) to process the acquired information into a format that is easy for the user to understand. The input is related information, and the output is the generated answer text.
[0314] Step 9:
[0315] The server recognizes the user's emotions and adjusts the response accordingly. The server uses an emotion engine (e.g., a Text Analytics API) to analyze the user's emotional state and adjust the response. The input is audio data, and the output is the adjusted response text. For example, if the user is agitated, a concise response is generated.
[0316] Step 10:
[0317] The server converts the response into audio data. The server uses a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert the generated response text into audio data. The input is the generated response text, and the output is audio data. Specifically, it makes an API call to convert the text into an audio file.
[0318] Step 11:
[0319] The server generates audio data and sends it to the terminal. The server then sends the generated audio data back to the terminal as an HTTP response. The input is the generated audio data, and the output is the byte stream audio data sent to the terminal.
[0320] Step 12:
[0321] The terminal plays audio data. The terminal decodes the received audio data and plays it through its built-in speaker. The input is audio data in byte stream format from the server, and the output is the audio that the user hears. Specifically, the audio is played using an audio device.
[0322] Example of a prompt:
[0323] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[0324] (Application Example 2)
[0325] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0326] In traditional sports viewing and entertainment events, users lacked real-time means of resolving questions, and the answers were not optimized according to the user's emotional state. This resulted in limitations on the user's viewing and event participation experience.
[0327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for adjusting the extracted information based on the user's emotional state to generate an answer, means for converting the answer into audio data, means for transmitting the audio data to the user terminal, means for playing the audio data on the user terminal, and means for implementation as a smartphone application. As a result, the user can resolve questions in real time, and the answer is optimized for their emotional state, significantly improving the viewing experience and event participation experience.
[0328] "Audio data" refers to data that digitizes audio information, converting user questions and statements into a format that can be processed by electronic devices.
[0329] "Means of acquisition" refers to a function for capturing the user's voice, which collects voice data using the smartphone's microphone or other voice input devices.
[0330] "Means of transmission" refers to communication functions for sending acquired audio data to a server via a network, and generally refers to the means of exchanging data over the internet.
[0331] "Means of converting to text" refers to a function that converts audio data into text information, using speech recognition technology to convert audio into text data.
[0332] "Means of analysis to understand user intent" refers to analyzing data converted into text using natural language processing technology to understand user questions and requests.
[0333] "Means for extracting relevant information from a database" refers to a function that searches for and extracts necessary information from a database based on the analysis results.
[0334] "Means for generating answers" refers to a function that constructs answers to user questions based on extracted information, thereby generating appropriate answers.
[0335] The "means of adjusting and generating responses" refer to a function for optimizing responses based on the user's emotional state. This function uses emotion recognition technology to determine the user's emotions and provide responses accordingly.
[0336] "Means of converting to audio data" refers to a function that converts text-generated responses back into audio data using speech synthesis technology.
[0337] "Means of sending to the user's terminal" refers to a communication function for sending the generated audio data to the user's smartphone or other devices.
[0338] "Means of playback" refers to the function of playing back audio data received on the user's device, outputting the sound through the smartphone's speaker or earphones.
[0339] "Implementation as a smartphone application" means that all of these functions are configured as a smartphone application and designed to be easily used by users.
[0340] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0341] System Configuration
[0342] The system consists of the following hardware and software components:
[0343] 1. Acquisition of audio data
[0344] Hardware: Smartphone microphone
[0345] Software: speech_recognition library
[0346] 2. Sending audio data
[0347] Hardware: Smartphone communication module
[0348] Software: Request Library
[0349] 3. Converting audio data to text
[0350] Software: Use the speech_recognition library to convert speech data into text.
[0351] 4. Text analysis and understanding of intent
[0352] Software: Natural Language Processing (NLP) Engine
[0353] 5. Extraction of relevant information
[0354] Software: Database query function
[0355] 6. Emotion Recognition and Response Generation
[0356] Software: emotion_recognition library
[0357] Adjust responses based on the user's emotional state.
[0358] 7. Voice conversion of the answer
[0359] Software: gTTS library
[0360] 8. Sending and playing audio data
[0361] Hardware: Smartphone speaker
[0362] Software: Use mpg321 to play audio files
[0363] Implementation method
[0364] Acquisition and transmission of audio data
[0365] When a user asks "What are the rules of this game?" through the application, the device captures the audio via the microphone. This audio data is then sent to a server over the internet.
[0366] Text conversion and analysis of audio data
[0367] The server receives the transmitted audio data and converts it into text using the speech_recognition library. It then analyzes this text using a natural language processing (NLP) engine to understand the user's intent.
[0368] Extraction of relevant information and generation of responses
[0369] Based on the analysis results, relevant information is extracted from the database. For example, if asked about the rules of a specific play, the information on those rules is extracted from the database. Meanwhile, the emotion recognition engine recognizes emotions from the user's voice and takes them into consideration in the response generation process.
[0370] Voice conversion, transmission, and playback of responses
[0371] The generated response is converted into audio data by the gTTS library and sent back from the server to the terminal. The terminal receives the audio data and plays the audio using mpg321.
[0372] Specific examples and prompt statements
[0373] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the smartphone app captures the audio and sends it to a server. The server then analyzes the audio data and generates an answer such as "This is the offside rule."
[0374] Example of a prompt:
[0375] User input: "What are the rules of this game?"
[0376] Emotional state: Excited
[0377] This allows users to receive quick and appropriate answers, and by providing information tailored to their emotional state, it is possible to significantly improve the viewing experience.
[0378] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0379] Step 1:
[0380] The user launches a smartphone application and, while watching a game, makes a voice input such as, "What are the rules of this play?" The voice data is captured through the smartphone's microphone. The input is the user's voice data, and the output is the captured voice data.
[0381] Step 2:
[0382] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request. The input is the audio data, and the output is the transmission of the byte stream to the server.
[0383] Step 3:
[0384] The server converts the received audio data into text using a speech recognition module (e.g., the speech_recognition library). The input is audio data, and the output after conversion is text data.
[0385] Step 4:
[0386] The server analyzes text data using a natural language processing (NLP) engine to understand the user's intent. The input is text data, and the output is the analysis result, which represents the content of the user's question.
[0387] Step 5:
[0388] The server extracts relevant information from the database based on the analysis results. For example, if a user asks, "What are the rules for this play?", the server queries for the relevant rules and retrieves that information. The input is the analysis results, and the output is the relevant information extracted from the database.
[0389] Step 6:
[0390] An emotion recognition engine (e.g., the emotion_recognition library) recognizes the user's emotional state from their voice data. The input is voice data, and the output is the user's emotional state (excitement, confusion, etc.).
[0391] Step 7:
[0392] The server generates a response based on the extracted information and the user's emotional state. For example, if the user is agitated, the generated response might be a concise statement such as "This is the offside rule!". The input is the relevant information and emotional state, and the output is the generated response text.
[0393] Step 8:
[0394] The server converts the generated responses into speech data using a speech synthesis module (e.g., the gTTS library). The input is the response text, and the output is the speech data.
[0395] Step 9:
[0396] The server sends the generated audio data back to the terminal. The input is audio data, and the output is transmission to the terminal.
[0397] Step 10:
[0398] The device plays the received audio data. The user can hear the audio response through the device's speaker. The input is the audio data, and the output is the audio response to the user.
[0399] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0400] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0401] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0402] [Second Embodiment]
[0403] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0404] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0405] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0406] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0407] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0409] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0410] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0411] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0412] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0413] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0414] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0415] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0416] User voice input
[0417] When a user has a question while watching a game, they can naturally ask it into their device, such as, "What are the rules of this play?" The device captures the user's voice through its microphone.
[0418] Sending audio data
[0419] The terminal sends the captured audio data to the server. During transmission, the audio data is converted into a byte stream and packaged as part of the HTTP request.
[0420] Speech recognition and analysis
[0421] The server passes the received audio data to a speech recognition module, which converts the audio data into text data. A highly accurate speech recognition engine is used for this conversion. The converted text data is then analyzed by a natural language processing engine to understand the intent of the question.
[0422] Database referencing and response generation
[0423] Based on the analysis results, the server queries a real-time database to retrieve relevant information. For example, if a user requests information about the rules of play, it extracts information about the relevant rules from the database. Based on this information, it generates an appropriate response. The generated response is in text format, which is then converted into audio data.
[0424] Sending and playing audio data
[0425] The server returns the generated audio data to the terminal, which receives the data and plays it back to the user. In this way, the user can get answers in real time, such as "This violates the offside rule."
[0426] Specific example
[0427] For example, let's say a user is watching a soccer match. If they have a question about a particular play and ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text, "What is the rule for this play?", and parses the text. Based on the parse results, the server searches its database and retrieves information about the rule, for example, "offside". It generates this as a text response, converts it back into audio data, and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0428] In this way, users can get real-time answers to their questions the moment they arise, allowing them to enjoy the viewing experience more deeply. Furthermore, this system can be applied to explanations of classical performing arts and translations of foreign films. By providing expert information in response to users' arbitrary questions, it enables them to enjoy various entertainment content from multiple perspectives.
[0429] The following describes the processing flow.
[0430] Step 1: Start voice input
[0431] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0432] Step 2: Acquire audio data
[0433] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0434] Step 3: Preprocessing of audio data
[0435] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0436] Step 4: Sending audio data
[0437] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0438] Step 5: Speech Recognition Processing
[0439] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[0440] Step 6: Natural Language Processing (NLP)
[0441] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0442] Step 7: Database query
[0443] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[0444] Step 8: Generating the answer
[0445] Server: Generates appropriate responses for the user based on information retrieved from the database. For example, it might create a response such as, "This violates the offside rule."
[0446] Step 9: Text to Speech
[0447] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[0448] Step 10: Sending audio data
[0449] Server: Sends the generated audio data to the terminal as an HTTP response.
[0450] Step 11: Receiving and Playing Audio
[0451] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This violates the offside rule."
[0452] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply.
[0453] (Example 1)
[0454] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0455] There is a need to provide real-time information on sports and entertainment events and quickly resolve user questions. However, conventional systems have low accuracy in speech recognition and natural language processing, making it difficult to accurately obtain the information users are looking for. In addition, when real-time is required, the overall processing speed becomes slow. As a result, users cannot enjoy the event and their satisfaction level is low.
[0456] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0457] In this invention, the server includes means for passing voice data to a speech recognition engine, means for converting the voice data into text data, means for passing the text data to a natural language processing engine for analysis, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the generated answer into voice data, and means for transmitting the voice data to a terminal. This enables the entire process from voice input to answering to be carried out quickly and accurately, allowing users to obtain event information in real time and enjoy a highly satisfying viewing experience.
[0458] "Audio data" refers to a digital representation of sound, including user voice input.
[0459] A "byte stream" is a sequence of consecutive bytes, and is a format for transferring data over a communication path.
[0460] A "speech recognition engine" is software or hardware used to convert speech data into text data.
[0461] "Text data" refers to sentence-format data extracted from audio data by a speech recognition engine.
[0462] A "natural language processing engine" is software or hardware that analyzes text data and understands its meaning and intent.
[0463] "Analysis results" refer to information about the meaning and intent of text data analyzed by a natural language processing engine.
[0464] A "database" is a system for storing and managing structured data, and is used by servers to extract relevant information.
[0465] "Answer" refers to the content of the response to a user's question, generated based on analysis results and information extracted from the database.
[0466] A "speech synthesis engine" is software or hardware used to convert text data into speech data.
[0467] A "terminal" is an electronic device used by a user to input voice data and play back audio data.
[0468] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0469] Users can ask questions to their devices while watching sporting or entertainment events. For example, if they have a question about a particular play during a soccer match, they can naturally ask, "What are the rules for this play?" The devices have built-in microphones that capture the user's voice in real time.
[0470] The captured audio data is converted into a byte stream on the device. The converted audio data is then sent to the server as an HTTP POST request. The device utilizes an internet connection during this process.
[0471] On the server side, the received audio data is passed to a high-precision speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio into text data. This speech recognition engine can quickly and accurately convert the audio data into text. The converted text data is then sent to a natural language processing engine (e.g., a generative AI model) for analysis. This natural language processing engine understands the intent of the user's question and performs the necessary analysis to generate an appropriate response.
[0472] Based on the analysis results, the server queries a database (e.g., a relational database) to retrieve relevant information. For example, if a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. The server then generates an appropriate answer based on the retrieved information and uses a speech synthesis engine (e.g., a speech synthesis API) to convert that answer into speech data.
[0473] The generated audio data is sent from the server to the terminal, and the terminal provides the answer to the user by playing back the received audio data. In this way, the user can obtain answers to questions in real time.
[0474] As a concrete example, let's say a user is watching a soccer match. If a question arises about a particular play and they ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text data, "What is the rule for this play?", and analyzes the text. Based on the analysis, the server retrieves rule information about "offside" from its database and generates it as the answer. Then, it uses a speech synthesis engine to convert it into audio data and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0475] The following are some examples of prompt statements.
[0476] User input: "What are the rules of this game?"
[0477] Prompts for generative AI models:
[0478] "The user is watching a soccer match and is asking about the rules regarding a specific play. Please explain the correct rule based on the following question: What are the rules for this play?"
[0479] Thus, the system of the present invention allows users to resolve their questions efficiently in real time and enjoy a richer viewing experience.
[0480] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0481] Step 1:
[0482] The user provides voice input. When a question arises during a game, the user can naturally ask it into the device, such as, "What are the rules of this play?" The device has a built-in microphone that captures the user's voice in real time.
[0483] Input: User's voice
[0484] Output: Captured audio data
[0485] Step 2:
[0486] The terminal converts the audio data into a byte stream. The terminal saves the captured audio data as digital audio data and converts it back into a byte stream.
[0487] Input: Captured audio data
[0488] Output: Audio data in byte stream format
[0489] Step 3:
[0490] The terminal sends audio data to the server. The terminal sends the converted byte stream of audio data to the server as an HTTP POST request. The terminal uses an internet connection during this process.
[0491] Input: Byte stream audio data
[0492] Output: HTTP request to the server
[0493] Step 4:
[0494] The server receives the audio data and passes it to the speech recognition engine. The server passes the received audio data to the speech recognition engine (general-purpose speech recognition API) and starts processing.
[0495] Input: Audio data in an HTTP request
[0496] Output: Audio data passed to the speech recognition engine
[0497] Step 5:
[0498] The server converts the audio data into text data. The speech recognition engine converts the audio data into text data. For example, the text data "What are the rules of this play?" is generated.
[0499] Input: Audio data passed to the speech recognition engine.
[0500] Output: Generated text data
[0501] Step 6:
[0502] The server passes the text data to a natural language processing engine for analysis. The server then passes the converted text data to the natural language processing engine (generative AI model) to analyze the intent of the question.
[0503] Input: Text data
[0504] Output: Data analyzing the intent behind the questions
[0505] Step 7:
[0506] The server queries the database based on the analysis results. The server queries the database (relational database) based on the analysis results and retrieves the relevant rules and information.
[0507] Input: Data analyzing the intent of the question
[0508] Output: Related information extracted from the database
[0509] Step 8:
[0510] The server generates the response and converts it into audio data. Based on the acquired information, the server generates a text-based response, and then converts that text into audio data using a speech synthesis engine (speech synthesis API).
[0511] Input: Related information extracted from the database
[0512] Output: Speech data generated by the speech synthesis engine
[0513] Step 9:
[0514] The server sends audio data to the terminal. The server then sends the generated audio data back to the terminal. During this process, the server utilizes an internet connection.
[0515] Input: Generated audio data
[0516] Output: Sending audio data to the terminal
[0517] Step 10:
[0518] The device receives and plays audio data. The device receives audio data sent from the server and plays it for the user. The user can receive real-time responses from the device, such as "This violates the offside rule."
[0519] Input: Audio data sent to the terminal
[0520] Output: Audio response played back to the user
[0521] (Application Example 1)
[0522] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0523] Within the factory, problems related to robot operation and troubleshooting frequently occur, and employees are unable to respond quickly, posing a challenge. Furthermore, the inability of employees to quickly obtain the information necessary for robot operation and problem solving poses a risk of decreased production efficiency. Traditional methods often involve manual information retrieval, which is time-consuming in finding appropriate solutions.
[0524] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0525] In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data to text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the answer to audio data, means for transmitting audio data to a user terminal, means for playing the audio data on the user terminal, means for capturing audio input to smart glasses, means for playing the audio data on the smart glasses, and means for querying a database for real-time answers to questions regarding the operation and troubleshooting of robots in the factory. This enables factory employees to obtain information regarding the operation and troubleshooting of robots in real time, facilitating rapid problem solving and improved production efficiency.
[0526] "Means for acquiring audio data" refers to a device or software that receives a user's speech as an audio signal and records that audio signal as digital data.
[0527] "Means for transmitting audio data to a server" refers to a device or software that transmits acquired audio data to a remote server via a network.
[0528] "Means for converting audio data to text" refers to a device or software that analyzes an audio signal and outputs its contents in text format.
[0529] "Means for analyzing text and understanding user intent" refers to a device or software that analyzes text data using natural language processing technology, understands its content, and identifies the intent behind the user's questions or instructions.
[0530] "Means for extracting relevant information from a database based on analysis results" refers to a device or software that searches for and extracts appropriate information from a database based on the analyzed user intent.
[0531] "Means for generating answers based on extracted information" refers to a device or software that automatically generates appropriate answers to user questions based on information extracted from a database.
[0532] "Means for converting responses into audio data" refers to a device or software that converts generated text-formatted responses into audio signals.
[0533] "Means for transmitting audio data to a user's terminal" refers to a device or software that transmits generated audio data to a user's terminal via a network.
[0534] "Means for playing audio data on a user terminal" refers to a device or software that plays received audio data and provides the user with an audio response.
[0535] "Means for capturing voice input on smart glasses" refers to a device or software that acquires the user's voice using a microphone installed in smart glasses.
[0536] "Means of playing audio data with smart glasses" refers to a device or software that plays audio data using speakers or bone conduction devices installed in smart glasses.
[0537] "Means of querying a database to provide real-time answers to questions regarding the operation and troubleshooting of robots in a factory" refers to a device or software that searches a database containing information on robot operation and errors in real time and obtains answers to questions.
[0538] This invention relates to a system for rapidly operating and troubleshooting robots in a factory. The system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0539] Hardware and software configuration
[0540] hardware
[0541] Smart glasses: These are devices that have a built-in microphone and speaker, capture the user's voice input, and play back the audio data.
[0542] Server: A computer device equipped with a high-performance processor and large-capacity memory, used for data processing and database retrieval.
[0543] software
[0544] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert user voice input into text data.
[0545] The system uses the OpenAI GPT-4 natural language processing engine to analyze text data and understand user intent.
[0546] Database: A cloud-based data management system (e.g., MySQL) is used to store information related to robot operation and troubleshooting.
[0547] Communication technology: The HTTP protocol is used to send and receive voice data and analysis results between the server and the smart glasses.
[0548] System operation
[0549] 1. Voice input capture
[0550] The user voice-inputs "Why did the robot stop working?" into the smart glasses. The smart glasses' microphone captures this voice.
[0551] 2. Sending audio data
[0552] The smart glasses convert the captured audio data into a byte stream and send it to the server as an HTTP request.
[0553] 3. Speech Recognition and Analysis
[0554] The server receives the audio data and converts it to text using Google Cloud Speech-to-Text. Then, OpenAI GPT-4 is used to analyze the text data and understand the user's intent.
[0555] 4. Database referencing and response generation
[0556] The server queries a cloud-based database based on the analysis results and extracts relevant information. Based on the information obtained, it generates an appropriate response.
[0557] 5. Converting and sending responses as audio data.
[0558] The server generates text-based responses, which are then converted into audio data and sent to the smart glasses.
[0559] 6. Playback of audio data
[0560] The smart glasses play back the received audio data and provide the user with a response such as, "There may be a sensor malfunction due to error code 123."
[0561] Examples of specific cases and prompt statements
[0562] For example, if a robot suddenly stops working in a factory, an employee might ask the smart glasses, "Why did the robot stop working?" This voice input is sent to a server, which performs speech recognition and text analysis, then queries a database to generate an answer. This answer is then sent as voice data to the smart glasses, allowing the employee to receive the cause of the problem and its solution in real time via voice.
[0563] Examples of prompt statements are as follows:
[0564] Employee question: Why did the robot stop working?
[0565] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0566] Step 1:
[0567] The user inputs a question by voice into the smart glasses. Specifically, the user says, "Why did the robot stop working?" and the smart glasses' microphone captures the voice. At this stage, the input is in the form of voice data.
[0568] Step 2:
[0569] The smart glasses convert the captured audio data into a byte stream. This processes the audio data into a format that can be transmitted over the network. The converted data is then sent to the server as an HTTP request. The output is audio data in byte stream format.
[0570] Step 3:
[0571] The server receives audio data and passes it to the Google Cloud Speech-to-Text speech recognition engine, which converts it into text data. The input is audio data in byte stream format, and the output is text data. Specifically, the speech recognition engine analyzes the audio signal and converts it from speech to text.
[0572] Step 4:
[0573] The server passes text data to OpenAI GPT-4, which performs natural language processing to analyze the user's intent. The input is text data converted from speech, and the output is text data including the analysis results. Specifically, GPT-4 analyzes the text and performs the process of understanding the user's question.
[0574] Step 5:
[0575] The server queries the database based on the analysis results to retrieve relevant information. The input is text data reflecting the analysis results, and the output is a database record containing the relevant information. Specifically, the server generates and executes an SQL query to retrieve the relevant information.
[0576] Step 6:
[0577] The server generates appropriate answers to user questions based on information retrieved from the database. The input is information retrieved from the database, and the output is answer text data. Specifically, the server combines the retrieved information to generate appropriate answer sentences.
[0578] Step 7:
[0579] The system converts the server-generated response text into audio data. The input is text data, and the output is audio data. Specifically, it uses a text-to-speech conversion engine to create a natural-sounding response.
[0580] Step 8:
[0581] The server converts the audio data and sends it to the smart glasses. The input is audio data, and the output is the audio data that reaches the smart glasses. Specifically, the server performs the process of sending the audio data as an HTTP response.
[0582] Step 9:
[0583] The smart glasses play back the received audio data to the user. The input is the received audio data, and the output is the audio information the user hears. Specifically, the smart glasses use their speakers or bone conduction device to play back the voice response.
[0584] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0585] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0586] User voice input
[0587] When a user has a question while watching a game, they can ask, "What are the rules for this play?" The device captures the user's voice via the microphone.
[0588] Sending audio data
[0589] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request.
[0590] Speech recognition and analysis
[0591] The server receives the audio data and converts it to text using a speech recognition module. This text is then analyzed by a natural language processing (NLP) engine to understand the user's intent.
[0592] Database referencing and response generation
[0593] The server queries a real-time database based on the analysis results to retrieve relevant information. For example, if a query is made about the rules for a specific play, it extracts the relevant rule information from the database. Based on that information, it generates an appropriate response.
[0594] Use of the emotion engine
[0595] The emotion engine recognizes emotions from the user's voice. This emotional information is taken into consideration during the response generation process. For example, if the user is excited, a concise and clear response will be provided. On the other hand, if the user is confused, a response with added detail will be provided.
[0596] Text to speech conversion
[0597] The generated responses are converted into speech data by a speech synthesis module. A Text-to-Speech (TTS) engine is used.
[0598] Sending and playing audio data
[0599] The server sends the generated audio data back to the terminal. The terminal receives the audio data and plays it back through its built-in speaker, providing the user with an appropriate response.
[0600] Specific example
[0601] If a user asks "What is the rule for this play?" while watching a soccer match, the device captures the audio and sends it to a server. The server converts the audio data into text and analyzes it with a natural language processing engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information, such as information about the "offside" rule.
[0602] In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear response. Something like, "This is the offside rule," is converted into audio data, sent to the device, and played back. This allows the user to receive a quick and appropriate response, enhancing their viewing experience.
[0603] The introduction of an emotion engine enables more personalized responses that reflect the user's emotional state, further improving system usability and user satisfaction.
[0604] The following describes the processing flow.
[0605] Step 1: Start voice input
[0606] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0607] Step 2: Acquire audio data
[0608] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0609] Step 3: Preprocessing of audio data
[0610] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0611] Step 4: Sending audio data
[0612] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0613] Step 5: Speech Recognition Processing
[0614] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[0615] Step 6: Natural Language Processing (NLP)
[0616] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0617] Step 7: Emotion Analysis
[0618] Server: Passes text and audio data to the emotion engine to recognize the user's emotions. For example, the emotion engine analyzes whether the user is excited or confused.
[0619] Step 8: Database query
[0620] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[0621] Step 9: Emotion-Based Response Adjustment
[0622] Server: Based on the results of the emotion engine, it generates responses that correspond to the user's emotional state. For example, if the user is excited, the response will be concise and clear.
[0623] Step 10: Generating the answer
[0624] Server: Generates appropriate responses to the user, including information retrieved from the database. For example, it might create a response such as, "This is the offside rule."
[0625] Step 11: Text to Speech
[0626] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[0627] Step 12: Sending audio data
[0628] Server: Sends the generated audio data to the terminal as an HTTP response.
[0629] Step 13: Receiving and Playing Audio
[0630] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This is an offside rule."
[0631] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply. Furthermore, the introduction of an emotion engine provides more personalized answers that reflect the user's emotional state, further improving the system's usability and user satisfaction.
[0632] (Example 2)
[0633] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0634] Conventional question-answering systems struggle to provide timely and appropriate answers to user questions, especially during sporting events and entertainment events where immediate explanations of rules and other information are required. Furthermore, they are unable to generate answers that take into account the user's emotional state, which hinders the improvement of the user experience.
[0635] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for recognizing the user's emotions from the voice data and reflecting that emotional information in the response generation process, and means for converting the response into voice data. This makes it possible to provide appropriate and emotionally personalized responses to user questions in real time.
[0636] "Audio data" refers to data that represents the voice information spoken by a user in digital format.
[0637] A "byte stream" is a data format that represents and transmits digital data in the form of consecutive bytes.
[0638] A "server" is a computer system used for processing, storing, and responding to requests from clients.
[0639] "Text" refers to data obtained by converting audio data into written information.
[0640] "Analyzing text" means analyzing text, which is character information, using natural language processing techniques to understand its meaning and intent.
[0641] "User intent" refers to the specific information or actions that the user is seeking, based on what they have said.
[0642] "Related information" refers to information retrieved from the database that is relevant to the user's intent.
[0643] A "database" is a system for efficiently storing, managing, and searching large amounts of data.
[0644] An "answer" is a piece of informational text generated based on a user's question.
[0645] "Emotional recognition" means analyzing voice and other data to identify the user's emotional state (e.g., excitement, confusion).
[0646] "Emotional information" refers to data about a user's emotional state obtained through the emotion recognition process.
[0647] A "Text-to-Speech (TTS) engine" is a technology used to convert text data into speech data.
[0648] "Playing audio data" refers to the act of playing converted audio data through sound equipment such as speakers.
[0649] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0650] This system includes the following components:
[0651] 1. Device (with microphone and speaker)
[0652] 2. Server
[0653] 3. Speech recognition module
[0654] 4. Natural Language Processing (NLP) Engine
[0655] 5. Real-time database
[0656] 6. Emotional Engine
[0657] 7. Text-to-Speech (TTS) Engine
[0658] User voice input
[0659] When a user has a question while watching a game, they might ask, "What are the rules for this play?" The device captures the user's voice via its microphone. The device uses the microphone as an audio input device to acquire the audio signal as digital data.
[0660] Sending audio data
[0661] The terminal converts the captured audio data into a byte stream and sends it to the server as part of an HTTP request. The terminal sends the audio data to the server using, for example, the Python requests library.
[0662] Speech recognition and analysis
[0663] The server receives the audio data and converts it to text using a speech recognition module (e.g., Google Speech-to-Text API). This text is then parsed using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API) to understand the user's intent.
[0664] Database referencing and response generation
[0665] The server queries a real-time database (e.g., Firebase or MySQL) based on the analysis results to retrieve relevant information. If a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. Then, using a template engine (e.g., Jinja2), it generates a grammatically correct answer based on the retrieved information.
[0666] Use of the emotion engine
[0667] The server uses an emotion engine (for example, Microsoft Azure's Text Analytics) to recognize emotions from the user's voice data. This recognized emotion is then reflected in the response generation process. For example, if the user is excited, the response is adjusted to be concise and clear. On the other hand, if the user is confused, more detailed explanations are added.
[0668] Text to speech conversion
[0669] The server passes the generated text response to a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert it into audio data. This audio data is saved as an audio file and used for subsequent processing.
[0670] Sending and playing audio data
[0671] The server sends the generated audio data back to the terminal, which receives the data and plays it back through its built-in speaker. This allows the user to receive quick and appropriate answers via voice.
[0672] Specific example
[0673] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the device captures the audio and sends it to the server. The server converts the audio data into text and analyzes it with an NLP engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information and extracts information about the "offside" rule. In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear answer. Something like "This is the offside rule" is converted into audio data, sent to the device, and played back. This allows the user to get a quick and appropriate answer in audio, improving the viewing experience.
[0674] Example of a prompt
[0675] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[0676] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0677] Step 1:
[0678] The user asks a question. When a user has a question while watching a game, they might ask, "What are the rules of this play?" The input is the user's voice, and the output is digital audio data captured by the microphone.
[0679] Step 2:
[0680] The device captures audio. The device's microphone captures the user's voice in real time. The input is the user's voice, and the output is digital audio data on the device. Specifically, the microphone device converts the audio signal into digital data.
[0681] Step 3:
[0682] The terminal converts the audio data into a byte stream and sends it to the server. The terminal encodes the audio data into a byte stream and sends it to the server as part of an HTTP request. The input is the captured digital audio data, and the output is the byte stream sent to the server. The Python requests library is used for this specific operation.
[0683] Step 4:
[0684] The server receives the audio data. The server receives an HTTP request and retrieves the audio data contained in the body. The input is audio data in byte stream format, and the output is digital audio data on the server.
[0685] Step 5:
[0686] The server converts audio data to text. The server uses a speech recognition module (e.g., Speech-to-Text API) to convert audio data to text. The input is digital audio data, and the output is text data. Specifically, it makes an API call to convert audio into text information.
[0687] Step 6:
[0688] The server analyzes the text and understands the user's intent. The server analyzes the text data using a natural language processing engine (e.g., Cloud Natural Language API) to identify the user's intent. The input is text data, and the output is information about the user's intent. For example, it analyzes the question "What are the rules of this game?" and understands the intent "I want to know the rules of the game."
[0689] Step 7:
[0690] The server references a database based on the analysis results. The server queries a real-time database (e.g., Firebase) based on the analysis results and retrieves relevant information. The input is user intent information, and the output is information from the relevant database. For example, it might retrieve rule information related to a specific play from the database.
[0691] Step 8:
[0692] The server generates an answer based on the information it extracts. The server uses a template engine (e.g., Jinja2) to process the acquired information into a format that is easy for the user to understand. The input is related information, and the output is the generated answer text.
[0693] Step 9:
[0694] The server recognizes the user's emotions and adjusts the response accordingly. The server uses an emotion engine (e.g., a Text Analytics API) to analyze the user's emotional state and adjust the response. The input is audio data, and the output is the adjusted response text. For example, if the user is agitated, a concise response is generated.
[0695] Step 10:
[0696] The server converts the response into audio data. The server uses a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert the generated response text into audio data. The input is the generated response text, and the output is audio data. Specifically, it makes an API call to convert the text into an audio file.
[0697] Step 11:
[0698] The server generates audio data and sends it to the terminal. The server then sends the generated audio data back to the terminal as an HTTP response. The input is the generated audio data, and the output is the byte stream audio data sent to the terminal.
[0699] Step 12:
[0700] The terminal plays audio data. The terminal decodes the received audio data and plays it through its built-in speaker. The input is audio data in byte stream format from the server, and the output is the audio that the user hears. Specifically, the audio is played using an audio device.
[0701] Example of a prompt:
[0702] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[0703] (Application Example 2)
[0704] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0705] In traditional sports viewing and entertainment events, users lacked real-time means of resolving questions, and the answers were not optimized according to the user's emotional state. This resulted in limitations on the user's viewing and event participation experience.
[0706] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for adjusting the extracted information based on the user's emotional state to generate an answer, means for converting the answer into audio data, means for transmitting the audio data to the user terminal, means for playing the audio data on the user terminal, and means for implementation as a smartphone application. As a result, the user can resolve questions in real time, and the answer is optimized for their emotional state, significantly improving the viewing experience and event participation experience.
[0707] "Audio data" refers to data that digitizes audio information, converting user questions and statements into a format that can be processed by electronic devices.
[0708] "Means of acquisition" refers to a function for capturing the user's voice, which collects voice data using the smartphone's microphone or other voice input devices.
[0709] "Means of transmission" refers to communication functions for sending acquired audio data to a server via a network, and generally refers to the means of exchanging data over the internet.
[0710] "Means of converting to text" refers to a function that converts audio data into text information, using speech recognition technology to convert audio into text data.
[0711] "Means of analysis to understand user intent" refers to analyzing data converted into text using natural language processing technology to understand user questions and requests.
[0712] "Means for extracting relevant information from a database" refers to a function that searches for and extracts necessary information from a database based on the analysis results.
[0713] "Means for generating answers" refers to a function that constructs answers to user questions based on extracted information, thereby generating appropriate answers.
[0714] The "means of adjusting and generating responses" refer to a function for optimizing responses based on the user's emotional state. This function uses emotion recognition technology to determine the user's emotions and provide responses accordingly.
[0715] "Means of converting to audio data" refers to a function that converts text-generated responses back into audio data using speech synthesis technology.
[0716] "Means of sending to the user's terminal" refers to a communication function for sending the generated audio data to the user's smartphone or other devices.
[0717] "Means of playback" refers to the function of playing back audio data received on the user's device, outputting the sound through the smartphone's speaker or earphones.
[0718] "Implementation as a smartphone application" means that all of these functions are configured as a smartphone application and designed to be easily used by users.
[0719] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0720] System Configuration
[0721] The system consists of the following hardware and software components:
[0722] 1. Acquisition of audio data
[0723] Hardware: Smartphone microphone
[0724] Software: speech_recognition library
[0725] 2. Sending audio data
[0726] Hardware: Smartphone communication module
[0727] Software: Request Library
[0728] 3. Converting audio data to text
[0729] Software: Use the speech_recognition library to convert speech data into text.
[0730] 4. Text analysis and understanding of intent
[0731] Software: Natural Language Processing (NLP) Engine
[0732] 5. Extraction of relevant information
[0733] Software: Database query function
[0734] 6. Emotion Recognition and Response Generation
[0735] Software: emotion_recognition library
[0736] Adjust responses based on the user's emotional state.
[0737] 7. Voice conversion of the answer
[0738] Software: gTTS library
[0739] 8. Sending and playing audio data
[0740] Hardware: Smartphone speaker
[0741] Software: Use mpg321 to play audio files
[0742] Implementation method
[0743] Acquisition and transmission of audio data
[0744] When a user asks "What are the rules of this game?" through the application, the device captures the audio via the microphone. This audio data is then sent to a server over the internet.
[0745] Text conversion and analysis of audio data
[0746] The server receives the transmitted audio data and converts it into text using the speech_recognition library. It then analyzes this text using a natural language processing (NLP) engine to understand the user's intent.
[0747] Extraction of relevant information and generation of responses
[0748] Based on the analysis results, relevant information is extracted from the database. For example, if asked about the rules of a specific play, the information on those rules is extracted from the database. Meanwhile, the emotion recognition engine recognizes emotions from the user's voice and takes them into consideration in the response generation process.
[0749] Voice conversion, transmission, and playback of responses
[0750] The generated response is converted into audio data by the gTTS library and sent back from the server to the terminal. The terminal receives the audio data and plays the audio using mpg321.
[0751] Specific examples and prompt statements
[0752] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the smartphone app captures the audio and sends it to a server. The server then analyzes the audio data and generates an answer such as "This is the offside rule."
[0753] Example of a prompt:
[0754] User input: "What are the rules of this game?"
[0755] Emotional state: Excited
[0756] This allows users to receive quick and appropriate answers, and by providing information tailored to their emotional state, it is possible to significantly improve the viewing experience.
[0757] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0758] Step 1:
[0759] The user launches a smartphone application and, while watching a game, makes a voice input such as, "What are the rules of this play?" The voice data is captured through the smartphone's microphone. The input is the user's voice data, and the output is the captured voice data.
[0760] Step 2:
[0761] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request. The input is the audio data, and the output is the transmission of the byte stream to the server.
[0762] Step 3:
[0763] The server converts the received audio data into text using a speech recognition module (e.g., the speech_recognition library). The input is audio data, and the output after conversion is text data.
[0764] Step 4:
[0765] The server analyzes text data using a natural language processing (NLP) engine to understand the user's intent. The input is text data, and the output is the analysis result, which represents the content of the user's question.
[0766] Step 5:
[0767] The server extracts relevant information from the database based on the analysis results. For example, if a user asks, "What are the rules for this play?", the server queries for the relevant rules and retrieves that information. The input is the analysis results, and the output is the relevant information extracted from the database.
[0768] Step 6:
[0769] An emotion recognition engine (e.g., the emotion_recognition library) recognizes the user's emotional state from their voice data. The input is voice data, and the output is the user's emotional state (excitement, confusion, etc.).
[0770] Step 7:
[0771] The server generates a response based on the extracted information and the user's emotional state. For example, if the user is agitated, the generated response might be a concise statement such as "This is the offside rule!". The input is the relevant information and emotional state, and the output is the generated response text.
[0772] Step 8:
[0773] The server converts the generated responses into speech data using a speech synthesis module (e.g., the gTTS library). The input is the response text, and the output is the speech data.
[0774] Step 9:
[0775] The server sends the generated audio data back to the terminal. The input is audio data, and the output is transmission to the terminal.
[0776] Step 10:
[0777] The device plays the received audio data. The user can hear the audio response through the device's speaker. The input is the audio data, and the output is the audio response to the user.
[0778] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0779] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0780] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0781] [Third Embodiment]
[0782] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0783] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0784] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0785] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0786] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0787] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0788] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0789] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0790] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0791] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0792] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0793] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0794] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0795] User voice input
[0796] When a user has a question while watching a game, they can naturally ask it into their device, such as, "What are the rules of this play?" The device captures the user's voice through its microphone.
[0797] Sending audio data
[0798] The terminal sends the captured audio data to the server. During transmission, the audio data is converted into a byte stream and packaged as part of the HTTP request.
[0799] Speech recognition and analysis
[0800] The server passes the received audio data to a speech recognition module, which converts the audio data into text data. A highly accurate speech recognition engine is used for this conversion. The converted text data is then analyzed by a natural language processing engine to understand the intent of the question.
[0801] Database referencing and response generation
[0802] Based on the analysis results, the server queries a real-time database to retrieve relevant information. For example, if a user requests information about the rules of play, it extracts information about the relevant rules from the database. Based on this information, it generates an appropriate response. The generated response is in text format, which is then converted into audio data.
[0803] Sending and playing audio data
[0804] The server returns the generated audio data to the terminal, which receives the data and plays it back to the user. In this way, the user can get answers in real time, such as "This violates the offside rule."
[0805] Specific example
[0806] For example, let's say a user is watching a soccer match. If they have a question about a particular play and ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text, "What is the rule for this play?", and parses the text. Based on the parse results, the server searches its database and retrieves information about the rule, for example, "offside". It generates this as a text response, converts it back into audio data, and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0807] In this way, users can get real-time answers to their questions the moment they arise, allowing them to enjoy the viewing experience more deeply. Furthermore, this system can be applied to explanations of classical performing arts and translations of foreign films. By providing expert information in response to users' arbitrary questions, it enables them to enjoy various entertainment content from multiple perspectives.
[0808] The following describes the processing flow.
[0809] Step 1: Start voice input
[0810] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0811] Step 2: Acquire audio data
[0812] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0813] Step 3: Preprocessing of audio data
[0814] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0815] Step 4: Sending audio data
[0816] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0817] Step 5: Speech Recognition Processing
[0818] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[0819] Step 6: Natural Language Processing (NLP)
[0820] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0821] Step 7: Database query
[0822] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[0823] Step 8: Generating the answer
[0824] Server: Generates appropriate responses for the user based on information retrieved from the database. For example, it might create a response such as, "This violates the offside rule."
[0825] Step 9: Text to Speech
[0826] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[0827] Step 10: Sending audio data
[0828] Server: Sends the generated audio data to the terminal as an HTTP response.
[0829] Step 11: Receiving and Playing Audio
[0830] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This violates the offside rule."
[0831] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply.
[0832] (Example 1)
[0833] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0834] There is a need to provide real-time information on sports and entertainment events and quickly resolve user questions. However, conventional systems have low accuracy in speech recognition and natural language processing, making it difficult to accurately obtain the information users are looking for. In addition, when real-time is required, the overall processing speed becomes slow. As a result, users cannot enjoy the event and their satisfaction level is low.
[0835] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0836] In this invention, the server includes means for passing voice data to a speech recognition engine, means for converting the voice data into text data, means for passing the text data to a natural language processing engine for analysis, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the generated answer into voice data, and means for transmitting the voice data to a terminal. This enables the entire process from voice input to answering to be carried out quickly and accurately, allowing users to obtain event information in real time and enjoy a highly satisfying viewing experience.
[0837] "Audio data" refers to a digital representation of sound, including user voice input.
[0838] A "byte stream" is a sequence of consecutive bytes, and is a format for transferring data over a communication path.
[0839] A "speech recognition engine" is software or hardware used to convert speech data into text data.
[0840] "Text data" refers to sentence-format data extracted from audio data by a speech recognition engine.
[0841] A "natural language processing engine" is software or hardware that analyzes text data and understands its meaning and intent.
[0842] "Analysis results" refer to information about the meaning and intent of text data analyzed by a natural language processing engine.
[0843] A "database" is a system for storing and managing structured data, and is used by servers to extract relevant information.
[0844] "Answer" refers to the content of the response to a user's question, generated based on analysis results and information extracted from the database.
[0845] A "speech synthesis engine" is software or hardware used to convert text data into speech data.
[0846] A "terminal" is an electronic device used by a user to input voice data and play back audio data.
[0847] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0848] Users can ask questions to their devices while watching sporting or entertainment events. For example, if they have a question about a particular play during a soccer match, they can naturally ask, "What are the rules for this play?" The devices have built-in microphones that capture the user's voice in real time.
[0849] The captured audio data is converted into a byte stream on the device. The converted audio data is then sent to the server as an HTTP POST request. The device utilizes an internet connection during this process.
[0850] On the server side, the received audio data is passed to a high-precision speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio into text data. This speech recognition engine can quickly and accurately convert the audio data into text. The converted text data is then sent to a natural language processing engine (e.g., a generative AI model) for analysis. This natural language processing engine understands the intent of the user's question and performs the necessary analysis to generate an appropriate response.
[0851] Based on the analysis results, the server queries a database (e.g., a relational database) to retrieve relevant information. For example, if a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. The server then generates an appropriate answer based on the retrieved information and uses a speech synthesis engine (e.g., a speech synthesis API) to convert that answer into speech data.
[0852] The generated audio data is sent from the server to the terminal, and the terminal provides the answer to the user by playing back the received audio data. In this way, the user can obtain answers to questions in real time.
[0853] As a concrete example, let's say a user is watching a soccer match. If a question arises about a particular play and they ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text data, "What is the rule for this play?", and analyzes the text. Based on the analysis, the server retrieves rule information about "offside" from its database and generates it as the answer. Then, it uses a speech synthesis engine to convert it into audio data and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[0854] The following are some examples of prompt statements.
[0855] User input: "What are the rules of this game?"
[0856] Prompts for generative AI models:
[0857] "The user is watching a soccer match and is asking about the rules regarding a specific play. Please explain the correct rule based on the following question: What are the rules for this play?"
[0858] Thus, the system of the present invention allows users to resolve their questions efficiently in real time and enjoy a richer viewing experience.
[0859] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0860] Step 1:
[0861] The user provides voice input. When a question arises during a game, the user can naturally ask it into the device, such as, "What are the rules of this play?" The device has a built-in microphone that captures the user's voice in real time.
[0862] Input: User's voice
[0863] Output: Captured audio data
[0864] Step 2:
[0865] The terminal converts the audio data into a byte stream. The terminal saves the captured audio data as digital audio data and converts it back into a byte stream.
[0866] Input: Captured audio data
[0867] Output: Audio data in byte stream format
[0868] Step 3:
[0869] The terminal sends audio data to the server. The terminal sends the converted byte stream of audio data to the server as an HTTP POST request. The terminal uses an internet connection during this process.
[0870] Input: Byte stream audio data
[0871] Output: HTTP request to the server
[0872] Step 4:
[0873] The server receives the audio data and passes it to the speech recognition engine. The server passes the received audio data to the speech recognition engine (general-purpose speech recognition API) and starts processing.
[0874] Input: Audio data in an HTTP request
[0875] Output: Audio data passed to the speech recognition engine
[0876] Step 5:
[0877] The server converts the audio data into text data. The speech recognition engine converts the audio data into text data. For example, the text data "What are the rules of this play?" is generated.
[0878] Input: Audio data passed to the speech recognition engine.
[0879] Output: Generated text data
[0880] Step 6:
[0881] The server passes the text data to a natural language processing engine for analysis. The server then passes the converted text data to the natural language processing engine (generative AI model) to analyze the intent of the question.
[0882] Input: Text data
[0883] Output: Data analyzing the intent behind the questions
[0884] Step 7:
[0885] The server queries the database based on the analysis results. The server queries the database (relational database) based on the analysis results and retrieves the relevant rules and information.
[0886] Input: Data analyzing the intent of the question
[0887] Output: Related information extracted from the database
[0888] Step 8:
[0889] The server generates the response and converts it into audio data. Based on the acquired information, the server generates a text-based response, and then converts that text into audio data using a speech synthesis engine (speech synthesis API).
[0890] Input: Related information extracted from the database
[0891] Output: Speech data generated by the speech synthesis engine
[0892] Step 9:
[0893] The server sends audio data to the terminal. The server then sends the generated audio data back to the terminal. During this process, the server utilizes an internet connection.
[0894] Input: Generated audio data
[0895] Output: Sending audio data to the terminal
[0896] Step 10:
[0897] The device receives and plays audio data. The device receives audio data sent from the server and plays it for the user. The user can receive real-time responses from the device, such as "This violates the offside rule."
[0898] Input: Audio data sent to the terminal
[0899] Output: Audio response played back to the user
[0900] (Application Example 1)
[0901] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0902] Within the factory, problems related to robot operation and troubleshooting frequently occur, and employees are unable to respond quickly, posing a challenge. Furthermore, the inability of employees to quickly obtain the information necessary for robot operation and problem solving poses a risk of decreased production efficiency. Traditional methods often involve manual information retrieval, which is time-consuming in finding appropriate solutions.
[0903] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0904] In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data to text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the answer to audio data, means for transmitting audio data to a user terminal, means for playing the audio data on the user terminal, means for capturing audio input to smart glasses, means for playing the audio data on the smart glasses, and means for querying a database for real-time answers to questions regarding the operation and troubleshooting of robots in the factory. This enables factory employees to obtain information regarding the operation and troubleshooting of robots in real time, facilitating rapid problem solving and improved production efficiency.
[0905] "Means for acquiring audio data" refers to a device or software that receives a user's speech as an audio signal and records that audio signal as digital data.
[0906] "Means for transmitting audio data to a server" refers to a device or software that transmits acquired audio data to a remote server via a network.
[0907] "Means for converting audio data to text" refers to a device or software that analyzes an audio signal and outputs its contents in text format.
[0908] "Means for analyzing text and understanding user intent" refers to a device or software that analyzes text data using natural language processing technology, understands its content, and identifies the intent behind the user's questions or instructions.
[0909] "Means for extracting relevant information from a database based on analysis results" refers to a device or software that searches for and extracts appropriate information from a database based on the analyzed user intent.
[0910] "Means for generating answers based on extracted information" refers to a device or software that automatically generates appropriate answers to user questions based on information extracted from a database.
[0911] "Means for converting responses into audio data" refers to a device or software that converts generated text-formatted responses into audio signals.
[0912] "Means for transmitting audio data to a user's terminal" refers to a device or software that transmits generated audio data to a user's terminal via a network.
[0913] "Means for playing audio data on a user terminal" refers to a device or software that plays received audio data and provides the user with an audio response.
[0914] "Means for capturing voice input on smart glasses" refers to a device or software that acquires the user's voice using a microphone installed in smart glasses.
[0915] "Means of playing audio data with smart glasses" refers to a device or software that plays audio data using speakers or bone conduction devices installed in smart glasses.
[0916] "Means of querying a database to provide real-time answers to questions regarding the operation and troubleshooting of robots in a factory" refers to a device or software that searches a database containing information on robot operation and errors in real time and obtains answers to questions.
[0917] This invention relates to a system for rapidly operating and troubleshooting robots in a factory. The system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[0918] Hardware and software configuration
[0919] hardware
[0920] Smart glasses: These are devices that have a built-in microphone and speaker, capture the user's voice input, and play back the audio data.
[0921] Server: A computer device equipped with a high-performance processor and large-capacity memory, used for data processing and database retrieval.
[0922] software
[0923] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert user voice input into text data.
[0924] The system uses the OpenAI GPT-4 natural language processing engine to analyze text data and understand user intent.
[0925] Database: A cloud-based data management system (e.g., MySQL) is used to store information related to robot operation and troubleshooting.
[0926] Communication technology: The HTTP protocol is used to send and receive voice data and analysis results between the server and the smart glasses.
[0927] System operation
[0928] 1. Voice input capture
[0929] The user voice-inputs "Why did the robot stop working?" into the smart glasses. The smart glasses' microphone captures this voice.
[0930] 2. Sending audio data
[0931] The smart glasses convert the captured audio data into a byte stream and send it to the server as an HTTP request.
[0932] 3. Speech Recognition and Analysis
[0933] The server receives the audio data and converts it to text using Google Cloud Speech-to-Text. Then, OpenAI GPT-4 is used to analyze the text data and understand the user's intent.
[0934] 4. Database referencing and response generation
[0935] The server queries a cloud-based database based on the analysis results and extracts relevant information. Based on the information obtained, it generates an appropriate response.
[0936] 5. Converting and sending responses as audio data.
[0937] The server generates text-based responses, which are then converted into audio data and sent to the smart glasses.
[0938] 6. Playback of audio data
[0939] The smart glasses play back the received audio data and provide the user with a response such as, "There may be a sensor malfunction due to error code 123."
[0940] Examples of specific cases and prompt statements
[0941] For example, if a robot suddenly stops working in a factory, an employee might ask the smart glasses, "Why did the robot stop working?" This voice input is sent to a server, which performs speech recognition and text analysis, then queries a database to generate an answer. This answer is then sent as voice data to the smart glasses, allowing the employee to receive the cause of the problem and its solution in real time via voice.
[0942] Examples of prompt statements are as follows:
[0943] Employee question: Why did the robot stop working?
[0944] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0945] Step 1:
[0946] The user inputs a question by voice into the smart glasses. Specifically, the user says, "Why did the robot stop working?" and the smart glasses' microphone captures the voice. At this stage, the input is in the form of voice data.
[0947] Step 2:
[0948] The smart glasses convert the captured audio data into a byte stream. This processes the audio data into a format that can be transmitted over the network. The converted data is then sent to the server as an HTTP request. The output is audio data in byte stream format.
[0949] Step 3:
[0950] The server receives audio data and passes it to the Google Cloud Speech-to-Text speech recognition engine, which converts it into text data. The input is audio data in byte stream format, and the output is text data. Specifically, the speech recognition engine analyzes the audio signal and converts it from speech to text.
[0951] Step 4:
[0952] The server passes text data to OpenAI GPT-4, which performs natural language processing to analyze the user's intent. The input is text data converted from speech, and the output is text data including the analysis results. Specifically, GPT-4 analyzes the text and performs the process of understanding the user's question.
[0953] Step 5:
[0954] The server queries the database based on the analysis results to retrieve relevant information. The input is text data reflecting the analysis results, and the output is a database record containing the relevant information. Specifically, the server generates and executes an SQL query to retrieve the relevant information.
[0955] Step 6:
[0956] The server generates appropriate answers to user questions based on information retrieved from the database. The input is information retrieved from the database, and the output is answer text data. Specifically, the server combines the retrieved information to generate appropriate answer sentences.
[0957] Step 7:
[0958] The system converts the server-generated response text into audio data. The input is text data, and the output is audio data. Specifically, it uses a text-to-speech conversion engine to create a natural-sounding response.
[0959] Step 8:
[0960] The server converts the audio data and sends it to the smart glasses. The input is audio data, and the output is the audio data that reaches the smart glasses. Specifically, the server performs the process of sending the audio data as an HTTP response.
[0961] Step 9:
[0962] The smart glasses play back the received audio data to the user. The input is the received audio data, and the output is the audio information the user hears. Specifically, the smart glasses use their speakers or bone conduction device to play back the voice response.
[0963] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0964] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[0965] User voice input
[0966] When a user has a question while watching a game, they can ask, "What are the rules for this play?" The device captures the user's voice via the microphone.
[0967] Sending audio data
[0968] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request.
[0969] Speech recognition and analysis
[0970] The server receives the audio data and converts it to text using a speech recognition module. This text is then analyzed by a natural language processing (NLP) engine to understand the user's intent.
[0971] Database referencing and response generation
[0972] The server queries a real-time database based on the analysis results to retrieve relevant information. For example, if a query is made about the rules for a specific play, it extracts the relevant rule information from the database. Based on that information, it generates an appropriate response.
[0973] Use of the emotion engine
[0974] The emotion engine recognizes emotions from the user's voice. This emotional information is taken into consideration during the response generation process. For example, if the user is excited, a concise and clear response will be provided. On the other hand, if the user is confused, a response with added detail will be provided.
[0975] Text to speech conversion
[0976] The generated responses are converted into speech data by a speech synthesis module. A Text-to-Speech (TTS) engine is used.
[0977] Sending and playing audio data
[0978] The server sends the generated audio data back to the terminal. The terminal receives the audio data and plays it back through its built-in speaker, providing the user with an appropriate response.
[0979] Specific example
[0980] If a user asks "What is the rule for this play?" while watching a soccer match, the device captures the audio and sends it to a server. The server converts the audio data into text and analyzes it with a natural language processing engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information, such as information about the "offside" rule.
[0981] In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear response. Something like, "This is the offside rule," is converted into audio data, sent to the device, and played back. This allows the user to receive a quick and appropriate response, enhancing their viewing experience.
[0982] The introduction of an emotion engine enables more personalized responses that reflect the user's emotional state, further improving system usability and user satisfaction.
[0983] The following describes the processing flow.
[0984] Step 1: Start voice input
[0985] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[0986] Step 2: Acquire audio data
[0987] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[0988] Step 3: Preprocessing of audio data
[0989] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[0990] Step 4: Sending audio data
[0991] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[0992] Step 5: Speech Recognition Processing
[0993] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[0994] Step 6: Natural Language Processing (NLP)
[0995] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[0996] Step 7: Emotion Analysis
[0997] Server: Passes text and audio data to the emotion engine to recognize the user's emotions. For example, the emotion engine analyzes whether the user is excited or confused.
[0998] Step 8: Database query
[0999] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[1000] Step 9: Emotion-Based Response Adjustment
[1001] Server: Based on the results of the emotion engine, it generates responses that correspond to the user's emotional state. For example, if the user is excited, the response will be concise and clear.
[1002] Step 10: Generating the answer
[1003] Server: Generates appropriate responses to the user, including information retrieved from the database. For example, it might create a response such as, "This is the offside rule."
[1004] Step 11: Text to Speech
[1005] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[1006] Step 12: Sending audio data
[1007] Server: Sends the generated audio data to the terminal as an HTTP response.
[1008] Step 13: Receiving and Playing Audio
[1009] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This is an offside rule."
[1010] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply. Furthermore, the introduction of an emotion engine provides more personalized answers that reflect the user's emotional state, further improving the system's usability and user satisfaction.
[1011] (Example 2)
[1012] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1013] Conventional question-answering systems struggle to provide timely and appropriate answers to user questions, especially during sporting events and entertainment events where immediate explanations of rules and other information are required. Furthermore, they are unable to generate answers that take into account the user's emotional state, which hinders the improvement of the user experience.
[1014] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for recognizing the user's emotions from the voice data and reflecting that emotional information in the response generation process, and means for converting the response into voice data. This makes it possible to provide appropriate and emotionally personalized responses to user questions in real time.
[1015] "Audio data" refers to data that represents the voice information spoken by a user in digital format.
[1016] A "byte stream" is a data format that represents and transmits digital data in the form of consecutive bytes.
[1017] A "server" is a computer system used for processing, storing, and responding to requests from clients.
[1018] "Text" refers to data obtained by converting audio data into written information.
[1019] "Analyzing text" means analyzing text, which is character information, using natural language processing techniques to understand its meaning and intent.
[1020] "User intent" refers to the specific information or actions that the user is seeking, based on what they have said.
[1021] "Related information" refers to information retrieved from the database that is relevant to the user's intent.
[1022] A "database" is a system for efficiently storing, managing, and searching large amounts of data.
[1023] An "answer" is a piece of informational text generated based on a user's question.
[1024] "Emotional recognition" means analyzing voice and other data to identify the user's emotional state (e.g., excitement, confusion).
[1025] "Emotional information" refers to data about a user's emotional state obtained through the emotion recognition process.
[1026] A "Text-to-Speech (TTS) engine" is a technology used to convert text data into speech data.
[1027] "Playing audio data" refers to the act of playing converted audio data through sound equipment such as speakers.
[1028] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[1029] This system includes the following components:
[1030] 1. Device (with microphone and speaker)
[1031] 2. Server
[1032] 3. Speech recognition module
[1033] 4. Natural Language Processing (NLP) Engine
[1034] 5. Real-time database
[1035] 6. Emotional Engine
[1036] 7. Text-to-Speech (TTS) Engine
[1037] User voice input
[1038] When a user has a question while watching a game, they might ask, "What are the rules for this play?" The device captures the user's voice via its microphone. The device uses the microphone as an audio input device to acquire the audio signal as digital data.
[1039] Sending audio data
[1040] The terminal converts the captured audio data into a byte stream and sends it to the server as part of an HTTP request. The terminal sends the audio data to the server using, for example, the Python requests library.
[1041] Speech recognition and analysis
[1042] The server receives the audio data and converts it to text using a speech recognition module (e.g., Google Speech-to-Text API). This text is then parsed using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API) to understand the user's intent.
[1043] Database referencing and response generation
[1044] The server queries a real-time database (e.g., Firebase or MySQL) based on the analysis results to retrieve relevant information. If a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. Then, using a template engine (e.g., Jinja2), it generates a grammatically correct answer based on the retrieved information.
[1045] Use of the emotion engine
[1046] The server uses an emotion engine (for example, Microsoft Azure's Text Analytics) to recognize emotions from the user's voice data. This recognized emotion is then reflected in the response generation process. For example, if the user is excited, the response is adjusted to be concise and clear. On the other hand, if the user is confused, more detailed explanations are added.
[1047] Text to speech conversion
[1048] The server passes the generated text response to a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert it into audio data. This audio data is saved as an audio file and used for subsequent processing.
[1049] Sending and playing audio data
[1050] The server sends the generated audio data back to the terminal, which receives the data and plays it back through its built-in speaker. This allows the user to receive quick and appropriate answers via voice.
[1051] Specific example
[1052] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the device captures the audio and sends it to the server. The server converts the audio data into text and analyzes it with an NLP engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information and extracts information about the "offside" rule. In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear answer. Something like "This is the offside rule" is converted into audio data, sent to the device, and played back. This allows the user to get a quick and appropriate answer in audio, improving the viewing experience.
[1053] Example of a prompt
[1054] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[1055] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1056] Step 1:
[1057] The user asks a question. When a user has a question while watching a game, they might ask, "What are the rules of this play?" The input is the user's voice, and the output is digital audio data captured by the microphone.
[1058] Step 2:
[1059] The device captures audio. The device's microphone captures the user's voice in real time. The input is the user's voice, and the output is digital audio data on the device. Specifically, the microphone device converts the audio signal into digital data.
[1060] Step 3:
[1061] The terminal converts the audio data into a byte stream and sends it to the server. The terminal encodes the audio data into a byte stream and sends it to the server as part of an HTTP request. The input is the captured digital audio data, and the output is the byte stream sent to the server. The Python requests library is used for this specific operation.
[1062] Step 4:
[1063] The server receives the audio data. The server receives an HTTP request and retrieves the audio data contained in the body. The input is audio data in byte stream format, and the output is digital audio data on the server.
[1064] Step 5:
[1065] The server converts audio data to text. The server uses a speech recognition module (e.g., Speech-to-Text API) to convert audio data to text. The input is digital audio data, and the output is text data. Specifically, it makes an API call to convert audio into text information.
[1066] Step 6:
[1067] The server analyzes the text and understands the user's intent. The server analyzes the text data using a natural language processing engine (e.g., Cloud Natural Language API) to identify the user's intent. The input is text data, and the output is information about the user's intent. For example, it analyzes the question "What are the rules of this game?" and understands the intent "I want to know the rules of the game."
[1068] Step 7:
[1069] The server references a database based on the analysis results. The server queries a real-time database (e.g., Firebase) based on the analysis results and retrieves relevant information. The input is user intent information, and the output is information from the relevant database. For example, it might retrieve rule information related to a specific play from the database.
[1070] Step 8:
[1071] The server generates an answer based on the information it extracts. The server uses a template engine (e.g., Jinja2) to process the acquired information into a format that is easy for the user to understand. The input is related information, and the output is the generated answer text.
[1072] Step 9:
[1073] The server recognizes the user's emotions and adjusts the response accordingly. The server uses an emotion engine (e.g., a Text Analytics API) to analyze the user's emotional state and adjust the response. The input is audio data, and the output is the adjusted response text. For example, if the user is agitated, a concise response is generated.
[1074] Step 10:
[1075] The server converts the response into audio data. The server uses a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert the generated response text into audio data. The input is the generated response text, and the output is audio data. Specifically, it makes an API call to convert the text into an audio file.
[1076] Step 11:
[1077] The server generates audio data and sends it to the terminal. The server then sends the generated audio data back to the terminal as an HTTP response. The input is the generated audio data, and the output is the byte stream audio data sent to the terminal.
[1078] Step 12:
[1079] The terminal plays audio data. The terminal decodes the received audio data and plays it through its built-in speaker. The input is audio data in byte stream format from the server, and the output is the audio that the user hears. Specifically, the audio is played using an audio device.
[1080] Example of a prompt:
[1081] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[1082] (Application Example 2)
[1083] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1084] In traditional sports viewing and entertainment events, users lacked real-time means of resolving questions, and the answers were not optimized according to the user's emotional state. This resulted in limitations on the user's viewing and event participation experience.
[1085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for adjusting the extracted information based on the user's emotional state to generate an answer, means for converting the answer into audio data, means for transmitting the audio data to the user terminal, means for playing the audio data on the user terminal, and means for implementation as a smartphone application. As a result, the user can resolve questions in real time, and the answer is optimized for their emotional state, significantly improving the viewing experience and event participation experience.
[1086] "Audio data" refers to data that digitizes audio information, converting user questions and statements into a format that can be processed by electronic devices.
[1087] "Means of acquisition" refers to a function for capturing the user's voice, which collects voice data using the smartphone's microphone or other voice input devices.
[1088] "Means of transmission" refers to communication functions for sending acquired audio data to a server via a network, and generally refers to the means of exchanging data over the internet.
[1089] "Means of converting to text" refers to a function that converts audio data into text information, using speech recognition technology to convert audio into text data.
[1090] "Means of analysis to understand user intent" refers to analyzing data converted into text using natural language processing technology to understand user questions and requests.
[1091] "Means for extracting relevant information from a database" refers to a function that searches for and extracts necessary information from a database based on the analysis results.
[1092] "Means for generating answers" refers to a function that constructs answers to user questions based on extracted information, thereby generating appropriate answers.
[1093] The "means of adjusting and generating responses" refer to a function for optimizing responses based on the user's emotional state. This function uses emotion recognition technology to determine the user's emotions and provide responses accordingly.
[1094] "Means of converting to audio data" refers to a function that converts text-generated responses back into audio data using speech synthesis technology.
[1095] "Means of sending to the user's terminal" refers to a communication function for sending the generated audio data to the user's smartphone or other devices.
[1096] "Means of playback" refers to the function of playing back audio data received on the user's device, outputting the sound through the smartphone's speaker or earphones.
[1097] "Implementation as a smartphone application" means that all of these functions are configured as a smartphone application and designed to be easily used by users.
[1098] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[1099] System Configuration
[1100] The system consists of the following hardware and software components:
[1101] 1. Acquisition of audio data
[1102] Hardware: Smartphone microphone
[1103] Software: speech_recognition library
[1104] 2. Sending audio data
[1105] Hardware: Smartphone communication module
[1106] Software: Request Library
[1107] 3. Converting audio data to text
[1108] Software: Use the speech_recognition library to convert speech data into text.
[1109] 4. Text analysis and understanding of intent
[1110] Software: Natural Language Processing (NLP) Engine
[1111] 5. Extraction of relevant information
[1112] Software: Database query function
[1113] 6. Emotion Recognition and Response Generation
[1114] Software: emotion_recognition library
[1115] Adjust responses based on the user's emotional state.
[1116] 7. Voice conversion of the answer
[1117] Software: gTTS library
[1118] 8. Sending and playing audio data
[1119] Hardware: Smartphone speaker
[1120] Software: Use mpg321 to play audio files
[1121] Implementation method
[1122] Acquisition and transmission of audio data
[1123] When a user asks "What are the rules of this game?" through the application, the device captures the audio via the microphone. This audio data is then sent to a server over the internet.
[1124] Text conversion and analysis of audio data
[1125] The server receives the transmitted audio data and converts it into text using the speech_recognition library. It then analyzes this text using a natural language processing (NLP) engine to understand the user's intent.
[1126] Extraction of relevant information and generation of responses
[1127] Based on the analysis results, relevant information is extracted from the database. For example, if asked about the rules of a specific play, the information on those rules is extracted from the database. Meanwhile, the emotion recognition engine recognizes emotions from the user's voice and takes them into consideration in the response generation process.
[1128] Voice conversion, transmission, and playback of responses
[1129] The generated response is converted into audio data by the gTTS library and sent back from the server to the terminal. The terminal receives the audio data and plays the audio using mpg321.
[1130] Specific examples and prompt statements
[1131] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the smartphone app captures the audio and sends it to a server. The server then analyzes the audio data and generates an answer such as "This is the offside rule."
[1132] Example of a prompt:
[1133] User input: "What are the rules of this game?"
[1134] Emotional state: Excited
[1135] This allows users to receive quick and appropriate answers, and by providing information tailored to their emotional state, it is possible to significantly improve the viewing experience.
[1136] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1137] Step 1:
[1138] The user launches a smartphone application and, while watching a game, makes a voice input such as, "What are the rules of this play?" The voice data is captured through the smartphone's microphone. The input is the user's voice data, and the output is the captured voice data.
[1139] Step 2:
[1140] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request. The input is the audio data, and the output is the transmission of the byte stream to the server.
[1141] Step 3:
[1142] The server converts the received audio data into text using a speech recognition module (e.g., the speech_recognition library). The input is audio data, and the output after conversion is text data.
[1143] Step 4:
[1144] The server analyzes text data using a natural language processing (NLP) engine to understand the user's intent. The input is text data, and the output is the analysis result, which represents the content of the user's question.
[1145] Step 5:
[1146] The server extracts relevant information from the database based on the analysis results. For example, if a user asks, "What are the rules for this play?", the server queries for the relevant rules and retrieves that information. The input is the analysis results, and the output is the relevant information extracted from the database.
[1147] Step 6:
[1148] An emotion recognition engine (e.g., the emotion_recognition library) recognizes the user's emotional state from their voice data. The input is voice data, and the output is the user's emotional state (excitement, confusion, etc.).
[1149] Step 7:
[1150] The server generates a response based on the extracted information and the user's emotional state. For example, if the user is agitated, the generated response might be a concise statement such as "This is the offside rule!". The input is the relevant information and emotional state, and the output is the generated response text.
[1151] Step 8:
[1152] The server converts the generated responses into speech data using a speech synthesis module (e.g., the gTTS library). The input is the response text, and the output is the speech data.
[1153] Step 9:
[1154] The server sends the generated audio data back to the terminal. The input is audio data, and the output is transmission to the terminal.
[1155] Step 10:
[1156] The device plays the received audio data. The user can hear the audio response through the device's speaker. The input is the audio data, and the output is the audio response to the user.
[1157] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1158] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1159] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1160] [Fourth Embodiment]
[1161] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1162] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1163] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1164] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1165] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1166] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1167] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1168] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1169] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1170] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1171] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1172] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1173] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1174] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[1175] User voice input
[1176] When a user has a question while watching a game, they can naturally ask it into their device, such as, "What are the rules of this play?" The device captures the user's voice through its microphone.
[1177] Sending audio data
[1178] The terminal sends the captured audio data to the server. During transmission, the audio data is converted into a byte stream and packaged as part of the HTTP request.
[1179] Speech recognition and analysis
[1180] The server passes the received audio data to a speech recognition module, which converts the audio data into text data. A highly accurate speech recognition engine is used for this conversion. The converted text data is then analyzed by a natural language processing engine to understand the intent of the question.
[1181] Database referencing and response generation
[1182] Based on the analysis results, the server queries a real-time database to retrieve relevant information. For example, if a user requests information about the rules of play, it extracts information about the relevant rules from the database. Based on this information, it generates an appropriate response. The generated response is in text format, which is then converted into audio data.
[1183] Sending and playing audio data
[1184] The server returns the generated audio data to the terminal, which receives the data and plays it back to the user. In this way, the user can get answers in real time, such as "This violates the offside rule."
[1185] Specific example
[1186] For example, let's say a user is watching a soccer match. If they have a question about a particular play and ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text, "What is the rule for this play?", and parses the text. Based on the parse results, the server searches its database and retrieves information about the rule, for example, "offside". It generates this as a text response, converts it back into audio data, and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[1187] In this way, users can get real-time answers to their questions the moment they arise, allowing them to enjoy the viewing experience more deeply. Furthermore, this system can be applied to explanations of classical performing arts and translations of foreign films. By providing expert information in response to users' arbitrary questions, it enables them to enjoy various entertainment content from multiple perspectives.
[1188] The following describes the processing flow.
[1189] Step 1: Start voice input
[1190] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[1191] Step 2: Acquire audio data
[1192] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[1193] Step 3: Preprocessing of audio data
[1194] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[1195] Step 4: Sending audio data
[1196] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[1197] Step 5: Speech Recognition Processing
[1198] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[1199] Step 6: Natural Language Processing (NLP)
[1200] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[1201] Step 7: Database query
[1202] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[1203] Step 8: Generating the answer
[1204] Server: Generates appropriate responses for the user based on information retrieved from the database. For example, it might create a response such as, "This violates the offside rule."
[1205] Step 9: Text to Speech
[1206] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[1207] Step 10: Sending audio data
[1208] Server: Sends the generated audio data to the terminal as an HTTP response.
[1209] Step 11: Receiving and Playing Audio
[1210] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This violates the offside rule."
[1211] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply.
[1212] (Example 1)
[1213] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1214] There is a need to provide real-time information on sports and entertainment events and quickly resolve user questions. However, conventional systems have low accuracy in speech recognition and natural language processing, making it difficult to accurately obtain the information users are looking for. In addition, when real-time is required, the overall processing speed becomes slow. As a result, users cannot enjoy the event and their satisfaction level is low.
[1215] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1216] In this invention, the server includes means for passing voice data to a speech recognition engine, means for converting the voice data into text data, means for passing the text data to a natural language processing engine for analysis, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the generated answer into voice data, and means for transmitting the voice data to a terminal. This enables the entire process from voice input to answering to be carried out quickly and accurately, allowing users to obtain event information in real time and enjoy a highly satisfying viewing experience.
[1217] "Audio data" refers to a digital representation of sound, including user voice input.
[1218] A "byte stream" is a sequence of consecutive bytes, and is a format for transferring data over a communication path.
[1219] A "speech recognition engine" is software or hardware used to convert speech data into text data.
[1220] "Text data" refers to sentence-format data extracted from audio data by a speech recognition engine.
[1221] A "natural language processing engine" is software or hardware that analyzes text data and understands its meaning and intent.
[1222] "Analysis results" refer to information about the meaning and intent of text data analyzed by a natural language processing engine.
[1223] A "database" is a system for storing and managing structured data, and is used by servers to extract relevant information.
[1224] "Answer" refers to the content of the response to a user's question, generated based on analysis results and information extracted from the database.
[1225] A "speech synthesis engine" is software or hardware used to convert text data into speech data.
[1226] A "terminal" is an electronic device used by a user to input voice data and play back audio data.
[1227] This invention relates to a system that allows users to obtain necessary information in real time while watching sports or attending entertainment events. This system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[1228] Users can ask questions to their devices while watching sporting or entertainment events. For example, if they have a question about a particular play during a soccer match, they can naturally ask, "What are the rules for this play?" The devices have built-in microphones that capture the user's voice in real time.
[1229] The captured audio data is converted into a byte stream on the device. The converted audio data is then sent to the server as an HTTP POST request. The device utilizes an internet connection during this process.
[1230] On the server side, the received audio data is passed to a high-precision speech recognition engine (e.g., a general-purpose speech recognition API) to convert the audio into text data. This speech recognition engine can quickly and accurately convert the audio data into text. The converted text data is then sent to a natural language processing engine (e.g., a generative AI model) for analysis. This natural language processing engine understands the intent of the user's question and performs the necessary analysis to generate an appropriate response.
[1231] Based on the analysis results, the server queries a database (e.g., a relational database) to retrieve relevant information. For example, if a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. The server then generates an appropriate answer based on the retrieved information and uses a speech synthesis engine (e.g., a speech synthesis API) to convert that answer into speech data.
[1232] The generated audio data is sent from the server to the terminal, and the terminal provides the answer to the user by playing back the received audio data. In this way, the user can obtain answers to questions in real time.
[1233] As a concrete example, let's say a user is watching a soccer match. If a question arises about a particular play and they ask, "What is the rule for this play?", the device sends the audio to the server. The server converts the audio data into text data, "What is the rule for this play?", and analyzes the text. Based on the analysis, the server retrieves rule information about "offside" from its database and generates it as the answer. Then, it uses a speech synthesis engine to convert it into audio data and sends it back to the device. The device plays the audio and provides the user with the answer, "This is an offside play."
[1234] The following are some examples of prompt statements.
[1235] User input: "What are the rules of this game?"
[1236] Prompts for generative AI models:
[1237] "The user is watching a soccer match and is asking about the rules regarding a specific play. Please explain the correct rule based on the following question: What are the rules for this play?"
[1238] Thus, the system of the present invention allows users to resolve their questions efficiently in real time and enjoy a richer viewing experience.
[1239] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1240] Step 1:
[1241] The user provides voice input. When a question arises during a game, the user can naturally ask it into the device, such as, "What are the rules of this play?" The device has a built-in microphone that captures the user's voice in real time.
[1242] Input: User's voice
[1243] Output: Captured audio data
[1244] Step 2:
[1245] The terminal converts the audio data into a byte stream. The terminal saves the captured audio data as digital audio data and converts it back into a byte stream.
[1246] Input: Captured audio data
[1247] Output: Audio data in byte stream format
[1248] Step 3:
[1249] The terminal sends audio data to the server. The terminal sends the converted byte stream of audio data to the server as an HTTP POST request. The terminal uses an internet connection during this process.
[1250] Input: Byte stream audio data
[1251] Output: HTTP request to the server
[1252] Step 4:
[1253] The server receives the audio data and passes it to the speech recognition engine. The server passes the received audio data to the speech recognition engine (general-purpose speech recognition API) and starts processing.
[1254] Input: Audio data in an HTTP request
[1255] Output: Audio data passed to the speech recognition engine
[1256] Step 5:
[1257] The server converts the audio data into text data. The speech recognition engine converts the audio data into text data. For example, the text data "What are the rules of this play?" is generated.
[1258] Input: Audio data passed to the speech recognition engine.
[1259] Output: Generated text data
[1260] Step 6:
[1261] The server passes the text data to a natural language processing engine for analysis. The server then passes the converted text data to the natural language processing engine (generative AI model) to analyze the intent of the question.
[1262] Input: Text data
[1263] Output: Data analyzing the intent behind the questions
[1264] Step 7:
[1265] The server queries the database based on the analysis results. The server queries the database (relational database) based on the analysis results and retrieves the relevant rules and information.
[1266] Input: Data analyzing the intent of the question
[1267] Output: Related information extracted from the database
[1268] Step 8:
[1269] The server generates the response and converts it into audio data. Based on the acquired information, the server generates a text-based response, and then converts that text into audio data using a speech synthesis engine (speech synthesis API).
[1270] Input: Related information extracted from the database
[1271] Output: Speech data generated by the speech synthesis engine
[1272] Step 9:
[1273] The server sends audio data to the terminal. The server then sends the generated audio data back to the terminal. During this process, the server utilizes an internet connection.
[1274] Input: Generated audio data
[1275] Output: Sending audio data to the terminal
[1276] Step 10:
[1277] The device receives and plays audio data. The device receives audio data sent from the server and plays it for the user. The user can receive real-time responses from the device, such as "This violates the offside rule."
[1278] Input: Audio data sent to the terminal
[1279] Output: Audio response played back to the user
[1280] (Application Example 1)
[1281] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1282] Within the factory, problems related to robot operation and troubleshooting frequently occur, and employees are unable to respond quickly, posing a challenge. Furthermore, the inability of employees to quickly obtain the information necessary for robot operation and problem solving poses a risk of decreased production efficiency. Traditional methods often involve manual information retrieval, which is time-consuming in finding appropriate solutions.
[1283] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1284] In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data to text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for generating an answer based on the extracted information, means for converting the answer to audio data, means for transmitting audio data to a user terminal, means for playing the audio data on the user terminal, means for capturing audio input to smart glasses, means for playing the audio data on the smart glasses, and means for querying a database for real-time answers to questions regarding the operation and troubleshooting of robots in the factory. This enables factory employees to obtain information regarding the operation and troubleshooting of robots in real time, facilitating rapid problem solving and improved production efficiency.
[1285] "Means for acquiring audio data" refers to a device or software that receives a user's speech as an audio signal and records that audio signal as digital data.
[1286] "Means for transmitting audio data to a server" refers to a device or software that transmits acquired audio data to a remote server via a network.
[1287] "Means for converting audio data to text" refers to a device or software that analyzes an audio signal and outputs its contents in text format.
[1288] "Means for analyzing text and understanding user intent" refers to a device or software that analyzes text data using natural language processing technology, understands its content, and identifies the intent behind the user's questions or instructions.
[1289] "Means for extracting relevant information from a database based on analysis results" refers to a device or software that searches for and extracts appropriate information from a database based on the analyzed user intent.
[1290] "Means for generating answers based on extracted information" refers to a device or software that automatically generates appropriate answers to user questions based on information extracted from a database.
[1291] "Means for converting responses into audio data" refers to a device or software that converts generated text-formatted responses into audio signals.
[1292] "Means for transmitting audio data to a user's terminal" refers to a device or software that transmits generated audio data to a user's terminal via a network.
[1293] "Means for playing audio data on a user terminal" refers to a device or software that plays received audio data and provides the user with an audio response.
[1294] "Means for capturing voice input on smart glasses" refers to a device or software that acquires the user's voice using a microphone installed in smart glasses.
[1295] "Means of playing audio data with smart glasses" refers to a device or software that plays audio data using speakers or bone conduction devices installed in smart glasses.
[1296] "Means of querying a database to provide real-time answers to questions regarding the operation and troubleshooting of robots in a factory" refers to a device or software that searches a database containing information on robot operation and errors in real time and obtains answers to questions.
[1297] This invention relates to a system for rapidly operating and troubleshooting robots in a factory. The system includes voice input, processing of voice data, real-time response generation, and provision of voice responses.
[1298] Hardware and software configuration
[1299] hardware
[1300] Smart glasses: These are devices that have a built-in microphone and speaker, capture the user's voice input, and play back the audio data.
[1301] Server: A computer device equipped with a high-performance processor and large-capacity memory, used for data processing and database retrieval.
[1302] software
[1303] Speech recognition engine: Uses Google Cloud Speech-to-Text to convert user voice input into text data.
[1304] The system uses the OpenAI GPT-4 natural language processing engine to analyze text data and understand user intent.
[1305] Database: A cloud-based data management system (e.g., MySQL) is used to store information related to robot operation and troubleshooting.
[1306] Communication technology: The HTTP protocol is used to send and receive voice data and analysis results between the server and the smart glasses.
[1307] System operation
[1308] 1. Voice input capture
[1309] The user voice-inputs "Why did the robot stop working?" into the smart glasses. The smart glasses' microphone captures this voice.
[1310] 2. Sending audio data
[1311] The smart glasses convert the captured audio data into a byte stream and send it to the server as an HTTP request.
[1312] 3. Speech Recognition and Analysis
[1313] The server receives the audio data and converts it to text using Google Cloud Speech-to-Text. Then, OpenAI GPT-4 is used to analyze the text data and understand the user's intent.
[1314] 4. Database referencing and response generation
[1315] The server queries a cloud-based database based on the analysis results and extracts relevant information. Based on the information obtained, it generates an appropriate response.
[1316] 5. Converting and sending responses as audio data.
[1317] The server generates text-based responses, which are then converted into audio data and sent to the smart glasses.
[1318] 6. Playback of audio data
[1319] The smart glasses play back the received audio data and provide the user with a response such as, "There may be a sensor malfunction due to error code 123."
[1320] Examples of specific cases and prompt statements
[1321] For example, if a robot suddenly stops working in a factory, an employee might ask the smart glasses, "Why did the robot stop working?" This voice input is sent to a server, which performs speech recognition and text analysis, then queries a database to generate an answer. This answer is then sent as voice data to the smart glasses, allowing the employee to receive the cause of the problem and its solution in real time via voice.
[1322] Examples of prompt statements are as follows:
[1323] Employee question: Why did the robot stop working?
[1324] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1325] Step 1:
[1326] The user inputs a question by voice into the smart glasses. Specifically, the user says, "Why did the robot stop working?" and the smart glasses' microphone captures the voice. At this stage, the input is in the form of voice data.
[1327] Step 2:
[1328] The smart glasses convert the captured audio data into a byte stream. This processes the audio data into a format that can be transmitted over the network. The converted data is then sent to the server as an HTTP request. The output is audio data in byte stream format.
[1329] Step 3:
[1330] The server receives audio data and passes it to the Google Cloud Speech-to-Text speech recognition engine, which converts it into text data. The input is audio data in byte stream format, and the output is text data. Specifically, the speech recognition engine analyzes the audio signal and converts it from speech to text.
[1331] Step 4:
[1332] The server passes text data to OpenAI GPT-4, which performs natural language processing to analyze the user's intent. The input is text data converted from speech, and the output is text data including the analysis results. Specifically, GPT-4 analyzes the text and performs the process of understanding the user's question.
[1333] Step 5:
[1334] The server queries the database based on the analysis results to retrieve relevant information. The input is text data reflecting the analysis results, and the output is a database record containing the relevant information. Specifically, the server generates and executes an SQL query to retrieve the relevant information.
[1335] Step 6:
[1336] The server generates appropriate answers to user questions based on information retrieved from the database. The input is information retrieved from the database, and the output is answer text data. Specifically, the server combines the retrieved information to generate appropriate answer sentences.
[1337] Step 7:
[1338] The system converts the server-generated response text into audio data. The input is text data, and the output is audio data. Specifically, it uses a text-to-speech conversion engine to create a natural-sounding response.
[1339] Step 8:
[1340] The server converts the audio data and sends it to the smart glasses. The input is audio data, and the output is the audio data that reaches the smart glasses. Specifically, the server performs the process of sending the audio data as an HTTP response.
[1341] Step 9:
[1342] The smart glasses play back the received audio data to the user. The input is the received audio data, and the output is the audio information the user hears. Specifically, the smart glasses use their speakers or bone conduction device to play back the voice response.
[1343] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1344] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[1345] User voice input
[1346] When a user has a question while watching a game, they can ask, "What are the rules for this play?" The device captures the user's voice via the microphone.
[1347] Sending audio data
[1348] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request.
[1349] Speech recognition and analysis
[1350] The server receives the audio data and converts it to text using a speech recognition module. This text is then analyzed by a natural language processing (NLP) engine to understand the user's intent.
[1351] Database referencing and response generation
[1352] The server queries a real-time database based on the analysis results to retrieve relevant information. For example, if a query is made about the rules for a specific play, it extracts the relevant rule information from the database. Based on that information, it generates an appropriate response.
[1353] Use of the emotion engine
[1354] The emotion engine recognizes emotions from the user's voice. This emotional information is taken into consideration during the response generation process. For example, if the user is excited, a concise and clear response will be provided. On the other hand, if the user is confused, a response with added detail will be provided.
[1355] Text to speech conversion
[1356] The generated responses are converted into speech data by a speech synthesis module. A Text-to-Speech (TTS) engine is used.
[1357] Sending and playing audio data
[1358] The server sends the generated audio data back to the terminal. The terminal receives the audio data and plays it back through its built-in speaker, providing the user with an appropriate response.
[1359] Specific example
[1360] If a user asks "What is the rule for this play?" while watching a soccer match, the device captures the audio and sends it to a server. The server converts the audio data into text and analyzes it with a natural language processing engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information, such as information about the "offside" rule.
[1361] In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear response. Something like, "This is the offside rule," is converted into audio data, sent to the device, and played back. This allows the user to receive a quick and appropriate response, enhancing their viewing experience.
[1362] The introduction of an emotion engine enables more personalized responses that reflect the user's emotional state, further improving system usability and user satisfaction.
[1363] The following describes the processing flow.
[1364] Step 1: Start voice input
[1365] User: While watching a sporting event, a question arises, and they ask, "What are the rules of this play?"
[1366] Step 2: Acquire audio data
[1367] Terminal: Captures user audio using the built-in microphone and saves it to a buffer.
[1368] Step 3: Preprocessing of audio data
[1369] Device: Improves the quality of audio data by performing noise cancellation, echo reduction, and other functions.
[1370] Step 4: Sending audio data
[1371] Terminal: Reads audio data from the buffer, converts it to a byte stream, and sends it to the server's API endpoint as an HTTP POST request.
[1372] Step 5: Speech Recognition Processing
[1373] Server: Converts the speech data received by the speech recognition module into text. For example, it uses the Google Cloud Speech-to-Text API or Microsoft Azure Cognitive Services.
[1374] Step 6: Natural Language Processing (NLP)
[1375] Server: Passes the text data obtained from speech recognition to a natural language processing engine to analyze the intent of the question. For example, it extracts keywords such as "play" and "rules" from the question.
[1376] Step 7: Emotion Analysis
[1377] Server: Passes text and audio data to the emotion engine to recognize the user's emotions. For example, the emotion engine analyzes whether the user is excited or confused.
[1378] Step 8: Database query
[1379] Server: Based on the analyzed keywords and intent, it searches a real-time database and retrieves relevant information.
[1380] Step 9: Emotion-Based Response Adjustment
[1381] Server: Based on the results of the emotion engine, it generates responses that correspond to the user's emotional state. For example, if the user is excited, the response will be concise and clear.
[1382] Step 10: Generating the answer
[1383] Server: Generates appropriate responses to the user, including information retrieved from the database. For example, it might create a response such as, "This is the offside rule."
[1384] Step 11: Text to Speech
[1385] Server: Passes the generated text response to the speech synthesis module and converts it into speech data. Uses a Text-to-Speech (TTS) engine.
[1386] Step 12: Sending audio data
[1387] Server: Sends the generated audio data to the terminal as an HTTP response.
[1388] Step 13: Receiving and Playing Audio
[1389] Terminal: Receives audio data, decodes it, and plays it back through the built-in speaker. The user can hear the response, "This is an offside rule."
[1390] This specific processing flow allows users to resolve their questions on the spot and enjoy sports viewing and entertainment events more deeply. Furthermore, the introduction of an emotion engine provides more personalized answers that reflect the user's emotional state, further improving the system's usability and user satisfaction.
[1391] (Example 2)
[1392] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1393] Conventional question-answering systems struggle to provide timely and appropriate answers to user questions, especially during sporting events and entertainment events where immediate explanations of rules and other information are required. Furthermore, they are unable to generate answers that take into account the user's emotional state, which hinders the improvement of the user experience.
[1394] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for recognizing the user's emotions from the voice data and reflecting that emotional information in the response generation process, and means for converting the response into voice data. This makes it possible to provide appropriate and emotionally personalized responses to user questions in real time.
[1395] "Audio data" refers to data that represents the voice information spoken by a user in digital format.
[1396] A "byte stream" is a data format that represents and transmits digital data in the form of consecutive bytes.
[1397] A "server" is a computer system used for processing, storing, and responding to requests from clients.
[1398] "Text" refers to data obtained by converting audio data into written information.
[1399] "Analyzing text" means analyzing text, which is character information, using natural language processing techniques to understand its meaning and intent.
[1400] "User intent" refers to the specific information or actions that the user is seeking, based on what they have said.
[1401] "Related information" refers to information retrieved from the database that is relevant to the user's intent.
[1402] A "database" is a system for efficiently storing, managing, and searching large amounts of data.
[1403] An "answer" is a piece of informational text generated based on a user's question.
[1404] "Emotional recognition" means analyzing voice and other data to identify the user's emotional state (e.g., excitement, confusion).
[1405] "Emotional information" refers to data about a user's emotional state obtained through the emotion recognition process.
[1406] A "Text-to-Speech (TTS) engine" is a technology used to convert text data into speech data.
[1407] "Playing audio data" refers to the act of playing converted audio data through sound equipment such as speakers.
[1408] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[1409] This system includes the following components:
[1410] 1. Device (with microphone and speaker)
[1411] 2. Server
[1412] 3. Speech recognition module
[1413] 4. Natural Language Processing (NLP) Engine
[1414] 5. Real-time database
[1415] 6. Emotional Engine
[1416] 7. Text-to-Speech (TTS) Engine
[1417] User voice input
[1418] When a user has a question while watching a game, they might ask, "What are the rules for this play?" The device captures the user's voice via its microphone. The device uses the microphone as an audio input device to acquire the audio signal as digital data.
[1419] Sending audio data
[1420] The terminal converts the captured audio data into a byte stream and sends it to the server as part of an HTTP request. The terminal sends the audio data to the server using, for example, the Python requests library.
[1421] Speech recognition and analysis
[1422] The server receives the audio data and converts it to text using a speech recognition module (e.g., Google Speech-to-Text API). This text is then parsed using a natural language processing (NLP) engine (e.g., Google Cloud Natural Language API) to understand the user's intent.
[1423] Database referencing and response generation
[1424] The server queries a real-time database (e.g., Firebase or MySQL) based on the analysis results to retrieve relevant information. If a user asks, "What are the rules of this play?", the server extracts information about the rules of that play from the database. Then, using a template engine (e.g., Jinja2), it generates a grammatically correct answer based on the retrieved information.
[1425] Use of the emotion engine
[1426] The server uses an emotion engine (for example, Microsoft Azure's Text Analytics) to recognize emotions from the user's voice data. This recognized emotion is then reflected in the response generation process. For example, if the user is excited, the response is adjusted to be concise and clear. On the other hand, if the user is confused, more detailed explanations are added.
[1427] Text to speech conversion
[1428] The server passes the generated text response to a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert it into audio data. This audio data is saved as an audio file and used for subsequent processing.
[1429] Sending and playing audio data
[1430] The server sends the generated audio data back to the terminal, which receives the data and plays it back through its built-in speaker. This allows the user to receive quick and appropriate answers via voice.
[1431] Specific example
[1432] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the device captures the audio and sends it to the server. The server converts the audio data into text and analyzes it with an NLP engine to understand the intent of the question. It then queries a real-time database to retrieve relevant information and extracts information about the "offside" rule. In parallel, if the emotion engine recognizes that the user is excited from their voice, the server generates a concise and clear answer. Something like "This is the offside rule" is converted into audio data, sent to the device, and played back. This allows the user to get a quick and appropriate answer in audio, improving the viewing experience.
[1433] Example of a prompt
[1434] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[1435] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1436] Step 1:
[1437] The user asks a question. When a user has a question while watching a game, they might ask, "What are the rules of this play?" The input is the user's voice, and the output is digital audio data captured by the microphone.
[1438] Step 2:
[1439] The device captures audio. The device's microphone captures the user's voice in real time. The input is the user's voice, and the output is digital audio data on the device. Specifically, the microphone device converts the audio signal into digital data.
[1440] Step 3:
[1441] The terminal converts the audio data into a byte stream and sends it to the server. The terminal encodes the audio data into a byte stream and sends it to the server as part of an HTTP request. The input is the captured digital audio data, and the output is the byte stream sent to the server. The Python requests library is used for this specific operation.
[1442] Step 4:
[1443] The server receives the audio data. The server receives an HTTP request and retrieves the audio data contained in the body. The input is audio data in byte stream format, and the output is digital audio data on the server.
[1444] Step 5:
[1445] The server converts audio data to text. The server uses a speech recognition module (e.g., Speech-to-Text API) to convert audio data to text. The input is digital audio data, and the output is text data. Specifically, it makes an API call to convert audio into text information.
[1446] Step 6:
[1447] The server analyzes the text and understands the user's intent. The server analyzes the text data using a natural language processing engine (e.g., Cloud Natural Language API) to identify the user's intent. The input is text data, and the output is information about the user's intent. For example, it analyzes the question "What are the rules of this game?" and understands the intent "I want to know the rules of the game."
[1448] Step 7:
[1449] The server references a database based on the analysis results. The server queries a real-time database (e.g., Firebase) based on the analysis results and retrieves relevant information. The input is user intent information, and the output is information from the relevant database. For example, it might retrieve rule information related to a specific play from the database.
[1450] Step 8:
[1451] The server generates an answer based on the information it extracts. The server uses a template engine (e.g., Jinja2) to process the acquired information into a format that is easy for the user to understand. The input is related information, and the output is the generated answer text.
[1452] Step 9:
[1453] The server recognizes the user's emotions and adjusts the response accordingly. The server uses an emotion engine (e.g., a Text Analytics API) to analyze the user's emotional state and adjust the response. The input is audio data, and the output is the adjusted response text. For example, if the user is agitated, a concise response is generated.
[1454] Step 10:
[1455] The server converts the response into audio data. The server uses a Text-to-Speech (TTS) engine (for example, Amazon Polly) to convert the generated response text into audio data. The input is the generated response text, and the output is audio data. Specifically, it makes an API call to convert the text into an audio file.
[1456] Step 11:
[1457] The server generates audio data and sends it to the terminal. The server then sends the generated audio data back to the terminal as an HTTP response. The input is the generated audio data, and the output is the byte stream audio data sent to the terminal.
[1458] Step 12:
[1459] The terminal plays audio data. The terminal decodes the received audio data and plays it through its built-in speaker. The input is audio data in byte stream format from the server, and the output is the audio that the user hears. Specifically, the audio is played using an audio device.
[1460] Example of a prompt:
[1461] Describe the process by which a user watching a soccer match asks about the rules of a specific play, and how the system understands the question and generates an appropriate answer in real time. Also, describe specific prompts for answer generation that take the user's emotional state into account.
[1462] (Application Example 2)
[1463] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1464] In traditional sports viewing and entertainment events, users lacked real-time means of resolving questions, and the answers were not optimized according to the user's emotional state. This resulted in limitations on the user's viewing and event participation experience.
[1465] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring audio data, means for transmitting audio data to the server, means for converting audio data into text, means for analyzing the text and understanding the user's intent, means for extracting relevant information from a database based on the analysis results, means for adjusting the extracted information based on the user's emotional state to generate an answer, means for converting the answer into audio data, means for transmitting the audio data to the user terminal, means for playing the audio data on the user terminal, and means for implementation as a smartphone application. As a result, the user can resolve questions in real time, and the answer is optimized for their emotional state, significantly improving the viewing experience and event participation experience.
[1466] "Audio data" refers to data that digitizes audio information, converting user questions and statements into a format that can be processed by electronic devices.
[1467] "Means of acquisition" refers to a function for capturing the user's voice, which collects voice data using the smartphone's microphone or other voice input devices.
[1468] "Means of transmission" refers to communication functions for sending acquired audio data to a server via a network, and generally refers to the means of exchanging data over the internet.
[1469] "Means of converting to text" refers to a function that converts audio data into text information, using speech recognition technology to convert audio into text data.
[1470] "Means of analysis to understand user intent" refers to analyzing data converted into text using natural language processing technology to understand user questions and requests.
[1471] "Means for extracting relevant information from a database" refers to a function that searches for and extracts necessary information from a database based on the analysis results.
[1472] "Means for generating answers" refers to a function that constructs answers to user questions based on extracted information, thereby generating appropriate answers.
[1473] The "means of adjusting and generating responses" refer to a function for optimizing responses based on the user's emotional state. This function uses emotion recognition technology to determine the user's emotions and provide responses accordingly.
[1474] "Means of converting to audio data" refers to a function that converts text-generated responses back into audio data using speech synthesis technology.
[1475] "Means of sending to the user's terminal" refers to a communication function for sending the generated audio data to the user's smartphone or other devices.
[1476] "Means of playback" refers to the function of playing back audio data received on the user's device, outputting the sound through the smartphone's speaker or earphones.
[1477] "Implementation as a smartphone application" means that all of these functions are configured as a smartphone application and designed to be easily used by users.
[1478] This invention relates to a system that allows users to resolve questions in real time while watching sports or attending entertainment events, and further adjusts the answers according to the user's emotional state. In addition to voice input, processing of voice data, real-time answer generation, and provision of voice responses, this system includes a function to recognize the user's emotions and adjust the answers accordingly.
[1479] System Configuration
[1480] The system consists of the following hardware and software components:
[1481] 1. Acquisition of audio data
[1482] Hardware: Smartphone microphone
[1483] Software: speech_recognition library
[1484] 2. Sending audio data
[1485] Hardware: Smartphone communication module
[1486] Software: Request Library
[1487] 3. Converting audio data to text
[1488] Software: Use the speech_recognition library to convert speech data into text.
[1489] 4. Text analysis and understanding of intent
[1490] Software: Natural Language Processing (NLP) Engine
[1491] 5. Extraction of relevant information
[1492] Software: Database query function
[1493] 6. Emotion Recognition and Response Generation
[1494] Software: emotion_recognition library
[1495] Adjust responses based on the user's emotional state.
[1496] 7. Voice conversion of the answer
[1497] Software: gTTS library
[1498] 8. Sending and playing audio data
[1499] Hardware: Smartphone speaker
[1500] Software: Use mpg321 to play audio files
[1501] Implementation method
[1502] Acquisition and transmission of audio data
[1503] When a user asks "What are the rules of this game?" through the application, the device captures the audio via the microphone. This audio data is then sent to a server over the internet.
[1504] Text conversion and analysis of audio data
[1505] The server receives the transmitted audio data and converts it into text using the speech_recognition library. It then analyzes this text using a natural language processing (NLP) engine to understand the user's intent.
[1506] Extraction of relevant information and generation of responses
[1507] Based on the analysis results, relevant information is extracted from the database. For example, if asked about the rules of a specific play, the information on those rules is extracted from the database. Meanwhile, the emotion recognition engine recognizes emotions from the user's voice and takes them into consideration in the response generation process.
[1508] Voice conversion, transmission, and playback of responses
[1509] The generated response is converted into audio data by the gTTS library and sent back from the server to the terminal. The terminal receives the audio data and plays the audio using mpg321.
[1510] Specific examples and prompt statements
[1511] For example, if a user asks "What is the rule of this play?" while watching a soccer match, the smartphone app captures the audio and sends it to a server. The server then analyzes the audio data and generates an answer such as "This is the offside rule."
[1512] Example of a prompt:
[1513] User input: "What are the rules of this game?"
[1514] Emotional state: Excited
[1515] This allows users to receive quick and appropriate answers, and by providing information tailored to their emotional state, it is possible to significantly improve the viewing experience.
[1516] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1517] Step 1:
[1518] The user launches a smartphone application and, while watching a game, makes a voice input such as, "What are the rules of this play?" The voice data is captured through the smartphone's microphone. The input is the user's voice data, and the output is the captured voice data.
[1519] Step 2:
[1520] The terminal sends the captured audio data to the server. The audio data is converted into a byte stream and sent to the server as part of an HTTP request. The input is the audio data, and the output is the transmission of the byte stream to the server.
[1521] Step 3:
[1522] The server converts the received audio data into text using a speech recognition module (e.g., the speech_recognition library). The input is audio data, and the output after conversion is text data.
[1523] Step 4:
[1524] The server analyzes text data using a natural language processing (NLP) engine to understand the user's intent. The input is text data, and the output is the analysis result, which represents the content of the user's question.
[1525] Step 5:
[1526] The server extracts relevant information from the database based on the analysis results. For example, if a user asks, "What are the rules for this play?", the server queries for the relevant rules and retrieves that information. The input is the analysis results, and the output is the relevant information extracted from the database.
[1527] Step 6:
[1528] An emotion recognition engine (e.g., the emotion_recognition library) recognizes the user's emotional state from their voice data. The input is voice data, and the output is the user's emotional state (excitement, confusion, etc.).
[1529] Step 7:
[1530] The server generates a response based on the extracted information and the user's emotional state. For example, if the user is agitated, the generated response might be a concise statement such as "This is the offside rule!". The input is the relevant information and emotional state, and the output is the generated response text.
[1531] Step 8:
[1532] The server converts the generated responses into speech data using a speech synthesis module (e.g., the gTTS library). The input is the response text, and the output is the speech data.
[1533] Step 9:
[1534] The server sends the generated audio data back to the terminal. The input is audio data, and the output is transmission to the terminal.
[1535] Step 10:
[1536] The device plays the received audio data. The user can hear the audio response through the device's speaker. The input is the audio data, and the output is the audio response to the user.
[1537] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1538] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1539] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1540] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1541] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1542] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1543] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1544] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1545] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1546] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1547] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1548] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1549] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1550] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1551] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1552] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1553] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1554] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1555] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1556] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1557] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1558] The following is further disclosed regarding the embodiments described above.
[1559] (Claim 1)
[1560] Means for acquiring audio data,
[1561] A means of sending audio data to a server,
[1562] A means of converting audio data into text,
[1563] A means of analyzing text and understanding user intent,
[1564] A means of extracting relevant information from a database based on the analysis results,
[1565] A means of generating an answer based on the extracted information,
[1566] A means of converting the answer into audio data,
[1567] A means of transmitting audio data to the user's terminal,
[1568] A means of playing audio data on the user's terminal,
[1569] A system that includes this.
[1570] (Claim 2)
[1571] The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
[1572] (Claim 3)
[1573] The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content.
[1574] "Example 1"
[1575] (Claim 1)
[1576] Means for acquiring audio data,
[1577] A means of converting audio data into a byte stream,
[1578] A means of sending audio data to a server,
[1579] A means of passing audio data to a speech recognition engine,
[1580] A means of converting audio data into text data,
[1581] A means of passing text data to a natural language processing engine for analysis,
[1582] A means of extracting relevant information from a database based on the analysis results,
[1583] A means of generating an answer based on the extracted information,
[1584] A means of converting the generated response into audio data,
[1585] A means of transmitting audio data to a terminal,
[1586] A means of playing audio data on a device,
[1587] A system that includes this.
[1588] (Claim 2)
[1589] The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
[1590] (Claim 3)
[1591] The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content.
[1592] "Application Example 1"
[1593] (Claim 1)
[1594] Means for acquiring audio data,
[1595] A means of sending audio data to a server,
[1596] A means of converting audio data into text,
[1597] A means of analyzing text and understanding user intent,
[1598] A means of extracting relevant information from a database based on the analysis results,
[1599] A means of generating an answer based on the extracted information,
[1600] A means of converting the answer into audio data,
[1601] A means of transmitting audio data to the user's terminal,
[1602] A means of playing audio data on the user's terminal,
[1603] A means of capturing voice input on smart glasses,
[1604] A method for playing audio data with smart glasses,
[1605] A means of querying a database to provide real-time answers to questions regarding the operation and troubleshooting of robots within a factory,
[1606] A system that includes this.
[1607] (Claim 2)
[1608] The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
[1609] (Claim 3)
[1610] The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content.
[1611] "Example 2 of combining an emotion engine"
[1612] (Claim 1)
[1613] Means for acquiring audio data,
[1614] A means of converting audio data into a byte stream and sending it to a server,
[1615] A means of converting audio data into text,
[1616] A means of analyzing text and understanding user intent,
[1617] A means of extracting relevant information from a database based on the analysis results,
[1618] A means of generating an answer based on the extracted information,
[1619] A means of recognizing the user's emotions from voice data and reflecting that emotional information in the response generation process,
[1620] A means of converting the answer into audio data,
[1621] A means of transmitting audio data to the user's terminal,
[1622] A means of playing audio data on the user's terminal,
[1623] A system that includes this.
[1624] (Claim 2)
[1625] The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
[1626] (Claim 3)
[1627] The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content.
[1628] "Application example 2 when combining with an emotional engine"
[1629] (Claim 1)
[1630] Means for acquiring audio data,
[1631] A means of sending audio data to a server,
[1632] A means of converting audio data into text,
[1633] A means of analyzing text and understanding user intent,
[1634] A means of extracting relevant information from a database based on the analysis results,
[1635] A means of generating a response by adjusting the extracted information based on the user's emotional state,
[1636] A means of converting the answer into audio data,
[1637] A means of transmitting audio data to the user's terminal,
[1638] A means of playing audio data on the user's terminal,
[1639] The means of implementation as a smartphone application,
[1640] A system that includes this.
[1641] (Claim 2)
[1642] The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
[1643] (Claim 3)
[1644] The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content. [Explanation of Symbols]
[1645] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring audio data, A means of sending audio data to a server, A means of converting audio data into text, A means of analyzing text and understanding user intent, A means of extracting relevant information from a database based on the analysis results, A means of generating an answer based on the extracted information, A means of converting the answer into audio data, A means of transmitting audio data to the user's terminal, A means of playing audio data on the user's terminal, A system that includes this.
2. The system according to claim 1, which answers questions about the rules and match status of a sports event in real time.
3. The system according to claim 1, which provides real-time commentary on entertainment events and translation of foreign language content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A