system
The system addresses long speeches in voice dialogue by converting voice to text, generating visual materials, and providing audio summaries, enhancing user understanding and interaction efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional voice dialogue systems provide rich knowledge in long speeches that are difficult for users to understand, leading to inefficiencies in information grasp.
A system that converts voice input to text, analyzes the text, generates visual materials, and provides audio summaries to facilitate quick and intuitive understanding of information.
Enhances interaction efficiency by enabling users to understand information quickly through both visual and auditory means, optimizing the layout of visual materials and adjusting information presentation based on user emotions.
Smart Images

Figure 2026068457000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0005] , , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a voice dialogue system, there is a problem that the rich knowledge provided by the generative AI becomes a long speech, which is difficult for the user to understand and burdensome. This problem makes it difficult for the user to immediately grasp information through the conversation and causes a decrease in the efficiency of the dialogue. The purpose of the present invention is to suppress such long speeches in voice dialogue and enable the user to quickly and intuitively understand important information.
Means for Solving the Problems
[0005] To solve this problem, the present invention provides a means for receiving voice input and converting it into text data. Furthermore, it includes means for analyzing the converted text data and retrieving relevant information. Based on the retrieved information, it automatically generates visual materials and outputs them to a display device, thereby assisting the user in visually understanding the information. In addition, by combining this with means for providing a summary of the visual materials in audio, the user can quickly grasp the information. Furthermore, by automatically optimizing the layout of the visual materials and simultaneously providing relevant audio information through the user interface, the efficiency of the interaction is enhanced.
[0006] "Voice input" is a method of conveying information by having the user speak directly to the system.
[0007] "Text data" refers to character information converted by speech recognition, and is in a format suitable for computer analysis and processing.
[0008] "Analysis" is the process of interpreting the content of input text data and identifying its intent and related information.
[0009] "Searching for information" is the process of retrieving relevant knowledge and data from databases and other resources based on analyzed data.
[0010] "Visual materials" are graphic or chart-based materials used to visually organize and present the key points of information.
[0011] A "display device" is a device used to present generated visual materials so that users can view them.
[0012] A "summary" is a shortened and concise summary of the main points of information, designed to allow users to quickly understand it.
[0013] "Providing information via audio" refers to a method of conveying information to users in audio format using synthesized speech or recordings.
[0014] "Automatically optimizing the layout" means that a computer autonomously adjusts the arrangement and design of visual materials to present information in the most effective and easy-to-understand way.
[0015] A "user interface" refers to the operating area or screen that a user uses to interact with a system, and is the part that handles information input and output. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] The system of this invention functions by combining multiple means to streamline information transmission in voice dialogue. When a user inputs a question by voice into the interface, it is converted into text data using speech recognition technology. At this stage, the terminal has the ability to receive the voice and transcribe it in real time.
[0038] Subsequently, the server retrieves the text data and performs analysis using natural language processing. In this analysis step, the server efficiently searches for relevant information that reflects the intent and theme of the user's question. Once relevant data is retrieved from the databases and external information sources used here, the foundational information for the visual materials is established.
[0039] Next, the server generates visual materials. Visual materials are visualizations, such as infographics, based on the acquired information, presenting the information in a format that is intuitively easy for the user to understand. This generation process includes optimizing the layout of the visual materials so that they are easy for the user to read and the key points can be clearly recognized.
[0040] Once the visual material is complete, the server sends it to the terminal, where it is displayed. The displayed visual material incorporates diagrams, graphs, text, and other elements to make the information easier to understand visually. Furthermore, if necessary, synthesized speech is used to provide audio summaries of the key points of the visual material, further supporting information comprehension.
[0041] As a concrete example, consider a scenario where a user asks, "Please tell me how to use solar energy." The terminal recognizes the speech and converts it into text data. The server analyzes this data and searches for information on how to use solar energy. It then generates relevant visual materials and sends them to the terminal. The terminal displays these materials on its screen and also provides an audio explanation of the main uses. The user can quickly understand the information using both sight and hearing. In this way, the system effectively supports the user's understanding of information.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The user inputs questions by voice using a microphone. The user asks the system questions by voice about a topic for which they want to obtain specific information.
[0045] Step 2:
[0046] The device uses speech recognition technology to convert the input speech into text data. The device then sends this converted text data to the next processing step.
[0047] Step 3:
[0048] The server analyzes the text data and uses natural language processing to identify the user's intent and related keywords. Based on the question, the server prepares to search for the most relevant data.
[0049] Step 4:
[0050] The server searches for relevant information from databases and external sources based on the analysis results. The server quickly retrieves accurate data in response to the user's questions.
[0051] Step 5:
[0052] Based on the information acquired by the server, visual materials are generated. The server creates infographics in a format that is easily understandable to the user and optimizes the layout of the visual materials.
[0053] Step 6:
[0054] The server sends the generated visual material to the terminal. The terminal prepares to present the received visual material to the user.
[0055] Step 7:
[0056] The device displays visual materials on its screen and provides users with summarized information using synthesized speech. The device is designed to facilitate visual and auditory comprehension of information.
[0057] Step 8:
[0058] Users can effectively understand information by reviewing visual materials and listening to audio summaries. Users can follow up with additional questions as needed.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] In systems that efficiently provide information through voice interaction, users need to quickly and intuitively understand large amounts of information. However, in conventional systems, information obtained from voice input is provided in text format, making it difficult to effectively convey information using both sight and hearing. A new method of information transmission is needed to solve this problem and facilitate user understanding.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for analyzing converted text data and searching for relevant information using natural language processing, means for a generation engine for generating visual materials using the relevant information, and means for visualizing the information in the generated visual materials. As a result, the user can receive not only audio but also visually organized information simultaneously, enabling quick and intuitive understanding of the information.
[0064] "Means for receiving voice input" refers to a device or technology that provides a function for taking in external voice information into a system.
[0065] "Methods for converting speech to text data" refers to technologies that perform a process of analyzing speech signals and converting them into corresponding text data.
[0066] "Means for analyzing text data and searching for related information using natural language processing" refers to technologies that analyze text data, understand the user's intent using natural language processing techniques, and have the functionality to search for related information from databases and external sources.
[0067] "Means equipped with a generation engine for generating visual materials" refers to a function equipped with an engine for generating visual materials based on acquired information, and is a technology that visualizes information as infographics or charts.
[0068] "Means of visualizing information in visual materials" refers to techniques for organizing generated materials and presenting them in a visually easy-to-understand format.
[0069] "A means of outputting visual materials to a display device and providing their key points in synthesized speech" refers to a technology that displays generated visual materials on a display device for user viewing and simultaneously explains the key points of the materials in synthesized speech.
[0070] This system combines multiple technologies to acquire information via voice and effectively convey it to the user. When a user provides voice input, the terminal first receives the voice. The terminal uses a microphone to convert the voice into digital data, performs speech recognition using that data, and then converts the voice back into text data. At this stage, speech recognition software such as Google® Speech-to-Text API is utilized.
[0071] Next, the server retrieves this text data. The server uses natural language processing techniques to analyze the text data. Based on the analyzed information, the server connects to internal databases and external information sources to retrieve relevant data. This process uses common data analysis tools and search algorithms.
[0072] The information obtained is then generated as visual materials through a generation engine on the server. Visual materials such as infographics and charts are optimized in layout so that the information is presented in a format that is intuitively easy for the user to understand. Visualization tools such as Infogram and Tableau are used during this generation process.
[0073] Finally, the generated visual materials are sent to the terminal. The terminal outputs them to a display device, allowing the user to view the materials. In addition, a speech synthesis engine installed in the terminal provides an audio summary of the visual materials.
[0074] As a concrete example, consider a scenario where a user asks a question via voice, "Please tell me how to use solar energy." In this case, the terminal converts the voice into text, the server collects and analyzes information about solar energy, generates visual materials, and sends them to the user's terminal. Through this entire process, the user can quickly understand the information through both voice and visual means.
[0075] An example of a prompt might be, "Generate a visual explanation of how to utilize solar energy." Such a prompt prompt would prompt the server to begin the process of visualizing the relevant information.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The user provides voice input through the interface. The input voice is captured by the terminal and converted into digital data via the microphone. The terminal then passes this digital voice data to a speech recognition engine, which generates text data in real time. During this process, voice signal processing is performed to reduce background noise and clarify speech.
[0079] Step 2:
[0080] The device sends text data to the server. Upon receiving the text data, the server analyzes it using natural language processing (NLP) techniques. Specifically, it extracts relevant keywords by performing data processing such as text tokenization, morphological analysis, and language model application to understand the intent of the user's question.
[0081] Step 3:
[0082] The server searches for relevant information based on the analysis results. It accesses databases and external information sources to retrieve data related to the keywords. Using information retrieval algorithms, it quickly selects the most relevant information. The retrieved information is then formatted into a dataset necessary for generating visual materials in the next step.
[0083] Step 4:
[0084] The server takes a dataset as input and generates visual materials using a generative AI model. This process involves visualizations such as charts, graphs, and infographics, organizing the data into a user-friendly format. Visual design optimization ensures that the generated materials are visually appealing.
[0085] Step 5:
[0086] The generated visual materials are sent from the server to the terminal. The terminal receives them and outputs them to the display device. Furthermore, the terminal's built-in speech synthesis function is used to provide the user with audio summaries of the key features and points of the visual materials, thereby supporting a deeper understanding of the information. In this state, the user can obtain information through both audio and visual means.
[0087] (Application Example 1)
[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0089] Traditional methods for customers to efficiently obtain information in physical stores are time-consuming, making it difficult to quickly acquire appropriate information. In particular, requests for detailed product usage instructions and characteristics require the assistance of store staff, leading to resource depletion and wasted time. There is a need to address these challenges and streamline the provision of information within stores.
[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0091] In this invention, the server includes a medium for receiving voice input, a medium for converting the received voice into text information, a medium for analyzing the converted text information and retrieving related information, and a medium for presenting the generated visual material to a portable visual aid device. This enables customers to ask questions by voice within the store and obtain information that is quick and easy to understand visually.
[0092] "Voice input" is the process of recognizing the voice signals spoken by a user and receiving them as digital data.
[0093] A "medium" is a means of processing, converting, displaying, or transmitting data such as audio signals, visual materials, and textual information.
[0094] "Textual information" refers to text data that is obtained by analyzing and converting audio signals.
[0095] "Analysis" is the process of deciphering input text information and understanding its meaning and intent.
[0096] "Relevant information" refers to the appropriate data and knowledge that are retrieved based on the user's questions or requests.
[0097] "Visual materials" refer to graphics, diagrams, text, and other materials created to present information in a way that is easy to understand visually.
[0098] A "display medium" refers to a device or part of a device used to present visual materials, allowing users to visually confirm information.
[0099] A "portable visual aid" is a device that a user can carry and use, and that has the ability to display visual information.
[0100] The system of this invention mainly consists of a server, a terminal, and a user. The user requests information through voice input, and that voice is converted into text information by the terminal. The terminal has speech recognition technology such as Google Cloud Speech-to-Text API installed, and has the ability to quickly convert voice data into text.
[0101] The server then acquires the text information and performs analysis using natural language processing. OpenAI's generative AI model and the BERT model are used for the analysis, ensuring that the intent of the user's question is accurately interpreted. Based on the analysis results, relevant information is retrieved from databases and external sources, such as Google FI® rebase and MongoDB, and the server uses this information to generate visual materials.
[0102] The visual materials are generated using D3.js or Tableau and are characterized by their user-friendly and interactive format. The server sends the generated visual materials to the terminal, which then presents them on the display of a portable visual aid, such as smart glasses, to provide the user with visual information.
[0103] Furthermore, to complement visual information, the server can use speech synthesis technologies such as Amazon Polly to provide the key points of the visual materials as audio.
[0104] As a concrete example, if a user asks "How do I use this product?" in a physical store, the terminal recognizes the voice, converts it into text, and sends it to the server. The server analyzes this information and searches its database for knowledge related to how to use the product. An infographic generated using D3.js is displayed on smart glasses, and an audio explanation is provided using Amazon Polly.
[0105] An example of a prompt to input into the generating AI model is: "The customer wants to know more details about the specifications of a specific product. Analyze this audio data to generate visual information and display it on the assistive device."
[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0107] Step 1:
[0108] The user provides voice input. The user speaks into the microphone of their smart glasses or device, asking for information. The input voice signal is sent to the device.
[0109] Step 2:
[0110] The device converts the audio signal into text information. Speech recognition software running on the device (e.g., Google Cloud Speech-to-Text API) outputs the audio signal as text in real time. Here, the input is an audio signal, and the output is text data.
[0111] Step 3:
[0112] The server acquires and analyzes text data. The server uses natural language processing (e.g., OpenAI GPT) to analyze the acquired text, understand the user's intent, and search for relevant information. The input is text data, and the output is relevant information.
[0113] Step 4:
[0114] The server generates visual materials based on relevant information. Visual material generation software (e.g., D3.js) creates an infographic based on the retrieved information. The input is relevant information, and the output is a visual material. In this step, the information is structured and visualized.
[0115] Step 5:
[0116] The server sends visual materials to the terminal for display. The terminal displays the received visual materials on the display of a portable visual aid (e.g., smart glasses). The input is the visual materials, and the output is the visual information on the display.
[0117] Step 6:
[0118] The server generates an audio summary of the visual information. Speech synthesis technology (e.g., Amazon Polly) converts the key points of the visual material into synthesized speech. The input is the key points of the visual material, and the output is synthesized speech.
[0119] Step 7:
[0120] The device plays synthesized speech to provide information to the user. The audio is played through the device's speaker, allowing the user to obtain information supplementarily through hearing. The input is synthesized speech data, and the output is audio playback.
[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0122] This invention aims to optimize the method of information presentation by recognizing user emotions through the integration of an emotion engine into a voice dialogue system. The system has the function of receiving user voice input via a terminal and converting it into text data through speech recognition technology. The converted text data is sent to a server and analyzed using natural language processing. Based on the analysis results, relevant information is retrieved from databases and external information sources.
[0123] However, a key feature of this invention is that it incorporates an emotion engine into the system that analyzes the user's emotions. The server recognizes the user's emotions from the voice data and adjusts the information presentation method based on the results. For example, if the system recognizes that the user is feeling anxious, it can provide information in a calm tone or organize visual materials concisely.
[0124] In generating visual materials, the server visually assembles the acquired information and optimizes it according to the user's emotional state. This material is then sent to the terminal and displayed to the user. The terminal displays the infographic on its screen and provides audio summaries as needed to aid user understanding.
[0125] As a concrete example, consider a scenario where a user asks, "Please tell me about the economic situation." The terminal converts the speech to text, and the emotion engine detects tension in the user's voice. As a result, the server provides a reassuring tone for the speech summary and presents visual materials more clearly and simply. In this way, it is possible to provide a conversational experience that responds to the user's emotions and improves the ease with which information can be received. This system utilizes both visual and audio information to provide flexible information tailored to the user's needs.
[0126] The following describes the processing flow.
[0127] Step 1:
[0128] The user inputs their question by voice using a microphone. For example, the user might ask the system for the latest weather information.
[0129] Step 2:
[0130] The device receives voice input and converts it into text data using speech recognition technology. The converted text is then sent to a server for analysis.
[0131] Step 3:
[0132] The server analyzes text data and user voice data, and evaluates the user's emotional state through an emotion engine. This information is used to ensure that users can receive information with confidence.
[0133] Step 4:
[0134] The server searches the database for relevant information based on the analysis results. The server then prepares to retrieve the most relevant information and provide it to the user.
[0135] Step 5:
[0136] The server generates visual materials based on the user's emotion recognition results. The layout of the visual materials is optimized according to the user's emotional state and designed to provide a sense of security and ease of understanding.
[0137] Step 6:
[0138] The generated visual materials are sent from the server to the terminal, which then displays the materials on its screen. Because the information is visually organized, users can easily understand it through their eyes.
[0139] Step 7:
[0140] Furthermore, the device uses synthesized speech to provide summarized information in a tone that is sensitive to the user's emotions. This audio is provided simultaneously with the infographic to assist the user in understanding the information.
[0141] Step 8:
[0142] Users understand the information by viewing visual materials and listening to audio summaries. Users can ask additional questions as needed, allowing the interaction to continue.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] In voice dialogue systems, there is a need to provide optimal information based on the user's emotions and situation. However, conventional systems have not adequately achieved flexible information delivery that takes user emotions into consideration. Therefore, technology is needed to analyze user emotions and provide information accordingly.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for analyzing the user's emotions, means for adjusting the information presentation method according to the user's emotional state, and means for generating video material and outputting it to a display device. This enables flexible information presentation that responds to the user's emotions.
[0148] "Voice input" refers to a method for recognizing and digitally processing voice signals from a user.
[0149] "Textual information" refers to data obtained by converting voice input into text format.
[0150] "Analysis" is the process of semantically understanding textual information and identifying related information.
[0151] "Visual materials" are digital content that visually represents acquired information.
[0152] A "display device" is a device that provides visual information to a user.
[0153] "Providing information via audio" refers to the process of conveying textual information and video materials to users through speech synthesis.
[0154] "Analyzing user emotions" refers to technology that identifies emotional states from voice input.
[0155] "Adjusting the method of information presentation" is the process of optimizing how information is delivered based on the user's emotional state.
[0156] This invention provides specific means for realizing flexible information presentation in a voice dialogue system that responds to the user's emotional state. This system can adjust the information presentation method by receiving voice input and analyzing the user's emotions.
[0157] When a user speaks a question or request into the device, the device uses its microphone to capture the audio as a digital signal. This captured audio data is then converted into text using Google Cloud Speech-to-Text or similar speech recognition technology.
[0158] The terminal then sends this converted text information to the server. The server uses natural language processing techniques to analyze the received text information. This process utilizes OpenAI's language model, a generative AI model, and other similar algorithms. Through this analysis, the server understands the user's intent and retrieves relevant information from databases and external sources.
[0159] The server then uses the voice input data to perform sentiment analysis. This involves techniques that analyze parameters such as voice tone, speed, and volume. Based on the sentiment analysis results, the server adjusts how information is presented. For example, if the user is feeling anxious, visual materials are designed to be simple and easy to understand, and the voice output is adjusted to a calm and reassuring tone.
[0160] In generating visual materials, the server uses tools to visually organize information. Visualization tools such as Tableau and Infogram are utilized to create video materials based on the acquired information. These materials are sent to the terminal, which displays them on its screen. The display includes infographics and statistical data, making the information easier for the user to understand visually.
[0161] As a concrete example, consider a scenario where a user speaks into the terminal saying, "Please tell me about the economic situation." This prompt is recognized as voice input and converted into text. If the emotion engine detects tension in the user's voice, the server uses this information to generate a simple infographic and create a calm voice summary. In this way, optimal information presentation tailored to the user's emotions is achieved.
[0162] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0163] Step 1:
[0164] The user makes a question or request to the terminal using their voice. The terminal uses a microphone to acquire this voice input and transmits it to the computer as digital audio data. The input is the user's voice signal, and the output is the digital data of that voice. This digital audio data is used as input for the next processing step.
[0165] Step 2:
[0166] The device converts the acquired digital audio data into text information using speech recognition software. For example, Google Cloud Speech-to-Text is used. The input is the digital audio data generated in step 1, and the output is the corresponding text data. Through this process, the user's voice input is converted into text information.
[0167] Step 3:
[0168] The terminal sends the converted character information to the server. The server analyzes the received character data using a generative AI model. Natural language processing techniques are used for the analysis, with OpenAI's language model being one example. The input is text data, and the output of the analysis is detailed information about the user's intent and the content of the inquiry. Instructions are generated to retrieve this information from a database or external sources.
[0169] Step 4:
[0170] The server performs sentiment analysis based on audio data and text information. It uses a sentiment analysis engine to evaluate the user's emotional state. Inputs are audio parameters (e.g., tone, speed) and text data, and output is the user's emotional state (e.g., tension, relief, etc.). Based on this output, instructions are generated to adjust how information is presented.
[0171] Step 5:
[0172] The server generates visual materials based on the acquired information and the results of sentiment analysis. Visualization tools (e.g., Tableau, Infogram) are used to generate these visual materials. The input consists of information acquired from the database and the results of sentiment analysis, and the output is a video material displayed to the user. This material is optimized according to the user's emotional state.
[0173] Step 6:
[0174] The server sends the generated visual materials and audio summaries to the terminal. The terminal presents this to the user on its display and, if necessary, uses speech synthesis technology to present the summaries aloud. The input is the visual materials and summary data from the server, and the output is the display on the user's screen and the audio output. This process allows the user to receive information in a way that is relevant to their emotions.
[0175] (Application Example 2)
[0176] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0177] Conventional voice dialogue systems fail to adequately consider individual emotional states when providing information to users, resulting in uniform information presentation and leading to inconsistencies in user understanding and acceptance. Furthermore, the lack of mechanisms to recognize user emotions in real time and provide information accordingly prevents the provision of a better customer experience.
[0178] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0179] In this invention, the server includes means for recognizing the user's emotions from voice data, means for adjusting the information presentation method based on the recognized emotions, and means for optimizing the layout of visual materials according to the emotional state. This enables information presentation tailored to each user's emotions, resulting in improved comprehension and a more satisfying conversational experience.
[0180] "Means for receiving voice input" refers to a device or technology that allows a system to receive voice signals emitted by a user.
[0181] "Means of converting to text data" refers to a device or technology for converting received audio into a digital written format.
[0182] "Means for analysis and retrieving relevant information" refers to devices or technologies for analyzing converted text data and identifying related information from databases or external sources.
[0183] "Means for generating visual materials" refers to devices or technologies that convert retrieved information into a visual format so that it can be presented visually.
[0184] "Means for outputting to a display device" refers to a device or technology for outputting generated visual material to a display medium such as a screen.
[0185] "Means of providing summaries in audio format" refers to a device or technology for concisely summarizing the content of visual materials and playing it back as audio.
[0186] "Means for recognizing a user's emotions from voice data" refers to a device or technology for analyzing the characteristics of voice to determine the emotional state of the speaker.
[0187] "Means for adjusting the method of information presentation" refers to devices or technologies for optimizing the method of information transmission and tone of presentation based on perceived emotions.
[0188] "Means for optimizing according to emotional state" refers to devices or technologies for adjusting the visual layout and information structure to match the user's emotions.
[0189] This system primarily consists of a user terminal, a server, and a visual display device. The terminal has an interface for receiving user voice input and uses speech recognition software to convert the input voice into text data. For example, a speech recognition API can be used for this process.
[0190] The server receives the converted text data and analyzes it using natural language processing techniques. This analysis system, for example, uses a natural language processing engine. Based on the analysis results, relevant information is retrieved from databases and external sources. The retrieved information is then converted into a visual format by a visual material generation module.
[0191] A key feature is that the server has a built-in emotion recognition engine. This engine analyzes the characteristics of the voice to determine the user's emotional state. Based on the results of the emotion recognition, the server adjusts the information presentation method and the layout of visual materials. For example, if the server determines that the user is nervous, the information will be presented in a calm tone, and the visual materials will be summarized concisely.
[0192] Finally, the generated visual materials and adjusted information are sent to the terminal and displayed on the screen. Users can view the visual materials through the screen and listen to an audio summary if needed. This audio summary is provided by playing the generated audio file.
[0193] As a concrete example, consider a scenario where a user asks, "Tell me about this week's popular products." If the emotion engine determines the user is excited during the process of the device converting speech to text and the server analyzing this information, the product information will be explained in a calm voice and presented on the screen in a visually easy-to-understand format. This allows the user to comfortably receive the information they are looking for.
[0194] Examples of prompts to input into a generative AI model include: "Please explain in detail how to provide product information displayed on a screen based on emotions."
[0195] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0196] Step 1:
[0197] The terminal receives voice input from the user. This voice input is the data that forms the basis for subsequent processing. The terminal uses speech recognition software to convert this voice data into text data. Here, the input is voice data and the output is string data. This conversion makes the voice into a format that can be analyzed.
[0198] Step 2:
[0199] The server analyzes the text data received from the terminal. This process uses natural language processing techniques to extract the user's intent from the text data and retrieve relevant information. The input is text data, and the output is a list of information. This information includes relevant data retrieved from a database.
[0200] Step 3:
[0201] The server uses an emotion recognition engine to recognize the user's emotions from the audio data. The input is the original audio data, and the output is data indicating the emotional state. This recognition result is used to determine the adjustment policy for information presentation.
[0202] Step 4:
[0203] The server uses a visual material generation module to create visually clear materials based on the acquired information and emotion recognition results. Inputs are information lists and emotion state data, and output is visual materials. The layout is adjusted according to the emotion state to provide information in the most optimal format for the user.
[0204] Step 5:
[0205] The server sends the generated visual materials and adjusted information to the terminal. The terminal outputs this material to a display device, making it available for the user to review. The input is the visual materials, and the output is the displayed material. In addition, if necessary, an audio summary of the information is provided as supplementary material.
[0206] Through the steps described above, the system can present users with information tailored to their individual emotional state.
[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0210] [Second Embodiment]
[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0223] The system of this invention functions by combining multiple means to streamline information transmission in voice dialogue. When a user inputs a question by voice into the interface, it is converted into text data using speech recognition technology. At this stage, the terminal has the ability to receive the voice and transcribe it in real time.
[0224] Subsequently, the server retrieves the text data and performs analysis using natural language processing. In this analysis step, the server efficiently searches for relevant information that reflects the intent and theme of the user's question. Once relevant data is retrieved from the databases and external information sources used here, the foundational information for the visual materials is established.
[0225] Next, the server generates visual materials. Visual materials are visualizations, such as infographics, based on the acquired information, presenting the information in a format that is intuitively easy for the user to understand. This generation process includes optimizing the layout of the visual materials so that they are easy for the user to read and the key points can be clearly recognized.
[0226] Once the visual material is complete, the server sends it to the terminal, where it is displayed. The displayed visual material incorporates diagrams, graphs, text, and other elements to make the information easier to understand visually. Furthermore, if necessary, synthesized speech is used to provide audio summaries of the key points of the visual material, further supporting information comprehension.
[0227] As a concrete example, consider a scenario where a user asks, "Please tell me how to use solar energy." The terminal recognizes the speech and converts it into text data. The server analyzes this data and searches for information on how to use solar energy. It then generates relevant visual materials and sends them to the terminal. The terminal displays these materials on its screen and also provides an audio explanation of the main uses. The user can quickly understand the information using both sight and hearing. In this way, the system effectively supports the user's understanding of information.
[0228] The following describes the processing flow.
[0229] Step 1:
[0230] The user inputs questions by voice using a microphone. The user asks the system questions by voice about a topic for which they want to obtain specific information.
[0231] Step 2:
[0232] The device uses speech recognition technology to convert the input speech into text data. The device then sends this converted text data to the next processing step.
[0233] Step 3:
[0234] The server analyzes the text data and uses natural language processing to identify the user's intent and related keywords. Based on the question, the server prepares to search for the most relevant data.
[0235] Step 4:
[0236] The server searches for relevant information from databases and external sources based on the analysis results. The server quickly retrieves accurate data in response to the user's questions.
[0237] Step 5:
[0238] Based on the information acquired by the server, visual materials are generated. The server creates infographics in a format that is easily understandable to the user and optimizes the layout of the visual materials.
[0239] Step 6:
[0240] The server sends the generated visual material to the terminal. The terminal prepares to present the received visual material to the user.
[0241] Step 7:
[0242] The device displays visual materials on its screen and provides users with summarized information using synthesized speech. The device is designed to facilitate visual and auditory comprehension of information.
[0243] Step 8:
[0244] Users can effectively understand information by reviewing visual materials and listening to audio summaries. Users can follow up with additional questions as needed.
[0245] (Example 1)
[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0247] In systems that efficiently provide information through voice interaction, users need to quickly and intuitively understand large amounts of information. However, in conventional systems, information obtained from voice input is provided in text format, making it difficult to effectively convey information using both sight and hearing. A new method of information transmission is needed to solve this problem and facilitate user understanding.
[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0249] In this invention, the server includes means for analyzing converted text data and searching for relevant information using natural language processing, means for a generation engine for generating visual materials using the relevant information, and means for visualizing the information in the generated visual materials. As a result, the user can receive not only audio but also visually organized information simultaneously, enabling quick and intuitive understanding of the information.
[0250] "Means for receiving voice input" refers to a device or technology that provides a function for taking in external voice information into a system.
[0251] "Methods for converting speech to text data" refers to technologies that perform a process of analyzing speech signals and converting them into corresponding text data.
[0252] "Means for analyzing text data and searching for related information using natural language processing" refers to technologies that analyze text data, understand the user's intent using natural language processing techniques, and have the functionality to search for related information from databases and external sources.
[0253] "Means equipped with a generation engine for generating visual materials" refers to a function equipped with an engine for generating visual materials based on acquired information, and is a technology that visualizes information as infographics or charts.
[0254] "Means of visualizing information in visual materials" refers to techniques for organizing generated materials and presenting them in a visually easy-to-understand format.
[0255] "A means of outputting visual materials to a display device and providing their key points in synthesized speech" refers to a technology that displays generated visual materials on a display device for user viewing and simultaneously explains the key points of the materials in synthesized speech.
[0256] This system combines multiple technologies to acquire information via voice and effectively convey it to the user. When a user provides voice input, the terminal first receives the voice. The terminal uses a microphone to convert the voice into digital data, performs speech recognition using that data, and then converts the voice back into text data. At this stage, speech recognition software such as the Google Speech-to-Text API is utilized.
[0257] Next, the server retrieves this text data. The server uses natural language processing techniques to analyze the text data. Based on the analyzed information, the server connects to internal databases and external information sources to retrieve relevant data. This process uses common data analysis tools and search algorithms.
[0258] The information obtained is then generated as visual materials through a generation engine on the server. Visual materials such as infographics and charts are optimized in layout so that the information is presented in a format that is intuitively easy for the user to understand. Visualization tools such as Infogram and Tableau are used during this generation process.
[0259] Finally, the generated visual materials are sent to the terminal. The terminal outputs them to a display device, allowing the user to view the materials. In addition, a speech synthesis engine installed in the terminal provides an audio summary of the visual materials.
[0260] As a concrete example, consider a scenario where a user asks a question via voice, "Please tell me how to use solar energy." In this case, the terminal converts the voice into text, the server collects and analyzes information about solar energy, generates visual materials, and sends them to the user's terminal. Through this entire process, the user can quickly understand the information through both voice and visual means.
[0261] An example of a prompt might be, "Generate a visual explanation of how to utilize solar energy." Such a prompt prompt would prompt the server to begin the process of visualizing the relevant information.
[0262] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0263] Step 1:
[0264] The user provides voice input through the interface. The input voice is captured by the terminal and converted into digital data via the microphone. The terminal then passes this digital voice data to a speech recognition engine, which generates text data in real time. During this process, voice signal processing is performed to reduce background noise and clarify speech.
[0265] Step 2:
[0266] The device sends text data to the server. Upon receiving the text data, the server analyzes it using natural language processing (NLP) techniques. Specifically, it extracts relevant keywords by performing data processing such as text tokenization, morphological analysis, and language model application to understand the intent of the user's question.
[0267] Step 3:
[0268] The server searches for relevant information based on the analysis results. It accesses databases and external information sources to retrieve data related to the keywords. Using information retrieval algorithms, it quickly selects the most relevant information. The retrieved information is then formatted into a dataset necessary for generating visual materials in the next step.
[0269] Step 4:
[0270] The server takes a dataset as input and generates visual materials using a generative AI model. This process involves visualizations such as charts, graphs, and infographics, organizing the data into a user-friendly format. Visual design optimization ensures that the generated materials are visually appealing.
[0271] Step 5:
[0272] The generated visual materials are sent from the server to the terminal. The terminal receives them and outputs them to the display device. Furthermore, the terminal's built-in speech synthesis function is used to provide the user with audio summaries of the key features and points of the visual materials, thereby supporting a deeper understanding of the information. In this state, the user can obtain information through both audio and visual means.
[0273] (Application Example 1)
[0274] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0275] Traditional methods for customers to efficiently obtain information in physical stores are time-consuming, making it difficult to quickly acquire appropriate information. In particular, requests for detailed product usage instructions and characteristics require the assistance of store staff, leading to resource depletion and wasted time. There is a need to address these challenges and streamline the provision of information within stores.
[0276] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0277] In this invention, the server includes a medium for receiving voice input, a medium for converting the received voice into text information, a medium for analyzing the converted text information and retrieving related information, and a medium for presenting the generated visual material to a portable visual aid device. This enables customers to ask questions by voice within the store and obtain information that is quick and easy to understand visually.
[0278] "Voice input" is the process of recognizing the voice signals spoken by a user and receiving them as digital data.
[0279] A "medium" is a means of processing, converting, displaying, or transmitting data such as audio signals, visual materials, and textual information.
[0280] "Textual information" refers to text data that is obtained by analyzing and converting audio signals.
[0281] "Analysis" is the process of deciphering input text information and understanding its meaning and intent.
[0282] "Relevant information" refers to the appropriate data and knowledge that are retrieved based on the user's questions or requests.
[0283] "Visual materials" refer to graphics, diagrams, text, and other materials created to present information in a way that is easy to understand visually.
[0284] The "display medium" is a device or a part of a device for presenting visual materials, through which users can visually confirm information.
[0285] The "portable visual assistance device" is a device that users can carry and use, and has the ability to display visual information.
[0286] The system of this invention is mainly composed of a server, a terminal, and a user. The user requests information through voice input, and the voice is converted into character information by the terminal. Google Cloud Speech-to-Text API and other voice recognition technologies are installed on the terminal, which has the ability to quickly convert voice data into text.
[0287] After that, the server obtains the character information and performs analysis using natural language processing. For the analysis, OpenAI's generative AI model or BERT model is used, and through these technologies, the intention of the user's question is accurately interpreted. Based on the analysis results, relevant information is retrieved from a database or external information sources, such as Google Firebase or MongoDB, and the server uses this information to generate visual materials.
[0288] The visual materials are generated by using D3.js or Tableau, and are characterized by being easy to view and in an interactive format. The server sends the generated visual materials to the terminal, and the terminal presents them on the display of a portable visual assistance device, such as smart glasses, to provide visual information to the user.
[0289] In addition, to complement the visual information, the server can use voice synthesis technology such as Amazon Polly to provide the key points of the visual materials as voice.
[0290] As a concrete example, if a user asks "How do I use this product?" in a physical store, the terminal recognizes the voice, converts it into text, and sends it to the server. The server analyzes this information and searches its database for knowledge related to how to use the product. An infographic generated using D3.js is displayed on smart glasses, and an audio explanation is provided using Amazon Polly.
[0291] An example of a prompt to input into the generating AI model is: "The customer wants to know more details about the specifications of a specific product. Analyze this audio data to generate visual information and display it on the assistive device."
[0292] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0293] Step 1:
[0294] The user provides voice input. The user speaks into the microphone of their smart glasses or device, asking for information. The input voice signal is sent to the device.
[0295] Step 2:
[0296] The device converts the audio signal into text information. Speech recognition software running on the device (e.g., Google Cloud Speech-to-Text API) outputs the audio signal as text in real time. Here, the input is an audio signal, and the output is text data.
[0297] Step 3:
[0298] The server acquires and analyzes text data. The server uses natural language processing (e.g., OpenAI GPT) to analyze the acquired text, understand the user's intent, and search for relevant information. The input is text data, and the output is relevant information.
[0299] Step 4:
[0300] The server generates visual materials based on relevant information. Visual material generation software (e.g., D3.js) creates infographics based on the retrieved information. The input is the relevant information, and the output is the visual materials. In this step, the structuring and visualization of information are performed.
[0301] Step 5:
[0302] The server sends the visual materials to the terminal for display. The terminal displays the received visual materials on the display of a portable visual assistance device (e.g., smart glasses). The input is the visual materials, and the output is the visual information on the display.
[0303] Step 6:
[0304] The server generates a summary of the visual information as audio. Speech synthesis technology (e.g., Amazon Polly) converts the key points of the visual materials into synthetic audio. The input is the key point information of the visual materials, and the output is the synthetic audio.
[0305] Step 7:
[0306] The terminal plays the synthetic audio to provide information to the user. By playing the audio through the terminal's speaker, the user can complementarily obtain information through hearing. The input is the synthetic audio data, and the output is the audio playback.
[0307] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0308] This invention aims to optimize the method of information presentation by recognizing user emotions through the integration of an emotion engine into a voice dialogue system. The system has the function of receiving user voice input via a terminal and converting it into text data through speech recognition technology. The converted text data is sent to a server and analyzed using natural language processing. Based on the analysis results, relevant information is retrieved from databases and external information sources.
[0309] However, a key feature of this invention is that it incorporates an emotion engine into the system that analyzes the user's emotions. The server recognizes the user's emotions from the voice data and adjusts the information presentation method based on the results. For example, if the system recognizes that the user is feeling anxious, it can provide information in a calm tone or organize visual materials concisely.
[0310] In generating visual materials, the server visually assembles the acquired information and optimizes it according to the user's emotional state. This material is then sent to the terminal and displayed to the user. The terminal displays the infographic on its screen and provides audio summaries as needed to aid user understanding.
[0311] As a concrete example, consider a scenario where a user asks, "Please tell me about the economic situation." The terminal converts the speech to text, and the emotion engine detects tension in the user's voice. As a result, the server provides a reassuring tone for the speech summary and presents visual materials more clearly and simply. In this way, it is possible to provide a conversational experience that responds to the user's emotions and improves the ease with which information can be received. This system utilizes both visual and audio information to provide flexible information tailored to the user's needs.
[0312] The following describes the processing flow.
[0313] Step 1:
[0314] The user inputs their question by voice using a microphone. For example, the user might ask the system for the latest weather information.
[0315] Step 2:
[0316] The device receives voice input and converts it into text data using speech recognition technology. The converted text is then sent to a server for analysis.
[0317] Step 3:
[0318] The server analyzes text data and user voice data, and evaluates the user's emotional state through an emotion engine. This information is used to ensure that users can receive information with confidence.
[0319] Step 4:
[0320] The server searches the database for relevant information based on the analysis results. The server then prepares to retrieve the most relevant information and provide it to the user.
[0321] Step 5:
[0322] The server generates visual materials based on the user's emotion recognition results. The layout of the visual materials is optimized according to the user's emotional state and designed to provide a sense of security and ease of understanding.
[0323] Step 6:
[0324] The generated visual materials are sent from the server to the terminal, which then displays the materials on its screen. Because the information is visually organized, users can easily understand it through their eyes.
[0325] Step 7:
[0326] Furthermore, the device uses synthesized speech to provide summarized information in a tone that is sensitive to the user's emotions. This audio is provided simultaneously with the infographic to assist the user in understanding the information.
[0327] Step 8:
[0328] Users understand the information by viewing visual materials and listening to audio summaries. Users can ask additional questions as needed, allowing the interaction to continue.
[0329] (Example 2)
[0330] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0331] In voice dialogue systems, there is a need to provide optimal information based on the user's emotions and situation. However, conventional systems have not adequately achieved flexible information delivery that takes user emotions into consideration. Therefore, technology is needed to analyze user emotions and provide information accordingly.
[0332] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0333] In this invention, the server includes means for analyzing the user's emotions, means for adjusting the information presentation method according to the user's emotional state, and means for generating video material and outputting it to a display device. This enables flexible information presentation that responds to the user's emotions.
[0334] "Voice input" refers to a method for recognizing and digitally processing voice signals from a user.
[0335] "Textual information" refers to data obtained by converting voice input into text format.
[0336] "Analysis" is the process of semantically understanding textual information and identifying related information.
[0337] "Visual materials" are digital content that visually represents acquired information.
[0338] A "display device" is a device that provides visual information to a user.
[0339] "Providing information via audio" refers to the process of conveying textual information and video materials to users through speech synthesis.
[0340] "Analyzing user emotions" refers to technology that identifies emotional states from voice input.
[0341] "Adjusting the method of information presentation" is the process of optimizing how information is delivered based on the user's emotional state.
[0342] This invention provides specific means for realizing flexible information presentation in a voice dialogue system that responds to the user's emotional state. This system can adjust the information presentation method by receiving voice input and analyzing the user's emotions.
[0343] When a user speaks a question or request into the device, the device uses its microphone to capture the audio as a digital signal. This captured audio data is then converted into text using Google Cloud Speech-to-Text or similar speech recognition technology.
[0344] The terminal then sends this converted text information to the server. The server uses natural language processing techniques to analyze the received text information. This process utilizes OpenAI's language model, a generative AI model, and other similar algorithms. Through this analysis, the server understands the user's intent and retrieves relevant information from databases and external sources.
[0345] The server then uses the voice input data to perform sentiment analysis. This involves techniques that analyze parameters such as voice tone, speed, and volume. Based on the sentiment analysis results, the server adjusts how information is presented. For example, if the user is feeling anxious, visual materials are designed to be simple and easy to understand, and the voice output is adjusted to a calm and reassuring tone.
[0346] In generating visual materials, the server uses tools to visually organize information. Visualization tools such as Tableau and Infogram are utilized to create video materials based on the acquired information. These materials are sent to the terminal, which displays them on its screen. The display includes infographics and statistical data, making the information easier for the user to understand visually.
[0347] As a concrete example, consider a scenario where a user speaks into the terminal saying, "Please tell me about the economic situation." This prompt is recognized as voice input and converted into text. If the emotion engine detects tension in the user's voice, the server uses this information to generate a simple infographic and create a calm voice summary. In this way, optimal information presentation tailored to the user's emotions is achieved.
[0348] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0349] Step 1:
[0350] The user makes a question or request to the terminal using their voice. The terminal uses a microphone to acquire this voice input and transmits it to the computer as digital audio data. The input is the user's voice signal, and the output is the digital data of that voice. This digital audio data is used as input for the next processing step.
[0351] Step 2:
[0352] The device converts the acquired digital audio data into text information using speech recognition software. For example, Google Cloud Speech-to-Text is used. The input is the digital audio data generated in step 1, and the output is the corresponding text data. Through this process, the user's voice input is converted into text information.
[0353] Step 3:
[0354] The terminal sends the converted character information to the server. The server analyzes the received character data using a generative AI model. Natural language processing techniques are used for the analysis, with OpenAI's language model being one example. The input is text data, and the output of the analysis is detailed information about the user's intent and the content of the inquiry. Instructions are generated to retrieve this information from a database or external sources.
[0355] Step 4:
[0356] The server performs sentiment analysis based on audio data and text information. It uses a sentiment analysis engine to evaluate the user's emotional state. Inputs are audio parameters (e.g., tone, speed) and text data, and output is the user's emotional state (e.g., tension, relief, etc.). Based on this output, instructions are generated to adjust how information is presented.
[0357] Step 5:
[0358] The server generates visual materials based on the acquired information and the results of sentiment analysis. Visualization tools (e.g., Tableau, Infogram) are used to generate these visual materials. The input consists of information acquired from the database and the results of sentiment analysis, and the output is a video material displayed to the user. This material is optimized according to the user's emotional state.
[0359] Step 6:
[0360] The server sends the generated visual materials and audio summaries to the terminal. The terminal presents this to the user on its display and, if necessary, uses speech synthesis technology to present the summaries aloud. The input is the visual materials and summary data from the server, and the output is the display on the user's screen and the audio output. This process allows the user to receive information in a way that is relevant to their emotions.
[0361] (Application Example 2)
[0362] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0363] Conventional voice dialogue systems fail to adequately consider individual emotional states when providing information to users, resulting in uniform information presentation and leading to inconsistencies in user understanding and acceptance. Furthermore, the lack of mechanisms to recognize user emotions in real time and provide information accordingly prevents the provision of a better customer experience.
[0364] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0365] In this invention, the server includes means for recognizing the user's emotions from voice data, means for adjusting the information presentation method based on the recognized emotions, and means for optimizing the layout of visual materials according to the emotional state. This enables information presentation tailored to each user's emotions, resulting in improved comprehension and a more satisfying conversational experience.
[0366] "Means for receiving voice input" refers to a device or technology that allows a system to receive voice signals emitted by a user.
[0367] "Means of converting to text data" refers to a device or technology for converting received audio into a digital written format.
[0368] "Means for analysis and retrieving relevant information" refers to devices or technologies for analyzing converted text data and identifying related information from databases or external sources.
[0369] "Means for generating visual materials" refers to devices or technologies that convert retrieved information into a visual format so that it can be presented visually.
[0370] "Means for outputting to a display device" refers to a device or technology for outputting generated visual material to a display medium such as a screen.
[0371] "Means of providing summaries in audio format" refers to a device or technology for concisely summarizing the content of visual materials and playing it back as audio.
[0372] "Means for recognizing a user's emotions from voice data" refers to a device or technology for analyzing the characteristics of voice to determine the emotional state of the speaker.
[0373] "Means for adjusting the method of information presentation" refers to devices or technologies for optimizing the method of information transmission and tone of presentation based on perceived emotions.
[0374] "Means for optimizing according to emotional state" refers to devices or technologies for adjusting the visual layout and information structure to match the user's emotions.
[0375] This system primarily consists of a user terminal, a server, and a visual display device. The terminal has an interface for receiving user voice input and uses speech recognition software to convert the input voice into text data. For example, a speech recognition API can be used for this process.
[0376] The server receives the converted text data and analyzes it using natural language processing techniques. This analysis system, for example, uses a natural language processing engine. Based on the analysis results, relevant information is retrieved from databases and external sources. The retrieved information is then converted into a visual format by a visual material generation module.
[0377] A key feature is that the server has a built-in emotion recognition engine. This engine analyzes the characteristics of the voice to determine the user's emotional state. Based on the results of the emotion recognition, the server adjusts the information presentation method and the layout of visual materials. For example, if the server determines that the user is nervous, the information will be presented in a calm tone, and the visual materials will be summarized concisely.
[0378] Finally, the generated visual materials and adjusted information are sent to the terminal and displayed on the screen. Users can view the visual materials through the screen and listen to an audio summary if needed. This audio summary is provided by playing the generated audio file.
[0379] As a concrete example, consider a scenario where a user asks, "Tell me about this week's popular products." If the emotion engine determines the user is excited during the process of the device converting speech to text and the server analyzing this information, the product information will be explained in a calm voice and presented on the screen in a visually easy-to-understand format. This allows the user to comfortably receive the information they are looking for.
[0380] Examples of prompts to input into a generative AI model include: "Please explain in detail how to provide product information displayed on a screen based on emotions."
[0381] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0382] Step 1:
[0383] The terminal receives voice input from the user. This voice input is the data that forms the basis for subsequent processing. The terminal uses speech recognition software to convert this voice data into text data. Here, the input is voice data and the output is string data. This conversion makes the voice into a format that can be analyzed.
[0384] Step 2:
[0385] The server analyzes the text data received from the terminal. This process uses natural language processing techniques to extract the user's intent from the text data and retrieve relevant information. The input is text data, and the output is a list of information. This information includes relevant data retrieved from a database.
[0386] Step 3:
[0387] The server uses an emotion recognition engine to recognize the user's emotions from the audio data. The input is the original audio data, and the output is data indicating the emotional state. This recognition result is used to determine the adjustment policy for information presentation.
[0388] Step 4:
[0389] The server uses a visual material generation module to create visually clear materials based on the acquired information and emotion recognition results. Inputs are information lists and emotion state data, and output is visual materials. The layout is adjusted according to the emotion state to provide information in the most optimal format for the user.
[0390] Step 5:
[0391] The server sends the generated visual materials and adjusted information to the terminal. The terminal outputs this material to a display device, making it available for the user to review. The input is the visual materials, and the output is the displayed material. In addition, if necessary, an audio summary of the information is provided as supplementary material.
[0392] Through the steps described above, the system can present users with information tailored to their individual emotional state.
[0393] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0396] [Third Embodiment]
[0397] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0398] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0400] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0404] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0405] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0406] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0407] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0408] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0409] The system of this invention functions by combining multiple means to streamline information transmission in voice dialogue. When a user inputs a question by voice into the interface, it is converted into text data using speech recognition technology. At this stage, the terminal has the ability to receive the voice and transcribe it in real time.
[0410] Subsequently, the server retrieves the text data and performs analysis using natural language processing. In this analysis step, the server efficiently searches for relevant information that reflects the intent and theme of the user's question. Once relevant data is retrieved from the databases and external information sources used here, the foundational information for the visual materials is established.
[0411] Next, the server generates visual materials. Visual materials are visualizations, such as infographics, based on the acquired information, presenting the information in a format that is intuitively easy for the user to understand. This generation process includes optimizing the layout of the visual materials so that they are easy for the user to read and the key points can be clearly recognized.
[0412] Once the visual material is complete, the server sends it to the terminal, where it is displayed. The displayed visual material incorporates diagrams, graphs, text, and other elements to make the information easier to understand visually. Furthermore, if necessary, synthesized speech is used to provide audio summaries of the key points of the visual material, further supporting information comprehension.
[0413] As a concrete example, consider a scenario where a user asks, "Please tell me how to use solar energy." The terminal recognizes the speech and converts it into text data. The server analyzes this data and searches for information on how to use solar energy. It then generates relevant visual materials and sends them to the terminal. The terminal displays these materials on its screen and also provides an audio explanation of the main uses. The user can quickly understand the information using both sight and hearing. In this way, the system effectively supports the user's understanding of information.
[0414] The following describes the processing flow.
[0415] Step 1:
[0416] The user inputs questions by voice using a microphone. The user asks the system questions by voice about a topic for which they want to obtain specific information.
[0417] Step 2:
[0418] The device uses speech recognition technology to convert the input speech into text data. The device then sends this converted text data to the next processing step.
[0419] Step 3:
[0420] The server analyzes the text data and uses natural language processing to identify the user's intent and related keywords. Based on the question, the server prepares to search for the most relevant data.
[0421] Step 4:
[0422] The server searches for relevant information from databases and external sources based on the analysis results. The server quickly retrieves accurate data in response to the user's questions.
[0423] Step 5:
[0424] Based on the information acquired by the server, visual materials are generated. The server creates infographics in a format that is easily understandable to the user and optimizes the layout of the visual materials.
[0425] Step 6:
[0426] The server sends the generated visual material to the terminal. The terminal prepares to present the received visual material to the user.
[0427] Step 7:
[0428] The device displays visual materials on its screen and provides users with summarized information using synthesized speech. The device is designed to facilitate visual and auditory comprehension of information.
[0429] Step 8:
[0430] Users can effectively understand information by reviewing visual materials and listening to audio summaries. Users can follow up with additional questions as needed.
[0431] (Example 1)
[0432] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0433] In systems that efficiently provide information through voice interaction, users need to quickly and intuitively understand large amounts of information. However, in conventional systems, information obtained from voice input is provided in text format, making it difficult to effectively convey information using both sight and hearing. A new method of information transmission is needed to solve this problem and facilitate user understanding.
[0434] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0435] In this invention, the server includes means for analyzing converted text data and searching for relevant information using natural language processing, means for a generation engine for generating visual materials using the relevant information, and means for visualizing the information in the generated visual materials. As a result, the user can receive not only audio but also visually organized information simultaneously, enabling quick and intuitive understanding of the information.
[0436] "Means for receiving voice input" refers to a device or technology that provides a function for taking in external voice information into a system.
[0437] "Methods for converting speech to text data" refers to technologies that perform a process of analyzing speech signals and converting them into corresponding text data.
[0438] "Means for analyzing text data and searching for related information using natural language processing" refers to technologies that analyze text data, understand the user's intent using natural language processing techniques, and have the functionality to search for related information from databases and external sources.
[0439] "Means equipped with a generation engine for generating visual materials" refers to a function equipped with an engine for generating visual materials based on acquired information, and is a technology that visualizes information as infographics or charts.
[0440] "Means of visualizing information in visual materials" refers to techniques for organizing generated materials and presenting them in a visually easy-to-understand format.
[0441] "A means of outputting visual materials to a display device and providing their key points in synthesized speech" refers to a technology that displays generated visual materials on a display device for user viewing and simultaneously explains the key points of the materials in synthesized speech.
[0442] This system combines multiple technologies to acquire information via voice and effectively convey it to the user. When a user provides voice input, the terminal first receives the voice. The terminal uses a microphone to convert the voice into digital data, performs speech recognition using that data, and then converts the voice back into text data. At this stage, speech recognition software such as the Google Speech-to-Text API is utilized.
[0443] Next, the server retrieves this text data. The server uses natural language processing techniques to analyze the text data. Based on the analyzed information, the server connects to internal databases and external information sources to retrieve relevant data. This process uses common data analysis tools and search algorithms.
[0444] The information obtained is then generated as visual materials through a generation engine on the server. Visual materials such as infographics and charts are optimized in layout so that the information is presented in a format that is intuitively easy for the user to understand. Visualization tools such as Infogram and Tableau are used during this generation process.
[0445] Finally, the generated visual materials are sent to the terminal. The terminal outputs them to a display device, allowing the user to view the materials. In addition, a speech synthesis engine installed in the terminal provides an audio summary of the visual materials.
[0446] As a concrete example, consider a scenario where a user asks a question via voice, "Please tell me how to use solar energy." In this case, the terminal converts the voice into text, the server collects and analyzes information about solar energy, generates visual materials, and sends them to the user's terminal. Through this entire process, the user can quickly understand the information through both voice and visual means.
[0447] An example of a prompt might be, "Generate a visual explanation of how to utilize solar energy." Such a prompt prompt would prompt the server to begin the process of visualizing the relevant information.
[0448] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0449] Step 1:
[0450] The user provides voice input through the interface. The input voice is captured by the terminal and converted into digital data via the microphone. The terminal then passes this digital voice data to a speech recognition engine, which generates text data in real time. During this process, voice signal processing is performed to reduce background noise and clarify speech.
[0451] Step 2:
[0452] The device sends text data to the server. Upon receiving the text data, the server analyzes it using natural language processing (NLP) techniques. Specifically, it extracts relevant keywords by performing data processing such as text tokenization, morphological analysis, and language model application to understand the intent of the user's question.
[0453] Step 3:
[0454] The server searches for relevant information based on the analysis results. It accesses databases and external information sources to retrieve data related to the keywords. Using information retrieval algorithms, it quickly selects the most relevant information. The retrieved information is then formatted into a dataset necessary for generating visual materials in the next step.
[0455] Step 4:
[0456] The server takes a dataset as input and generates visual materials using a generative AI model. This process involves visualizations such as charts, graphs, and infographics, organizing the data into a user-friendly format. Visual design optimization ensures that the generated materials are visually appealing.
[0457] Step 5:
[0458] The generated visual materials are sent from the server to the terminal. The terminal receives them and outputs them to the display device. Furthermore, the terminal's built-in speech synthesis function is used to provide the user with audio summaries of the key features and points of the visual materials, thereby supporting a deeper understanding of the information. In this state, the user can obtain information through both audio and visual means.
[0459] (Application Example 1)
[0460] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0461] Traditional methods for customers to efficiently obtain information in physical stores are time-consuming, making it difficult to quickly acquire appropriate information. In particular, requests for detailed product usage instructions and characteristics require the assistance of store staff, leading to resource depletion and wasted time. There is a need to address these challenges and streamline the provision of information within stores.
[0462] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0463] In this invention, the server includes a medium for receiving voice input, a medium for converting the received voice into text information, a medium for analyzing the converted text information and retrieving related information, and a medium for presenting the generated visual material to a portable visual aid device. This enables customers to ask questions by voice within the store and obtain information that is quick and easy to understand visually.
[0464] "Voice input" is the process of recognizing the voice signals spoken by a user and receiving them as digital data.
[0465] A "medium" is a means of processing, converting, displaying, or transmitting data such as audio signals, visual materials, and textual information.
[0466] "Textual information" refers to text data that is obtained by analyzing and converting audio signals.
[0467] "Analysis" is the process of deciphering input text information and understanding its meaning and intent.
[0468] "Relevant information" refers to the appropriate data and knowledge that are retrieved based on the user's questions or requests.
[0469] "Visual materials" refer to graphics, diagrams, text, and other materials created to present information in a way that is easy to understand visually.
[0470] A "display medium" refers to a device or part of a device used to present visual materials, allowing users to visually confirm information.
[0471] A "portable visual aid" is a device that a user can carry and use, and that has the ability to display visual information.
[0472] The system of this invention mainly consists of a server, a terminal, and a user. The user requests information through voice input, and that voice is converted into text information by the terminal. The terminal has speech recognition technology such as Google Cloud Speech-to-Text API installed, and has the ability to quickly convert voice data into text.
[0473] The server then retrieves the text information and performs analysis using natural language processing. OpenAI's generative AI model and the BERT model are used for the analysis, ensuring that the intent of the user's question is accurately interpreted. Based on the analysis results, relevant information is retrieved from databases and external sources, such as Google Firebase and MongoDB, and the server uses this information to generate visual materials.
[0474] The visual materials are generated using D3.js or Tableau and are characterized by their user-friendly and interactive format. The server sends the generated visual materials to the terminal, which then presents them on the display of a portable visual aid, such as smart glasses, to provide the user with visual information.
[0475] Furthermore, to complement visual information, the server can use speech synthesis technologies such as Amazon Polly to provide the key points of the visual materials as audio.
[0476] As a concrete example, if a user asks "How do I use this product?" in a physical store, the terminal recognizes the voice, converts it into text, and sends it to the server. The server analyzes this information and searches its database for knowledge related to how to use the product. An infographic generated using D3.js is displayed on smart glasses, and an audio explanation is provided using Amazon Polly.
[0477] An example of a prompt to input into the generating AI model is: "The customer wants to know more details about the specifications of a specific product. Analyze this audio data to generate visual information and display it on the assistive device."
[0478] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0479] Step 1:
[0480] The user provides voice input. The user speaks into the microphone of their smart glasses or device, asking for information. The input voice signal is sent to the device.
[0481] Step 2:
[0482] The device converts the audio signal into text information. Speech recognition software running on the device (e.g., Google Cloud Speech-to-Text API) outputs the audio signal as text in real time. Here, the input is an audio signal, and the output is text data.
[0483] Step 3:
[0484] The server acquires and analyzes text data. The server uses natural language processing (e.g., OpenAI GPT) to analyze the acquired text, understand the user's intent, and search for relevant information. The input is text data, and the output is relevant information.
[0485] Step 4:
[0486] The server generates visual materials based on relevant information. Visual material generation software (e.g., D3.js) creates an infographic based on the retrieved information. The input is relevant information, and the output is a visual material. In this step, the information is structured and visualized.
[0487] Step 5:
[0488] The server sends visual materials to the terminal for display. The terminal displays the received visual materials on the display of a portable visual aid (e.g., smart glasses). The input is the visual materials, and the output is the visual information on the display.
[0489] Step 6:
[0490] The server generates an audio summary of the visual information. Speech synthesis technology (e.g., Amazon Polly) converts the key points of the visual material into synthesized speech. The input is the key points of the visual material, and the output is synthesized speech.
[0491] Step 7:
[0492] The device plays synthesized speech to provide information to the user. The audio is played through the device's speaker, allowing the user to obtain information supplementarily through hearing. The input is synthesized speech data, and the output is audio playback.
[0493] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0494] This invention aims to optimize the method of information presentation by recognizing user emotions through the integration of an emotion engine into a voice dialogue system. The system has the function of receiving user voice input via a terminal and converting it into text data through speech recognition technology. The converted text data is sent to a server and analyzed using natural language processing. Based on the analysis results, relevant information is retrieved from databases and external information sources.
[0495] However, a key feature of this invention is that it incorporates an emotion engine into the system that analyzes the user's emotions. The server recognizes the user's emotions from the voice data and adjusts the information presentation method based on the results. For example, if the system recognizes that the user is feeling anxious, it can provide information in a calm tone or organize visual materials concisely.
[0496] In generating visual materials, the server visually assembles the acquired information and optimizes it according to the user's emotional state. This material is then sent to the terminal and displayed to the user. The terminal displays the infographic on its screen and provides audio summaries as needed to aid user understanding.
[0497] As a concrete example, consider a scenario where a user asks, "Please tell me about the economic situation." The terminal converts the speech to text, and the emotion engine detects tension in the user's voice. As a result, the server provides a reassuring tone for the speech summary and presents visual materials more clearly and simply. In this way, it is possible to provide a conversational experience that responds to the user's emotions and improves the ease with which information can be received. This system utilizes both visual and audio information to provide flexible information tailored to the user's needs.
[0498] The following describes the processing flow.
[0499] Step 1:
[0500] The user inputs their question by voice using a microphone. For example, the user might ask the system for the latest weather information.
[0501] Step 2:
[0502] The device receives voice input and converts it into text data using speech recognition technology. The converted text is then sent to a server for analysis.
[0503] Step 3:
[0504] The server analyzes text data and user voice data, and evaluates the user's emotional state through an emotion engine. This information is used to ensure that users can receive information with confidence.
[0505] Step 4:
[0506] The server searches the database for relevant information based on the analysis results. The server then prepares to retrieve the most relevant information and provide it to the user.
[0507] Step 5:
[0508] The server generates visual materials based on the user's emotion recognition results. The layout of the visual materials is optimized according to the user's emotional state and designed to provide a sense of security and ease of understanding.
[0509] Step 6:
[0510] The generated visual materials are sent from the server to the terminal, which then displays the materials on its screen. Because the information is visually organized, users can easily understand it through their eyes.
[0511] Step 7:
[0512] Furthermore, the device uses synthesized speech to provide summarized information in a tone that is sensitive to the user's emotions. This audio is provided simultaneously with the infographic to assist the user in understanding the information.
[0513] Step 8:
[0514] Users understand the information by viewing visual materials and listening to audio summaries. Users can ask additional questions as needed, allowing the interaction to continue.
[0515] (Example 2)
[0516] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0517] In voice dialogue systems, there is a need to provide optimal information based on the user's emotions and situation. However, conventional systems have not adequately achieved flexible information delivery that takes user emotions into consideration. Therefore, technology is needed to analyze user emotions and provide information accordingly.
[0518] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0519] In this invention, the server includes means for analyzing the user's emotions, means for adjusting the information presentation method according to the user's emotional state, and means for generating video material and outputting it to a display device. This enables flexible information presentation that responds to the user's emotions.
[0520] "Voice input" refers to a method for recognizing and digitally processing voice signals from a user.
[0521] "Textual information" refers to data obtained by converting voice input into text format.
[0522] "Analysis" is the process of semantically understanding textual information and identifying related information.
[0523] "Visual materials" are digital content that visually represents acquired information.
[0524] A "display device" is a device that provides visual information to a user.
[0525] "Providing information via audio" refers to the process of conveying textual information and video materials to users through speech synthesis.
[0526] "Analyzing user emotions" refers to technology that identifies emotional states from voice input.
[0527] "Adjusting the method of information presentation" is the process of optimizing how information is delivered based on the user's emotional state.
[0528] This invention provides specific means for realizing flexible information presentation in a voice dialogue system that responds to the user's emotional state. This system can adjust the information presentation method by receiving voice input and analyzing the user's emotions.
[0529] When a user speaks a question or request into the device, the device uses its microphone to capture the audio as a digital signal. This captured audio data is then converted into text using Google Cloud Speech-to-Text or similar speech recognition technology.
[0530] The terminal then sends this converted text information to the server. The server uses natural language processing techniques to analyze the received text information. This process utilizes OpenAI's language model, a generative AI model, and other similar algorithms. Through this analysis, the server understands the user's intent and retrieves relevant information from databases and external sources.
[0531] The server then uses the voice input data to perform sentiment analysis. This involves techniques that analyze parameters such as voice tone, speed, and volume. Based on the sentiment analysis results, the server adjusts how information is presented. For example, if the user is feeling anxious, visual materials are designed to be simple and easy to understand, and the voice output is adjusted to a calm and reassuring tone.
[0532] In generating visual materials, the server uses tools to visually organize information. Visualization tools such as Tableau and Infogram are utilized to create video materials based on the acquired information. These materials are sent to the terminal, which displays them on its screen. The display includes infographics and statistical data, making the information easier for the user to understand visually.
[0533] As a concrete example, consider a scenario where a user speaks into the terminal saying, "Please tell me about the economic situation." This prompt is recognized as voice input and converted into text. If the emotion engine detects tension in the user's voice, the server uses this information to generate a simple infographic and create a calm voice summary. In this way, optimal information presentation tailored to the user's emotions is achieved.
[0534] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0535] Step 1:
[0536] The user makes a question or request to the terminal using their voice. The terminal uses a microphone to acquire this voice input and transmits it to the computer as digital audio data. The input is the user's voice signal, and the output is the digital data of that voice. This digital audio data is used as input for the next processing step.
[0537] Step 2:
[0538] The device converts the acquired digital audio data into text information using speech recognition software. For example, Google Cloud Speech-to-Text is used. The input is the digital audio data generated in step 1, and the output is the corresponding text data. Through this process, the user's voice input is converted into text information.
[0539] Step 3:
[0540] The terminal sends the converted character information to the server. The server analyzes the received character data using a generative AI model. Natural language processing techniques are used for the analysis, with OpenAI's language model being one example. The input is text data, and the output of the analysis is detailed information about the user's intent and the content of the inquiry. Instructions are generated to retrieve this information from a database or external sources.
[0541] Step 4:
[0542] The server performs sentiment analysis based on audio data and text information. It uses a sentiment analysis engine to evaluate the user's emotional state. Inputs are audio parameters (e.g., tone, speed) and text data, and output is the user's emotional state (e.g., tension, relief, etc.). Based on this output, instructions are generated to adjust how information is presented.
[0543] Step 5:
[0544] The server generates visual materials based on the acquired information and the results of sentiment analysis. Visualization tools (e.g., Tableau, Infogram) are used to generate these visual materials. The input consists of information acquired from the database and the results of sentiment analysis, and the output is a video material displayed to the user. This material is optimized according to the user's emotional state.
[0545] Step 6:
[0546] The server sends the generated visual materials and audio summaries to the terminal. The terminal presents this to the user on its display and, if necessary, uses speech synthesis technology to present the summaries aloud. The input is the visual materials and summary data from the server, and the output is the display on the user's screen and the audio output. This process allows the user to receive information in a way that is relevant to their emotions.
[0547] (Application Example 2)
[0548] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0549] Conventional voice dialogue systems fail to adequately consider individual emotional states when providing information to users, resulting in uniform information presentation and leading to inconsistencies in user understanding and acceptance. Furthermore, the lack of mechanisms to recognize user emotions in real time and provide information accordingly prevents the provision of a better customer experience.
[0550] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0551] In this invention, the server includes means for recognizing the user's emotions from voice data, means for adjusting the information presentation method based on the recognized emotions, and means for optimizing the layout of visual materials according to the emotional state. This enables information presentation tailored to each user's emotions, resulting in improved comprehension and a more satisfying conversational experience.
[0552] "Means for receiving voice input" refers to a device or technology that allows a system to receive voice signals emitted by a user.
[0553] "Means of converting to text data" refers to a device or technology for converting received audio into a digital written format.
[0554] "Means for analysis and retrieving relevant information" refers to devices or technologies for analyzing converted text data and identifying related information from databases or external sources.
[0555] "Means for generating visual materials" refers to devices or technologies that convert retrieved information into a visual format so that it can be presented visually.
[0556] "Means for outputting to a display device" refers to a device or technology for outputting generated visual material to a display medium such as a screen.
[0557] "Means of providing summaries in audio format" refers to a device or technology for concisely summarizing the content of visual materials and playing it back as audio.
[0558] "Means for recognizing a user's emotions from voice data" refers to a device or technology for analyzing the characteristics of voice to determine the emotional state of the speaker.
[0559] "Means for adjusting the method of information presentation" refers to devices or technologies for optimizing the method of information transmission and tone of presentation based on perceived emotions.
[0560] "Means for optimizing according to emotional state" refers to devices or technologies for adjusting the visual layout and information structure to match the user's emotions.
[0561] This system primarily consists of a user terminal, a server, and a visual display device. The terminal has an interface for receiving user voice input and uses speech recognition software to convert the input voice into text data. For example, a speech recognition API can be used for this process.
[0562] The server receives the converted text data and analyzes it using natural language processing techniques. This analysis system, for example, uses a natural language processing engine. Based on the analysis results, relevant information is retrieved from databases and external sources. The retrieved information is then converted into a visual format by a visual material generation module.
[0563] A key feature is that the server has a built-in emotion recognition engine. This engine analyzes the characteristics of the voice to determine the user's emotional state. Based on the results of the emotion recognition, the server adjusts the information presentation method and the layout of visual materials. For example, if the server determines that the user is nervous, the information will be presented in a calm tone, and the visual materials will be summarized concisely.
[0564] Finally, the generated visual materials and adjusted information are sent to the terminal and displayed on the screen. Users can view the visual materials through the screen and listen to an audio summary if needed. This audio summary is provided by playing the generated audio file.
[0565] As a concrete example, consider a scenario where a user asks, "Tell me about this week's popular products." If the emotion engine determines the user is excited during the process of the device converting speech to text and the server analyzing this information, the product information will be explained in a calm voice and presented on the screen in a visually easy-to-understand format. This allows the user to comfortably receive the information they are looking for.
[0566] Examples of prompts to input into a generative AI model include: "Please explain in detail how to provide product information displayed on a screen based on emotions."
[0567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0568] Step 1:
[0569] The terminal receives voice input from the user. This voice input is the data that forms the basis for subsequent processing. The terminal uses speech recognition software to convert this voice data into text data. Here, the input is voice data and the output is string data. This conversion makes the voice into a format that can be analyzed.
[0570] Step 2:
[0571] The server analyzes the text data received from the terminal. This process uses natural language processing techniques to extract the user's intent from the text data and retrieve relevant information. The input is text data, and the output is a list of information. This information includes relevant data retrieved from a database.
[0572] Step 3:
[0573] The server uses an emotion recognition engine to recognize the user's emotions from the audio data. The input is the original audio data, and the output is data indicating the emotional state. This recognition result is used to determine the adjustment policy for information presentation.
[0574] Step 4:
[0575] The server uses a visual material generation module to create visually clear materials based on the acquired information and emotion recognition results. Inputs are information lists and emotion state data, and output is visual materials. The layout is adjusted according to the emotion state to provide information in the most optimal format for the user.
[0576] Step 5:
[0577] The server sends the generated visual materials and adjusted information to the terminal. The terminal outputs this material to a display device, making it available for the user to review. The input is the visual materials, and the output is the displayed material. In addition, if necessary, an audio summary of the information is provided as supplementary material.
[0578] Through the steps described above, the system can present users with information tailored to their individual emotional state.
[0579] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0580] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0581] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0582] [Fourth Embodiment]
[0583] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0584] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0585] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0586] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0587] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0588] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0589] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0590] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0591] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0592] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0593] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0594] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0595] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0596] The system of this invention functions by combining multiple means to streamline information transmission in voice dialogue. When a user inputs a question by voice into the interface, it is converted into text data using speech recognition technology. At this stage, the terminal has the ability to receive the voice and transcribe it in real time.
[0597] Subsequently, the server retrieves the text data and performs analysis using natural language processing. In this analysis step, the server efficiently searches for relevant information that reflects the intent and theme of the user's question. Once relevant data is retrieved from the databases and external information sources used here, the foundational information for the visual materials is established.
[0598] Next, the server generates visual materials. Visual materials are visualizations, such as infographics, based on the acquired information, presenting the information in a format that is intuitively easy for the user to understand. This generation process includes optimizing the layout of the visual materials so that they are easy for the user to read and the key points can be clearly recognized.
[0599] Once the visual material is complete, the server sends it to the terminal, where it is displayed. The displayed visual material incorporates diagrams, graphs, text, and other elements to make the information easier to understand visually. Furthermore, if necessary, synthesized speech is used to provide audio summaries of the key points of the visual material, further supporting information comprehension.
[0600] As a concrete example, consider a scenario where a user asks, "Please tell me how to use solar energy." The terminal recognizes the speech and converts it into text data. The server analyzes this data and searches for information on how to use solar energy. It then generates relevant visual materials and sends them to the terminal. The terminal displays these materials on its screen and also provides an audio explanation of the main uses. The user can quickly understand the information using both sight and hearing. In this way, the system effectively supports the user's understanding of information.
[0601] The following describes the processing flow.
[0602] Step 1:
[0603] The user inputs questions by voice using a microphone. The user asks the system questions by voice about a topic for which they want to obtain specific information.
[0604] Step 2:
[0605] The device uses speech recognition technology to convert the input speech into text data. The device then sends this converted text data to the next processing step.
[0606] Step 3:
[0607] The server analyzes the text data and uses natural language processing to identify the user's intent and related keywords. Based on the question, the server prepares to search for the most relevant data.
[0608] Step 4:
[0609] The server searches for relevant information from databases and external sources based on the analysis results. The server quickly retrieves accurate data in response to the user's questions.
[0610] Step 5:
[0611] Based on the information acquired by the server, visual materials are generated. The server creates infographics in a format that is easily understandable to the user and optimizes the layout of the visual materials.
[0612] Step 6:
[0613] The server sends the generated visual material to the terminal. The terminal prepares to present the received visual material to the user.
[0614] Step 7:
[0615] The device displays visual materials on its screen and provides users with summarized information using synthesized speech. The device is designed to facilitate visual and auditory comprehension of information.
[0616] Step 8:
[0617] Users can effectively understand information by reviewing visual materials and listening to audio summaries. Users can follow up with additional questions as needed.
[0618] (Example 1)
[0619] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0620] In systems that efficiently provide information through voice interaction, users need to quickly and intuitively understand large amounts of information. However, in conventional systems, information obtained from voice input is provided in text format, making it difficult to effectively convey information using both sight and hearing. A new method of information transmission is needed to solve this problem and facilitate user understanding.
[0621] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0622] In this invention, the server includes means for analyzing converted text data and searching for relevant information using natural language processing, means for a generation engine for generating visual materials using the relevant information, and means for visualizing the information in the generated visual materials. As a result, the user can receive not only audio but also visually organized information simultaneously, enabling quick and intuitive understanding of the information.
[0623] "Means for receiving voice input" refers to a device or technology that provides a function for taking in external voice information into a system.
[0624] "Methods for converting speech to text data" refers to technologies that perform a process of analyzing speech signals and converting them into corresponding text data.
[0625] "Means for analyzing text data and searching for related information using natural language processing" refers to technologies that analyze text data, understand the user's intent using natural language processing techniques, and have the functionality to search for related information from databases and external sources.
[0626] "Means equipped with a generation engine for generating visual materials" refers to a function equipped with an engine for generating visual materials based on acquired information, and is a technology that visualizes information as infographics or charts.
[0627] "Means of visualizing information in visual materials" refers to techniques for organizing generated materials and presenting them in a visually easy-to-understand format.
[0628] "A means of outputting visual materials to a display device and providing their key points in synthesized speech" refers to a technology that displays generated visual materials on a display device for user viewing and simultaneously explains the key points of the materials in synthesized speech.
[0629] This system combines multiple technologies to acquire information via voice and effectively convey it to the user. When a user provides voice input, the terminal first receives the voice. The terminal uses a microphone to convert the voice into digital data, performs speech recognition using that data, and then converts the voice back into text data. At this stage, speech recognition software such as the Google Speech-to-Text API is utilized.
[0630] Next, the server retrieves this text data. The server uses natural language processing techniques to analyze the text data. Based on the analyzed information, the server connects to internal databases and external information sources to retrieve relevant data. This process uses common data analysis tools and search algorithms.
[0631] The information obtained is then generated as visual materials through a generation engine on the server. Visual materials such as infographics and charts are optimized in layout so that the information is presented in a format that is intuitively easy for the user to understand. Visualization tools such as Infogram and Tableau are used during this generation process.
[0632] Finally, the generated visual materials are sent to the terminal. The terminal outputs them to a display device, allowing the user to view the materials. In addition, a speech synthesis engine installed in the terminal provides an audio summary of the visual materials.
[0633] As a concrete example, consider a scenario where a user asks a question via voice, "Please tell me how to use solar energy." In this case, the terminal converts the voice into text, the server collects and analyzes information about solar energy, generates visual materials, and sends them to the user's terminal. Through this entire process, the user can quickly understand the information through both voice and visual means.
[0634] An example of a prompt might be, "Generate a visual explanation of how to utilize solar energy." Such a prompt prompt would prompt the server to begin the process of visualizing the relevant information.
[0635] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0636] Step 1:
[0637] The user provides voice input through the interface. The input voice is captured by the terminal and converted into digital data via the microphone. The terminal then passes this digital voice data to a speech recognition engine, which generates text data in real time. During this process, voice signal processing is performed to reduce background noise and clarify speech.
[0638] Step 2:
[0639] The device sends text data to the server. Upon receiving the text data, the server analyzes it using natural language processing (NLP) techniques. Specifically, it extracts relevant keywords by performing data processing such as text tokenization, morphological analysis, and language model application to understand the intent of the user's question.
[0640] Step 3:
[0641] The server searches for relevant information based on the analysis results. It accesses databases and external information sources to retrieve data related to the keywords. Using information retrieval algorithms, it quickly selects the most relevant information. The retrieved information is then formatted into a dataset necessary for generating visual materials in the next step.
[0642] Step 4:
[0643] The server takes a dataset as input and generates visual materials using a generative AI model. This process involves visualizations such as charts, graphs, and infographics, organizing the data into a user-friendly format. Visual design optimization ensures that the generated materials are visually appealing.
[0644] Step 5:
[0645] The generated visual materials are sent from the server to the terminal. The terminal receives them and outputs them to the display device. Furthermore, the terminal's built-in speech synthesis function is used to provide the user with audio summaries of the key features and points of the visual materials, thereby supporting a deeper understanding of the information. In this state, the user can obtain information through both audio and visual means.
[0646] (Application Example 1)
[0647] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0648] Traditional methods for customers to efficiently obtain information in physical stores are time-consuming, making it difficult to quickly acquire appropriate information. In particular, requests for detailed product usage instructions and characteristics require the assistance of store staff, leading to resource depletion and wasted time. There is a need to address these challenges and streamline the provision of information within stores.
[0649] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0650] In this invention, the server includes a medium for receiving voice input, a medium for converting the received voice into text information, a medium for analyzing the converted text information and retrieving related information, and a medium for presenting the generated visual material to a portable visual aid device. This enables customers to ask questions by voice within the store and obtain information that is quick and easy to understand visually.
[0651] "Voice input" is the process of recognizing the voice signals spoken by a user and receiving them as digital data.
[0652] A "medium" is a means of processing, converting, displaying, or transmitting data such as audio signals, visual materials, and textual information.
[0653] "Textual information" refers to text data that is obtained by analyzing and converting audio signals.
[0654] "Analysis" is the process of deciphering input text information and understanding its meaning and intent.
[0655] "Relevant information" refers to the appropriate data and knowledge that are retrieved based on the user's questions or requests.
[0656] "Visual materials" refer to graphics, diagrams, text, and other materials created to present information in a way that is easy to understand visually.
[0657] A "display medium" refers to a device or part of a device used to present visual materials, allowing users to visually confirm information.
[0658] A "portable visual aid" is a device that a user can carry and use, and that has the ability to display visual information.
[0659] The system of this invention mainly consists of a server, a terminal, and a user. The user requests information through voice input, and that voice is converted into text information by the terminal. The terminal has speech recognition technology such as Google Cloud Speech-to-Text API installed, and has the ability to quickly convert voice data into text.
[0660] The server then retrieves the text information and performs analysis using natural language processing. OpenAI's generative AI model and the BERT model are used for the analysis, ensuring that the intent of the user's question is accurately interpreted. Based on the analysis results, relevant information is retrieved from databases and external sources, such as Google Firebase and MongoDB, and the server uses this information to generate visual materials.
[0661] The visual materials are generated using D3.js or Tableau and are characterized by their user-friendly and interactive format. The server sends the generated visual materials to the terminal, which then presents them on the display of a portable visual aid, such as smart glasses, to provide the user with visual information.
[0662] Furthermore, to complement visual information, the server can use speech synthesis technologies such as Amazon Polly to provide the key points of the visual materials as audio.
[0663] As a concrete example, if a user asks "How do I use this product?" in a physical store, the terminal recognizes the voice, converts it into text, and sends it to the server. The server analyzes this information and searches its database for knowledge related to how to use the product. An infographic generated using D3.js is displayed on smart glasses, and an audio explanation is provided using Amazon Polly.
[0664] An example of a prompt to input into the generating AI model is: "The customer wants to know more details about the specifications of a specific product. Analyze this audio data to generate visual information and display it on the assistive device."
[0665] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0666] Step 1:
[0667] The user provides voice input. The user speaks into the microphone of their smart glasses or device, asking for information. The input voice signal is sent to the device.
[0668] Step 2:
[0669] The device converts the audio signal into text information. Speech recognition software running on the device (e.g., Google Cloud Speech-to-Text API) outputs the audio signal as text in real time. Here, the input is an audio signal, and the output is text data.
[0670] Step 3:
[0671] The server acquires and analyzes text data. The server uses natural language processing (e.g., OpenAI GPT) to analyze the acquired text, understand the user's intent, and search for relevant information. The input is text data, and the output is relevant information.
[0672] Step 4:
[0673] The server generates visual materials based on relevant information. Visual material generation software (e.g., D3.js) creates an infographic based on the retrieved information. The input is relevant information, and the output is a visual material. In this step, the information is structured and visualized.
[0674] Step 5:
[0675] The server sends visual materials to the terminal for display. The terminal displays the received visual materials on the display of a portable visual aid (e.g., smart glasses). The input is the visual materials, and the output is the visual information on the display.
[0676] Step 6:
[0677] The server generates an audio summary of the visual information. Speech synthesis technology (e.g., Amazon Polly) converts the key points of the visual material into synthesized speech. The input is the key points of the visual material, and the output is synthesized speech.
[0678] Step 7:
[0679] The device plays synthesized speech to provide information to the user. The audio is played through the device's speaker, allowing the user to obtain information supplementarily through hearing. The input is synthesized speech data, and the output is audio playback.
[0680] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0681] This invention aims to optimize the method of information presentation by recognizing user emotions through the integration of an emotion engine into a voice dialogue system. The system has the function of receiving user voice input via a terminal and converting it into text data through speech recognition technology. The converted text data is sent to a server and analyzed using natural language processing. Based on the analysis results, relevant information is retrieved from databases and external information sources.
[0682] However, a key feature of this invention is that it incorporates an emotion engine into the system that analyzes the user's emotions. The server recognizes the user's emotions from the voice data and adjusts the information presentation method based on the results. For example, if the system recognizes that the user is feeling anxious, it can provide information in a calm tone or organize visual materials concisely.
[0683] In generating visual materials, the server visually assembles the acquired information and optimizes it according to the user's emotional state. This material is then sent to the terminal and displayed to the user. The terminal displays the infographic on its screen and provides audio summaries as needed to aid user understanding.
[0684] As a concrete example, consider a scenario where a user asks, "Please tell me about the economic situation." The terminal converts the speech to text, and the emotion engine detects tension in the user's voice. As a result, the server provides a reassuring tone for the speech summary and presents visual materials more clearly and simply. In this way, it is possible to provide a conversational experience that responds to the user's emotions and improves the ease with which information can be received. This system utilizes both visual and audio information to provide flexible information tailored to the user's needs.
[0685] The following describes the processing flow.
[0686] Step 1:
[0687] The user inputs their question by voice using a microphone. For example, the user might ask the system for the latest weather information.
[0688] Step 2:
[0689] The device receives voice input and converts it into text data using speech recognition technology. The converted text is then sent to a server for analysis.
[0690] Step 3:
[0691] The server analyzes text data and user voice data, and evaluates the user's emotional state through an emotion engine. This information is used to ensure that users can receive information with confidence.
[0692] Step 4:
[0693] The server searches the database for relevant information based on the analysis results. The server then prepares to retrieve the most relevant information and provide it to the user.
[0694] Step 5:
[0695] The server generates visual materials based on the user's emotion recognition results. The layout of the visual materials is optimized according to the user's emotional state and designed to provide a sense of security and ease of understanding.
[0696] Step 6:
[0697] The generated visual materials are sent from the server to the terminal, which then displays the materials on its screen. Because the information is visually organized, users can easily understand it through their eyes.
[0698] Step 7:
[0699] Furthermore, the device uses synthesized speech to provide summarized information in a tone that is sensitive to the user's emotions. This audio is provided simultaneously with the infographic to assist the user in understanding the information.
[0700] Step 8:
[0701] Users understand the information by viewing visual materials and listening to audio summaries. Users can ask additional questions as needed, allowing the interaction to continue.
[0702] (Example 2)
[0703] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0704] In voice dialogue systems, there is a need to provide optimal information based on the user's emotions and situation. However, conventional systems have not adequately achieved flexible information delivery that takes user emotions into consideration. Therefore, technology is needed to analyze user emotions and provide information accordingly.
[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0706] In this invention, the server includes means for analyzing the user's emotions, means for adjusting the information presentation method according to the user's emotional state, and means for generating video material and outputting it to a display device. This enables flexible information presentation that responds to the user's emotions.
[0707] "Voice input" refers to a method for recognizing and digitally processing voice signals from a user.
[0708] "Textual information" refers to data obtained by converting voice input into text format.
[0709] "Analysis" is the process of semantically understanding textual information and identifying related information.
[0710] "Visual materials" are digital content that visually represents acquired information.
[0711] A "display device" is a device that provides visual information to a user.
[0712] "Providing information via audio" refers to the process of conveying textual information and video materials to users through speech synthesis.
[0713] "Analyzing user emotions" refers to technology that identifies emotional states from voice input.
[0714] "Adjusting the method of information presentation" is the process of optimizing how information is delivered based on the user's emotional state.
[0715] This invention provides specific means for realizing flexible information presentation in a voice dialogue system that responds to the user's emotional state. This system can adjust the information presentation method by receiving voice input and analyzing the user's emotions.
[0716] When a user speaks a question or request into the device, the device uses its microphone to capture the audio as a digital signal. This captured audio data is then converted into text using Google Cloud Speech-to-Text or similar speech recognition technology.
[0717] The terminal then sends this converted text information to the server. The server uses natural language processing techniques to analyze the received text information. This process utilizes OpenAI's language model, a generative AI model, and other similar algorithms. Through this analysis, the server understands the user's intent and retrieves relevant information from databases and external sources.
[0718] The server then uses the voice input data to perform sentiment analysis. This involves techniques that analyze parameters such as voice tone, speed, and volume. Based on the sentiment analysis results, the server adjusts how information is presented. For example, if the user is feeling anxious, visual materials are designed to be simple and easy to understand, and the voice output is adjusted to a calm and reassuring tone.
[0719] In generating visual materials, the server uses tools to visually organize information. Visualization tools such as Tableau and Infogram are utilized to create video materials based on the acquired information. These materials are sent to the terminal, which displays them on its screen. The display includes infographics and statistical data, making the information easier for the user to understand visually.
[0720] As a concrete example, consider a scenario where a user speaks into the terminal saying, "Please tell me about the economic situation." This prompt is recognized as voice input and converted into text. If the emotion engine detects tension in the user's voice, the server uses this information to generate a simple infographic and create a calm voice summary. In this way, optimal information presentation tailored to the user's emotions is achieved.
[0721] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0722] Step 1:
[0723] The user makes a question or request to the terminal using their voice. The terminal uses a microphone to acquire this voice input and transmits it to the computer as digital audio data. The input is the user's voice signal, and the output is the digital data of that voice. This digital audio data is used as input for the next processing step.
[0724] Step 2:
[0725] The device converts the acquired digital audio data into text information using speech recognition software. For example, Google Cloud Speech-to-Text is used. The input is the digital audio data generated in step 1, and the output is the corresponding text data. Through this process, the user's voice input is converted into text information.
[0726] Step 3:
[0727] The terminal sends the converted character information to the server. The server analyzes the received character data using a generative AI model. Natural language processing techniques are used for the analysis, with OpenAI's language model being one example. The input is text data, and the output of the analysis is detailed information about the user's intent and the content of the inquiry. Instructions are generated to retrieve this information from a database or external sources.
[0728] Step 4:
[0729] The server performs sentiment analysis based on audio data and text information. It uses a sentiment analysis engine to evaluate the user's emotional state. Inputs are audio parameters (e.g., tone, speed) and text data, and output is the user's emotional state (e.g., tension, relief, etc.). Based on this output, instructions are generated to adjust how information is presented.
[0730] Step 5:
[0731] The server generates visual materials based on the acquired information and the results of sentiment analysis. Visualization tools (e.g., Tableau, Infogram) are used to generate these visual materials. The input consists of information acquired from the database and the results of sentiment analysis, and the output is a video material displayed to the user. This material is optimized according to the user's emotional state.
[0732] Step 6:
[0733] The server sends the generated visual materials and audio summaries to the terminal. The terminal presents this to the user on its display and, if necessary, uses speech synthesis technology to present the summaries aloud. The input is the visual materials and summary data from the server, and the output is the display on the user's screen and the audio output. This process allows the user to receive information in a way that is relevant to their emotions.
[0734] (Application Example 2)
[0735] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0736] Conventional voice dialogue systems fail to adequately consider individual emotional states when providing information to users, resulting in uniform information presentation and leading to inconsistencies in user understanding and acceptance. Furthermore, the lack of mechanisms to recognize user emotions in real time and provide information accordingly prevents the provision of a better customer experience.
[0737] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0738] In this invention, the server includes means for recognizing the user's emotions from voice data, means for adjusting the information presentation method based on the recognized emotions, and means for optimizing the layout of visual materials according to the emotional state. This enables information presentation tailored to each user's emotions, resulting in improved comprehension and a more satisfying conversational experience.
[0739] "Means for receiving voice input" refers to a device or technology that allows a system to receive voice signals emitted by a user.
[0740] "Means of converting to text data" refers to a device or technology for converting received audio into a digital written format.
[0741] "Means for analysis and retrieving relevant information" refers to devices or technologies for analyzing converted text data and identifying related information from databases or external sources.
[0742] "Means for generating visual materials" refers to devices or technologies that convert retrieved information into a visual format so that it can be presented visually.
[0743] "Means for outputting to a display device" refers to a device or technology for outputting generated visual material to a display medium such as a screen.
[0744] "Means of providing summaries in audio format" refers to a device or technology for concisely summarizing the content of visual materials and playing it back as audio.
[0745] "Means for recognizing a user's emotions from voice data" refers to a device or technology for analyzing the characteristics of voice to determine the emotional state of the speaker.
[0746] "Means for adjusting the method of information presentation" refers to devices or technologies for optimizing the method of information transmission and tone of presentation based on perceived emotions.
[0747] "Means for optimizing according to emotional state" refers to devices or technologies for adjusting the visual layout and information structure to match the user's emotions.
[0748] This system primarily consists of a user terminal, a server, and a visual display device. The terminal has an interface for receiving user voice input and uses speech recognition software to convert the input voice into text data. For example, a speech recognition API can be used for this process.
[0749] The server receives the converted text data and analyzes it using natural language processing techniques. This analysis system, for example, uses a natural language processing engine. Based on the analysis results, relevant information is retrieved from databases and external sources. The retrieved information is then converted into a visual format by a visual material generation module.
[0750] A key feature is that the server has a built-in emotion recognition engine. This engine analyzes the characteristics of the voice to determine the user's emotional state. Based on the results of the emotion recognition, the server adjusts the information presentation method and the layout of visual materials. For example, if the server determines that the user is nervous, the information will be presented in a calm tone, and the visual materials will be summarized concisely.
[0751] Finally, the generated visual materials and adjusted information are sent to the terminal and displayed on the screen. Users can view the visual materials through the screen and listen to an audio summary if needed. This audio summary is provided by playing the generated audio file.
[0752] As a concrete example, consider a scenario where a user asks, "Tell me about this week's popular products." If the emotion engine determines the user is excited during the process of the device converting speech to text and the server analyzing this information, the product information will be explained in a calm voice and presented on the screen in a visually easy-to-understand format. This allows the user to comfortably receive the information they are looking for.
[0753] Examples of prompts to input into a generative AI model include: "Please explain in detail how to provide product information displayed on a screen based on emotions."
[0754] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0755] Step 1:
[0756] The terminal receives voice input from the user. This voice input is the data that forms the basis for subsequent processing. The terminal uses speech recognition software to convert this voice data into text data. Here, the input is voice data and the output is string data. This conversion makes the voice into a format that can be analyzed.
[0757] Step 2:
[0758] The server analyzes the text data received from the terminal. This process uses natural language processing techniques to extract the user's intent from the text data and retrieve relevant information. The input is text data, and the output is a list of information. This information includes relevant data retrieved from a database.
[0759] Step 3:
[0760] The server uses an emotion recognition engine to recognize the user's emotions from the audio data. The input is the original audio data, and the output is data indicating the emotional state. This recognition result is used to determine the adjustment policy for information presentation.
[0761] Step 4:
[0762] The server uses a visual material generation module to create visually clear materials based on the acquired information and emotion recognition results. Inputs are information lists and emotion state data, and output is visual materials. The layout is adjusted according to the emotion state to provide information in the most optimal format for the user.
[0763] Step 5:
[0764] The server sends the generated visual materials and adjusted information to the terminal. The terminal outputs this material to a display device, making it available for the user to review. The input is the visual materials, and the output is the displayed material. In addition, if necessary, an audio summary of the information is provided as supplementary material.
[0765] Through the steps described above, the system can present users with information tailored to their individual emotional state.
[0766] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0767] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0768] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0769] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0770] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0771] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0772] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0773] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0774] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0775] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0776] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0777] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0778] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0779] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0780] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0781] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0782] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0783] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0784] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0785] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0786] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0787] The following is further disclosed regarding the embodiments described above.
[0788] (Claim 1)
[0789] A means of accepting voice input,
[0790] A means of converting received audio into text data,
[0791] A means of analyzing the converted text data and searching for related information,
[0792] A means for generating visual materials based on retrieved information,
[0793] A means for outputting the generated visual material to a display device,
[0794] A means of providing summaries of visual materials in audio format,
[0795] A system that includes this.
[0796] (Claim 2)
[0797] The system according to claim 1, comprising means for automatically optimizing the layout of visual materials.
[0798] (Claim 3)
[0799] The system according to claim 1, comprising means for simultaneously providing audio information related to visual material in a user interface.
[0800] "Example 1"
[0801] (Claim 1)
[0802] A means of accepting voice input,
[0803] A means of converting received audio into text data,
[0804] A means for analyzing converted text data and searching for related information using natural language processing,
[0805] A means comprising a generation engine for generating visual materials using related information,
[0806] Means for visualizing information in generated visual materials,
[0807] A means of outputting visual materials to a display device and providing their key points in synthesized speech,
[0808] A system that includes this.
[0809] (Claim 2)
[0810] The system according to claim 1, comprising means for automatically optimizing the layout of visual materials.
[0811] (Claim 3)
[0812] The system according to claim 1, comprising means for simultaneously providing audio information related to visual materials in the user interface of a user terminal.
[0813] "Application Example 1"
[0814] (Claim 1)
[0815] A medium that accepts voice input,
[0816] A medium that converts received audio into text information,
[0817] A medium that analyzes converted text information and searches for related information,
[0818] A medium that generates visual materials based on searched information,
[0819] A medium that outputs the generated visual materials to a display medium,
[0820] A medium that provides summaries of visual materials in audio format,
[0821] A medium for presenting visual materials to a portable visual aid,
[0822] A system that includes this.
[0823] (Claim 2)
[0824] The system according to claim 1, comprising a medium for automatically optimizing the arrangement of visual materials.
[0825] (Claim 3)
[0826] The system according to claim 1, comprising a medium that simultaneously provides audio information related to visual material at the user interface.
[0827] "Example 2 of combining an emotion engine"
[0828] (Claim 1)
[0829] A means of accepting voice input,
[0830] A means of converting received audio into text information,
[0831] A means for analyzing the converted character information and obtaining related information,
[0832] A means for generating video materials based on acquired information,
[0833] A means for outputting the generated video material to a display device,
[0834] A means of providing an audio summary of video materials,
[0835] A means of analyzing user emotions,
[0836] A means of adjusting the way information is presented according to the user's emotional state,
[0837] A system that includes this.
[0838] (Claim 2)
[0839] The system according to claim 1, comprising means for automatically optimizing the layout of video materials.
[0840] (Claim 3)
[0841] The system according to claim 1, comprising means for simultaneously providing audio information related to video material in a user interface.
[0842] "Application example 2 of combining emotional engines"
[0843] (Claim 1)
[0844] A means of accepting voice input,
[0845] A means of converting received audio into text data,
[0846] A means of analyzing the converted text data and searching for related information,
[0847] A means for generating visual materials based on retrieved information,
[0848] A means for outputting the generated visual material to a display device,
[0849] A means of providing summaries of visual materials in audio format,
[0850] A means of recognizing the user's emotions from voice data,
[0851] Means for adjusting information presentation methods based on perceived emotions,
[0852] A system that includes this.
[0853] (Claim 2)
[0854] The system according to claim 1, comprising means for optimizing the layout of visual materials according to emotional state.
[0855] (Claim 3)
[0856] The system according to claim 1, comprising means for simultaneously providing audio information related to visual materials in an emotionally appropriate tone within a user interface. [Explanation of Symbols]
[0857] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of accepting voice input, A means of converting received audio into text data, A means of analyzing the converted text data and searching for related information, A means for generating visual materials based on retrieved information, A means for outputting the generated visual material to a display device, A means of providing summaries of visual materials in audio format, A system that includes this.
2. The system according to claim 1, comprising means for automatically optimizing the layout of visual materials.
3. The system according to claim 1, comprising means for simultaneously providing audio information related to visual material in a user interface.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A