system

The system addresses integration of diverse information forms by processing voice input into text, analyzing user intent, and generating emotionally sensitive responses, enhancing accessibility and user experience for visually and hearing-impaired individuals.

JP2026070963APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional technologies fail to integrate the processing of diverse information forms such as voice, text, and image, leading to limitations in accessing information for visually and hearing-impaired individuals, and face challenges in voice recognition accuracy, intent analysis, and natural response generation.

Method used

A system comprising an audio input means, conversion means, analysis means, information acquisition means, and presentation means to process voice input into text, analyze user intent, acquire relevant information, and generate appropriate responses, incorporating emotion recognition to tailor responses to user emotions.

Benefits of technology

Enables efficient, accurate, and emotionally sensitive information retrieval and response generation, improving accessibility for visually and hearing-impaired users and enhancing user experience through natural and timely information delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070963000001_ABST
    Figure 2026070963000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Audio input means for acquiring audio signals, A conversion means for converting the aforementioned audio signal into a text representation, An analysis means for analyzing the aforementioned text expression to identify user intent, Information acquisition means for acquiring relevant information based on the user intent, A response generation means that generates a response based on the acquired information, A system including means for presenting the aforementioned response to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern information society, it is required to integrally utilize information in different forms such as voice, text, and image, and provide information quickly and appropriately based on the intention of the user. However, in the conventional technology, means for consistently processing these information forms are not sufficiently provided, and there is a problem that there are restrictions on information access for visually and hearing-impaired persons in particular.

Means for Solving the Problems

[0005] The present invention provides information that comprehensively processes the format of information and responds to the diverse needs of users by comprising: an audio input means for acquiring an audio signal; a conversion means for converting the audio signal into a text representation; an analysis means for analyzing the text representation to identify user intent; an information acquisition means for acquiring relevant information based on the user intent; a response generation means for generating a response based on the acquired information; and a presentation means for presenting the response to the user.

[0006] "Voice input means" refers to a device or function for acquiring voice signals emitted by a user.

[0007] "Conversion means" refers to a device or function that converts an acquired audio signal into text format.

[0008] "Analysis means" refers to a device or function that analyzes text-based data to identify the user's intent.

[0009] "Information acquisition means" refers to a device or function that acquires necessary information from an external database or web service according to the user's intent.

[0010] "Response generation means" refers to a device or function that constructs a response to the user based on acquired information.

[0011] "Presentation means" refers to a device or function for displaying or playing back the generated response to the user. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] The system of the present invention appropriately acquires the user's voice input, efficiently converts it into text to analyze the user's intent, acquires necessary information from external sources based on that intent, and ultimately generates an appropriate response for the user.

[0034] Users access the system using devices such as smartphones and tablets. These devices are equipped with voice input capabilities, which are used to capture the user's voice as a digital signal. The captured voice signal is then sent from the device to the server via the network.

[0035] The server converts audio data into text using a speech recognition API. This speech recognition process utilizes advanced language models to analyze speech patterns and convert them into human-readable text. This converted text is then used for subsequent analysis.

[0036] The analysis engine utilizes natural language processing techniques to extract intent from the user's utterances. For example, if a user says, "Tell me the weather for tomorrow," the analysis engine identifies the need to search for weather information. Based on this information request, the server retrieves the necessary data from external databases and web services.

[0037] The acquired information is processed by a response generation module, which uses commonly used text generation and speech synthesis technologies to construct a response for the user. For example, if it's weather information, it will generate specific information such as, "Tomorrow in Tokyo it will be sunny, with a high of 25 degrees Celsius."

[0038] The final generated response is returned to the terminal. The terminal receives this response data and displays it as text on the screen or plays it back as audio through the speaker, allowing the user to intuitively obtain the requested information. This enables the user to quickly obtain information and achieve their goals.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The user activates the device's voice input function and speaks their question or request. The device captures the voice using the microphone and temporarily stores it as digital data.

[0042] Step 2:

[0043] The terminal converts the acquired audio signal into a format that is easy to process and sends this audio data to the server. The audio data is delivered via the network to the API endpoint specified by the server.

[0044] Step 3:

[0045] The server passes the received audio data to a speech recognition API, which converts the audio into text. This conversion process uses a language model to analyze the audio signal and generate corresponding text information.

[0046] Step 4:

[0047] The server inputs text data into a natural language processing engine to analyze the intent of the user's request. This analysis process uses techniques such as keyword extraction and sentiment analysis to identify the information the user is looking for.

[0048] Step 5:

[0049] Based on the analysis results, the server collects the information the user needs from the internet and external databases. For example, it might call web services such as weather information APIs to obtain relevant information.

[0050] Step 6:

[0051] The server uses the collected information to generate a response to present to the user. The response generation module constructs the information in a natural way, utilizing text and speech synthesis.

[0052] Step 7:

[0053] The server sends the generated response to the terminal. The terminal presents this response data to the user using methods such as audio output or display. The user obtains the answer to their question through the presented information.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] In recent years, voice assistants and natural language interfaces have become widely used, but they suffer from problems such as low voice recognition accuracy, inability to accurately analyze intent from speech, and unnatural responses to users. These problems hinder the improvement of information retrieval and efficiency in daily life using voice interfaces. The present invention aims to address these problems and provide a voice interface that enables more natural, rapid, and accurate information retrieval and response generation.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes a conversion device for converting voice signals into text information, an analysis device for analyzing the text information and extracting intent, and an information acquisition device for obtaining relevant information based on the intent. This enables the user to efficiently acquire information via voice and obtain appropriate and natural responses.

[0059] An "audio signal" is data that represents sound in digital or analog format.

[0060] An "input device" is a device that receives voice information from a user and acquires it as a digital signal.

[0061] A "conversion device" is a device used to convert acquired audio signals into text information.

[0062] "Textual information" refers to data obtained by converting audio signals into a text format that can be analyzed.

[0063] An "analysis device" is a device that analyzes textual information and extracts the user's intended meaning.

[0064] "Intent" refers to a specific request for information or a directive that can be inferred from the user's statements.

[0065] An "information acquisition device" is a device that acquires relevant information from an external source based on the user's intent.

[0066] A "generation device" is a device that creates an appropriate response based on acquired information.

[0067] A "display device" is a device that visually presents the generated response to the user as text.

[0068] A "sound playback device" is a device that presents the generated response to the user audibly as sound.

[0069] This invention provides a system that offers an interface for users to acquire information via voice and receive appropriate responses. The user uses a device such as a smartphone or tablet to acquire voice signals using a voice input device. The device converts these voice signals into a digital format and transmits them to a server via a network.

[0070] The server converts the audio signal into text information using a speech recognition API (for example, using general speech recognition technology) to analyze the audio signal. This converted text information is then processed by an analysis device using natural language processing technology to extract the user's intent. The analysis device uses generative AI model technology (for example, a general natural language processing engine) to identify the user's intent.

[0071] Based on the identified user intent, the information acquisition device accesses external information sources and retrieves relevant information. These external information sources could include various information service providers and databases. The acquired information is processed by a generator on the server and converted into a natural language response to the user. Here, generative AI modeling technology is used to improve the naturalness and accuracy of the response.

[0072] Finally, the generated response is sent to the terminal, which then presents this response to the user. Using a display device or audio playback device, intuitive feedback can be provided to the user by displaying it as text or playing it aloud.

[0073] As a concrete example, a user might use voice input via their device, saying, "Tell me the weather for tomorrow." In this case, the server converts the voice signal into text, analyzes the intent of "Tell me the weather for tomorrow," and retrieves weather information by referring to an appropriate external database. The generated response would then be provided in the form of, "Tomorrow will be sunny, with a high of 25 degrees Celsius."

[0074] Examples of prompt phrases include "Voice input: Tell me the weather tomorrow" and "Voice input: Tell me the latest news."

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The user inputs an audio signal via the terminal's audio input device. The terminal converts this audio into a digital signal and transmits it to the server via the internet. The input is raw audio, and the output is encoded digital data.

[0078] Step 2:

[0079] The server passes the received digital audio signal to a speech recognition API, which converts the audio into text. In this process, the API analyzes the frequency pattern of the audio and generates the corresponding text. The input is a digital audio signal, and the output is text data.

[0080] Step 3:

[0081] The analysis device on the server analyzes the acquired text data using natural language processing techniques to extract the user's intent. Here, it understands keywords and context within the text to identify the information or commands the user is seeking. The input is text data, and the output is a data structure representing the intent.

[0082] Step 4:

[0083] The server's information retrieval device accesses external information sources based on the user's intent and retrieves the necessary data. Specifically, it sends requests to appropriate APIs or databases and receives the information. At this stage, the input is the user's intent data structure, and the output is the retrieved information data.

[0084] Step 5:

[0085] The server's generation device generates natural and easy-to-understand responses based on acquired information data. This process utilizes a generative AI model to construct text, and in some cases, also performs speech synthesis. The input is information data, and the output is response text or audio data.

[0086] Step 6:

[0087] The server sends the final generated response to the terminal. The terminal receives this response and presents it to the user via a display device or audio playback device. For example, information is provided to the user by displaying text on the screen or playing audio through a speaker. In this step, the input is the response data, and the output is the visual or auditory feedback to the user.

[0088] (Application Example 1)

[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0090] Using voice commands to acquire information within a mobile vehicle allows users to concentrate on driving while quickly and accurately obtaining necessary information. However, conventional technologies have faced challenges in ensuring sufficient driver safety and comfort, such as low voice recognition accuracy and the failure to provide necessary information in a timely manner. Furthermore, delays in acquiring information from external sources have also been a factor that detracts from the user experience.

[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0092] In this invention, the server includes an input device means for acquiring voice information, a conversion device means for converting the voice information into a string, an analysis device means for analyzing the string and extracting the user's request, a data acquisition device means for acquiring relevant information based on the user's request, and a response generation device means for generating a response based on the acquired relevant information. This makes it possible to recognize voice instructions with high accuracy and to quickly acquire and provide necessary information from external information sources.

[0093] An "input device for acquiring voice information" is a device that captures voice signals from a user and inputs them into the system as digital signals.

[0094] "The conversion device for converting the aforementioned audio information into a string" refers to a device that has the function of analyzing the acquired audio signal and converting it into a corresponding string format.

[0095] "The analysis device for analyzing the aforementioned string and extracting the user's request" is a device for identifying the user's intent from the converted string and determining the necessary actions.

[0096] "Data acquisition device for obtaining relevant information based on the user's request" refers to a device that has the function of obtaining relevant information corresponding to the user's request from external information sources such as the Internet.

[0097] A "response generation device for generating responses based on acquired relevant information" is a device that generates appropriate responses to present to the user based on information obtained from a data acquisition device.

[0098] An "in-vehicle voice assistant device with driving assistance functions that recognizes voice and presents information in a moving vehicle" is a device that has a voice assistant function that assists the driver by recognizing voice commands in a moving vehicle and quickly presenting necessary information.

[0099] This invention improves driver safety and convenience by implementing a voice-controlled information acquisition system in a mobile device.

[0100] The server uses an input device to acquire voice information and collects the user's voice commands from a microphone inside the mobile device. Next, it uses a conversion device, specifically a speech recognition API (e.g., Google® Cloud Speech-to-Text), to convert the collected voice information into a string.

[0101] This string is analyzed by an analysis device to identify the user's request. This analysis device uses a natural language processing library (e.g., spaCy) to extract the user's intent. After the intent is identified, a data acquisition device retrieves relevant information from an external source (e.g., Google Places API), and based on this, a response generation device generates an appropriate response. The response generation device uses speech synthesis technology (e.g., gTTS) to convert the text response into speech.

[0102] The terminal presents the generated voice response to the user through the in-car speaker. This entire process allows the driver to safely obtain necessary information while on the move.

[0103] For example, if a driver asks, "Where's the next gas station?", the server will collect location information for the nearest gas station and generate a voice response saying, "The next gas station is 2 kilometers away." This allows the driver to obtain the necessary information without taking their eyes off the road.

[0104] As an example of a prompt, we will use the instruction, "When the user asks 'Where is the next gas station?', retrieve the location of the most suitable gas station."

[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0106] Step 1:

[0107] The user speaks aloud while in a moving vehicle. This voice is captured by the device's microphone. The input is the user's voice, and the output is a digital audio signal. The device prepares to send this audio signal to the server.

[0108] Step 2:

[0109] The server processes the received audio signal with a converter and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is a digital audio signal, and the output is text data. The server analyzes the audio pattern and converts what the user says into text.

[0110] Step 3:

[0111] The server receives the string converted by the parsing device and analyzes it using a natural language processing library (e.g., spaCy). The input is text data, and the output is the analysis result indicating the user's intent. The server identifies the user's request from this analysis result.

[0112] Step 4:

[0113] The server uses a data acquisition device to access external information sources (e.g., Google Places API) based on user requests. The input is the parsed result representing the user's request, and the output is the data containing the relevant information. The server quickly retrieves the necessary information and prepares it for use within the system.

[0114] Step 5:

[0115] The server generates a response to the user based on information obtained using a response generation device. Text is converted to speech using speech synthesis technology (e.g., gTTS). The input is data containing relevant information, and the output is audio data. The server creates voice guidance and prepares it to be clearly communicated to the user.

[0116] Step 6:

[0117] The terminal processes audio data received from the server using a playback device and presents it to the user through a speaker. The input is audio data, and the output is the audio heard by the user. The terminal conveys information intuitively without attracting the driver's attention.

[0118] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0119] The system of the present invention receives voice input from a user, effectively analyzes its content, and generates an appropriate response based on the obtained information. Furthermore, it recognizes the user's emotional state and fine-tunes the content and presentation method of the response to match that emotion, thereby improving the user experience.

[0120] The user inputs voice commands into the system using a mobile device such as a smartphone or tablet. The device captures the user's voice and sends the voice signal to the server. The server uses a speech recognition API to convert the voice to text and then analyzes the resulting text.

[0121] As an analysis tool, the server incorporates an emotion recognition engine that identifies emotions from the user's speech content and tone of voice. This emotion information, along with the results of text analysis, is incorporated as an important element in information retrieval and response generation.

[0122] The emotion recognition engine analyzes the user's tone of voice and speaking patterns, classifying their emotions into categories such as positive, negative, and neutral. For example, when a user says, "I'm a little tired," the engine can sense fatigue not only from the content of the statement but also from the tone of voice, resulting in the recognition of a negative emotion.

[0123] Recognized emotions are taken into consideration when generating responses, enabling responses in an emotionally appropriate tone, such as incorporating words of encouragement. By not only providing information but also responding in a way that is sensitive to the user's emotions, more natural and approachable communication is achieved.

[0124] Ultimately, the terminal receives the generated response and presents it to the user. This presentation is done through audio output or a visual display, and by using special displays and sounds that respond to emotions, the user can experience deeper satisfaction. This makes the system both technically advanced and emotionally perceptually valuable.

[0125] The following describes the processing flow.

[0126] Step 1:

[0127] Users use the device's voice input function to ask questions or make requests by voice. The device captures the voice and temporarily stores it as digital data.

[0128] Step 2:

[0129] The device sends the saved audio data to the server. The server uses a speech recognition API to convert the audio data into text.

[0130] Step 3:

[0131] The server passes the converted text to the analysis engine, which then performs the analysis. During this analysis, the emotion recognition engine identifies the user's emotions from the text and audio features. For example, it determines whether the user's voice sounds calm or anxious.

[0132] Step 4:

[0133] The server collects relevant information in response to user requests based on analysis results and sentiment recognition data. It utilizes databases and external web services to obtain the necessary information.

[0134] Step 5:

[0135] The server generates a response based on the acquired information and the results of sentiment analysis. During response generation, the tone and content are adjusted specifically according to the user's emotions; for example, an encouraging message might be included for a depressed user.

[0136] Step 6:

[0137] The server sends the generated response data to the terminal. The terminal presents this data to the user. This presentation is done through text display on a visual display or speech synthesis. Specific examples include displaying text in soft colors on the screen and outputting speech in a gentle voice.

[0138] Through the above process, the system goes beyond simply providing information and delivers responses that are sensitive to the user's emotions.

[0139] (Example 2)

[0140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0141] In voice-input information systems, conventional technologies have struggled to generate responses that take user emotions into account, limiting the ability to achieve natural and user-friendly communication. To address this challenge, there is a need for technology that can accurately identify user emotions and generate appropriate responses based on them.

[0142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0143] In this invention, the server includes an input means for acquiring an audio signal, a conversion means for converting the audio signal into a written expression, and an analysis means for analyzing the written expression to identify the user's emotions. This enables the natural generation of responses that take the user's emotions into consideration, resulting in a deeper communication experience.

[0144] An "audio signal" is a signal used to electronically transmit information in the form of sound.

[0145] "Input means" refers to a device or technology that has the function of acquiring a user's voice signal and receiving it for subsequent processing.

[0146] "Conversion means" refers to the process or technology of converting audio signals into written expressions.

[0147] "Textual representation" refers to data in text format converted from speech, which is used for analysis and response generation.

[0148] "Analysis means" refers to a technology or process for processing written expressions and identifying the user's emotions and intentions.

[0149] "Generative means" refer to technologies and functions that produce appropriate responses based on identified emotions or intentions.

[0150] "Presentation means" refers to technologies or methods that show the generated response to the user visually or audibly.

[0151] A "computational model" is a mathematical or statistical model used to process and analyze information.

[0152] An "external information source" is a database or service that exists outside the system and is accessed to retrieve relevant information.

[0153] This invention is a system that generates and presents emotionally responsive responses based on voice input from a user. Users can input voice commands using devices such as smartphones or tablets. The terminal acquires this voice signal and transmits it to a server via a network.

[0154] The server converts the audio signal into text data using speech recognition software such as the Google Cloud Speech-to-Text API. The resulting text is then analyzed using an emotion recognition engine. The emotion recognition engine analyzes the user's speech content and tone of voice to classify emotions. This analysis utilizes machine learning models to categorize emotions as positive, negative, or neutral.

[0155] Based on the analysis results, the server generates a response using a generative AI model. This generative AI model includes commonly used natural language processing models. This process can use prompts that take the user's emotions into account, forming a response with an appropriate tone. For example, it might use a prompt such as, "If the user says, 'I'm feeling a little down today,' generate an encouraging message that is empathetic to the user's feelings."

[0156] Finally, the generated response is sent to the terminal and presented to the user via voice output or visual display. Using speech synthesis technology, the response is played back aloud and presented in a way that is appropriate to the user's emotions. This system allows users to receive not only information but also emotional support. As a result, the user experience is technologically advanced and becomes more humane.

[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0158] Step 1:

[0159] Users input voice commands using a smartphone or tablet. The input voice signals are captured by the device's microphone. The device converts these voice signals into digital data and sends it to a server over the network.

[0160] Step 2:

[0161] The server converts audio signals received over the network into text data using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, the input audio signal is analyzed by the speech recognition engine, and a written representation is obtained as output. This conversion uses phonological analysis and statistical models.

[0162] Step 3:

[0163] The server passes the obtained textual expression to the emotion recognition engine, which analyzes the user's utterances and tone of voice. Based on the analysis, the user's emotions are classified into categories such as positive, negative, and neutral. The input is textual expression, and the output is the recognized emotion category. Specifically, the system analyzes keywords within the text and the intensity of the voice.

[0164] Step 4:

[0165] The server uses the analysis results to issue prompts to the generative AI model, generating appropriate responses. In this process, instructions in the form of prompt sentences are used as input, and the output is a natural, contextualized response that is sensitive to emotions. The generative AI model performs advanced language generation using, for example, OpenAI's (registered trademark) natural language processing tools.

[0166] Step 5:

[0167] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology and a visual display to present the response to the user. The input is the generated text response, and the output is the response appropriately presented to the user. The specific operation of speech synthesis includes processes such as converting text into a speech waveform and outputting it through a speaker.

[0168] Through these processing steps, the system provides the user with an emotionally responsive and interactive experience.

[0169] (Application Example 2)

[0170] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0171] Modern security systems lack the functionality to accurately assess the urgency of voice alerts. This creates a challenge: potentially delays in prompt and appropriate responses during emergencies.

[0172] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0173] In this invention, the server includes a receiving means for acquiring an acoustic signal, a conversion device for converting the acoustic signal into a text representation, and an emotion recognition device for identifying the emotional state from the user's voice and determining the degree of urgency. This enables immediate determination of the degree of urgency based on the user's voice input, allowing for a quick and appropriate response.

[0174] An "acoustic signal" is an electrical signal obtained by converting the waveform of sound traveling through space into an electrical signal, and is treated as audio data.

[0175] A "receiving means" is a device for acquiring acoustic signals, and its role is to convert audio input into electrical signals and deliver them to the system.

[0176] A "conversion device" is a device that converts acoustic signals into written representations, and has the function of converting audio data into text data.

[0177] An "analysis device" is a device that analyzes text data to identify the user's intent and is used to understand the intent behind the information.

[0178] An "information acquisition device" is a device for acquiring relevant information based on the user's intent, and is equipped with the function of acquiring necessary information from external resources.

[0179] A "response generation device" is a device that generates a response based on acquired information, and is intended to form an appropriate response for the user.

[0180] A "presentation means" is a means of presenting a generated response to the user, and has the role of displaying or playing information in the form of sound or visual means.

[0181] An "emotion recognition device" is a device that identifies an emotional state from a user's voice and determines the degree of urgency, playing a role in estimating emotions from the tone and content of the voice.

[0182] The system that implements this application generates appropriate responses based on user voice input and takes emergency action as needed. The system operates using the following hardware and software:

[0183] First, the user's voice is acquired by the smartphone's receiving device and processed as an acoustic signal. The acquired acoustic signal is then converted into text through a conversion device. It is common to use speech recognition APIs such as Google Cloud Speech-to-Text for speech recognition.

[0184] Next, the server uses an analysis device to analyze the text representation and understand the user's intent. At this stage, natural language processing techniques are applied to identify specific requests from the user's speech. The analysis results are sent to an information retrieval device that accesses external information sources and network services to obtain relevant information. For example, an API is used to retrieve necessary information from the web.

[0185] Subsequently, the server uses a response generation device to generate an appropriate response to the user based on the acquired information. During this process, an emotion recognition device is used to identify the user's emotional state and adjust the tone and urgency of the response accordingly. Emotion recognition engines such as Aurora and IBM Watson® are often utilized.

[0186] Finally, the response is sent to the user through a presentation method. On smartphones, information is provided via voice output or a visual interface. If the situation is deemed urgent, the security company and close family members are immediately notified.

[0187] For example, if a user speaks into their smartphone at night saying, "There might be an intruder," and their tone of voice indicates anxiety, the system will immediately detect negative emotions and initiate an emergency call.

[0188] An example of a prompt message might be something like, "We will use our emotion recognition engine to classify the user's emotions from their voice and provide immediate notification if it is an emergency."

[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0190] Step 1:

[0191] The device acquires an acoustic signal.

[0192] When a user speaks into their smartphone, the device's microphone receives the sound as an acoustic signal. This acoustic signal is then acquired as input data.

[0193] Step 2:

[0194] The device converts the audio signal into text.

[0195] The device converts the acquired acoustic signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). It analyzes the acoustic signal and generates the corresponding text as output.

[0196] Step 3:

[0197] The server analyzes the text data to identify the user's intent.

[0198] The server receives text data sent from the terminal and analyzes it using natural language processing technology. The user's intent, identified as a result of the text analysis, is output, and this intent forms the basis for generating related responses and retrieving information.

[0199] Step 4:

[0200] The server uses an emotion recognition device to identify the user's emotional state.

[0201] The analyzed text data and voice tone are then analyzed by an emotion recognition engine on the server (e.g., Aurora or IBM Watson). This process outputs emotion data classified as positive, negative, or neutral, which is used as an important element in response generation.

[0202] Step 5:

[0203] The server retrieves the relevant information.

[0204] Based on the user's intent, the system retrieves necessary information by accessing external information sources and network services. The server collects relevant information using an information retrieval API and uses that information as input for the next response generation step.

[0205] Step 6:

[0206] The server generates a response.

[0207] Based on the acquired relevant information and emotional state, the server constructs the optimal response using a response generator. The generated response becomes the output, and is adjusted according to the tone of voice and urgency.

[0208] Step 7:

[0209] The terminal presents the generated response to the user.

[0210] The terminal, upon receiving the response from the server, presents it via audio output or visually on its display. In emergencies, it can also contact security services or notify family members. This step provides the user with final information.

[0211] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0212] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0213] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0214] [Second Embodiment]

[0215] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0216] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0217] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0218] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0219] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0221] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0222] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0223] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0224] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0225] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0226] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0227] The system of the present invention appropriately acquires the user's voice input, efficiently converts it into text to analyze the user's intent, acquires necessary information from external sources based on that intent, and ultimately generates an appropriate response for the user.

[0228] Users access the system using devices such as smartphones and tablets. These devices are equipped with voice input capabilities, which are used to capture the user's voice as a digital signal. The captured voice signal is then sent from the device to the server via the network.

[0229] The server converts audio data into text using a speech recognition API. This speech recognition process utilizes advanced language models to analyze speech patterns and convert them into human-readable text. This converted text is then used for subsequent analysis.

[0230] The analysis engine utilizes natural language processing techniques to extract intent from the user's utterances. For example, if a user says, "Tell me the weather for tomorrow," the analysis engine identifies the need to search for weather information. Based on this information request, the server retrieves the necessary data from external databases and web services.

[0231] The acquired information is processed by a response generation module, which uses commonly used text generation and speech synthesis technologies to construct a response for the user. For example, if it's weather information, it will generate specific information such as, "Tomorrow in Tokyo it will be sunny, with a high of 25 degrees Celsius."

[0232] The final generated response is returned to the terminal. The terminal receives this response data and displays it as text on the screen or plays it back as audio through the speaker, allowing the user to intuitively obtain the requested information. This enables the user to quickly obtain information and achieve their goals.

[0233] The following describes the processing flow.

[0234] Step 1:

[0235] The user activates the device's voice input function and speaks their question or request. The device captures the voice using the microphone and temporarily stores it as digital data.

[0236] Step 2:

[0237] The terminal converts the acquired audio signal into a format that is easy to process and sends this audio data to the server. The audio data is delivered via the network to the API endpoint specified by the server.

[0238] Step 3:

[0239] The server passes the received audio data to a speech recognition API, which converts the audio into text. This conversion process uses a language model to analyze the audio signal and generate corresponding text information.

[0240] Step 4:

[0241] The server inputs text data into a natural language processing engine to analyze the intent of the user's request. This analysis process uses techniques such as keyword extraction and sentiment analysis to identify the information the user is looking for.

[0242] Step 5:

[0243] Based on the analysis results, the server collects the information the user needs from the internet and external databases. For example, it might call web services such as weather information APIs to obtain relevant information.

[0244] Step 6:

[0245] The server uses the collected information to generate a response to present to the user. The response generation module constructs the information in a natural way, utilizing text and speech synthesis.

[0246] Step 7:

[0247] The server sends the generated response to the terminal. The terminal presents this response data to the user using methods such as audio output or display. The user obtains the answer to their question through the presented information.

[0248] (Example 1)

[0249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0250] In recent years, voice assistants and natural language interfaces have become widely used, but they suffer from problems such as low voice recognition accuracy, inability to accurately analyze intent from speech, and unnatural responses to users. These problems hinder the improvement of information retrieval and efficiency in daily life using voice interfaces. The present invention aims to address these problems and provide a voice interface that enables more natural, rapid, and accurate information retrieval and response generation.

[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0252] In this invention, the server includes a conversion device for converting voice signals into text information, an analysis device for analyzing the text information and extracting intent, and an information acquisition device for obtaining relevant information based on the intent. This enables the user to efficiently acquire information via voice and obtain appropriate and natural responses.

[0253] An "audio signal" is data that represents sound in digital or analog format.

[0254] An "input device" is a device that receives voice information from a user and acquires it as a digital signal.

[0255] A "conversion device" is a device used to convert acquired audio signals into text information.

[0256] "Textual information" refers to data obtained by converting audio signals into a text format that can be analyzed.

[0257] An "analysis device" is a device that analyzes textual information and extracts the user's intended meaning.

[0258] "Intent" refers to a specific request for information or a directive that can be inferred from the user's statements.

[0259] An "information acquisition device" is a device that acquires relevant information from an external source based on the user's intent.

[0260] A "generation device" is a device that creates an appropriate response based on acquired information.

[0261] A "display device" is a device that visually presents the generated response to the user as text.

[0262] A "sound playback device" is a device that presents the generated response to the user audibly as sound.

[0263] This invention provides a system that offers an interface for users to acquire information via voice and receive appropriate responses. The user uses a device such as a smartphone or tablet to acquire voice signals using a voice input device. The device converts these voice signals into a digital format and transmits them to a server via a network.

[0264] The server converts the audio signal into text information using a speech recognition API (for example, using general speech recognition technology) to analyze the audio signal. This converted text information is then processed by an analysis device using natural language processing technology to extract the user's intent. The analysis device uses generative AI model technology (for example, a general natural language processing engine) to identify the user's intent.

[0265] Based on the identified user intent, the information acquisition device accesses external information sources and retrieves relevant information. These external information sources could include various information service providers and databases. The acquired information is processed by a generator on the server and converted into a natural language response to the user. Here, generative AI modeling technology is used to improve the naturalness and accuracy of the response.

[0266] Finally, the generated response is sent to the terminal, which then presents this response to the user. Using a display device or audio playback device, intuitive feedback can be provided to the user by displaying it as text or playing it aloud.

[0267] As a concrete example, a user might use voice input via their device, saying, "Tell me the weather for tomorrow." In this case, the server converts the voice signal into text, analyzes the intent of "Tell me the weather for tomorrow," and retrieves weather information by referring to an appropriate external database. The generated response would then be provided in the form of, "Tomorrow will be sunny, with a high of 25 degrees Celsius."

[0268] Examples of prompt phrases include "Voice input: Tell me the weather tomorrow" and "Voice input: Tell me the latest news."

[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0270] Step 1:

[0271] The user inputs an audio signal via the terminal's audio input device. The terminal converts this audio into a digital signal and transmits it to the server via the internet. The input is raw audio, and the output is encoded digital data.

[0272] Step 2:

[0273] The server passes the received digital audio signal to a speech recognition API, which converts the audio into text. In this process, the API analyzes the frequency pattern of the audio and generates the corresponding text. The input is a digital audio signal, and the output is text data.

[0274] Step 3:

[0275] The analysis device on the server analyzes the acquired text data using natural language processing techniques to extract the user's intent. Here, it understands keywords and context within the text to identify the information or commands the user is seeking. The input is text data, and the output is a data structure representing the intent.

[0276] Step 4:

[0277] The server's information retrieval device accesses external information sources based on the user's intent and retrieves the necessary data. Specifically, it sends requests to appropriate APIs or databases and receives the information. At this stage, the input is the user's intent data structure, and the output is the retrieved information data.

[0278] Step 5:

[0279] The server's generation device generates natural and easy-to-understand responses based on acquired information data. This process utilizes a generative AI model to construct text, and in some cases, also performs speech synthesis. The input is information data, and the output is response text or audio data.

[0280] Step 6:

[0281] The server sends the final generated response to the terminal. The terminal receives this response and presents it to the user via a display device or audio playback device. For example, information is provided to the user by displaying text on the screen or playing audio through a speaker. In this step, the input is the response data, and the output is the visual or auditory feedback to the user.

[0282] (Application Example 1)

[0283] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0284] Obtaining information within a moving body using voice instructions enables the user to obtain necessary information quickly and accurately while concentrating on driving. However, in conventional technologies, there were problems such as low voice recognition accuracy and failure to provide necessary information in a timely manner, making it difficult to sufficiently ensure the safety and comfort of drivers. In addition, the acquisition of information from external information sources was sometimes delayed, which was a factor that impaired the user experience.

[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0286] In this invention, the server includes input device means for acquiring voice information, conversion device means for converting the voice information into a character string, analysis device means for analyzing the character string and extracting the user's request, data acquisition device means for acquiring related information based on the user's request, and response generation device means for generating a response based on the acquired related information. Thereby, it becomes possible to accurately recognize voice instructions and quickly acquire and provide necessary information from an external information source.

[0287] The "input device for acquiring voice information" is a device for capturing a voice signal from a user and inputting it into the system as a digital signal.

[0288] The "conversion device for converting the voice information into a character string" is a device having a function of analyzing the acquired voice signal and converting it into a corresponding character string format.

[0289] "The analysis device for analyzing the aforementioned string and extracting the user's request" is a device for identifying the user's intent from the converted string and determining the necessary actions.

[0290] "Data acquisition device for obtaining relevant information based on the user's request" refers to a device that has the function of obtaining relevant information corresponding to the user's request from external information sources such as the Internet.

[0291] A "response generation device for generating responses based on acquired relevant information" is a device that generates appropriate responses to present to the user based on information obtained from a data acquisition device.

[0292] An "in-vehicle voice assistant device with driving assistance functions that recognizes voice and presents information in a moving vehicle" is a device that has a voice assistant function that assists the driver by recognizing voice commands in a moving vehicle and quickly presenting necessary information.

[0293] This invention improves driver safety and convenience by implementing a voice-controlled information acquisition system in a mobile device.

[0294] The server uses an input device to acquire voice information and collects the user's voice commands from a microphone inside the mobile device. Next, it uses a conversion device, specifically a speech recognition API (e.g., Google Cloud Speech-to-Text), to convert the collected voice information into a string.

[0295] This string is analyzed by an analysis device to identify the user's request. This analysis device uses a natural language processing library (e.g., spaCy) to extract the user's intent. After the intent is identified, a data acquisition device retrieves relevant information from an external source (e.g., Google Places API), and based on this, a response generation device generates an appropriate response. The response generation device uses speech synthesis technology (e.g., gTTS) to convert the text response into speech.

[0296] The terminal presents the generated voice response to the user through the in-car speaker. This entire process allows the driver to safely obtain necessary information while on the move.

[0297] For example, if a driver asks, "Where's the next gas station?", the server will collect location information for the nearest gas station and generate a voice response saying, "The next gas station is 2 kilometers away." This allows the driver to obtain the necessary information without taking their eyes off the road.

[0298] As an example of a prompt, we will use the instruction, "When the user asks 'Where is the next gas station?', retrieve the location of the most suitable gas station."

[0299] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0300] Step 1:

[0301] The user speaks aloud while in a moving vehicle. This voice is captured by the device's microphone. The input is the user's voice, and the output is a digital audio signal. The device prepares to send this audio signal to the server.

[0302] Step 2:

[0303] The server processes the received audio signal with a converter and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is a digital audio signal, and the output is text data. The server analyzes the audio pattern and converts what the user says into text.

[0304] Step 3:

[0305] The server receives the string converted by the analysis device and analyzes it using a natural language processing library (e.g., spaCy). The input is text data, and the output is an analysis result indicating the user's intention. The server identifies the user's request from this analysis result.

[0306] Step 4:

[0307] The server uses a data acquisition device to access an external information source (e.g., Google Places API) based on the user's request. The input is the analysis result indicating the user's request, and the output is data of relevant information. The server quickly acquires the necessary information and prepares it for use within the system.

[0308] Step 5:

[0309] The server uses a response generation device to generate a response to the user based on the acquired information. The text is converted to audio by voice synthesis technology (e.g., gTTS). The input is data of relevant information, and the output is audio data. The server creates an audio guide and prepares to convey it to the user in an easy-to-understand manner.

[0310] Step 6:

[0311] The terminal processes the audio data received from the server with a playback device and presents it to the user through a speaker. The input is audio data, and the output is the audio that enters the user's ear. The terminal conveys the information intuitively without attracting the driver's attention.

[0312] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.

[0313] The system of the present invention receives voice input from a user, effectively analyzes its content, and generates an appropriate response based on the obtained information. Furthermore, it recognizes the user's emotional state and fine-tunes the content and presentation method of the response to match that emotion, thereby improving the user experience.

[0314] The user inputs voice commands into the system using a mobile device such as a smartphone or tablet. The device captures the user's voice and sends the voice signal to the server. The server uses a speech recognition API to convert the voice to text and then analyzes the resulting text.

[0315] As an analysis tool, the server incorporates an emotion recognition engine that identifies emotions from the user's speech content and tone of voice. This emotion information, along with the results of text analysis, is incorporated as an important element in information retrieval and response generation.

[0316] The emotion recognition engine analyzes the user's tone of voice and speaking patterns, classifying their emotions into categories such as positive, negative, and neutral. For example, when a user says, "I'm a little tired," the engine can sense fatigue not only from the content of the statement but also from the tone of voice, resulting in the recognition of a negative emotion.

[0317] Recognized emotions are taken into consideration when generating responses, enabling responses in an emotionally appropriate tone, such as incorporating words of encouragement. By not only providing information but also responding in a way that is sensitive to the user's emotions, more natural and approachable communication is achieved.

[0318] Ultimately, the terminal receives the generated response and presents it to the user. This presentation is done through audio output or a visual display, and by using special displays and sounds that respond to emotions, the user can experience deeper satisfaction. This makes the system both technically advanced and emotionally perceptually valuable.

[0319] The following describes the processing flow.

[0320] Step 1:

[0321] Users use the device's voice input function to ask questions or make requests by voice. The device captures the voice and temporarily stores it as digital data.

[0322] Step 2:

[0323] The device sends the saved audio data to the server. The server uses a speech recognition API to convert the audio data into text.

[0324] Step 3:

[0325] The server passes the converted text to the analysis engine, which then performs the analysis. During this analysis, the emotion recognition engine identifies the user's emotions from the text and audio features. For example, it determines whether the user's voice sounds calm or anxious.

[0326] Step 4:

[0327] The server collects relevant information in response to user requests based on analysis results and sentiment recognition data. It utilizes databases and external web services to obtain the necessary information.

[0328] Step 5:

[0329] The server generates a response based on the acquired information and the results of sentiment analysis. During response generation, the tone and content are adjusted specifically according to the user's emotions; for example, an encouraging message might be included for a depressed user.

[0330] Step 6:

[0331] The server sends the generated response data to the terminal. The terminal presents this data to the user. This presentation is done through text display on a visual display or speech synthesis. Specific examples include displaying text in soft colors on the screen and outputting speech in a gentle voice.

[0332] Through the above process, the system goes beyond simply providing information and delivers responses that are sensitive to the user's emotions.

[0333] (Example 2)

[0334] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0335] In voice-input information systems, conventional technologies have struggled to generate responses that take user emotions into account, limiting the ability to achieve natural and user-friendly communication. To address this challenge, there is a need for technology that can accurately identify user emotions and generate appropriate responses based on them.

[0336] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0337] In this invention, the server includes an input means for acquiring an audio signal, a conversion means for converting the audio signal into a written expression, and an analysis means for analyzing the written expression to identify the user's emotions. This enables the natural generation of responses that take the user's emotions into consideration, resulting in a deeper communication experience.

[0338] An "audio signal" is a signal used to electronically transmit information in the form of sound.

[0339] "Input means" refers to a device or technology that has the function of acquiring a user's voice signal and receiving it for subsequent processing.

[0340] "Conversion means" refers to the process or technology of converting audio signals into written expressions.

[0341] "Textual representation" refers to data in text format converted from speech, which is used for analysis and response generation.

[0342] "Analysis means" refers to a technology or process for processing written expressions and identifying the user's emotions and intentions.

[0343] "Generative means" refer to technologies and functions that produce appropriate responses based on identified emotions or intentions.

[0344] "Presentation means" refers to technologies or methods that show the generated response to the user visually or audibly.

[0345] A "computational model" is a mathematical or statistical model used to process and analyze information.

[0346] An "external information source" is a database or service that exists outside the system and is accessed to retrieve relevant information.

[0347] This invention is a system that generates and presents emotionally responsive responses based on voice input from a user. Users can input voice commands using devices such as smartphones or tablets. The terminal acquires this voice signal and transmits it to a server via a network.

[0348] The server converts the audio signal into text data using speech recognition software such as the Google Cloud Speech-to-Text API. The resulting text is then analyzed using an emotion recognition engine. The emotion recognition engine analyzes the user's speech content and tone of voice to classify emotions. This analysis utilizes machine learning models to categorize emotions as positive, negative, or neutral.

[0349] Based on the analysis results, the server generates a response using a generative AI model. This generative AI model includes commonly used natural language processing models. This process can use prompts that take the user's emotions into account, forming a response with an appropriate tone. For example, it might use a prompt such as, "If the user says, 'I'm feeling a little down today,' generate an encouraging message that is empathetic to the user's feelings."

[0350] Finally, the generated response is sent to the terminal and presented to the user via voice output or visual display. Using speech synthesis technology, the response is played back aloud and presented in a way that is appropriate to the user's emotions. This system allows users to receive not only information but also emotional support. As a result, the user experience is technologically advanced and becomes more humane.

[0351] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0352] Step 1:

[0353] Users input voice commands using a smartphone or tablet. The input voice signals are captured by the device's microphone. The device converts these voice signals into digital data and sends it to a server over the network.

[0354] Step 2:

[0355] The server converts audio signals received over the network into text data using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, the input audio signal is analyzed by the speech recognition engine, and a written representation is obtained as output. This conversion uses phonological analysis and statistical models.

[0356] Step 3:

[0357] The server passes the obtained textual expression to the emotion recognition engine, which analyzes the user's utterances and tone of voice. Based on the analysis, the user's emotions are classified into categories such as positive, negative, and neutral. The input is textual expression, and the output is the recognized emotion category. Specifically, the system analyzes keywords within the text and the intensity of the voice.

[0358] Step 4:

[0359] The server uses the analysis results to issue prompts to the generative AI model, generating appropriate responses. In this process, prompt statements serve as input, and the output is a natural, contextualized response that aligns with emotions. The generative AI model performs advanced language generation using, for example, OpenAI's natural language processing tools.

[0360] Step 5:

[0361] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology and a visual display to present the response to the user. The input is the generated text response, and the output is the response appropriately presented to the user. The specific operation of speech synthesis includes processes such as converting text into a speech waveform and outputting it through a speaker.

[0362] Through these processing steps, the system provides the user with an emotionally responsive and interactive experience.

[0363] (Application Example 2)

[0364] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0365] Modern security systems lack the functionality to accurately assess the urgency of voice alerts. This creates a challenge: potentially delays in prompt and appropriate responses during emergencies.

[0366] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0367] In this invention, the server includes a receiving means for acquiring an acoustic signal, a conversion device for converting the acoustic signal into a text representation, and an emotion recognition device for identifying the emotional state from the user's voice and determining the degree of urgency. This enables immediate determination of the degree of urgency based on the user's voice input, allowing for a quick and appropriate response.

[0368] An "acoustic signal" is an electrical signal obtained by converting the waveform of sound traveling through space into an electrical signal, and is treated as audio data.

[0369] A "receiving means" is a device for acquiring acoustic signals, and its role is to convert audio input into electrical signals and deliver them to the system.

[0370] A "conversion device" is a device that converts acoustic signals into written representations, and has the function of converting audio data into text data.

[0371] An "analysis device" is a device that analyzes text data to identify the user's intent and is used to understand the intent behind the information.

[0372] An "information acquisition device" is a device for acquiring relevant information based on the user's intent, and is equipped with the function of acquiring necessary information from external resources.

[0373] A "response generation device" is a device that generates a response based on acquired information, and is intended to form an appropriate response for the user.

[0374] A "presentation means" is a means of presenting a generated response to the user, and has the role of displaying or playing information in the form of sound or visual means.

[0375] An "emotion recognition device" is a device that identifies an emotional state from a user's voice and determines the degree of urgency, playing a role in estimating emotions from the tone and content of the voice.

[0376] The system that implements this application generates appropriate responses based on user voice input and takes emergency action as needed. The system operates using the following hardware and software:

[0377] First, the user's voice is acquired by the smartphone's receiving device and processed as an acoustic signal. The acquired acoustic signal is then converted into text through a conversion device. It is common to use speech recognition APIs such as Google Cloud Speech-to-Text for speech recognition.

[0378] Next, the server uses an analysis device to analyze the text representation and understand the user's intent. At this stage, natural language processing techniques are applied to identify specific requests from the user's speech. The analysis results are sent to an information retrieval device that accesses external information sources and network services to obtain relevant information. For example, an API is used to retrieve necessary information from the web.

[0379] Subsequently, the server uses a response generation device to generate an appropriate response to the user based on the acquired information. During this process, an emotion recognition device is used to identify the user's emotional state and adjust the tone and urgency of the response accordingly. Emotion recognition engines such as Aurora and IBM Watson are often utilized.

[0380] Finally, the response is sent to the user through a presentation method. On smartphones, information is provided via voice output or a visual interface. If the situation is deemed urgent, the security company and close family members are immediately notified.

[0381] For example, if a user speaks into their smartphone at night saying, "There might be an intruder," and their tone of voice indicates anxiety, the system will immediately detect negative emotions and initiate an emergency call.

[0382] An example of a prompt message might be something like, "We will use our emotion recognition engine to classify the user's emotions from their voice and provide immediate notification if it is an emergency."

[0383] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0384] Step 1:

[0385] The device acquires an acoustic signal.

[0386] When a user speaks into their smartphone, the device's microphone receives the sound as an acoustic signal. This acoustic signal is then acquired as input data.

[0387] Step 2:

[0388] The device converts the audio signal into text.

[0389] The device converts the acquired acoustic signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). It analyzes the acoustic signal and generates the corresponding text as output.

[0390] Step 3:

[0391] The server analyzes the text data to identify the user's intent.

[0392] The server receives text data sent from the terminal and analyzes it using natural language processing technology. The user's intent, identified as a result of the text analysis, is output, and this intent forms the basis for generating related responses and retrieving information.

[0393] Step 4:

[0394] The server uses an emotion recognition device to identify the user's emotional state.

[0395] The analyzed text data and voice tone are then analyzed by an emotion recognition engine on the server (e.g., Aurora or IBM Watson). This process outputs emotion data classified as positive, negative, or neutral, which is used as an important element in response generation.

[0396] Step 5:

[0397] The server retrieves the relevant information.

[0398] Based on the user's intent, the system retrieves necessary information by accessing external information sources and network services. The server collects relevant information using an information retrieval API and uses that information as input for the next response generation step.

[0399] Step 6:

[0400] The server generates a response.

[0401] Based on the acquired relevant information and emotional state, the server constructs the optimal response using a response generator. The generated response becomes the output, and is adjusted according to the tone of voice and urgency.

[0402] Step 7:

[0403] The terminal presents the generated response to the user.

[0404] The terminal, upon receiving the response from the server, presents it via audio output or visually on its display. In emergencies, it can also contact security services or notify family members. This step provides the user with final information.

[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0408] [Third Embodiment]

[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0421] The system of the present invention appropriately acquires the user's voice input, efficiently converts it into text to analyze the user's intent, acquires necessary information from external sources based on that intent, and ultimately generates an appropriate response for the user.

[0422] Users access the system using devices such as smartphones and tablets. These devices are equipped with voice input capabilities, which are used to capture the user's voice as a digital signal. The captured voice signal is then sent from the device to the server via the network.

[0423] The server converts audio data into text using a speech recognition API. This speech recognition process utilizes advanced language models to analyze speech patterns and convert them into human-readable text. This converted text is then used for subsequent analysis.

[0424] The analysis engine utilizes natural language processing techniques to extract intent from the user's utterances. For example, if a user says, "Tell me the weather for tomorrow," the analysis engine identifies the need to search for weather information. Based on this information request, the server retrieves the necessary data from external databases and web services.

[0425] The acquired information is processed by a response generation module, which uses commonly used text generation and speech synthesis technologies to construct a response for the user. For example, if it's weather information, it will generate specific information such as, "Tomorrow in Tokyo it will be sunny, with a high of 25 degrees Celsius."

[0426] The final generated response is returned to the terminal. The terminal receives this response data and displays it as text on the screen or plays it back as audio through the speaker, allowing the user to intuitively obtain the requested information. This enables the user to quickly obtain information and achieve their goals.

[0427] The following describes the processing flow.

[0428] Step 1:

[0429] The user activates the device's voice input function and speaks their question or request. The device captures the voice using the microphone and temporarily stores it as digital data.

[0430] Step 2:

[0431] The terminal converts the acquired audio signal into a format that is easy to process and sends this audio data to the server. The audio data is delivered via the network to the API endpoint specified by the server.

[0432] Step 3:

[0433] The server passes the received audio data to a speech recognition API, which converts the audio into text. This conversion process uses a language model to analyze the audio signal and generate corresponding text information.

[0434] Step 4:

[0435] The server inputs text data into a natural language processing engine to analyze the intent of the user's request. This analysis process uses techniques such as keyword extraction and sentiment analysis to identify the information the user is looking for.

[0436] Step 5:

[0437] Based on the analysis results, the server collects the information the user needs from the internet and external databases. For example, it might call web services such as weather information APIs to obtain relevant information.

[0438] Step 6:

[0439] The server uses the collected information to generate a response to present to the user. The response generation module constructs the information in a natural way, utilizing text and speech synthesis.

[0440] Step 7:

[0441] The server sends the generated response to the terminal. The terminal presents this response data to the user using methods such as audio output or display. The user obtains the answer to their question through the presented information.

[0442] (Example 1)

[0443] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0444] In recent years, voice assistants and natural language interfaces have become widely used, but they suffer from problems such as low voice recognition accuracy, inability to accurately analyze intent from speech, and unnatural responses to users. These problems hinder the improvement of information retrieval and efficiency in daily life using voice interfaces. The present invention aims to address these problems and provide a voice interface that enables more natural, rapid, and accurate information retrieval and response generation.

[0445] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0446] In this invention, the server includes a conversion device for converting voice signals into text information, an analysis device for analyzing the text information and extracting intent, and an information acquisition device for obtaining relevant information based on the intent. This enables the user to efficiently acquire information via voice and obtain appropriate and natural responses.

[0447] An "audio signal" is data that represents sound in digital or analog format.

[0448] An "input device" is a device that receives voice information from a user and acquires it as a digital signal.

[0449] A "conversion device" is a device used to convert acquired audio signals into text information.

[0450] "Textual information" refers to data obtained by converting audio signals into a text format that can be analyzed.

[0451] An "analysis device" is a device that analyzes textual information and extracts the user's intended meaning.

[0452] "Intent" refers to a specific request for information or a directive that can be inferred from the user's statements.

[0453] An "information acquisition device" is a device that acquires relevant information from an external source based on the user's intent.

[0454] A "generation device" is a device that creates an appropriate response based on acquired information.

[0455] A "display device" is a device that visually presents the generated response to the user as text.

[0456] A "sound playback device" is a device that presents the generated response to the user audibly as sound.

[0457] This invention provides a system that offers an interface for users to acquire information via voice and receive appropriate responses. The user uses a device such as a smartphone or tablet to acquire voice signals using a voice input device. The device converts these voice signals into a digital format and transmits them to a server via a network.

[0458] The server converts the audio signal into text information using a speech recognition API (for example, using general speech recognition technology) to analyze the audio signal. This converted text information is then processed by an analysis device using natural language processing technology to extract the user's intent. The analysis device uses generative AI model technology (for example, a general natural language processing engine) to identify the user's intent.

[0459] Based on the identified user intent, the information acquisition device accesses external information sources and retrieves relevant information. These external information sources could include various information service providers and databases. The acquired information is processed by a generator on the server and converted into a natural language response to the user. Here, generative AI modeling technology is used to improve the naturalness and accuracy of the response.

[0460] Finally, the generated response is sent to the terminal, which then presents this response to the user. Using a display device or audio playback device, intuitive feedback can be provided to the user by displaying it as text or playing it aloud.

[0461] As a concrete example, a user might use voice input via their device, saying, "Tell me the weather for tomorrow." In this case, the server converts the voice signal into text, analyzes the intent of "Tell me the weather for tomorrow," and retrieves weather information by referring to an appropriate external database. The generated response would then be provided in the form of, "Tomorrow will be sunny, with a high of 25 degrees Celsius."

[0462] Examples of prompt phrases include "Voice input: Tell me the weather tomorrow" and "Voice input: Tell me the latest news."

[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0464] Step 1:

[0465] The user inputs an audio signal via the terminal's audio input device. The terminal converts this audio into a digital signal and transmits it to the server via the internet. The input is raw audio, and the output is encoded digital data.

[0466] Step 2:

[0467] The server passes the received digital audio signal to a speech recognition API, which converts the audio into text. In this process, the API analyzes the frequency pattern of the audio and generates the corresponding text. The input is a digital audio signal, and the output is text data.

[0468] Step 3:

[0469] The analysis device on the server analyzes the acquired text data using natural language processing techniques to extract the user's intent. Here, it understands keywords and context within the text to identify the information or commands the user is seeking. The input is text data, and the output is a data structure representing the intent.

[0470] Step 4:

[0471] The server's information retrieval device accesses external information sources based on the user's intent and retrieves the necessary data. Specifically, it sends requests to appropriate APIs or databases and receives the information. At this stage, the input is the user's intent data structure, and the output is the retrieved information data.

[0472] Step 5:

[0473] The server's generation device generates natural and easy-to-understand responses based on acquired information data. This process utilizes a generative AI model to construct text, and in some cases, also performs speech synthesis. The input is information data, and the output is response text or audio data.

[0474] Step 6:

[0475] The server sends the final generated response to the terminal. The terminal receives this response and presents it to the user via a display device or audio playback device. For example, information is provided to the user by displaying text on the screen or playing audio through a speaker. In this step, the input is the response data, and the output is the visual or auditory feedback to the user.

[0476] (Application Example 1)

[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0478] Using voice commands to acquire information within a mobile vehicle allows users to concentrate on driving while quickly and accurately obtaining necessary information. However, conventional technologies have faced challenges in ensuring sufficient driver safety and comfort, such as low voice recognition accuracy and the failure to provide necessary information in a timely manner. Furthermore, delays in acquiring information from external sources have also been a factor that detracts from the user experience.

[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0480] In this invention, the server includes an input device means for acquiring voice information, a conversion device means for converting the voice information into a string, an analysis device means for analyzing the string and extracting the user's request, a data acquisition device means for acquiring relevant information based on the user's request, and a response generation device means for generating a response based on the acquired relevant information. This makes it possible to recognize voice instructions with high accuracy and to quickly acquire and provide necessary information from external information sources.

[0481] An "input device for acquiring voice information" is a device that captures voice signals from a user and inputs them into the system as digital signals.

[0482] "The conversion device for converting the aforementioned audio information into a string" refers to a device that has the function of analyzing the acquired audio signal and converting it into a corresponding string format.

[0483] "The analysis device for analyzing the aforementioned string and extracting the user's request" is a device for identifying the user's intent from the converted string and determining the necessary actions.

[0484] "Data acquisition device for obtaining relevant information based on the user's request" refers to a device that has the function of obtaining relevant information corresponding to the user's request from external information sources such as the Internet.

[0485] A "response generation device for generating responses based on acquired relevant information" is a device that generates appropriate responses to present to the user based on information obtained from a data acquisition device.

[0486] An "in-vehicle voice assistant device with driving assistance functions that recognizes voice and presents information in a moving vehicle" is a device that has a voice assistant function that assists the driver by recognizing voice commands in a moving vehicle and quickly presenting necessary information.

[0487] This invention improves driver safety and convenience by implementing a voice-controlled information acquisition system in a mobile device.

[0488] The server uses an input device to acquire voice information and collects the user's voice commands from a microphone inside the mobile device. Next, it uses a conversion device, specifically a speech recognition API (e.g., Google Cloud Speech-to-Text), to convert the collected voice information into a string.

[0489] This string is analyzed by an analysis device to identify the user's request. This analysis device uses a natural language processing library (e.g., spaCy) to extract the user's intent. After the intent is identified, a data acquisition device retrieves relevant information from an external source (e.g., Google Places API), and based on this, a response generation device generates an appropriate response. The response generation device uses speech synthesis technology (e.g., gTTS) to convert the text response into speech.

[0490] The terminal presents the generated voice response to the user through the in-car speaker. This entire process allows the driver to safely obtain necessary information while on the move.

[0491] For example, if a driver asks, "Where's the next gas station?", the server will collect location information for the nearest gas station and generate a voice response saying, "The next gas station is 2 kilometers away." This allows the driver to obtain the necessary information without taking their eyes off the road.

[0492] As an example of a prompt, we will use the instruction, "When the user asks 'Where is the next gas station?', retrieve the location of the most suitable gas station."

[0493] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0494] Step 1:

[0495] The user speaks aloud while in a moving vehicle. This voice is captured by the device's microphone. The input is the user's voice, and the output is a digital audio signal. The device prepares to send this audio signal to the server.

[0496] Step 2:

[0497] The server processes the received audio signal with a converter and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is a digital audio signal, and the output is text data. The server analyzes the audio pattern and converts what the user says into text.

[0498] Step 3:

[0499] The server receives the string converted by the parsing device and analyzes it using a natural language processing library (e.g., spaCy). The input is text data, and the output is the analysis result indicating the user's intent. The server identifies the user's request from this analysis result.

[0500] Step 4:

[0501] The server uses a data acquisition device to access external information sources (e.g., Google Places API) based on user requests. The input is the parsed result representing the user's request, and the output is the data containing the relevant information. The server quickly retrieves the necessary information and prepares it for use within the system.

[0502] Step 5:

[0503] The server generates a response to the user based on information obtained using a response generation device. Text is converted to speech using speech synthesis technology (e.g., gTTS). The input is data containing relevant information, and the output is audio data. The server creates voice guidance and prepares it to be clearly communicated to the user.

[0504] Step 6:

[0505] The terminal processes audio data received from the server using a playback device and presents it to the user through a speaker. The input is audio data, and the output is the audio heard by the user. The terminal conveys information intuitively without attracting the driver's attention.

[0506] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0507] The system of the present invention receives voice input from a user, effectively analyzes its content, and generates an appropriate response based on the obtained information. Furthermore, it recognizes the user's emotional state and fine-tunes the content and presentation method of the response to match that emotion, thereby improving the user experience.

[0508] The user inputs voice commands into the system using a mobile device such as a smartphone or tablet. The device captures the user's voice and sends the voice signal to the server. The server uses a speech recognition API to convert the voice to text and then analyzes the resulting text.

[0509] As an analysis tool, the server incorporates an emotion recognition engine that identifies emotions from the user's speech content and tone of voice. This emotion information, along with the results of text analysis, is incorporated as an important element in information retrieval and response generation.

[0510] The emotion recognition engine analyzes the user's tone of voice and speaking patterns, classifying their emotions into categories such as positive, negative, and neutral. For example, when a user says, "I'm a little tired," the engine can sense fatigue not only from the content of the statement but also from the tone of voice, resulting in the recognition of a negative emotion.

[0511] Recognized emotions are taken into consideration when generating responses, enabling responses in an emotionally appropriate tone, such as incorporating words of encouragement. By not only providing information but also responding in a way that is sensitive to the user's emotions, more natural and approachable communication is achieved.

[0512] Ultimately, the terminal receives the generated response and presents it to the user. This presentation is done through audio output or a visual display, and by using special displays and sounds that respond to emotions, the user can experience deeper satisfaction. This makes the system both technically advanced and emotionally perceptually valuable.

[0513] The following describes the processing flow.

[0514] Step 1:

[0515] Users use the device's voice input function to ask questions or make requests by voice. The device captures the voice and temporarily stores it as digital data.

[0516] Step 2:

[0517] The device sends the saved audio data to the server. The server uses a speech recognition API to convert the audio data into text.

[0518] Step 3:

[0519] The server passes the converted text to the analysis engine, which then performs the analysis. During this analysis, the emotion recognition engine identifies the user's emotions from the text and audio features. For example, it determines whether the user's voice sounds calm or anxious.

[0520] Step 4:

[0521] The server collects relevant information in response to user requests based on analysis results and sentiment recognition data. It utilizes databases and external web services to obtain the necessary information.

[0522] Step 5:

[0523] The server generates a response based on the acquired information and the results of sentiment analysis. During response generation, the tone and content are adjusted specifically according to the user's emotions; for example, an encouraging message might be included for a depressed user.

[0524] Step 6:

[0525] The server sends the generated response data to the terminal. The terminal presents this data to the user. This presentation is done through text display on a visual display or speech synthesis. Specific examples include displaying text in soft colors on the screen and outputting speech in a gentle voice.

[0526] Through the above process, the system goes beyond simply providing information and delivers responses that are sensitive to the user's emotions.

[0527] (Example 2)

[0528] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] In voice-input information systems, conventional technologies have struggled to generate responses that take user emotions into account, limiting the ability to achieve natural and user-friendly communication. To address this challenge, there is a need for technology that can accurately identify user emotions and generate appropriate responses based on them.

[0530] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0531] In this invention, the server includes an input means for acquiring an audio signal, a conversion means for converting the audio signal into a written expression, and an analysis means for analyzing the written expression to identify the user's emotions. This enables the natural generation of responses that take the user's emotions into consideration, resulting in a deeper communication experience.

[0532] An "audio signal" is a signal used to electronically transmit information in the form of sound.

[0533] "Input means" refers to a device or technology that has the function of acquiring a user's voice signal and receiving it for subsequent processing.

[0534] "Conversion means" refers to the process or technology of converting audio signals into written expressions.

[0535] "Textual representation" refers to data in text format converted from speech, which is used for analysis and response generation.

[0536] "Analysis means" refers to a technology or process for processing written expressions and identifying the user's emotions and intentions.

[0537] "Generative means" refer to technologies and functions that produce appropriate responses based on identified emotions or intentions.

[0538] "Presentation means" refers to technologies or methods that show the generated response to the user visually or audibly.

[0539] A "computational model" is a mathematical or statistical model used to process and analyze information.

[0540] An "external information source" is a database or service that exists outside the system and is accessed to retrieve relevant information.

[0541] This invention is a system that generates and presents emotionally responsive responses based on voice input from a user. Users can input voice commands using devices such as smartphones or tablets. The terminal acquires this voice signal and transmits it to a server via a network.

[0542] The server converts the audio signal into text data using speech recognition software such as the Google Cloud Speech-to-Text API. The resulting text is then analyzed using an emotion recognition engine. The emotion recognition engine analyzes the user's speech content and tone of voice to classify emotions. This analysis utilizes machine learning models to categorize emotions as positive, negative, or neutral.

[0543] Based on the analysis results, the server generates a response using a generative AI model. This generative AI model includes commonly used natural language processing models. This process can use prompts that take the user's emotions into account, forming a response with an appropriate tone. For example, it might use a prompt such as, "If the user says, 'I'm feeling a little down today,' generate an encouraging message that is empathetic to the user's feelings."

[0544] Finally, the generated response is sent to the terminal and presented to the user via voice output or visual display. Using speech synthesis technology, the response is played back aloud and presented in a way that is appropriate to the user's emotions. This system allows users to receive not only information but also emotional support. As a result, the user experience is technologically advanced and becomes more humane.

[0545] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0546] Step 1:

[0547] Users input voice commands using a smartphone or tablet. The input voice signals are captured by the device's microphone. The device converts these voice signals into digital data and sends it to a server over the network.

[0548] Step 2:

[0549] The server converts audio signals received over the network into text data using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, the input audio signal is analyzed by the speech recognition engine, and a written representation is obtained as output. This conversion uses phonological analysis and statistical models.

[0550] Step 3:

[0551] The server passes the obtained textual expression to the emotion recognition engine, which analyzes the user's utterances and tone of voice. Based on the analysis, the user's emotions are classified into categories such as positive, negative, and neutral. The input is textual expression, and the output is the recognized emotion category. Specifically, the system analyzes keywords within the text and the intensity of the voice.

[0552] Step 4:

[0553] The server uses the analysis results to issue prompts to the generative AI model, generating appropriate responses. In this process, prompt statements serve as input, and the output is a natural, contextualized response that aligns with emotions. The generative AI model performs advanced language generation using, for example, OpenAI's natural language processing tools.

[0554] Step 5:

[0555] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology and a visual display to present the response to the user. The input is the generated text response, and the output is the response appropriately presented to the user. The specific operation of speech synthesis includes processes such as converting text into a speech waveform and outputting it through a speaker.

[0556] Through these processing steps, the system provides the user with an emotionally responsive and interactive experience.

[0557] (Application Example 2)

[0558] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0559] Modern security systems lack the functionality to accurately assess the urgency of voice alerts. This creates a challenge: potentially delays in prompt and appropriate responses during emergencies.

[0560] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0561] In this invention, the server includes a receiving means for acquiring an acoustic signal, a conversion device for converting the acoustic signal into a text representation, and an emotion recognition device for identifying the emotional state from the user's voice and determining the degree of urgency. This enables immediate determination of the degree of urgency based on the user's voice input, allowing for a quick and appropriate response.

[0562] An "acoustic signal" is an electrical signal obtained by converting the waveform of sound traveling through space into an electrical signal, and is treated as audio data.

[0563] A "receiving means" is a device for acquiring acoustic signals, and its role is to convert audio input into electrical signals and deliver them to the system.

[0564] A "conversion device" is a device that converts acoustic signals into written representations, and has the function of converting audio data into text data.

[0565] An "analysis device" is a device that analyzes text data to identify the user's intent and is used to understand the intent behind the information.

[0566] An "information acquisition device" is a device for acquiring relevant information based on the user's intent, and is equipped with the function of acquiring necessary information from external resources.

[0567] A "response generation device" is a device that generates a response based on acquired information, and is intended to form an appropriate response for the user.

[0568] A "presentation means" is a means of presenting a generated response to the user, and has the role of displaying or playing information in the form of sound or visual means.

[0569] An "emotion recognition device" is a device that identifies an emotional state from a user's voice and determines the degree of urgency, playing a role in estimating emotions from the tone and content of the voice.

[0570] The system that implements this application generates appropriate responses based on user voice input and takes emergency action as needed. The system operates using the following hardware and software:

[0571] First, the user's voice is acquired by the smartphone's receiving device and processed as an acoustic signal. The acquired acoustic signal is then converted into text through a conversion device. It is common to use speech recognition APIs such as Google Cloud Speech-to-Text for speech recognition.

[0572] Next, the server uses an analysis device to analyze the text representation and understand the user's intent. At this stage, natural language processing techniques are applied to identify specific requests from the user's speech. The analysis results are sent to an information retrieval device that accesses external information sources and network services to obtain relevant information. For example, an API is used to retrieve necessary information from the web.

[0573] Subsequently, the server uses a response generation device to generate an appropriate response to the user based on the acquired information. During this process, an emotion recognition device is used to identify the user's emotional state and adjust the tone and urgency of the response accordingly. Emotion recognition engines such as Aurora and IBM Watson are often utilized.

[0574] Finally, the response is sent to the user through a presentation method. On smartphones, information is provided via voice output or a visual interface. If the situation is deemed urgent, the security company and close family members are immediately notified.

[0575] For example, if a user speaks into their smartphone at night saying, "There might be an intruder," and their tone of voice indicates anxiety, the system will immediately detect negative emotions and initiate an emergency call.

[0576] An example of a prompt message might be something like, "We will use our emotion recognition engine to classify the user's emotions from their voice and provide immediate notification if it is an emergency."

[0577] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0578] Step 1:

[0579] The device acquires an acoustic signal.

[0580] When a user speaks into their smartphone, the device's microphone receives the sound as an acoustic signal. This acoustic signal is then acquired as input data.

[0581] Step 2:

[0582] The device converts the audio signal into text.

[0583] The device converts the acquired acoustic signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). It analyzes the acoustic signal and generates the corresponding text as output.

[0584] Step 3:

[0585] The server analyzes the text data to identify the user's intent.

[0586] The server receives text data sent from the terminal and analyzes it using natural language processing technology. The user's intent, identified as a result of the text analysis, is output, and this intent forms the basis for generating related responses and retrieving information.

[0587] Step 4:

[0588] The server uses an emotion recognition device to identify the user's emotional state.

[0589] The analyzed text data and voice tone are then analyzed by an emotion recognition engine on the server (e.g., Aurora or IBM Watson). This process outputs emotion data classified as positive, negative, or neutral, which is used as an important element in response generation.

[0590] Step 5:

[0591] The server retrieves the relevant information.

[0592] Based on the user's intent, the system retrieves necessary information by accessing external information sources and network services. The server collects relevant information using an information retrieval API and uses that information as input for the next response generation step.

[0593] Step 6:

[0594] The server generates a response.

[0595] Based on the acquired relevant information and emotional state, the server constructs the optimal response using a response generator. The generated response becomes the output, and is adjusted according to the tone of voice and urgency.

[0596] Step 7:

[0597] The terminal presents the generated response to the user.

[0598] The terminal, upon receiving the response from the server, presents it via audio output or visually on its display. In emergencies, it can also contact security services or notify family members. This step provides the user with final information.

[0599] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0600] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0601] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0602] [Fourth Embodiment]

[0603] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0604] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0605] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0606] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0607] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0608] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0609] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0610] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0611] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0612] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0613] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0614] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0615] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0616] The system of the present invention appropriately acquires the user's voice input, efficiently converts it into text to analyze the user's intent, acquires necessary information from external sources based on that intent, and ultimately generates an appropriate response for the user.

[0617] Users access the system using devices such as smartphones and tablets. These devices are equipped with voice input capabilities, which are used to capture the user's voice as a digital signal. The captured voice signal is then sent from the device to the server via the network.

[0618] The server converts audio data into text using a speech recognition API. This speech recognition process utilizes advanced language models to analyze speech patterns and convert them into human-readable text. This converted text is then used for subsequent analysis.

[0619] The analysis engine utilizes natural language processing techniques to extract intent from the user's utterances. For example, if a user says, "Tell me the weather for tomorrow," the analysis engine identifies the need to search for weather information. Based on this information request, the server retrieves the necessary data from external databases and web services.

[0620] The acquired information is processed by a response generation module, which uses commonly used text generation and speech synthesis technologies to construct a response for the user. For example, if it's weather information, it will generate specific information such as, "Tomorrow in Tokyo it will be sunny, with a high of 25 degrees Celsius."

[0621] The final generated response is returned to the terminal. The terminal receives this response data and displays it as text on the screen or plays it back as audio through the speaker, allowing the user to intuitively obtain the requested information. This enables the user to quickly obtain information and achieve their goals.

[0622] The following describes the processing flow.

[0623] Step 1:

[0624] The user activates the device's voice input function and speaks their question or request. The device captures the voice using the microphone and temporarily stores it as digital data.

[0625] Step 2:

[0626] The terminal converts the acquired audio signal into a format that is easy to process and sends this audio data to the server. The audio data is delivered via the network to the API endpoint specified by the server.

[0627] Step 3:

[0628] The server passes the received audio data to a speech recognition API, which converts the audio into text. This conversion process uses a language model to analyze the audio signal and generate corresponding text information.

[0629] Step 4:

[0630] The server inputs text data into a natural language processing engine to analyze the intent of the user's request. This analysis process uses techniques such as keyword extraction and sentiment analysis to identify the information the user is looking for.

[0631] Step 5:

[0632] Based on the analysis results, the server collects the information the user needs from the internet and external databases. For example, it might call web services such as weather information APIs to obtain relevant information.

[0633] Step 6:

[0634] The server uses the collected information to generate a response to present to the user. The response generation module constructs the information in a natural way, utilizing text and speech synthesis.

[0635] Step 7:

[0636] The server sends the generated response to the terminal. The terminal presents this response data to the user using methods such as audio output or display. The user obtains the answer to their question through the presented information.

[0637] (Example 1)

[0638] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0639] In recent years, voice assistants and natural language interfaces have become widely used, but they suffer from problems such as low voice recognition accuracy, inability to accurately analyze intent from speech, and unnatural responses to users. These problems hinder the improvement of information retrieval and efficiency in daily life using voice interfaces. The present invention aims to address these problems and provide a voice interface that enables more natural, rapid, and accurate information retrieval and response generation.

[0640] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0641] In this invention, the server includes a conversion device for converting voice signals into text information, an analysis device for analyzing the text information and extracting intent, and an information acquisition device for obtaining relevant information based on the intent. This enables the user to efficiently acquire information via voice and obtain appropriate and natural responses.

[0642] An "audio signal" is data that represents sound in digital or analog format.

[0643] An "input device" is a device that receives voice information from a user and acquires it as a digital signal.

[0644] A "conversion device" is a device used to convert acquired audio signals into text information.

[0645] "Textual information" refers to data obtained by converting audio signals into a text format that can be analyzed.

[0646] An "analysis device" is a device that analyzes textual information and extracts the user's intended meaning.

[0647] "Intent" refers to a specific request for information or a directive that can be inferred from the user's statements.

[0648] An "information acquisition device" is a device that acquires relevant information from an external source based on the user's intent.

[0649] A "generation device" is a device that creates an appropriate response based on acquired information.

[0650] A "display device" is a device that visually presents the generated response to the user as text.

[0651] A "sound playback device" is a device that presents the generated response to the user audibly as sound.

[0652] This invention provides a system that offers an interface for users to acquire information via voice and receive appropriate responses. The user uses a device such as a smartphone or tablet to acquire voice signals using a voice input device. The device converts these voice signals into a digital format and transmits them to a server via a network.

[0653] The server converts the audio signal into text information using a speech recognition API (for example, using general speech recognition technology) to analyze the audio signal. This converted text information is then processed by an analysis device using natural language processing technology to extract the user's intent. The analysis device uses generative AI model technology (for example, a general natural language processing engine) to identify the user's intent.

[0654] Based on the identified user intent, the information acquisition device accesses external information sources and retrieves relevant information. These external information sources could include various information service providers and databases. The acquired information is processed by a generator on the server and converted into a natural language response to the user. Here, generative AI modeling technology is used to improve the naturalness and accuracy of the response.

[0655] Finally, the generated response is sent to the terminal, which then presents this response to the user. Using a display device or audio playback device, intuitive feedback can be provided to the user by displaying it as text or playing it aloud.

[0656] As a concrete example, a user might use voice input via their device, saying, "Tell me the weather for tomorrow." In this case, the server converts the voice signal into text, analyzes the intent of "Tell me the weather for tomorrow," and retrieves weather information by referring to an appropriate external database. The generated response would then be provided in the form of, "Tomorrow will be sunny, with a high of 25 degrees Celsius."

[0657] Examples of prompt phrases include "Voice input: Tell me the weather tomorrow" and "Voice input: Tell me the latest news."

[0658] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0659] Step 1:

[0660] The user inputs an audio signal via the terminal's audio input device. The terminal converts this audio into a digital signal and transmits it to the server via the internet. The input is raw audio, and the output is encoded digital data.

[0661] Step 2:

[0662] The server passes the received digital audio signal to a speech recognition API, which converts the audio into text. In this process, the API analyzes the frequency pattern of the audio and generates the corresponding text. The input is a digital audio signal, and the output is text data.

[0663] Step 3:

[0664] The analysis device on the server analyzes the acquired text data using natural language processing techniques to extract the user's intent. Here, it understands keywords and context within the text to identify the information or commands the user is seeking. The input is text data, and the output is a data structure representing the intent.

[0665] Step 4:

[0666] The server's information retrieval device accesses external information sources based on the user's intent and retrieves the necessary data. Specifically, it sends requests to appropriate APIs or databases and receives the information. At this stage, the input is the user's intent data structure, and the output is the retrieved information data.

[0667] Step 5:

[0668] The server's generation device generates natural and easy-to-understand responses based on acquired information data. This process utilizes a generative AI model to construct text, and in some cases, also performs speech synthesis. The input is information data, and the output is response text or audio data.

[0669] Step 6:

[0670] The server sends the final generated response to the terminal. The terminal receives this response and presents it to the user via a display device or audio playback device. For example, information is provided to the user by displaying text on the screen or playing audio through a speaker. In this step, the input is the response data, and the output is the visual or auditory feedback to the user.

[0671] (Application Example 1)

[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0673] Using voice commands to acquire information within a mobile vehicle allows users to concentrate on driving while quickly and accurately obtaining necessary information. However, conventional technologies have faced challenges in ensuring sufficient driver safety and comfort, such as low voice recognition accuracy and the failure to provide necessary information in a timely manner. Furthermore, delays in acquiring information from external sources have also been a factor that detracts from the user experience.

[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0675] In this invention, the server includes an input device means for acquiring voice information, a conversion device means for converting the voice information into a string, an analysis device means for analyzing the string and extracting the user's request, a data acquisition device means for acquiring relevant information based on the user's request, and a response generation device means for generating a response based on the acquired relevant information. This makes it possible to recognize voice instructions with high accuracy and to quickly acquire and provide necessary information from external information sources.

[0676] An "input device for acquiring voice information" is a device that captures voice signals from a user and inputs them into the system as digital signals.

[0677] "The conversion device for converting the aforementioned audio information into a string" refers to a device that has the function of analyzing the acquired audio signal and converting it into a corresponding string format.

[0678] "The analysis device for analyzing the aforementioned string and extracting the user's request" is a device for identifying the user's intent from the converted string and determining the necessary actions.

[0679] "Data acquisition device for obtaining relevant information based on the user's request" refers to a device that has the function of obtaining relevant information corresponding to the user's request from external information sources such as the Internet.

[0680] A "response generation device for generating responses based on acquired relevant information" is a device that generates appropriate responses to present to the user based on information obtained from a data acquisition device.

[0681] An "in-vehicle voice assistant device with driving assistance functions that recognizes voice and presents information in a moving vehicle" is a device that has a voice assistant function that assists the driver by recognizing voice commands in a moving vehicle and quickly presenting necessary information.

[0682] This invention improves driver safety and convenience by implementing a voice-controlled information acquisition system in a mobile device.

[0683] The server uses an input device to acquire voice information and collects the user's voice commands from a microphone inside the mobile device. Next, it uses a conversion device, specifically a speech recognition API (e.g., Google Cloud Speech-to-Text), to convert the collected voice information into a string.

[0684] This string is analyzed by an analysis device to identify the user's request. This analysis device uses a natural language processing library (e.g., spaCy) to extract the user's intent. After the intent is identified, a data acquisition device retrieves relevant information from an external source (e.g., Google Places API), and based on this, a response generation device generates an appropriate response. The response generation device uses speech synthesis technology (e.g., gTTS) to convert the text response into speech.

[0685] The terminal presents the generated voice response to the user through the in-car speaker. This entire process allows the driver to safely obtain necessary information while on the move.

[0686] For example, if a driver asks, "Where's the next gas station?", the server will collect location information for the nearest gas station and generate a voice response saying, "The next gas station is 2 kilometers away." This allows the driver to obtain the necessary information without taking their eyes off the road.

[0687] As an example of a prompt, we will use the instruction, "When the user asks 'Where is the next gas station?', retrieve the location of the most suitable gas station."

[0688] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0689] Step 1:

[0690] The user speaks aloud while in a moving vehicle. This voice is captured by the device's microphone. The input is the user's voice, and the output is a digital audio signal. The device prepares to send this audio signal to the server.

[0691] Step 2:

[0692] The server processes the received audio signal with a converter and converts it into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). The input is a digital audio signal, and the output is text data. The server analyzes the audio pattern and converts what the user says into text.

[0693] Step 3:

[0694] The server receives the string converted by the parsing device and analyzes it using a natural language processing library (e.g., spaCy). The input is text data, and the output is the analysis result indicating the user's intent. The server identifies the user's request from this analysis result.

[0695] Step 4:

[0696] The server uses a data acquisition device to access external information sources (e.g., Google Places API) based on user requests. The input is the parsed result representing the user's request, and the output is the data containing the relevant information. The server quickly retrieves the necessary information and prepares it for use within the system.

[0697] Step 5:

[0698] The server generates a response to the user based on information obtained using a response generation device. Text is converted to speech using speech synthesis technology (e.g., gTTS). The input is data containing relevant information, and the output is audio data. The server creates voice guidance and prepares it to be clearly communicated to the user.

[0699] Step 6:

[0700] The terminal processes audio data received from the server using a playback device and presents it to the user through a speaker. The input is audio data, and the output is the audio heard by the user. The terminal conveys information intuitively without attracting the driver's attention.

[0701] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0702] The system of the present invention receives voice input from a user, effectively analyzes its content, and generates an appropriate response based on the obtained information. Furthermore, it recognizes the user's emotional state and fine-tunes the content and presentation method of the response to match that emotion, thereby improving the user experience.

[0703] The user inputs voice commands into the system using a mobile device such as a smartphone or tablet. The device captures the user's voice and sends the voice signal to the server. The server uses a speech recognition API to convert the voice to text and then analyzes the resulting text.

[0704] As an analysis tool, the server incorporates an emotion recognition engine that identifies emotions from the user's speech content and tone of voice. This emotion information, along with the results of text analysis, is incorporated as an important element in information retrieval and response generation.

[0705] The emotion recognition engine analyzes the user's tone of voice and speaking patterns, classifying their emotions into categories such as positive, negative, and neutral. For example, when a user says, "I'm a little tired," the engine can sense fatigue not only from the content of the statement but also from the tone of voice, resulting in the recognition of a negative emotion.

[0706] Recognized emotions are taken into consideration when generating responses, enabling responses in an emotionally appropriate tone, such as incorporating words of encouragement. By not only providing information but also responding in a way that is sensitive to the user's emotions, more natural and approachable communication is achieved.

[0707] Ultimately, the terminal receives the generated response and presents it to the user. This presentation is done through audio output or a visual display, and by using special displays and sounds that respond to emotions, the user can experience deeper satisfaction. This makes the system both technically advanced and emotionally perceptually valuable.

[0708] The following describes the processing flow.

[0709] Step 1:

[0710] Users use the device's voice input function to ask questions or make requests by voice. The device captures the voice and temporarily stores it as digital data.

[0711] Step 2:

[0712] The device sends the saved audio data to the server. The server uses a speech recognition API to convert the audio data into text.

[0713] Step 3:

[0714] The server passes the converted text to the analysis engine, which then performs the analysis. During this analysis, the emotion recognition engine identifies the user's emotions from the text and audio features. For example, it determines whether the user's voice sounds calm or anxious.

[0715] Step 4:

[0716] The server collects relevant information in response to user requests based on analysis results and sentiment recognition data. It utilizes databases and external web services to obtain the necessary information.

[0717] Step 5:

[0718] The server generates a response based on the acquired information and the results of sentiment analysis. During response generation, the tone and content are adjusted specifically according to the user's emotions; for example, an encouraging message might be included for a depressed user.

[0719] Step 6:

[0720] The server sends the generated response data to the terminal. The terminal presents this data to the user. This presentation is done through text display on a visual display or speech synthesis. Specific examples include displaying text in soft colors on the screen and outputting speech in a gentle voice.

[0721] Through the above process, the system goes beyond simply providing information and delivers responses that are sensitive to the user's emotions.

[0722] (Example 2)

[0723] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0724] In voice-input information systems, conventional technologies have struggled to generate responses that take user emotions into account, limiting the ability to achieve natural and user-friendly communication. To address this challenge, there is a need for technology that can accurately identify user emotions and generate appropriate responses based on them.

[0725] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0726] In this invention, the server includes an input means for acquiring an audio signal, a conversion means for converting the audio signal into a written expression, and an analysis means for analyzing the written expression to identify the user's emotions. This enables the natural generation of responses that take the user's emotions into consideration, resulting in a deeper communication experience.

[0727] An "audio signal" is a signal used to electronically transmit information in the form of sound.

[0728] "Input means" refers to a device or technology that has the function of acquiring a user's voice signal and receiving it for subsequent processing.

[0729] "Conversion means" refers to the process or technology of converting audio signals into written expressions.

[0730] "Textual representation" refers to data in text format converted from speech, which is used for analysis and response generation.

[0731] "Analysis means" refers to a technology or process for processing written expressions and identifying the user's emotions and intentions.

[0732] "Generative means" refer to technologies and functions that produce appropriate responses based on identified emotions or intentions.

[0733] "Presentation means" refers to technologies or methods that show the generated response to the user visually or audibly.

[0734] A "computational model" is a mathematical or statistical model used to process and analyze information.

[0735] An "external information source" is a database or service that exists outside the system and is accessed to retrieve relevant information.

[0736] This invention is a system that generates and presents emotionally responsive responses based on voice input from a user. Users can input voice commands using devices such as smartphones or tablets. The terminal acquires this voice signal and transmits it to a server via a network.

[0737] The server converts the audio signal into text data using speech recognition software such as the Google Cloud Speech-to-Text API. The resulting text is then analyzed using an emotion recognition engine. The emotion recognition engine analyzes the user's speech content and tone of voice to classify emotions. This analysis utilizes machine learning models to categorize emotions as positive, negative, or neutral.

[0738] Based on the analysis results, the server generates a response using a generative AI model. This generative AI model includes commonly used natural language processing models. This process can use prompts that take the user's emotions into account, forming a response with an appropriate tone. For example, it might use a prompt such as, "If the user says, 'I'm feeling a little down today,' generate an encouraging message that is empathetic to the user's feelings."

[0739] Finally, the generated response is sent to the terminal and presented to the user via voice output or visual display. Using speech synthesis technology, the response is played back aloud and presented in a way that is appropriate to the user's emotions. This system allows users to receive not only information but also emotional support. As a result, the user experience is technologically advanced and becomes more humane.

[0740] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0741] Step 1:

[0742] Users input voice commands using a smartphone or tablet. The input voice signals are captured by the device's microphone. The device converts these voice signals into digital data and sends it to a server over the network.

[0743] Step 2:

[0744] The server converts audio signals received over the network into text data using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, the input audio signal is analyzed by the speech recognition engine, and a written representation is obtained as output. This conversion uses phonological analysis and statistical models.

[0745] Step 3:

[0746] The server passes the obtained textual expression to the emotion recognition engine, which analyzes the user's utterances and tone of voice. Based on the analysis, the user's emotions are classified into categories such as positive, negative, and neutral. The input is textual expression, and the output is the recognized emotion category. Specifically, the system analyzes keywords within the text and the intensity of the voice.

[0747] Step 4:

[0748] The server uses the analysis results to issue prompts to the generative AI model, generating appropriate responses. In this process, prompt statements serve as input, and the output is a natural, contextualized response that aligns with emotions. The generative AI model performs advanced language generation using, for example, OpenAI's natural language processing tools.

[0749] Step 5:

[0750] The generated response is sent from the server to the terminal. The terminal uses speech synthesis technology and a visual display to present the response to the user. The input is the generated text response, and the output is the response appropriately presented to the user. The specific operation of speech synthesis includes processes such as converting text into a speech waveform and outputting it through a speaker.

[0751] Through these processing steps, the system provides the user with an emotionally responsive and interactive experience.

[0752] (Application Example 2)

[0753] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0754] Modern security systems lack the functionality to accurately assess the urgency of voice alerts. This creates a challenge: potentially delays in prompt and appropriate responses during emergencies.

[0755] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0756] In this invention, the server includes a receiving means for acquiring an acoustic signal, a conversion device for converting the acoustic signal into a text representation, and an emotion recognition device for identifying the emotional state from the user's voice and determining the degree of urgency. This enables immediate determination of the degree of urgency based on the user's voice input, allowing for a quick and appropriate response.

[0757] An "acoustic signal" is an electrical signal obtained by converting the waveform of sound traveling through space into an electrical signal, and is treated as audio data.

[0758] A "receiving means" is a device for acquiring acoustic signals, and its role is to convert audio input into electrical signals and deliver them to the system.

[0759] A "conversion device" is a device that converts acoustic signals into written representations, and has the function of converting audio data into text data.

[0760] An "analysis device" is a device that analyzes text data to identify the user's intent and is used to understand the intent behind the information.

[0761] An "information acquisition device" is a device for acquiring relevant information based on the user's intent, and is equipped with the function of acquiring necessary information from external resources.

[0762] A "response generation device" is a device that generates a response based on acquired information, and is intended to form an appropriate response for the user.

[0763] A "presentation means" is a means of presenting a generated response to the user, and has the role of displaying or playing information in the form of sound or visual means.

[0764] An "emotion recognition device" is a device that identifies an emotional state from a user's voice and determines the degree of urgency, playing a role in estimating emotions from the tone and content of the voice.

[0765] The system that implements this application generates appropriate responses based on user voice input and takes emergency action as needed. The system operates using the following hardware and software:

[0766] First, the user's voice is acquired by the smartphone's receiving device and processed as an acoustic signal. The acquired acoustic signal is then converted into text through a conversion device. It is common to use speech recognition APIs such as Google Cloud Speech-to-Text for speech recognition.

[0767] Next, the server uses an analysis device to analyze the text representation and understand the user's intent. At this stage, natural language processing techniques are applied to identify specific requests from the user's speech. The analysis results are sent to an information retrieval device that accesses external information sources and network services to obtain relevant information. For example, an API is used to retrieve necessary information from the web.

[0768] Subsequently, the server uses a response generation device to generate an appropriate response to the user based on the acquired information. During this process, an emotion recognition device is used to identify the user's emotional state and adjust the tone and urgency of the response accordingly. Emotion recognition engines such as Aurora and IBM Watson are often utilized.

[0769] Finally, the response is sent to the user through a presentation method. On smartphones, information is provided via voice output or a visual interface. If the situation is deemed urgent, the security company and close family members are immediately notified.

[0770] For example, if a user speaks into their smartphone at night saying, "There might be an intruder," and their tone of voice indicates anxiety, the system will immediately detect negative emotions and initiate an emergency call.

[0771] An example of a prompt message might be something like, "We will use our emotion recognition engine to classify the user's emotions from their voice and provide immediate notification if it is an emergency."

[0772] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0773] Step 1:

[0774] The device acquires an acoustic signal.

[0775] When a user speaks into their smartphone, the device's microphone receives the sound as an acoustic signal. This acoustic signal is then acquired as input data.

[0776] Step 2:

[0777] The device converts the audio signal into text.

[0778] The device converts the acquired acoustic signal into text data using a speech recognition API (e.g., Google Cloud Speech-to-Text). It analyzes the acoustic signal and generates the corresponding text as output.

[0779] Step 3:

[0780] The server analyzes the text data to identify the user's intent.

[0781] The server receives text data sent from the terminal and analyzes it using natural language processing technology. The user's intent, identified as a result of the text analysis, is output, and this intent forms the basis for generating related responses and retrieving information.

[0782] Step 4:

[0783] The server uses an emotion recognition device to identify the user's emotional state.

[0784] The analyzed text data and voice tone are then analyzed by an emotion recognition engine on the server (e.g., Aurora or IBM Watson). This process outputs emotion data classified as positive, negative, or neutral, which is used as an important element in response generation.

[0785] Step 5:

[0786] The server retrieves the relevant information.

[0787] Based on the user's intent, the system retrieves necessary information by accessing external information sources and network services. The server collects relevant information using an information retrieval API and uses that information as input for the next response generation step.

[0788] Step 6:

[0789] The server generates a response.

[0790] Based on the acquired relevant information and emotional state, the server constructs the optimal response using a response generator. The generated response becomes the output, and is adjusted according to the tone of voice and urgency.

[0791] Step 7:

[0792] The terminal presents the generated response to the user.

[0793] The terminal, upon receiving the response from the server, presents it via audio output or visually on its display. In emergencies, it can also contact security services or notify family members. This step provides the user with final information.

[0794] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0795] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0796] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0797] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0798] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0799] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0800] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0801] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0802] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0803] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0804] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0805] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0806] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0807] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0808] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0809] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0810] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0811] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0812] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0813] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0814] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0815] The following is further disclosed regarding the embodiments described above.

[0816] (Claim 1)

[0817] Audio input means for acquiring audio signals,

[0818] A conversion means for converting the aforementioned audio signal into a text representation,

[0819] An analysis means for analyzing the aforementioned text expression to identify user intent,

[0820] Information acquisition means for acquiring relevant information based on the user intent,

[0821] A response generation means that generates a response based on the acquired information,

[0822] A system including means for presenting the aforementioned response to the user.

[0823] (Claim 2)

[0824] The system according to claim 1, wherein the presentation means has the function of presenting the response to the user as an audio playback.

[0825] (Claim 3)

[0826] The system according to claim 1, wherein the information acquisition means has a function to access an external database or web service to acquire relevant information.

[0827] "Example 1"

[0828] (Claim 1)

[0829] An input device for acquiring audio signals,

[0830] A conversion device for converting the aforementioned audio signal into text information,

[0831] An analysis device for analyzing the aforementioned textual information and extracting intent,

[0832] An information acquisition device for acquiring relevant information based on the aforementioned intent,

[0833] A generation device for generating a response based on acquired information,

[0834] A system including a display device or audio playback device for presenting the aforementioned response.

[0835] (Claim 2)

[0836] The system according to claim 1, wherein the presentation device has a function of reproducing the response as sound.

[0837] (Claim 3)

[0838] The system according to claim 1, wherein the information acquisition device has a function to access an external information source and acquire relevant information.

[0839] "Application Example 1"

[0840] (Claim 1)

[0841] An input device means for acquiring audio information,

[0842] A conversion device means for converting the aforementioned audio information into a string,

[0843] An analysis device means for analyzing the aforementioned string and extracting the user's request,

[0844] A data acquisition device means for obtaining relevant information based on the user's request,

[0845] A response generation device means for generating a response based on acquired related information,

[0846] A presentation device means for conveying the aforementioned response to the user,

[0847] An in-vehicle voice assistant device means that has a driving assistance function and recognizes voice in a moving vehicle to present information,

[0848] A system that includes this.

[0849] (Claim 2)

[0850] The system according to claim 1, wherein the presentation device has a function of transmitting the response to the user as voice.

[0851] (Claim 3)

[0852] The system according to claim 1, wherein the data acquisition device has a function to access an external storage device or a web information source to acquire relevant information.

[0853] "Example 2 of combining an emotion engine"

[0854] (Claim 1)

[0855] An input means for acquiring audio signals,

[0856] A conversion means for converting the aforementioned audio signal into written text,

[0857] An analytical means for analyzing the aforementioned textual expression to identify the user's emotions,

[0858] Means for using a computational model that includes generation means for generating an appropriate response based on the aforementioned emotion,

[0859] A system including a presentation means that includes a method for presenting the aforementioned response to a user.

[0860] (Claim 2)

[0861] The system according to claim 1, wherein the presentation means has the function of presenting the response to the user as an audio playback.

[0862] (Claim 3)

[0863] The system according to claim 1, wherein the analysis means has a function to access an external information source and obtain relevant information based on the user's emotions.

[0864] "Application example 2 when combining with an emotional engine"

[0865] (Claim 1)

[0866] A receiving means for acquiring an acoustic signal,

[0867] A conversion device that converts the aforementioned acoustic signal into a character representation,

[0868] An analysis device that analyzes the aforementioned textual expression and identifies the user's intent,

[0869] An information acquisition device that acquires relevant information based on the user's intent,

[0870] A response generation device that generates a response based on acquired information,

[0871] A means for presenting the aforementioned response to the user,

[0872] A system including an emotion recognition device that identifies the emotional state from the user's voice and determines the degree of urgency.

[0873] (Claim 2)

[0874] The system according to claim 1, wherein the presentation means has a function of presenting the response to the user as an audio output and determining a response according to the urgency.

[0875] (Claim 3)

[0876] The system according to claim 1, wherein the information acquisition device has a function to access an external information source or network service to acquire relevant information and, in the case of an emergency, to notify the responding agency. [Explanation of Symbols]

[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Audio input means for acquiring audio signals, A conversion means for converting the aforementioned audio signal into a text representation, An analysis means for analyzing the aforementioned text expression to identify user intent, Information acquisition means for acquiring relevant information based on the user intent, A response generation means that generates a response based on the acquired information, A system including means for presenting the aforementioned response to the user.

2. The system according to claim 1, wherein the presentation means has the function of presenting the response to the user as audio playback.

3. The system according to claim 1, wherein the information acquisition means has a function to access an external database or web service to acquire relevant information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A