system
The system addresses the complexity of existing voice recognition systems by integrating voice and visual feedback for intuitive schedule management and information retrieval, improving user experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Existing voice recognition systems and calendar-linked applications require complex operations and lack integrated text and voice interaction, making them difficult for users to efficiently manage schedules and obtain information.
A system that integrates voice recognition, natural language processing, speech synthesis, and calendar integration, allowing users to interact intuitively through voice and visual feedback, providing schedule notifications and information.
Enables users to easily manage schedules and obtain information through natural voice dialogue and visual feedback without complex operations, enhancing user convenience and satisfaction.
Smart Images

Figure 2026038090000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In everyday life and business, users need support to efficiently obtain information and manage their schedules. However, many existing voice recognition systems and calendar-linked applications require complex operations, making them difficult for users to use. Furthermore, these systems have incomplete integration of text and voice, making it difficult to provide the experience users desire. The present invention aims to solve these problems and provide a desktop secretary system that users can operate easily and intuitively. [Means for solving the problem]
[0005] The present invention provides a system that solves the above-mentioned problems through the following means.
[0006] First, it has a means for recognizing the user's voice and generating voice data. A standard microphone can be used for this. Next, it has a means for converting the generated voice data into text data, using a voice recognition API for this conversion. It also includes a means for analyzing the text data and generating answer data, using natural language processing technology for this analysis. It also has a means for outputting the generated answer data as voice data, using voice synthesis technology. Finally, it has a means for acquiring calendar data and generating and outputting alerts containing schedule information. This calendar integration allows users to receive timely schedule notifications.
[0007] Furthermore, it also includes a means for visualizing the above means on a display device, thereby realizing a system that can provide information intuitively to users. By combining these means, users can not only ask questions by voice, but also receive schedule notifications by voice and check information visually, greatly improving convenience.
[0008] A "user" is a person who operates the system and provides voice input.
[0009] "Voice Data" means data that is recorded and stored in digital form based on a user's voice input.
[0010] "Text data" is data in the form of a character string converted by speech recognition.
[0011] The "analysis means" is a means for interpreting text data and generating appropriate answer data based on the user's intentions.
[0012] "Answer data" is text data that expresses an answer to a user's question.
[0013] The "voice synthesis means" is a means for converting text data into voice data and outputting the voice to the user.
[0014] "Alert" is a function that provides notifications or warnings to the user by voice or visually.
[0015] "Calendar data" is data that includes information about schedules and appointments.
[0016] "Schedule information" is detailed information about a particular event contained within the calendar data.
[0017] A "display device" is a hardware device for providing visual information to a user.
[0018] A "voice recognition API" is a program interface for converting voice data into text data.
[0019] "Natural language processing technology" is a computer technology for understanding and analyzing human language.
[0020] "Speech synthesis technology" is a technology for converting text data into natural speech.
[0021] "Interactive" means having the property of being able to respond intuitively to user operations. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] The present invention aims to realize a desktop secretary system that can intuitively interact with a user, provide information, and manage schedules. The program processing of this system will be explained below.
[0044] The user speaks
[0045] The user asks questions and gives instructions to the device in natural language. For example, say, "What's the weather like tomorrow?"
[0046] Voice Recognition
[0047] The device collects the user's voice using a built-in microphone, saves it as voice data, and then sends the voice data to a voice recognition API (e.g., a publicly available voice recognition API) to convert the voice data into text data.
[0048] Text Analysis
[0049] The device sends the converted text data to a natural language processing API (e.g., a publicly available natural language processing API) to analyze the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, based on the analysis result of "What's the weather going to be like tomorrow?", it generates answer data such as "Tomorrow's weather will be sunny."
[0050] Audio response
[0051] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API) to convert the text data into voice data. The converted voice data is played back to the user through the speaker. As a result, the device responds to the user by voice, saying, "Tomorrow's weather will be sunny."
[0052] Calendar import and integration
[0053] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. The calendar data contains events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The schedule information is periodically checked, and an alert message is generated when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API and notified to the user through the speaker. This causes the device to alert the user by voice, saying, "You have a meeting in 15 minutes."
[0054] Visualization by display device
[0055] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move, blink, or other animations may be displayed when the user speaks or responds. This provides a more intuitive and enjoyable operating experience for the user.
[0056] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, calendar integration, and visual feedback. This allows the user to effectively obtain information and manage their schedule through natural voice dialogue and intuitive interaction without the need for complex operations.
[0057] The processing flow will be explained below.
[0058] Step 1:
[0059] The user speaks to the device, for example, "What's the weather like tomorrow?"
[0060] Step 2:
[0061] The device collects the user's voice using a built-in microphone and saves it as voice data.
[0062] Step 3:
[0063] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[0064] Step 4:
[0065] The server receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends it to the natural language processing API.
[0066] Step 5:
[0067] A natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[0068] Step 6:
[0069] The server receives the generated answer data and transmits it to the terminal.
[0070] Step 7:
[0071] The device sends the answer data to the speech synthesis API, which converts the answer data into voice data.
[0072] Step 8:
[0073] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[0074] Step 9:
[0075] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[0076] Step 10:
[0077] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[0078] Step 11:
[0079] The server sends the generated alert information to the terminal.
[0080] Step 12:
[0081] The device sends the alert information to a speech synthesis API, which converts it into voice data.
[0082] Step 13:
[0083] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[0084] Step 14:
[0085] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[0086] Step 15:
[0087] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] In modern society, there is a demand for systems that users can operate intuitively and that allow them to acquire information and manage schedules without requiring complex operations. However, existing systems are complex to operate, and many users are unable to master them. In addition, they lack an intuitive operating experience and visual feedback through user interaction, making them insufficient to improve user satisfaction.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for analyzing the text data and generating answer data, means for outputting the generated answer data as voice data, means for acquiring calendar data and generating and outputting an alert including schedule information, means for visualizing the execution of the above means on a display device, and means for generating and displaying dynamic character animation to provide an intuitive operating experience through dialogue with the user. This allows the user to easily acquire information and manage their schedule through natural voice dialogue, and further enjoy an intuitive operating experience through visual feedback.
[0093] "User" refers to a person who uses this system to obtain information by voice and perform operations.
[0094] "Voice data" refers to data that is a digital recording of a user's speech.
[0095] "Text data" refers to data in the form of a character string converted from audio data.
[0096] "Answer data" refers to data that contains answers to questions and instructions from the user, which are obtained by analyzing text data.
[0097] The term "voice input device" refers to a device for collecting a user's voice and recording it as voice data.
[0098] "Speech conversion API" refers to an application program interface for converting voice data into text data.
[0099] "Calendar data" refers to data containing a user's schedule information.
[0100] An "alert" refers to a notification that notifies a user that a scheduled time is approaching.
[0101] "Display device" refers to a device that visually presents the execution status of the system and the contents of the dialogue.
[0102] "Dynamic character animation" refers to a moving character that is displayed on a display device to visually represent interaction with a user.
[0103] The present invention relates to a desktop secretary system that intuitively interacts with users, provides information, and manages schedules. This system integrates speech recognition, natural language processing, and speech synthesis technologies to provide users with information efficiently and effectively. Specific embodiments of this system are described below.
[0104] The user speaks
[0105] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather going to be like tomorrow?" This speech is collected through the device's microphone.
[0106] Voice Recognition
[0107] The device collects the user's voice using a built-in microphone and saves it as voice data. The voice data is then sent to an existing voice recognition API (e.g., a publicly available voice recognition API) and converted into text data. Specifically, the voice recognition API used is something like the Google (registered trademark) Speech-to-Text API.
[0108] Text Analysis
[0109] The device sends the converted text data to a natural language processing API (for example, a publicly available natural language processing API) and analyzes the intent of the user's question or instruction. Appropriate answer data is generated based on the analysis results. For example, if the question is "What's the weather like tomorrow?", the answer data generated based on the analysis results will be "Tomorrow's weather will be sunny." An example of a natural language processing API that can be used is "Google Dialogflow."
[0110] Audio response
[0111] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API), which converts the text data into voice data. The converted voice data is played back to the user through the speaker. This causes the device to respond to the user by voice, saying, "Tomorrow's weather will be sunny." The speech synthesis API used is something like the "Google Text-to-Speech API."
[0112] Calendar import and integration
[0113] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. Calendar data includes events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The device periodically checks the schedule information, and generates an alert message when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API, and notified to the user through the speaker. For example, an alert message saying "There is a meeting in 15 minutes" may be played.
[0114] Visualization by display device
[0115] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move or its eyes may blink when the user speaks or responds. This provides the user with a more intuitive and enjoyable operating experience.
[0116] Specific examples
[0117] For example, consider the following prompt:
[0118] "Look at my calendar, what's next?"
[0119] "Please tell me what the weather will be like tomorrow."
[0120] "I have a meeting at 3 PM. Set a reminder."
[0121] As a result, the system of the present invention not only allows users to obtain information and manage their schedules through natural voice dialogue, but also allows them to enjoy an intuitive operating experience through visual feedback.
[0122] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0123] Step 1:
[0124] The user speaks questions or commands.
[0125] Specific behavior:
[0126] The user speaks to the terminal, "Please tell me what the weather will be like tomorrow."
[0127] Input: User voice input
[0128] Output: Audio data collected by the device's microphone
[0129] Step 2:
[0130] The device stores the voice data and sends it to the voice recognition API.
[0131] Specific behavior:
[0132] The device's microphone records the user's speech as voice data, temporarily stores it in memory, and then sends the voice data to an existing speech recognition API (e.g., Google Speech-to-Text API).
[0133] Input: Audio data
[0134] Output: Audio data sent to the API
[0135] Step 3:
[0136] The server converts the voice data into text data and sends it back to the terminal.
[0137] Specific behavior:
[0138] The speech recognition API analyzes the voice data and converts it into text data such as "Please tell me what the weather will be like tomorrow." The text data is then sent from the server to the device.
[0139] Input: Audio data
[0140] Output: Text data
[0141] Step 4:
[0142] The device sends the text data to a natural language processing API.
[0143] Specific behavior:
[0144] The device sends the received text data to a natural language processing API (e.g., Google Dialogflow).
[0145] Input: Text data
[0146] Output: Text data sent to the API
[0147] Step 5:
[0148] The server analyzes the text data and generates answer data.
[0149] Specific behavior:
[0150] The natural language processing API analyzes the text data, understands the intent, and generates the answer data, "The weather will be sunny tomorrow."
[0151] Input: Text data
[0152] Output: Answer data
[0153] Step 6:
[0154] The device sends the answer data to the speech synthesis API and outputs it as voice data.
[0155] Specific behavior:
[0156] The device sends the generated answer data to a speech synthesis API (for example, Google Text-to-Speech API), which generates speech data saying, "Tomorrow's weather will be sunny."
[0157] Input: Answer data
[0158] Output: Audio data
[0159] Step 7:
[0160] The terminal plays the audio data to the user through the speaker.
[0161] Specific behavior:
[0162] The terminal receives the voice data and responds to the user through the speaker by voice, "Tomorrow's weather will be sunny."
[0163] Input: Audio data
[0164] Output: The audio response that is played to the user
[0165] Step 8:
[0166] The device periodically calls the calendar API to obtain the user's schedule data.
[0167] Specific behavior:
[0168] The device accesses a calendar API (for example, Google Calendar API) to obtain the user's schedule data. For example, it obtains information such as "Meeting tomorrow at 10:00 AM."
[0169] Input: None (periodic processing)
[0170] Output: Retrieved calendar data
[0171] Step 9:
[0172] The device checks the schedule information, and when the scheduled time approaches, it generates an alert message and converts it into voice data.
[0173] Specific behavior:
[0174] The device analyzes the calendar data it has acquired, generates an alert message 15 minutes before the scheduled meeting saying "You have a meeting in 15 minutes," and converts it into voice data using a speech synthesis API.
[0175] Input: Calendar data
[0176] Output: Audio data of the alert message
[0177] Step 10:
[0178] The terminal notifies the user of the alert audio data through the speaker.
[0179] Specific behavior:
[0180] The device plays the generated alert audio data through the speaker and notifies the user that "there is a meeting in 15 minutes."
[0181] Input: Alert audio data
[0182] Output: The alert notification that is played to the user
[0183] Step 11:
[0184] The terminal visualizes the content of the dialogue with the user on a display device and generates and displays dynamic character animations.
[0185] Specific behavior:
[0186] In order for the terminal to provide interactive responses and animations, animations such as the character's mouth movements and blinking eyes are displayed on the display device in accordance with the content of the user's dialogue.
[0187] Input: Dialogue
[0188] Output: Animation displayed on a display device
[0189] (Application example 1)
[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0191] Conventional voice recognition systems provide information based on user voice input, but this is limited to schedule management and general information provision. This makes it difficult to support users in more specific actions and decision-making in their daily lives. In particular, the lack of support for financial management, such as daily payment history and coupon information, is inconvenient for users and is an issue that needs to be resolved.
[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0193] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to start with a user's voice input, output the answer data as voice, obtain payment history data, generate alerts, and visualize them.
[0194] "Voice data" refers to data obtained by converting a user's voice into digital form.
[0195] "Text data" refers to data obtained by converting voice data into character information.
[0196] "Answer data" is data of the answer content generated based on the analysis results of the text data.
[0197] "Payment history data" is data that records a series of information regarding payments made by a user in the past.
[0198] An "alert" is a message that conveys information that requires the user's attention or notification.
[0199] A "display device" is a device for visually presenting information to a user.
[0200] The present invention is a system for recognizing a user's voice and generating, analyzing, and displaying voice data. This system is configured using a variety of hardware and software, and provides intuitive and useful information and schedule management to the user. Detailed embodiments of the system are described below.
[0201] Hardware and Software Configuration
[0202] This system uses the following hardware and software:
[0203] 1. Voice input device
[0204] Hardware: Smartphone microphone
[0205] Description: A device that allows a user to enter verbal instructions or questions.
[0206] 2. Voice Recognition Software
[0207] Software: Google Cloud Speech-to-Text API
[0208] Description: Software for converting audio data into text data.
[0209] 3. Natural Language Processing Software
[0210] Software: Natural language processing APIs such as OpenAI (registered trademark) GPT-3 (registered trademark)
[0211] Description: Software for analyzing text data and generating answer data.
[0212] 4. Text-to-Speech Software
[0213] Software: Microsoft® Azure® Text-to-Speech API
[0214] Description: Software that converts answer data into audio data and provides auditory feedback to the user.
[0215] 5. Payment History Acquisition Software
[0216] Software: Stripe API
[0217] Description: Software for acquiring payment history data and managing user payment history information.
[0218] 6. Display device
[0219] Hardware: Smartphone screen
[0220] Description: A device for displaying visual information to a user.
[0221] System processing procedure
[0222] We will now explain how the system of the present invention works.
[0223] 1. User voice input
[0224] The user gives voice instructions to the smartphone, for example, saying, "Tell me my latest payment history."
[0225] 2. Speech-to-text
[0226] The smartphone sends the voice data to the Google Cloud Speech-to-Text API, which converts it into text data.
[0227] 3. Text Data Analysis
[0228] The text data is analyzed using OpenAI GPT-3 to understand the user's intent. For example, the intent "Check payment history" can be recognized.
[0229] 4. Obtaining the necessary data
[0230] Call the Stripe API to get the latest payment history data.
[0231] 5. Generating Audio Data
[0232] Using the Microsoft Azure Text-to-Speech API, the answer provided to the user is converted into audio data and played back to the user.
[0233] 6. Visual Feedback
[0234] In addition to providing voice responses, payment history and related information will be displayed on the smartphone screen.
[0235] Examples of concrete examples and prompts
[0236] Examples:
[0237] When a user says to their smartphone, "Tell me my recent payment history," the system converts the speech into text, analyzes it, and recognizes the intent of "check payment history." It then calls the Stripe API to retrieve the payment history, provides that information to the user via voice, and displays the payment history on the smartphone screen.
[0238] Example prompt sentence:
[0239] Text sent to the natural language processing API: "What is your recent payment history?"
[0240] Expected intent: "Check payment history"
[0241] Required API call: Stripe API
[0242] Output data: Latest payment history information
[0243] Example of synthesized voice: "The most recent payment history is a payment of XXXX yen on March 15th."
[0244] Display example: "March 15th, XXXX yen" on a smartphone screen.
[0245] This format allows users to intuitively obtain payment history information through voice, making daily financial management more convenient.
[0246] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0247] Step 1:
[0248] The user speaks into the smartphone's microphone, for example, saying, "Tell me my recent payment history."
[0249] Input: User's voice
[0250] Output: Audio data
[0251] Specific operation: The user's voice is collected by the smartphone's microphone and saved as audio data within the application.
[0252] Step 2:
[0253] The device sends the collected voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data.
[0254] Input: Audio data
[0255] Output: Text data
[0256] Specific operation: The application on the device sends the voice data to the cloud service and receives and stores the returned text data.
[0257] Step 3:
[0258] The device then sends the converted text data to a natural language processing API such as OpenAI GPT-3 for analysis. Based on the analysis results, the device understands the user's intent and generates appropriate answer data.
[0259] Input: Text data
[0260] Output: Answer data
[0261] Specific operation: The device sends text data to a natural language processing API, understands the instruction "Tell me my payment history," and generates data to respond.
[0262] Step 4:
[0263] Based on the generated answer data, the terminal calls the Stripe API to obtain payment history data.
[0264] Input: Answer data
[0265] Output: Payment history data
[0266] Specific operation: The terminal calls the Stripe API based on the answer data and obtains the user's latest payment history information.
[0267] Step 5:
[0268] The terminal uses the acquired payment history data to generate a response and converts it into voice data using the Microsoft Azure Text-to-Speech API.
[0269] Input: Payment history data
[0270] Output: Audio data
[0271] Specific operation: The terminal prepares the payment history in text format and sends it to the speech synthesis API to obtain voice data.
[0272] Step 6:
[0273] The terminal reproduces the generated voice data to the user through a speaker.
[0274] Input: Audio data
[0275] Output: The audio the user hears
[0276] Specific operation: The device plays audio data and notifies the user by voice that "The most recent payment history is a payment of XXXX yen on March 15th."
[0277] Step 7:
[0278] At the same time, the terminal displays the acquired payment history information on the smartphone screen.
[0279] Input: Payment history data
[0280] Output: Visual display
[0281] Specific operation: The terminal application reads the payment history data and displays the latest payment history to the user through a visual interface.
[0282] This allows users to intuitively obtain payment history information through both voice input and visual display.
[0283] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0284] The present invention is a system that combines a desktop secretary system that intuitively interacts with a user, provides information, and manages schedules with an emotion engine that recognizes the user's emotions. The program processing of this system is explained below.
[0285] The user speaks
[0286] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather like tomorrow?"
[0287] Voice Recognition
[0288] The device collects the user's voice using a built-in microphone and saves it as voice data, which is then sent to a voice recognition API and converted into text data.
[0289] Text Analysis
[0290] The device sends the converted text data to a natural language processing API, which analyzes the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, in response to the input "What's the weather like tomorrow?", the device generates the answer "Tomorrow's weather will be sunny."
[0291] Audio response
[0292] The device sends the generated answer data to a speech synthesis API, which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the device may respond with a voice saying, "Tomorrow's weather will be sunny."
[0293] Emotion recognition
[0294] The device sends the voice data to the emotion engine to recognize the user's emotions. The emotion engine identifies the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the tone, speed, and strength of the voice.
[0295] Regulating responses based on emotions
[0296] The server then adjusts the response data to be more appropriate based on the emotion recognition results obtained from the emotion engine. For example, if the user expresses anger, the server adjusts the response data to say, "I'm sorry, I'll try my best to help you."
[0297] Calendar import and integration
[0298] The device periodically calls the calendar API to obtain the user's schedule data. From the obtained calendar data, information about the next appointment is saved in the database. For example, a "meeting tomorrow at 10:00 AM" registered by the user is included in the calendar data.
[0299] Alert generation and output
[0300] The server analyzes the calendar data and generates an alert message when it detects an approaching scheduled time. This alert message is also tailored based on emotion. For example, "You have a meeting in 15 minutes" may be supplemented with an emotionally-based message of encouragement depending on the importance and situation, such as "You have a meeting in 15 minutes, good luck!"
[0301] Visualization by display device
[0302] The device is equipped with a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character is used, and animations such as the character's mouth moving and blinking are displayed when the user speaks or responds. The results of user emotion recognition are also visualized. For example, the character's facial expression changes depending on the user's emotion.
[0303] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, by adjusting responses based on the user's emotions, a more satisfying user experience can be provided.
[0304] The processing flow will be explained below.
[0305] Step 1:
[0306] The user speaks to the device, for example, "What's the weather like tomorrow?"
[0307] Step 2:
[0308] The device collects the user's voice using a built-in microphone and saves it as voice data.
[0309] Step 3:
[0310] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[0311] Step 4:
[0312] The device receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends this data to the natural language processing API.
[0313] Step 5:
[0314] The natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[0315] Step 6:
[0316] The device receives the answer data generated by the natural language processing API and sends it to the speech synthesis API, which converts the answer data into voice data.
[0317] Step 7:
[0318] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[0319] Step 8:
[0320] The device sends the voice data to the emotion engine, which analyzes the voice data and recognizes the user's emotions (e.g., joy, anger, sadness, surprise, etc.).
[0321] Step 9:
[0322] The terminal transmits the emotion recognition result obtained from the emotion engine to the server.
[0323] Step 10:
[0324] The server adjusts the answer data based on the emotion recognition results. For example, if the user expresses anger, it generates an additional message saying, "I'm sorry, I'll try to help you."
[0325] Step 11:
[0326] The device then sends the adjusted answer data back to the speech synthesis API, where it is converted into voice data.
[0327] Step 12:
[0328] The device plays the adjusted voice data to the user through the speaker, for example, responding with a voice message saying, "The weather will be sunny tomorrow. Sorry, we'll do our best to help you."
[0329] Step 13:
[0330] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[0331] Step 14:
[0332] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[0333] Step 15:
[0334] The device sends the generated alert information to a speech synthesis API, which converts it into voice data.
[0335] Step 16:
[0336] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[0337] Step 17:
[0338] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[0339] Step 18:
[0340] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[0341] Example 2
[0342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0343] Conventional voice recognition systems and digital assistants provide information and schedule management based on voice input, but they do not provide responses that take the user's emotions into account, making it difficult to provide a satisfying user experience. Furthermore, schedule notifications are monotonous, requiring flexible responses that respond to the user's emotions and circumstances. The present invention aims to solve these problems by providing a system that provides integrated responses that take the user's emotions into account and schedule management.
[0344] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to provide information based on the user's voice input. Furthermore, it includes means for recognizing a user's emotion, means for adjusting the answer data based on the emotion recognition result, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This makes it possible to provide an integrated response and schedule management that takes the user's emotion into consideration.
[0345] "Voice Data" refers to information that is stored and processed digitally from a user's voice.
[0346] "Text data" refers to voice data converted into character string information, and refers to user questions and instructions that are handled as character information.
[0347] "Answer data" refers to answer information to a user's question or instruction, which is generated based on the analyzed text data.
[0348] A "voice recognition program" refers to software or algorithms for converting voice data into text data.
[0349] An "emotion engine" refers to a program or algorithm that analyzes and recognizes a user's emotional state from voice and text data.
[0350] "Calendar data" refers to digital data containing a user's schedule information.
[0351] The term "display device" refers to an output device, such as a monitor or display, for visually interacting with a user or providing information.
[0352] An "alert message" refers to a message that is sent based on the user's schedule information, and the content of the message is adjusted according to the user's emotional state and situation.
[0353] "User" refers to a person who uses this system to input voice and receive information provision and schedule management services.
[0354] "Terminal" refers to a device that receives a user's voice input, generates voice data, generates text data, plays back the voice data, recognizes emotions, and visualizes the data on a display device.
[0355] "Server" refers to a computer system that has the function of adjusting answer data based on emotion recognition results and the function of acquiring and analyzing calendar data.
[0356] This invention relates to a desktop secretary system that provides information and manages schedules through voice dialogue with a user. The system also has the function of recognizing the user's emotions and adjusting responses based on the results.
[0357] This system operates using the following hardware and software.
[0358] Hardware used
[0359] Device: Desktop or laptop computer.
[0360] Sound recording device (microphone): A device for collecting the user's voice.
[0361] Display Device: A monitor or display for visualizing user interactions.
[0362] Speaker: A device for playing back generated audio data.
[0363] Software used
[0364] Speech recognition program (API): For example, the Google Speech-to-Text API, which is used to convert voice data into text data.
[0365] Natural Language Processing Program (API): For example, the OpenAI GPT-4 (registered trademark) API used to analyze text data and generate answer data.
[0366] Speech synthesis program (API): For example, Google Text-to-Speech API, which is used to convert the generated answer data into audio data.
[0367] Emotion recognition engine: for example IBM Watson® Tone Analyzer used to recognize user emotions from voice data.
[0368] Calendar management program (API): For example, use the Google Calendar API to retrieve and manage calendar data.
[0369] System Operation Overview
[0370] The user asks questions or gives instructions to the device in natural language. For example, they might say, "What's the weather going to be like tomorrow?" This voice input is collected as voice data through the device's built-in sound recording device (microphone).
[0371] The device sends this voice data to a voice recognition program (Google Speech-to-Text API), which converts the voice data into text data. The converted text data is then sent to a natural language processing program (OpenAI GPT-4 API), which analyzes the intent of the user's question or instruction. Based on the results of this analysis, for example, in response to an input such as "What's the weather like tomorrow?", the answer "Tomorrow's weather will be sunny" is generated.
[0372] This answer data is sent to a speech synthesis program (Google Text-to-Speech API), which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the answer may be "Tomorrow's weather will be sunny."
[0373] The device sends the voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. Based on the emotion recognition results, the server adjusts the response data to be more appropriate. For example, if the user expresses anger, the response may be adjusted to say, "I'm sorry, I'll try my best to help you."
[0374] The device also periodically calls the calendar management program (Google Calendar API) to obtain the user's schedule data. This data is analyzed, and information about upcoming appointments is saved in the database. For example, if the calendar data includes a "meeting tomorrow at 10:00 AM," this information is saved.
[0375] The server analyzes calendar data and generates alert messages when an appointment time approaches, which are also tailored based on emotion: for example, "You have a meeting in 15 minutes," plus a message like "Good luck!" depending on the importance and situation.
[0376] The device uses a display device to visualize the dialogue with the user. Interactive responses and animations are provided, such as the character's mouth moving or blinking when the user speaks or responds. Furthermore, the user's emotion recognition results are also visualized, and the character's facial expression changes depending on the user's emotion.
[0377] Examples of prompt statements
[0378] When a user says, "What's the weather going to be like tomorrow?", the device performs voice recognition and generates text data saying, "What's the weather going to be like tomorrow?"
[0379] Next, the terminal analyzes the text data and generates answer data such as "Tomorrow's weather will be sunny."
[0380] The answer data is converted into voice data via a voice synthesis program, and the answer "Tomorrow's weather will be sunny" is spoken.
[0381] If the emotion recognition engine determines that the user is excited, the device will respond with a tailored response: "Looks like it's going to be a great day, have fun!"
[0382] As described above, this invention starts with a user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, adjusting responses based on the user's emotions can provide a more satisfying user experience.
[0383] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0384] Step 1:
[0385] The user speaks a question or instruction to the device in natural language. For example, they might say, "What's the weather like tomorrow?" This voice input becomes the initial input to the system.
[0386] Input: User's voice input ("What's the weather like tomorrow?")
[0387] Output: None (pass to next step)
[0388] Step 2:
[0389] The device collects the user's voice using a built-in audio recording device (microphone) and saves it as voice data.The device then sends the voice data to a voice recognition program (Google Speech-to-Text API) and converts it into text data.
[0390] Input: User's voice data
[0391] Output: Text data ("What's the weather like tomorrow?")
[0392] Specific operation: The user's voice is collected by the microphone and sent to the speech recognition API to obtain text data such as "What's the weather like tomorrow?"
[0393] Step 3:
[0394] The device sends the converted text data to a natural language processing program (OpenAI GPT-4 API) and analyzes the information the user is looking for. Based on the analysis results, the device calls an external weather information API (such as the OpenWeatherMap API) to obtain weather data. Answer data is generated based on the analysis results and the obtained data.
[0395] Input: Text data ("What's the weather like tomorrow?")
[0396] Output: Answer data ("Tomorrow's weather will be sunny")
[0397] Specific operation: The text data is analyzed using a natural language processing API to analyze the intent, a weather information API is called to obtain weather data, and the answer data "Tomorrow's weather will be sunny" is generated.
[0398] Step 4:
[0399] The device sends the generated answer data to a speech synthesis program (Google Text-to-Speech API), which converts the text data into audio data, which is then played back to the user through the speaker.
[0400] Input: Answer data ("Tomorrow's weather will be sunny")
[0401] Output: Voice data (voice response: "Tomorrow's weather will be sunny")
[0402] Specific operation: The answer data is converted into audio data using a speech synthesis API, and the speaker responds by saying, "Tomorrow's weather will be sunny."
[0403] Step 5:
[0404] The device sends the voice data collected during the conversation with the user to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion recognition engine identifies the user's emotions from the tone, speed, and strength of the voice.
[0405] Input: User's voice data
[0406] Output: Emotion recognition data ("User is excited")
[0407] Specific operation: Sends voice data to an emotion recognition engine to recognize the user's state of excitement.
[0408] Step 6:
[0409] The server then adjusts the answer data to be more appropriate based on the emotion recognition results obtained from the emotion recognition engine. For example, if the user is excited, the server adjusts the response to something like, "It looks like it's going to be a great day, have fun!"
[0410] Input: Emotion recognition data ("user is excited"), Answer data
[0411] Output: Adjusted answer data ("It looks like you're having a great day, have fun!")
[0412] Specific operation: Adjust answer data based on emotion recognition data to generate a response that takes into account the user's emotions.
[0413] Step 7:
[0414] The device periodically calls the calendar program (Google Calendar API) to obtain the user's schedule data. The obtained calendar data is analyzed and information about the next schedule is saved in the database. For example, if a user schedules a meeting at 10:00 AM tomorrow, that information is saved.
[0415] Input: None (or a periodic timer condition)
[0416] Output: Calendar data ("Meeting tomorrow at 10 AM")
[0417] Specific behavior: Calls the Google Calendar API to obtain new schedule data.
[0418] Step 8:
[0419] The server analyzes the calendar data and generates an alert message when it detects an approaching appointment time, which is also tailored based on emotion.
[0420] Input: Calendar data ("Meeting tomorrow at 10 AM")
[0421] Output: Alert message ("You have a meeting in 15 minutes. Good luck!")
[0422] Specific behavior: Check the appointment time and generate an alert message to notify you.
[0423] Step 9:
[0424] The device visualizes conversations with the user and schedule information on a display device. When the user speaks or responds, interactive responses such as character animations are displayed. Emotion recognition results are also displayed.
[0425] Input: User emotion recognition data, dialogue data
[0426] Output: Visually displayed animations and character facial expressions
[0427] Specific operation: Animation is displayed on the display device, and the character's facial expression changes based on the user's emotions and the content of the dialogue.
[0428] (Application example 2)
[0429] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0430] While conventional factory robot assistant systems can provide natural voice interaction with workers and schedule management, they lack the ability to recognize workers' emotions and optimize responses based on them, which hinders improvements in user experience and work efficiency. Additionally, changes to work processes and alert notifications within the factory are difficult to intuitively understand. This can increase the burden on workers and ultimately reduce production efficiency.
[0431] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for performing emotion recognition and generating emotion data, means for adjusting answer data based on the emotion data, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This not only enables natural voice dialogue with the worker, but also optimizes responses based on emotion recognition and makes it possible to provide intuitively understandable alert notifications and work processes.
[0432] "Speech recognition" is a technology that converts a user's voice into text data.
[0433] "Voice data" is digital sound information generated by voice recognition.
[0434] "Text data" is character information obtained by analyzing voice data.
[0435] "Answer data" is answer information to questions and instructions generated by analyzing text data.
[0436] "Speech synthesis" is a technology that converts text data into voice data.
[0437] "Emotion recognition" is a technology that generates emotional data from a user's voice or text.
[0438] "Emotion data" is information indicating the emotional state of the user obtained by emotion recognition.
[0439] "Schedule information" is time information for tasks and plans obtained from calendar data.
[0440] An "alert" is a notification message that is generated based on schedule information.
[0441] A "display device" is a device for providing information visually.
[0442] The present invention relates to a smart factory robot assistant system that intuitively interacts with workers in a factory to improve work efficiency. This system has speech recognition and emotion recognition functions, and provides schedule management and work instructions. A specific example of the system is shown below.
[0443] Hardware Configuration
[0444] The system of the present invention consists of smart glasses worn by a user (worker), a microphone, an internal speaker, and a display device. The server also includes multiple APIs and data storage.
[0445] Software Configuration
[0446] The software used is as follows:
[0447] Speech recognition API (e.g. Google Speech-to-Text)
[0448] Natural Language Processing API (e.g. Google NLP API)
[0449] Text-to-speech API (e.g. Google Text-to-Speech)
[0450] Emotion recognition API (e.g. Affectiva)
[0451] Calendar API (e.g. Google Calendar API)
[0452] System Operation
[0453] The system operates in the following manner.
[0454] 1. Voice Recognition
[0455] When a user speaks to the smart glasses, the built-in microphone collects the voice, which is then sent to a speech recognition API and converted into text data.
[0456] 2. Text Analysis
[0457] The text data is analyzed using a natural language processing API and used to understand the intent of the user's questions and instructions. Appropriate answer data is generated based on the analysis results.
[0458] 3. Emotional Recognition
[0459] The created text data and voice tone are sent to an emotion recognition API to generate user emotion data, which indicates the user's emotions (joy, anger, sadness, etc.).
[0460] 4. Adjusting response data
[0461] Based on the emotion data, the generated answer data is adjusted. For example, if the user expresses anger, an apology or an encouraging message is added to the response.
[0462] 5. Audio and visual representation
[0463] The adjusted answer data is converted into audio data using a speech synthesis API and played back through the smart glasses' speakers, while visual feedback such as work progress and alerts is displayed on the smart glasses' display.
[0464] 6. Schedule Management
[0465] Schedule information is retrieved using the calendar API, and alerts are generated when important events or deadlines are approaching, with audio and visual notifications.
[0466] Specific examples
[0467] For example, if a worker says, "Please tell me the next work step," their intention is analyzed using voice recognition and natural language processing technology. Based on the analysis results, answer data that provides the next work step is generated. Furthermore, emotion recognition is used to determine the worker's emotions, so if the worker is tired, an appropriate response will be given according to their emotions, such as, "The next task is △△. You seem tired, let's take a short break."
[0468] Prompt Sentence Examples
[0469] An example of a prompt sentence to input to the generative AI model is as follows:
[0470] If the question "Please tell me the next work step" is input, please output the answer text as "The next work step is ____." Also, if the questioner is feeling sad, please add a message of encouragement.
[0471] As described above, the system of the present invention integrates voice input, voice recognition, natural language processing, voice synthesis, emotion recognition, emotion-based response adjustment, schedule linkage, and visual feedback. This allows users to effectively acquire information through natural voice dialogue and improve work efficiency. In addition, responses that take the user's emotions into consideration can provide a more comfortable work environment.
[0472] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0473] Step 1:
[0474] The user speaks to the smart glasses. The built-in microphone of the smart glasses collects the user's voice and transmits it to the terminal as voice data. The input is the user's voice, and the output is voice data.
[0475] Step 2:
[0476] The device sends voice data to the voice recognition API. The API analyzes the voice data and converts it into text data. In this process, the voice data is analyzed and text data is generated as character information. The input is voice data and the output is text data.
[0477] Step 3:
[0478] The device sends the generated text data to a natural language processing API. The natural language processing API analyzes the text data and understands the intent of the user's question or instruction. This process generates appropriate answer data based on the analyzed intent. The input is text data, and the output is answer data.
[0479] Step 4:
[0480] The device sends the generated answer data to the emotion recognition API, and simultaneously analyzes the worker's voice tone. The emotion recognition API generates emotion data and identifies the user's emotion (e.g., joy, anger, sadness, etc.). The input is the answer data and voice tone, and the output is emotion data.
[0481] Step 5:
[0482] The server adjusts the answer data based on the emotion data. For example, if the user expresses anger, it adds an apology or comforting message to the answer data. The input is the answer data and emotion data, and the output is the adjusted answer data.
[0483] Step 6:
[0484] The device sends the adjusted answer data to a speech synthesis API, which converts the text data into speech data. In this process, the adjusted answer data is converted into speech. The input is the adjusted answer data, and the output is speech data.
[0485] Step 7:
[0486] The terminal plays the generated voice data through the speaker of the smart glasses, allowing the user to receive a voice response. The input is voice data, and the output is voice output.
[0487] Step 8:
[0488] The server periodically calls the calendar API to retrieve calendar data. The retrieved calendar data is parsed and schedule information is extracted. The input is the calendar data, and the output is the schedule information.
[0489] Step 9:
[0490] The server generates alerts based on schedule information and sends the alert data to the terminal. Alert messages are generated when important events or work deadlines approach. The input is schedule information and the output is alert data.
[0491] Step 10:
[0492] The device sends the generated alert data to the speech synthesis API, which converts the alert message into voice data. It also generates data to display the alert content on the smart glasses display. The input is the alert data, and the output is voice data and display data.
[0493] Step 11:
[0494] The device notifies the user of the alert through the smart glasses' speaker and displays the alert content on the display. The input is audio data and display data, and the output is audio and display notifications.
[0495] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0496] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0497] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0498] [Second embodiment]
[0499] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0500] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0501] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0502] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0503] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0504] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0505] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0506] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0507] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0508] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0509] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0510] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0511] The present invention aims to realize a desktop secretary system that can intuitively interact with a user, provide information, and manage schedules. The program processing of this system will be explained below.
[0512] The user speaks
[0513] The user asks questions and gives instructions to the device in natural language. For example, say, "What's the weather like tomorrow?"
[0514] Voice Recognition
[0515] The device collects the user's voice using a built-in microphone, saves it as voice data, and then sends the voice data to a voice recognition API (e.g., a publicly available voice recognition API) to convert the voice data into text data.
[0516] Text Analysis
[0517] The device sends the converted text data to a natural language processing API (e.g., a publicly available natural language processing API) to analyze the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, based on the analysis result of "What's the weather going to be like tomorrow?", it generates answer data such as "Tomorrow's weather will be sunny."
[0518] Audio response
[0519] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API) to convert the text data into voice data. The converted voice data is played back to the user through the speaker. As a result, the device responds to the user by voice, saying, "Tomorrow's weather will be sunny."
[0520] Calendar import and integration
[0521] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. The calendar data contains events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The schedule information is periodically checked, and an alert message is generated when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API and notified to the user through the speaker. This causes the device to alert the user by voice, saying, "You have a meeting in 15 minutes."
[0522] Visualization by display device
[0523] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move, blink, or other animations may be displayed when the user speaks or responds. This provides a more intuitive and enjoyable operating experience for the user.
[0524] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, calendar integration, and visual feedback. This allows the user to effectively obtain information and manage their schedule through natural voice dialogue and intuitive interaction without the need for complex operations.
[0525] The processing flow will be explained below.
[0526] Step 1:
[0527] The user speaks to the device, for example, "What's the weather like tomorrow?"
[0528] Step 2:
[0529] The device collects the user's voice using a built-in microphone and saves it as voice data.
[0530] Step 3:
[0531] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[0532] Step 4:
[0533] The server receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends it to the natural language processing API.
[0534] Step 5:
[0535] A natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[0536] Step 6:
[0537] The server receives the generated answer data and transmits it to the terminal.
[0538] Step 7:
[0539] The device sends the answer data to the speech synthesis API, which converts the answer data into voice data.
[0540] Step 8:
[0541] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[0542] Step 9:
[0543] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[0544] Step 10:
[0545] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[0546] Step 11:
[0547] The server sends the generated alert information to the terminal.
[0548] Step 12:
[0549] The device sends the alert information to a speech synthesis API, which converts it into voice data.
[0550] Step 13:
[0551] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[0552] Step 14:
[0553] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[0554] Step 15:
[0555] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[0556] Example 1
[0557] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0558] In modern society, there is a demand for systems that users can operate intuitively and that allow them to acquire information and manage schedules without requiring complex operations. However, existing systems are complex to operate, and many users are unable to master them. In addition, they lack an intuitive operating experience and visual feedback through user interaction, making them insufficient to improve user satisfaction.
[0559] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0560] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for analyzing the text data and generating answer data, means for outputting the generated answer data as voice data, means for acquiring calendar data and generating and outputting an alert including schedule information, means for visualizing the execution of the above means on a display device, and means for generating and displaying dynamic character animation to provide an intuitive operating experience through dialogue with the user. This allows the user to easily acquire information and manage their schedule through natural voice dialogue, and further enjoy an intuitive operating experience through visual feedback.
[0561] "User" refers to a person who uses this system to obtain information by voice and perform operations.
[0562] "Voice data" refers to data that is a digital recording of a user's speech.
[0563] "Text data" refers to data in the form of a character string converted from audio data.
[0564] "Answer data" refers to data that contains answers to questions and instructions from the user, which are obtained by analyzing text data.
[0565] The term "voice input device" refers to a device for collecting a user's voice and recording it as voice data.
[0566] "Speech conversion API" refers to an application program interface for converting voice data into text data.
[0567] "Calendar data" refers to data containing a user's schedule information.
[0568] An "alert" refers to a notification that notifies a user that a scheduled time is approaching.
[0569] "Display device" refers to a device that visually presents the execution status of the system and the contents of the dialogue.
[0570] "Dynamic character animation" refers to a moving character that is displayed on a display device to visually represent interaction with a user.
[0571] The present invention relates to a desktop secretary system that intuitively interacts with users, provides information, and manages schedules. This system integrates speech recognition, natural language processing, and speech synthesis technologies to provide users with information efficiently and effectively. Specific embodiments of this system are described below.
[0572] The user speaks
[0573] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather going to be like tomorrow?" This speech is collected through the device's microphone.
[0574] Voice Recognition
[0575] The device collects the user's voice using a built-in microphone and saves it as voice data. The voice data is then sent to an existing voice recognition API (for example, a publicly available voice recognition API) and converted into text data. Specifically, the voice recognition API used is something like the Google Speech-to-Text API.
[0576] Text Analysis
[0577] The device sends the converted text data to a natural language processing API (for example, a publicly available natural language processing API) and analyzes the intent of the user's question or instruction. Appropriate answer data is generated based on the analysis results. For example, if the question is "What's the weather like tomorrow?", the answer data generated based on the analysis results will be "Tomorrow's weather will be sunny." An example of a natural language processing API that can be used is "Google Dialogflow."
[0578] Audio response
[0579] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API), which converts the text data into voice data. The converted voice data is played back to the user through the speaker. This causes the device to respond to the user by voice, saying, "Tomorrow's weather will be sunny." The speech synthesis API used is something like the "Google Text-to-Speech API."
[0580] Calendar import and integration
[0581] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. Calendar data includes events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The device periodically checks the schedule information, and generates an alert message when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API, and notified to the user through the speaker. For example, an alert message saying "There is a meeting in 15 minutes" may be played.
[0582] Visualization by display device
[0583] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move or its eyes may blink when the user speaks or responds. This provides the user with a more intuitive and enjoyable operating experience.
[0584] Specific examples
[0585] For example, consider the following prompt:
[0586] "Look at my calendar, what's next?"
[0587] "Please tell me what the weather will be like tomorrow."
[0588] "I have a meeting at 3 PM. Set a reminder."
[0589] As a result, the system of the present invention not only allows users to obtain information and manage their schedules through natural voice dialogue, but also allows them to enjoy an intuitive operating experience through visual feedback.
[0590] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0591] Step 1:
[0592] The user speaks questions or commands.
[0593] Specific behavior:
[0594] The user speaks to the terminal, "Please tell me what the weather will be like tomorrow."
[0595] Input: User voice input
[0596] Output: Audio data collected by the device's microphone
[0597] Step 2:
[0598] The device stores the voice data and sends it to the voice recognition API.
[0599] Specific behavior:
[0600] The device's microphone records the user's speech as voice data, temporarily stores it in memory, and then sends the voice data to an existing speech recognition API (e.g., Google Speech-to-Text API).
[0601] Input: Audio data
[0602] Output: Audio data sent to the API
[0603] Step 3:
[0604] The server converts the voice data into text data and sends it back to the terminal.
[0605] Specific behavior:
[0606] The speech recognition API analyzes the voice data and converts it into text data such as "Please tell me what the weather will be like tomorrow." The text data is then sent from the server to the device.
[0607] Input: Audio data
[0608] Output: Text data
[0609] Step 4:
[0610] The device sends the text data to a natural language processing API.
[0611] Specific behavior:
[0612] The device sends the received text data to a natural language processing API (e.g., Google Dialogflow).
[0613] Input: Text data
[0614] Output: Text data sent to the API
[0615] Step 5:
[0616] The server analyzes the text data and generates answer data.
[0617] Specific behavior:
[0618] The natural language processing API analyzes the text data, understands the intent, and generates the answer data, "The weather will be sunny tomorrow."
[0619] Input: Text data
[0620] Output: Answer data
[0621] Step 6:
[0622] The device sends the answer data to the speech synthesis API and outputs it as voice data.
[0623] Specific behavior:
[0624] The device sends the generated answer data to a speech synthesis API (for example, Google Text-to-Speech API), which generates speech data saying, "Tomorrow's weather will be sunny."
[0625] Input: Answer data
[0626] Output: Audio data
[0627] Step 7:
[0628] The terminal plays the audio data to the user through the speaker.
[0629] Specific behavior:
[0630] The terminal receives the voice data and responds to the user through the speaker by voice, "Tomorrow's weather will be sunny."
[0631] Input: Audio data
[0632] Output: The audio response that is played to the user
[0633] Step 8:
[0634] The device periodically calls the calendar API to obtain the user's schedule data.
[0635] Specific behavior:
[0636] The device accesses a calendar API (for example, Google Calendar API) to obtain the user's schedule data. For example, it obtains information such as "Meeting tomorrow at 10:00 AM."
[0637] Input: None (periodic processing)
[0638] Output: Retrieved calendar data
[0639] Step 9:
[0640] The device checks the schedule information, and when the scheduled time approaches, it generates an alert message and converts it into voice data.
[0641] Specific behavior:
[0642] The device analyzes the calendar data it has acquired, generates an alert message 15 minutes before the scheduled meeting saying "You have a meeting in 15 minutes," and converts it into voice data using a speech synthesis API.
[0643] Input: Calendar data
[0644] Output: Audio data of the alert message
[0645] Step 10:
[0646] The terminal notifies the user of the alert audio data through the speaker.
[0647] Specific behavior:
[0648] The device plays the generated alert audio data through the speaker and notifies the user that "there is a meeting in 15 minutes."
[0649] Input: Alert audio data
[0650] Output: The alert notification that is played to the user
[0651] Step 11:
[0652] The terminal visualizes the content of the dialogue with the user on a display device and generates and displays dynamic character animations.
[0653] Specific behavior:
[0654] In order for the terminal to provide interactive responses and animations, animations such as the character's mouth movements and blinking eyes are displayed on the display device in accordance with the content of the user's dialogue.
[0655] Input: Dialogue
[0656] Output: Animation displayed on a display device
[0657] (Application example 1)
[0658] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0659] Conventional voice recognition systems provide information based on user voice input, but this is limited to schedule management and general information provision. This makes it difficult to support users in more specific actions and decision-making in their daily lives. In particular, the lack of support for financial management, such as daily payment history and coupon information, is inconvenient for users and is an issue that needs to be resolved.
[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0661] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to start with a user's voice input, output the answer data as voice, obtain payment history data, generate alerts, and visualize them.
[0662] "Voice data" refers to data obtained by converting a user's voice into digital form.
[0663] "Text data" refers to data obtained by converting voice data into character information.
[0664] "Answer data" is data of the answer content generated based on the analysis results of the text data.
[0665] "Payment history data" is data that records a series of information regarding payments made by a user in the past.
[0666] An "alert" is a message that conveys information that requires the user's attention or notification.
[0667] A "display device" is a device for visually presenting information to a user.
[0668] The present invention is a system for recognizing a user's voice and generating, analyzing, and displaying voice data. This system is configured using a variety of hardware and software, and provides intuitive and useful information and schedule management to the user. Detailed embodiments of the system are described below.
[0669] Hardware and Software Configuration
[0670] This system uses the following hardware and software:
[0671] 1. Voice input device
[0672] Hardware: Smartphone microphone
[0673] Description: A device that allows a user to enter verbal instructions or questions.
[0674] 2. Voice Recognition Software
[0675] Software: Google Cloud Speech-to-Text API
[0676] Description: Software for converting audio data into text data.
[0677] 3. Natural Language Processing Software
[0678] Software: Natural Language Processing APIs such as OpenAI GPT-3
[0679] Description: Software for analyzing text data and generating answer data.
[0680] 4. Text-to-Speech Software
[0681] Software: Microsoft Azure Text-to-Speech API
[0682] Description: Software that converts answer data into audio data and provides auditory feedback to the user.
[0683] 5. Payment History Acquisition Software
[0684] Software: Stripe API
[0685] Description: Software for acquiring payment history data and managing user payment history information.
[0686] 6. Display device
[0687] Hardware: Smartphone screen
[0688] Description: A device for displaying visual information to a user.
[0689] System processing procedure
[0690] We will now explain how the system of the present invention works.
[0691] 1. User voice input
[0692] The user gives voice instructions to the smartphone, for example, saying, "Tell me my latest payment history."
[0693] 2. Speech-to-text
[0694] The smartphone sends the voice data to the Google Cloud Speech-to-Text API, which converts it into text data.
[0695] 3. Text Data Analysis
[0696] The text data is analyzed using OpenAI GPT-3 to understand the user's intent. For example, the intent "Check payment history" can be recognized.
[0697] 4. Obtaining the necessary data
[0698] Call the Stripe API to get the latest payment history data.
[0699] 5. Generating Audio Data
[0700] Using the Microsoft Azure Text-to-Speech API, the answer provided to the user is converted into audio data and played back to the user.
[0701] 6. Visual Feedback
[0702] In addition to providing voice responses, payment history and related information will be displayed on the smartphone screen.
[0703] Examples of concrete examples and prompts
[0704] Examples:
[0705] When a user says to their smartphone, "Tell me my recent payment history," the system converts the speech into text, analyzes it, and recognizes the intent of "check payment history." It then calls the Stripe API to retrieve the payment history, provides that information to the user via voice, and displays the payment history on the smartphone screen.
[0706] Example prompt sentence:
[0707] Text sent to the natural language processing API: "What is your recent payment history?"
[0708] Expected intent: "Check payment history"
[0709] Required API call: Stripe API
[0710] Output data: Latest payment history information
[0711] Example of synthesized voice: "The most recent payment history is a payment of XXXX yen on March 15th."
[0712] Display example: "March 15th, XXXX yen" on a smartphone screen.
[0713] This format allows users to intuitively obtain payment history information through voice, making daily financial management more convenient.
[0714] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0715] Step 1:
[0716] The user speaks into the smartphone's microphone, for example, saying, "Tell me my recent payment history."
[0717] Input: User's voice
[0718] Output: Audio data
[0719] Specific operation: The user's voice is collected by the smartphone's microphone and saved as audio data within the application.
[0720] Step 2:
[0721] The device sends the collected voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data.
[0722] Input: Audio data
[0723] Output: Text data
[0724] Specific operation: The application on the device sends the voice data to the cloud service and receives and stores the returned text data.
[0725] Step 3:
[0726] The device then sends the converted text data to a natural language processing API such as OpenAI GPT-3 for analysis. Based on the analysis results, the device understands the user's intent and generates appropriate answer data.
[0727] Input: Text data
[0728] Output: Answer data
[0729] Specific operation: The device sends text data to a natural language processing API, understands the instruction "Tell me my payment history," and generates data to respond.
[0730] Step 4:
[0731] Based on the generated answer data, the terminal calls the Stripe API to obtain payment history data.
[0732] Input: Answer data
[0733] Output: Payment history data
[0734] Specific operation: The terminal calls the Stripe API based on the answer data and obtains the user's latest payment history information.
[0735] Step 5:
[0736] The terminal uses the acquired payment history data to generate a response and converts it into voice data using the Microsoft Azure Text-to-Speech API.
[0737] Input: Payment history data
[0738] Output: Audio data
[0739] Specific operation: The terminal prepares the payment history in text format and sends it to the speech synthesis API to obtain voice data.
[0740] Step 6:
[0741] The terminal reproduces the generated voice data to the user through a speaker.
[0742] Input: Audio data
[0743] Output: The audio the user hears
[0744] Specific operation: The device plays audio data and notifies the user by voice that "The most recent payment history is a payment of XXXX yen on March 15th."
[0745] Step 7:
[0746] At the same time, the terminal displays the acquired payment history information on the smartphone screen.
[0747] Input: Payment history data
[0748] Output: Visual display
[0749] Specific operation: The terminal application reads the payment history data and displays the latest payment history to the user through a visual interface.
[0750] This allows users to intuitively obtain payment history information through both voice input and visual display.
[0751] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0752] The present invention is a system that combines a desktop secretary system that intuitively interacts with a user, provides information, and manages schedules with an emotion engine that recognizes the user's emotions. The program processing of this system is explained below.
[0753] The user speaks
[0754] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather like tomorrow?"
[0755] Voice Recognition
[0756] The device collects the user's voice using a built-in microphone and saves it as voice data, which is then sent to a voice recognition API and converted into text data.
[0757] Text Analysis
[0758] The device sends the converted text data to a natural language processing API, which analyzes the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, in response to the input "What's the weather like tomorrow?", the device generates the answer "Tomorrow's weather will be sunny."
[0759] Audio response
[0760] The device sends the generated answer data to a speech synthesis API, which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the device may respond with a voice saying, "Tomorrow's weather will be sunny."
[0761] Emotion recognition
[0762] The device sends the voice data to the emotion engine to recognize the user's emotions. The emotion engine identifies the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the tone, speed, and strength of the voice.
[0763] Regulating responses based on emotions
[0764] The server then adjusts the response data to be more appropriate based on the emotion recognition results obtained from the emotion engine. For example, if the user expresses anger, the server adjusts the response data to say, "I'm sorry, I'll try my best to help you."
[0765] Calendar import and integration
[0766] The device periodically calls the calendar API to obtain the user's schedule data. From the obtained calendar data, information about the next appointment is saved in the database. For example, a "meeting tomorrow at 10:00 AM" registered by the user is included in the calendar data.
[0767] Alert generation and output
[0768] The server analyzes the calendar data and generates an alert message when it detects an approaching scheduled time. This alert message is also tailored based on emotion. For example, "You have a meeting in 15 minutes" may be supplemented with an emotionally-based message of encouragement depending on the importance and situation, such as "You have a meeting in 15 minutes, good luck!"
[0769] Visualization by display device
[0770] The device is equipped with a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character is used, and animations such as the character's mouth moving and blinking are displayed when the user speaks or responds. The results of user emotion recognition are also visualized. For example, the character's facial expression changes depending on the user's emotion.
[0771] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, by adjusting responses based on the user's emotions, a more satisfying user experience can be provided.
[0772] The processing flow will be explained below.
[0773] Step 1:
[0774] The user speaks to the device, for example, "What's the weather like tomorrow?"
[0775] Step 2:
[0776] The device collects the user's voice using a built-in microphone and saves it as voice data.
[0777] Step 3:
[0778] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[0779] Step 4:
[0780] The device receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends this data to the natural language processing API.
[0781] Step 5:
[0782] The natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[0783] Step 6:
[0784] The device receives the answer data generated by the natural language processing API and sends it to the speech synthesis API, which converts the answer data into voice data.
[0785] Step 7:
[0786] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[0787] Step 8:
[0788] The device sends the voice data to the emotion engine, which analyzes the voice data and recognizes the user's emotions (e.g., joy, anger, sadness, surprise, etc.).
[0789] Step 9:
[0790] The terminal transmits the emotion recognition result obtained from the emotion engine to the server.
[0791] Step 10:
[0792] The server adjusts the answer data based on the emotion recognition results. For example, if the user expresses anger, it generates an additional message saying, "I'm sorry, I'll try to help you."
[0793] Step 11:
[0794] The device then sends the adjusted answer data back to the speech synthesis API, where it is converted into voice data.
[0795] Step 12:
[0796] The device plays the adjusted voice data to the user through the speaker, for example, responding with a voice message saying, "The weather will be sunny tomorrow. Sorry, we'll do our best to help you."
[0797] Step 13:
[0798] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[0799] Step 14:
[0800] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[0801] Step 15:
[0802] The device sends the generated alert information to a speech synthesis API, which converts it into voice data.
[0803] Step 16:
[0804] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[0805] Step 17:
[0806] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[0807] Step 18:
[0808] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[0809] Example 2
[0810] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0811] Conventional voice recognition systems and digital assistants provide information and schedule management based on voice input, but they do not provide responses that take the user's emotions into account, making it difficult to provide a satisfying user experience. Furthermore, schedule notifications are monotonous, requiring flexible responses that respond to the user's emotions and circumstances. The present invention aims to solve these problems by providing a system that provides integrated responses that take the user's emotions into account and schedule management.
[0812] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to provide information based on the user's voice input. Furthermore, it includes means for recognizing a user's emotion, means for adjusting the answer data based on the emotion recognition result, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This makes it possible to provide an integrated response and schedule management that takes the user's emotion into consideration.
[0813] "Voice Data" refers to information that is stored and processed digitally from a user's voice.
[0814] "Text data" refers to voice data converted into character string information, and refers to user questions and instructions that are handled as character information.
[0815] "Answer data" refers to answer information to a user's question or instruction, which is generated based on the analyzed text data.
[0816] A "voice recognition program" refers to software or algorithms for converting voice data into text data.
[0817] An "emotion engine" refers to a program or algorithm that analyzes and recognizes a user's emotional state from voice and text data.
[0818] "Calendar data" refers to digital data containing a user's schedule information.
[0819] The term "display device" refers to an output device, such as a monitor or display, for visually interacting with a user or providing information.
[0820] An "alert message" refers to a message that is sent based on the user's schedule information, and the content of the message is adjusted according to the user's emotional state and situation.
[0821] "User" refers to a person who uses this system to input voice and receive information provision and schedule management services.
[0822] "Terminal" refers to a device that receives a user's voice input, generates voice data, generates text data, plays back the voice data, recognizes emotions, and visualizes the data on a display device.
[0823] "Server" refers to a computer system that has the function of adjusting answer data based on emotion recognition results and the function of acquiring and analyzing calendar data.
[0824] This invention relates to a desktop secretary system that provides information and manages schedules through voice dialogue with a user. The system also has the function of recognizing the user's emotions and adjusting responses based on the results.
[0825] This system operates using the following hardware and software.
[0826] Hardware used
[0827] Device: Desktop or laptop computer.
[0828] Sound recording device (microphone): A device for collecting the user's voice.
[0829] Display Device: A monitor or display for visualizing user interactions.
[0830] Speaker: A device for playing back generated audio data.
[0831] Software used
[0832] Speech recognition program (API): For example, the Google Speech-to-Text API, which is used to convert voice data into text data.
[0833] Natural language processing program (API): For example, the OpenAI GPT-4 API used to analyze text data and generate answer data.
[0834] Speech synthesis program (API): For example, Google Text-to-Speech API, which is used to convert the generated answer data into audio data.
[0835] Emotion recognition engine: For example, IBM Watson Tone Analyzer, which is used to recognize user emotions from voice data.
[0836] Calendar management program (API): For example, use the Google Calendar API to retrieve and manage calendar data.
[0837] System Operation Overview
[0838] The user asks questions or gives instructions to the device in natural language. For example, they might say, "What's the weather going to be like tomorrow?" This voice input is collected as voice data through the device's built-in sound recording device (microphone).
[0839] The device sends this voice data to a voice recognition program (Google Speech-to-Text API), which converts the voice data into text data. The converted text data is then sent to a natural language processing program (OpenAI GPT-4 API), which analyzes the intent of the user's question or instruction. Based on the results of this analysis, for example, in response to an input such as "What's the weather like tomorrow?", the answer "Tomorrow's weather will be sunny" is generated.
[0840] This answer data is sent to a speech synthesis program (Google Text-to-Speech API), which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the answer may be "Tomorrow's weather will be sunny."
[0841] The device sends the voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. Based on the emotion recognition results, the server adjusts the response data to be more appropriate. For example, if the user expresses anger, the response may be adjusted to say, "I'm sorry, I'll try my best to help you."
[0842] The device also periodically calls the calendar management program (Google Calendar API) to obtain the user's schedule data. This data is analyzed, and information about upcoming appointments is saved in the database. For example, if the calendar data includes a "meeting tomorrow at 10:00 AM," this information is saved.
[0843] The server analyzes calendar data and generates alert messages when an appointment time approaches, which are also tailored based on emotion: for example, "You have a meeting in 15 minutes," plus a message like "Good luck!" depending on the importance and situation.
[0844] The device uses a display device to visualize the dialogue with the user. Interactive responses and animations are provided, such as the character's mouth moving or blinking when the user speaks or responds. Furthermore, the user's emotion recognition results are also visualized, and the character's facial expression changes depending on the user's emotion.
[0845] Examples of prompt statements
[0846] When a user says, "What's the weather going to be like tomorrow?", the device performs voice recognition and generates text data saying, "What's the weather going to be like tomorrow?"
[0847] Next, the terminal analyzes the text data and generates answer data such as "Tomorrow's weather will be sunny."
[0848] The answer data is converted into voice data via a voice synthesis program, and the answer "Tomorrow's weather will be sunny" is spoken.
[0849] If the emotion recognition engine determines that the user is excited, the device will respond with a tailored response: "Looks like it's going to be a great day, have fun!"
[0850] As described above, this invention starts with a user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, adjusting responses based on the user's emotions can provide a more satisfying user experience.
[0851] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0852] Step 1:
[0853] The user speaks a question or instruction to the device in natural language. For example, they might say, "What's the weather like tomorrow?" This voice input becomes the initial input to the system.
[0854] Input: User's voice input ("What's the weather like tomorrow?")
[0855] Output: None (pass to next step)
[0856] Step 2:
[0857] The device collects the user's voice using a built-in audio recording device (microphone) and saves it as voice data.The device then sends the voice data to a voice recognition program (Google Speech-to-Text API) and converts it into text data.
[0858] Input: User's voice data
[0859] Output: Text data ("What's the weather like tomorrow?")
[0860] Specific operation: The user's voice is collected by the microphone and sent to the speech recognition API to obtain text data such as "What's the weather like tomorrow?"
[0861] Step 3:
[0862] The device sends the converted text data to a natural language processing program (OpenAI GPT-4 API) and analyzes the information the user is looking for. Based on the analysis results, the device calls an external weather information API (such as the OpenWeatherMap API) to obtain weather data. Answer data is generated based on the analysis results and the obtained data.
[0863] Input: Text data ("What's the weather like tomorrow?")
[0864] Output: Answer data ("Tomorrow's weather will be sunny")
[0865] Specific operation: The text data is analyzed using a natural language processing API to analyze the intent, a weather information API is called to obtain weather data, and the answer data "Tomorrow's weather will be sunny" is generated.
[0866] Step 4:
[0867] The device sends the generated answer data to a speech synthesis program (Google Text-to-Speech API), which converts the text data into audio data, which is then played back to the user through the speaker.
[0868] Input: Answer data ("Tomorrow's weather will be sunny")
[0869] Output: Voice data (voice response: "Tomorrow's weather will be sunny")
[0870] Specific operation: The answer data is converted into audio data using a speech synthesis API, and the speaker responds by saying, "Tomorrow's weather will be sunny."
[0871] Step 5:
[0872] The device sends the voice data collected during the conversation with the user to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion recognition engine identifies the user's emotions from the tone, speed, and strength of the voice.
[0873] Input: User's voice data
[0874] Output: Emotion recognition data ("User is excited")
[0875] Specific operation: Sends voice data to an emotion recognition engine to recognize the user's state of excitement.
[0876] Step 6:
[0877] The server then adjusts the answer data to be more appropriate based on the emotion recognition results obtained from the emotion recognition engine. For example, if the user is excited, the server adjusts the response to something like, "It looks like it's going to be a great day, have fun!"
[0878] Input: Emotion recognition data ("user is excited"), Answer data
[0879] Output: Adjusted answer data ("It looks like you're having a great day, have fun!")
[0880] Specific operation: Adjust answer data based on emotion recognition data to generate a response that takes into account the user's emotions.
[0881] Step 7:
[0882] The device periodically calls the calendar program (Google Calendar API) to obtain the user's schedule data. The obtained calendar data is analyzed and information about the next schedule is saved in the database. For example, if a user schedules a meeting at 10:00 AM tomorrow, that information is saved.
[0883] Input: None (or a periodic timer condition)
[0884] Output: Calendar data ("Meeting tomorrow at 10 AM")
[0885] Specific behavior: Calls the Google Calendar API to obtain new schedule data.
[0886] Step 8:
[0887] The server analyzes the calendar data and generates an alert message when it detects an approaching appointment time, which is also tailored based on emotion.
[0888] Input: Calendar data ("Meeting tomorrow at 10 AM")
[0889] Output: Alert message ("You have a meeting in 15 minutes. Good luck!")
[0890] Specific behavior: Check the appointment time and generate an alert message to notify you.
[0891] Step 9:
[0892] The device visualizes conversations with the user and schedule information on a display device. When the user speaks or responds, interactive responses such as character animations are displayed. Emotion recognition results are also displayed.
[0893] Input: User emotion recognition data, dialogue data
[0894] Output: Visually displayed animations and character facial expressions
[0895] Specific operation: Animation is displayed on the display device, and the character's facial expression changes based on the user's emotions and the content of the dialogue.
[0896] (Application example 2)
[0897] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0898] While conventional factory robot assistant systems can provide natural voice interaction with workers and schedule management, they lack the ability to recognize workers' emotions and optimize responses based on them, which hinders improvements in user experience and work efficiency. Additionally, changes to work processes and alert notifications within the factory are difficult to intuitively understand. This can increase the burden on workers and ultimately reduce production efficiency.
[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for performing emotion recognition and generating emotion data, means for adjusting answer data based on the emotion data, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This not only enables natural voice dialogue with the worker, but also optimizes responses based on emotion recognition and makes it possible to provide intuitively understandable alert notifications and work processes.
[0900] "Speech recognition" is a technology that converts a user's voice into text data.
[0901] "Voice data" is digital sound information generated by voice recognition.
[0902] "Text data" is character information obtained by analyzing voice data.
[0903] "Answer data" is answer information to questions and instructions generated by analyzing text data.
[0904] "Speech synthesis" is a technology that converts text data into voice data.
[0905] "Emotion recognition" is a technology that generates emotional data from a user's voice or text.
[0906] "Emotion data" is information indicating the emotional state of the user obtained by emotion recognition.
[0907] "Schedule information" is time information for tasks and plans obtained from calendar data.
[0908] An "alert" is a notification message that is generated based on schedule information.
[0909] A "display device" is a device for providing information visually.
[0910] The present invention relates to a smart factory robot assistant system that intuitively interacts with workers in a factory to improve work efficiency. This system has speech recognition and emotion recognition functions, and provides schedule management and work instructions. A specific example of the system is shown below.
[0911] Hardware Configuration
[0912] The system of the present invention consists of smart glasses worn by a user (worker), a microphone, an internal speaker, and a display device. The server also includes multiple APIs and data storage.
[0913] Software Configuration
[0914] The software used is as follows:
[0915] Speech recognition API (e.g. Google Speech-to-Text)
[0916] Natural Language Processing API (e.g. Google NLP API)
[0917] Text-to-speech API (e.g. Google Text-to-Speech)
[0918] Emotion recognition API (e.g. Affectiva)
[0919] Calendar API (e.g. Google Calendar API)
[0920] System Operation
[0921] The system operates in the following manner.
[0922] 1. Voice Recognition
[0923] When a user speaks to the smart glasses, the built-in microphone collects the voice, which is then sent to a speech recognition API and converted into text data.
[0924] 2. Text Analysis
[0925] The text data is analyzed using a natural language processing API and used to understand the intent of the user's questions and instructions. Appropriate answer data is generated based on the analysis results.
[0926] 3. Emotional Recognition
[0927] The created text data and voice tone are sent to an emotion recognition API to generate user emotion data, which indicates the user's emotions (joy, anger, sadness, etc.).
[0928] 4. Adjusting response data
[0929] Based on the emotion data, the generated answer data is adjusted. For example, if the user expresses anger, an apology or an encouraging message is added to the response.
[0930] 5. Audio and visual representation
[0931] The adjusted answer data is converted into audio data using a speech synthesis API and played back through the smart glasses' speakers, while visual feedback such as work progress and alerts is displayed on the smart glasses' display.
[0932] 6. Schedule Management
[0933] Schedule information is retrieved using the calendar API, and alerts are generated when important events or deadlines are approaching, with audio and visual notifications.
[0934] Specific examples
[0935] For example, if a worker says, "Please tell me the next work step," their intention is analyzed using voice recognition and natural language processing technology. Based on the analysis results, answer data that provides the next work step is generated. Furthermore, emotion recognition is used to determine the worker's emotions, so if the worker is tired, an appropriate response will be given according to their emotions, such as, "The next task is △△. You seem tired, let's take a short break."
[0936] Prompt Sentence Examples
[0937] An example of a prompt sentence to input to the generative AI model is as follows:
[0938] If the question "Please tell me the next work step" is input, please output the answer text as "The next work step is ____." Also, if the questioner is feeling sad, please add a message of encouragement.
[0939] As described above, the system of the present invention integrates voice input, voice recognition, natural language processing, voice synthesis, emotion recognition, emotion-based response adjustment, schedule linkage, and visual feedback. This allows users to effectively acquire information through natural voice dialogue and improve work efficiency. In addition, responses that take the user's emotions into consideration can provide a more comfortable work environment.
[0940] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0941] Step 1:
[0942] The user speaks to the smart glasses. The built-in microphone of the smart glasses collects the user's voice and transmits it to the terminal as voice data. The input is the user's voice, and the output is voice data.
[0943] Step 2:
[0944] The device sends voice data to the voice recognition API. The API analyzes the voice data and converts it into text data. In this process, the voice data is analyzed and text data is generated as character information. The input is voice data and the output is text data.
[0945] Step 3:
[0946] The device sends the generated text data to a natural language processing API. The natural language processing API analyzes the text data and understands the intent of the user's question or instruction. This process generates appropriate answer data based on the analyzed intent. The input is text data, and the output is answer data.
[0947] Step 4:
[0948] The device sends the generated answer data to the emotion recognition API, and simultaneously analyzes the worker's voice tone. The emotion recognition API generates emotion data and identifies the user's emotion (e.g., joy, anger, sadness, etc.). The input is the answer data and voice tone, and the output is emotion data.
[0949] Step 5:
[0950] The server adjusts the answer data based on the emotion data. For example, if the user expresses anger, it adds an apology or comforting message to the answer data. The input is the answer data and emotion data, and the output is the adjusted answer data.
[0951] Step 6:
[0952] The device sends the adjusted answer data to a speech synthesis API, which converts the text data into speech data. In this process, the adjusted answer data is converted into speech. The input is the adjusted answer data, and the output is speech data.
[0953] Step 7:
[0954] The terminal plays the generated voice data through the speaker of the smart glasses, allowing the user to receive a voice response. The input is voice data, and the output is voice output.
[0955] Step 8:
[0956] The server periodically calls the calendar API to retrieve calendar data. The retrieved calendar data is parsed and schedule information is extracted. The input is the calendar data, and the output is the schedule information.
[0957] Step 9:
[0958] The server generates alerts based on schedule information and sends the alert data to the terminal. Alert messages are generated when important events or work deadlines approach. The input is schedule information and the output is alert data.
[0959] Step 10:
[0960] The device sends the generated alert data to the speech synthesis API, which converts the alert message into voice data. It also generates data to display the alert content on the smart glasses display. The input is the alert data, and the output is voice data and display data.
[0961] Step 11:
[0962] The device notifies the user of the alert through the smart glasses' speaker and displays the alert content on the display. The input is audio data and display data, and the output is audio and display notifications.
[0963] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0964] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0965] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0966] [Third embodiment]
[0967] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0968] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0969] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0970] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0971] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0972] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0973] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0974] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0975] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0976] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0977] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0978] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0979] The present invention aims to realize a desktop secretary system that can intuitively interact with a user, provide information, and manage schedules. The program processing of this system will be explained below.
[0980] The user speaks
[0981] The user asks questions and gives instructions to the device in natural language. For example, say, "What's the weather like tomorrow?"
[0982] Voice Recognition
[0983] The device collects the user's voice using a built-in microphone, saves it as voice data, and then sends the voice data to a voice recognition API (e.g., a publicly available voice recognition API) to convert the voice data into text data.
[0984] Text Analysis
[0985] The device sends the converted text data to a natural language processing API (e.g., a publicly available natural language processing API) to analyze the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, based on the analysis result of "What's the weather going to be like tomorrow?", it generates answer data such as "Tomorrow's weather will be sunny."
[0986] Audio response
[0987] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API) to convert the text data into voice data. The converted voice data is played back to the user through the speaker. As a result, the device responds to the user by voice, saying, "Tomorrow's weather will be sunny."
[0988] Calendar import and integration
[0989] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. The calendar data contains events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The schedule information is periodically checked, and an alert message is generated when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API and notified to the user through the speaker. This causes the device to alert the user by voice, saying, "You have a meeting in 15 minutes."
[0990] Visualization by display device
[0991] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move, blink, or other animations may be displayed when the user speaks or responds. This provides a more intuitive and enjoyable operating experience for the user.
[0992] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, calendar integration, and visual feedback. This allows the user to effectively obtain information and manage their schedule through natural voice dialogue and intuitive interaction without the need for complex operations.
[0993] The processing flow will be explained below.
[0994] Step 1:
[0995] The user speaks to the device, for example, "What's the weather like tomorrow?"
[0996] Step 2:
[0997] The device collects the user's voice using a built-in microphone and saves it as voice data.
[0998] Step 3:
[0999] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[1000] Step 4:
[1001] The server receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends it to the natural language processing API.
[1002] Step 5:
[1003] A natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[1004] Step 6:
[1005] The server receives the generated answer data and transmits it to the terminal.
[1006] Step 7:
[1007] The device sends the answer data to the speech synthesis API, which converts the answer data into voice data.
[1008] Step 8:
[1009] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[1010] Step 9:
[1011] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[1012] Step 10:
[1013] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[1014] Step 11:
[1015] The server sends the generated alert information to the terminal.
[1016] Step 12:
[1017] The device sends the alert information to a speech synthesis API, which converts it into voice data.
[1018] Step 13:
[1019] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[1020] Step 14:
[1021] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[1022] Step 15:
[1023] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[1024] Example 1
[1025] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1026] In modern society, there is a demand for systems that users can operate intuitively and that allow them to acquire information and manage schedules without requiring complex operations. However, existing systems are complex to operate, and many users are unable to master them. In addition, they lack an intuitive operating experience and visual feedback through user interaction, making them insufficient to improve user satisfaction.
[1027] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1028] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for analyzing the text data and generating answer data, means for outputting the generated answer data as voice data, means for acquiring calendar data and generating and outputting an alert including schedule information, means for visualizing the execution of the above means on a display device, and means for generating and displaying dynamic character animation to provide an intuitive operating experience through dialogue with the user. This allows the user to easily acquire information and manage their schedule through natural voice dialogue, and further enjoy an intuitive operating experience through visual feedback.
[1029] "User" refers to a person who uses this system to obtain information by voice and perform operations.
[1030] "Voice data" refers to data that is a digital recording of a user's speech.
[1031] "Text data" refers to data in the form of a character string converted from audio data.
[1032] "Answer data" refers to data that contains answers to questions and instructions from the user, which are obtained by analyzing text data.
[1033] The term "voice input device" refers to a device for collecting a user's voice and recording it as voice data.
[1034] "Speech conversion API" refers to an application program interface for converting voice data into text data.
[1035] "Calendar data" refers to data containing a user's schedule information.
[1036] An "alert" refers to a notification that notifies a user that a scheduled time is approaching.
[1037] "Display device" refers to a device that visually presents the execution status of the system and the contents of the dialogue.
[1038] "Dynamic character animation" refers to a moving character that is displayed on a display device to visually represent interaction with a user.
[1039] The present invention relates to a desktop secretary system that intuitively interacts with users, provides information, and manages schedules. This system integrates speech recognition, natural language processing, and speech synthesis technologies to provide users with information efficiently and effectively. Specific embodiments of this system are described below.
[1040] The user speaks
[1041] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather going to be like tomorrow?" This speech is collected through the device's microphone.
[1042] Voice Recognition
[1043] The device collects the user's voice using a built-in microphone and saves it as voice data. The voice data is then sent to an existing voice recognition API (for example, a publicly available voice recognition API) and converted into text data. Specifically, the voice recognition API used is something like the Google Speech-to-Text API.
[1044] Text Analysis
[1045] The device sends the converted text data to a natural language processing API (for example, a publicly available natural language processing API) and analyzes the intent of the user's question or instruction. Appropriate answer data is generated based on the analysis results. For example, if the question is "What's the weather like tomorrow?", the answer data generated based on the analysis results will be "Tomorrow's weather will be sunny." An example of a natural language processing API that can be used is "Google Dialogflow."
[1046] Audio response
[1047] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API), which converts the text data into voice data. The converted voice data is played back to the user through the speaker. This causes the device to respond to the user by voice, saying, "Tomorrow's weather will be sunny." The speech synthesis API used is something like the "Google Text-to-Speech API."
[1048] Calendar import and integration
[1049] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. Calendar data includes events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The device periodically checks the schedule information, and generates an alert message when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API, and notified to the user through the speaker. For example, an alert message saying "There is a meeting in 15 minutes" may be played.
[1050] Visualization by display device
[1051] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move or its eyes may blink when the user speaks or responds. This provides the user with a more intuitive and enjoyable operating experience.
[1052] Specific examples
[1053] For example, consider the following prompt:
[1054] "Look at my calendar, what's next?"
[1055] "Please tell me what the weather will be like tomorrow."
[1056] "I have a meeting at 3 PM. Set a reminder."
[1057] As a result, the system of the present invention not only allows users to obtain information and manage their schedules through natural voice dialogue, but also allows them to enjoy an intuitive operating experience through visual feedback.
[1058] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1059] Step 1:
[1060] The user speaks questions or commands.
[1061] Specific behavior:
[1062] The user speaks to the terminal, "Please tell me what the weather will be like tomorrow."
[1063] Input: User voice input
[1064] Output: Audio data collected by the device's microphone
[1065] Step 2:
[1066] The device stores the voice data and sends it to the voice recognition API.
[1067] Specific behavior:
[1068] The device's microphone records the user's speech as voice data, temporarily stores it in memory, and then sends the voice data to an existing speech recognition API (e.g., Google Speech-to-Text API).
[1069] Input: Audio data
[1070] Output: Audio data sent to the API
[1071] Step 3:
[1072] The server converts the voice data into text data and sends it back to the terminal.
[1073] Specific behavior:
[1074] The speech recognition API analyzes the voice data and converts it into text data such as "Please tell me what the weather will be like tomorrow." The text data is then sent from the server to the device.
[1075] Input: Audio data
[1076] Output: Text data
[1077] Step 4:
[1078] The device sends the text data to a natural language processing API.
[1079] Specific behavior:
[1080] The device sends the received text data to a natural language processing API (e.g., Google Dialogflow).
[1081] Input: Text data
[1082] Output: Text data sent to the API
[1083] Step 5:
[1084] The server analyzes the text data and generates answer data.
[1085] Specific behavior:
[1086] The natural language processing API analyzes the text data, understands the intent, and generates the answer data, "The weather will be sunny tomorrow."
[1087] Input: Text data
[1088] Output: Answer data
[1089] Step 6:
[1090] The device sends the answer data to the speech synthesis API and outputs it as voice data.
[1091] Specific behavior:
[1092] The device sends the generated answer data to a speech synthesis API (for example, Google Text-to-Speech API), which generates speech data saying, "Tomorrow's weather will be sunny."
[1093] Input: Answer data
[1094] Output: Audio data
[1095] Step 7:
[1096] The terminal plays the audio data to the user through the speaker.
[1097] Specific behavior:
[1098] The terminal receives the voice data and responds to the user through the speaker by voice, "Tomorrow's weather will be sunny."
[1099] Input: Audio data
[1100] Output: The audio response that is played to the user
[1101] Step 8:
[1102] The device periodically calls the calendar API to obtain the user's schedule data.
[1103] Specific behavior:
[1104] The device accesses a calendar API (for example, Google Calendar API) to obtain the user's schedule data. For example, it obtains information such as "Meeting tomorrow at 10:00 AM."
[1105] Input: None (periodic processing)
[1106] Output: Retrieved calendar data
[1107] Step 9:
[1108] The device checks the schedule information, and when the scheduled time approaches, it generates an alert message and converts it into voice data.
[1109] Specific behavior:
[1110] The device analyzes the calendar data it has acquired, generates an alert message 15 minutes before the scheduled meeting saying "You have a meeting in 15 minutes," and converts it into voice data using a speech synthesis API.
[1111] Input: Calendar data
[1112] Output: Audio data of the alert message
[1113] Step 10:
[1114] The terminal notifies the user of the alert audio data through the speaker.
[1115] Specific behavior:
[1116] The device plays the generated alert audio data through the speaker and notifies the user that "there is a meeting in 15 minutes."
[1117] Input: Alert audio data
[1118] Output: The alert notification that is played to the user
[1119] Step 11:
[1120] The terminal visualizes the content of the dialogue with the user on a display device and generates and displays dynamic character animations.
[1121] Specific behavior:
[1122] In order for the terminal to provide interactive responses and animations, animations such as the character's mouth movements and blinking eyes are displayed on the display device in accordance with the content of the user's dialogue.
[1123] Input: Dialogue
[1124] Output: Animation displayed on a display device
[1125] (Application example 1)
[1126] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1127] Conventional voice recognition systems provide information based on user voice input, but this is limited to schedule management and general information provision. This makes it difficult to support users in more specific actions and decision-making in their daily lives. In particular, the lack of support for financial management, such as daily payment history and coupon information, is inconvenient for users and is an issue that needs to be resolved.
[1128] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1129] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to start with a user's voice input, output the answer data as voice, obtain payment history data, generate alerts, and visualize them.
[1130] "Voice data" refers to data obtained by converting a user's voice into digital form.
[1131] "Text data" refers to data obtained by converting voice data into character information.
[1132] "Answer data" is data of the answer content generated based on the analysis results of the text data.
[1133] "Payment history data" is data that records a series of information regarding payments made by a user in the past.
[1134] An "alert" is a message that conveys information that requires the user's attention or notification.
[1135] A "display device" is a device for visually presenting information to a user.
[1136] The present invention is a system for recognizing a user's voice and generating, analyzing, and displaying voice data. This system is configured using a variety of hardware and software, and provides intuitive and useful information and schedule management to the user. Detailed embodiments of the system are described below.
[1137] Hardware and Software Configuration
[1138] This system uses the following hardware and software:
[1139] 1. Voice input device
[1140] Hardware: Smartphone microphone
[1141] Description: A device that allows a user to enter verbal instructions or questions.
[1142] 2. Voice Recognition Software
[1143] Software: Google Cloud Speech-to-Text API
[1144] Description: Software for converting audio data into text data.
[1145] 3. Natural Language Processing Software
[1146] Software: Natural Language Processing APIs such as OpenAI GPT-3
[1147] Description: Software for analyzing text data and generating answer data.
[1148] 4. Text-to-Speech Software
[1149] Software: Microsoft Azure Text-to-Speech API
[1150] Description: Software that converts answer data into audio data and provides auditory feedback to the user.
[1151] 5. Payment History Acquisition Software
[1152] Software: Stripe API
[1153] Description: Software for acquiring payment history data and managing user payment history information.
[1154] 6. Display device
[1155] Hardware: Smartphone screen
[1156] Description: A device for displaying visual information to a user.
[1157] System processing procedure
[1158] We will now explain how the system of the present invention works.
[1159] 1. User voice input
[1160] The user gives voice instructions to the smartphone, for example, saying, "Tell me my latest payment history."
[1161] 2. Speech-to-text
[1162] The smartphone sends the voice data to the Google Cloud Speech-to-Text API, which converts it into text data.
[1163] 3. Text Data Analysis
[1164] The text data is analyzed using OpenAI GPT-3 to understand the user's intent. For example, the intent "Check payment history" can be recognized.
[1165] 4. Obtaining the necessary data
[1166] Call the Stripe API to get the latest payment history data.
[1167] 5. Generating Audio Data
[1168] Using the Microsoft Azure Text-to-Speech API, the answer provided to the user is converted into audio data and played back to the user.
[1169] 6. Visual Feedback
[1170] In addition to providing voice responses, payment history and related information will be displayed on the smartphone screen.
[1171] Examples of concrete examples and prompts
[1172] Examples:
[1173] When a user says to their smartphone, "Tell me my recent payment history," the system converts the speech into text, analyzes it, and recognizes the intent of "check payment history." It then calls the Stripe API to retrieve the payment history, provides that information to the user via voice, and displays the payment history on the smartphone screen.
[1174] Example prompt sentence:
[1175] Text sent to the natural language processing API: "What is your recent payment history?"
[1176] Expected intent: "Check payment history"
[1177] Required API call: Stripe API
[1178] Output data: Latest payment history information
[1179] Example of synthesized voice: "The most recent payment history is a payment of XXXX yen on March 15th."
[1180] Display example: "March 15th, XXXX yen" on a smartphone screen.
[1181] This format allows users to intuitively obtain payment history information through voice, making daily financial management more convenient.
[1182] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1183] Step 1:
[1184] The user speaks into the smartphone's microphone, for example, saying, "Tell me my recent payment history."
[1185] Input: User's voice
[1186] Output: Audio data
[1187] Specific operation: The user's voice is collected by the smartphone's microphone and saved as audio data within the application.
[1188] Step 2:
[1189] The device sends the collected voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data.
[1190] Input: Audio data
[1191] Output: Text data
[1192] Specific operation: The application on the device sends the voice data to the cloud service and receives and stores the returned text data.
[1193] Step 3:
[1194] The device then sends the converted text data to a natural language processing API such as OpenAI GPT-3 for analysis. Based on the analysis results, the device understands the user's intent and generates appropriate answer data.
[1195] Input: Text data
[1196] Output: Answer data
[1197] Specific operation: The device sends text data to a natural language processing API, understands the instruction "Tell me my payment history," and generates data to respond.
[1198] Step 4:
[1199] Based on the generated answer data, the terminal calls the Stripe API to obtain payment history data.
[1200] Input: Answer data
[1201] Output: Payment history data
[1202] Specific operation: The terminal calls the Stripe API based on the answer data and obtains the user's latest payment history information.
[1203] Step 5:
[1204] The terminal uses the acquired payment history data to generate a response and converts it into voice data using the Microsoft Azure Text-to-Speech API.
[1205] Input: Payment history data
[1206] Output: Audio data
[1207] Specific operation: The terminal prepares the payment history in text format and sends it to the speech synthesis API to obtain voice data.
[1208] Step 6:
[1209] The terminal reproduces the generated voice data to the user through a speaker.
[1210] Input: Audio data
[1211] Output: The audio the user hears
[1212] Specific operation: The device plays audio data and notifies the user by voice that "The most recent payment history is a payment of XXXX yen on March 15th."
[1213] Step 7:
[1214] At the same time, the terminal displays the acquired payment history information on the smartphone screen.
[1215] Input: Payment history data
[1216] Output: Visual display
[1217] Specific operation: The terminal application reads the payment history data and displays the latest payment history to the user through a visual interface.
[1218] This allows users to intuitively obtain payment history information through both voice input and visual display.
[1219] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1220] The present invention is a system that combines a desktop secretary system that intuitively interacts with a user, provides information, and manages schedules with an emotion engine that recognizes the user's emotions. The program processing of this system is explained below.
[1221] The user speaks
[1222] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather like tomorrow?"
[1223] Voice Recognition
[1224] The device collects the user's voice using a built-in microphone and saves it as voice data, which is then sent to a voice recognition API and converted into text data.
[1225] Text Analysis
[1226] The device sends the converted text data to a natural language processing API, which analyzes the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, in response to the input "What's the weather like tomorrow?", the device generates the answer "Tomorrow's weather will be sunny."
[1227] Audio response
[1228] The device sends the generated answer data to a speech synthesis API, which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the device may respond with a voice saying, "Tomorrow's weather will be sunny."
[1229] Emotion recognition
[1230] The device sends the voice data to the emotion engine to recognize the user's emotions. The emotion engine identifies the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the tone, speed, and strength of the voice.
[1231] Regulating responses based on emotions
[1232] The server then adjusts the response data to be more appropriate based on the emotion recognition results obtained from the emotion engine. For example, if the user expresses anger, the server adjusts the response data to say, "I'm sorry, I'll try my best to help you."
[1233] Calendar import and integration
[1234] The device periodically calls the calendar API to obtain the user's schedule data. From the obtained calendar data, information about the next appointment is saved in the database. For example, a "meeting tomorrow at 10:00 AM" registered by the user is included in the calendar data.
[1235] Alert generation and output
[1236] The server analyzes the calendar data and generates an alert message when it detects an approaching scheduled time. This alert message is also tailored based on emotion. For example, "You have a meeting in 15 minutes" may be supplemented with an emotionally-based message of encouragement depending on the importance and situation, such as "You have a meeting in 15 minutes, good luck!"
[1237] Visualization by display device
[1238] The device is equipped with a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character is used, and animations such as the character's mouth moving and blinking are displayed when the user speaks or responds. The results of user emotion recognition are also visualized. For example, the character's facial expression changes depending on the user's emotion.
[1239] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, by adjusting responses based on the user's emotions, a more satisfying user experience can be provided.
[1240] The processing flow will be explained below.
[1241] Step 1:
[1242] The user speaks to the device, for example, "What's the weather like tomorrow?"
[1243] Step 2:
[1244] The device collects the user's voice using a built-in microphone and saves it as voice data.
[1245] Step 3:
[1246] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[1247] Step 4:
[1248] The device receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends this data to the natural language processing API.
[1249] Step 5:
[1250] The natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[1251] Step 6:
[1252] The device receives the answer data generated by the natural language processing API and sends it to the speech synthesis API, which converts the answer data into voice data.
[1253] Step 7:
[1254] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[1255] Step 8:
[1256] The device sends the voice data to the emotion engine, which analyzes the voice data and recognizes the user's emotions (e.g., joy, anger, sadness, surprise, etc.).
[1257] Step 9:
[1258] The terminal transmits the emotion recognition result obtained from the emotion engine to the server.
[1259] Step 10:
[1260] The server adjusts the answer data based on the emotion recognition results. For example, if the user expresses anger, it generates an additional message saying, "I'm sorry, I'll try to help you."
[1261] Step 11:
[1262] The device then sends the adjusted answer data back to the speech synthesis API, where it is converted into voice data.
[1263] Step 12:
[1264] The device plays the adjusted voice data to the user through the speaker, for example, responding with a voice message saying, "The weather will be sunny tomorrow. Sorry, we'll do our best to help you."
[1265] Step 13:
[1266] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[1267] Step 14:
[1268] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[1269] Step 15:
[1270] The device sends the generated alert information to a speech synthesis API, which converts it into voice data.
[1271] Step 16:
[1272] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[1273] Step 17:
[1274] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[1275] Step 18:
[1276] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[1277] Example 2
[1278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1279] Conventional voice recognition systems and digital assistants provide information and schedule management based on voice input, but they do not provide responses that take the user's emotions into account, making it difficult to provide a satisfying user experience. Furthermore, schedule notifications are monotonous, requiring flexible responses that respond to the user's emotions and circumstances. The present invention aims to solve these problems by providing a system that provides integrated responses that take the user's emotions into account and schedule management.
[1280] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to provide information based on the user's voice input. Furthermore, it includes means for recognizing a user's emotion, means for adjusting the answer data based on the emotion recognition result, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This makes it possible to provide an integrated response and schedule management that takes the user's emotion into consideration.
[1281] "Voice Data" refers to information that is stored and processed digitally from a user's voice.
[1282] "Text data" refers to voice data converted into character string information, and refers to user questions and instructions that are handled as character information.
[1283] "Answer data" refers to answer information to a user's question or instruction, which is generated based on the analyzed text data.
[1284] A "voice recognition program" refers to software or algorithms for converting voice data into text data.
[1285] An "emotion engine" refers to a program or algorithm that analyzes and recognizes a user's emotional state from voice and text data.
[1286] "Calendar data" refers to digital data containing a user's schedule information.
[1287] The term "display device" refers to an output device, such as a monitor or display, for visually interacting with a user or providing information.
[1288] An "alert message" refers to a message that is sent based on the user's schedule information, and the content of the message is adjusted according to the user's emotional state and situation.
[1289] "User" refers to a person who uses this system to input voice and receive information provision and schedule management services.
[1290] "Terminal" refers to a device that receives a user's voice input, generates voice data, generates text data, plays back the voice data, recognizes emotions, and visualizes the data on a display device.
[1291] "Server" refers to a computer system that has the function of adjusting answer data based on emotion recognition results and the function of acquiring and analyzing calendar data.
[1292] This invention relates to a desktop secretary system that provides information and manages schedules through voice dialogue with a user. The system also has the function of recognizing the user's emotions and adjusting responses based on the results.
[1293] This system operates using the following hardware and software.
[1294] Hardware used
[1295] Device: Desktop or laptop computer.
[1296] Sound recording device (microphone): A device for collecting the user's voice.
[1297] Display Device: A monitor or display for visualizing user interactions.
[1298] Speaker: A device for playing back generated audio data.
[1299] Software used
[1300] Speech recognition program (API): For example, the Google Speech-to-Text API, which is used to convert voice data into text data.
[1301] Natural language processing program (API): For example, the OpenAI GPT-4 API used to analyze text data and generate answer data.
[1302] Speech synthesis program (API): For example, Google Text-to-Speech API, which is used to convert the generated answer data into audio data.
[1303] Emotion recognition engine: For example, IBM Watson Tone Analyzer, which is used to recognize user emotions from voice data.
[1304] Calendar management program (API): For example, use the Google Calendar API to retrieve and manage calendar data.
[1305] System Operation Overview
[1306] The user asks questions or gives instructions to the device in natural language. For example, they might say, "What's the weather going to be like tomorrow?" This voice input is collected as voice data through the device's built-in sound recording device (microphone).
[1307] The device sends this voice data to a voice recognition program (Google Speech-to-Text API), which converts the voice data into text data. The converted text data is then sent to a natural language processing program (OpenAI GPT-4 API), which analyzes the intent of the user's question or instruction. Based on the results of this analysis, for example, in response to an input such as "What's the weather like tomorrow?", the answer "Tomorrow's weather will be sunny" is generated.
[1308] This answer data is sent to a speech synthesis program (Google Text-to-Speech API), which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the answer may be "Tomorrow's weather will be sunny."
[1309] The device sends the voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. Based on the emotion recognition results, the server adjusts the response data to be more appropriate. For example, if the user expresses anger, the response may be adjusted to say, "I'm sorry, I'll try my best to help you."
[1310] The device also periodically calls the calendar management program (Google Calendar API) to obtain the user's schedule data. This data is analyzed, and information about upcoming appointments is saved in the database. For example, if the calendar data includes a "meeting tomorrow at 10:00 AM," this information is saved.
[1311] The server analyzes calendar data and generates alert messages when an appointment time approaches, which are also tailored based on emotion: for example, "You have a meeting in 15 minutes," plus a message like "Good luck!" depending on the importance and situation.
[1312] The device uses a display device to visualize the dialogue with the user. Interactive responses and animations are provided, such as the character's mouth moving or blinking when the user speaks or responds. Furthermore, the user's emotion recognition results are also visualized, and the character's facial expression changes depending on the user's emotion.
[1313] Examples of prompt statements
[1314] When a user says, "What's the weather going to be like tomorrow?", the device performs voice recognition and generates text data saying, "What's the weather going to be like tomorrow?"
[1315] Next, the terminal analyzes the text data and generates answer data such as "Tomorrow's weather will be sunny."
[1316] The answer data is converted into voice data via a voice synthesis program, and the answer "Tomorrow's weather will be sunny" is spoken.
[1317] If the emotion recognition engine determines that the user is excited, the device will respond with a tailored response: "Looks like it's going to be a great day, have fun!"
[1318] As described above, this invention starts with a user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, adjusting responses based on the user's emotions can provide a more satisfying user experience.
[1319] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1320] Step 1:
[1321] The user speaks a question or instruction to the device in natural language. For example, they might say, "What's the weather like tomorrow?" This voice input becomes the initial input to the system.
[1322] Input: User's voice input ("What's the weather like tomorrow?")
[1323] Output: None (pass to next step)
[1324] Step 2:
[1325] The device collects the user's voice using a built-in audio recording device (microphone) and saves it as voice data.The device then sends the voice data to a voice recognition program (Google Speech-to-Text API) and converts it into text data.
[1326] Input: User's voice data
[1327] Output: Text data ("What's the weather like tomorrow?")
[1328] Specific operation: The user's voice is collected by the microphone and sent to the speech recognition API to obtain text data such as "What's the weather like tomorrow?"
[1329] Step 3:
[1330] The device sends the converted text data to a natural language processing program (OpenAI GPT-4 API) and analyzes the information the user is looking for. Based on the analysis results, the device calls an external weather information API (such as the OpenWeatherMap API) to obtain weather data. Answer data is generated based on the analysis results and the obtained data.
[1331] Input: Text data ("What's the weather like tomorrow?")
[1332] Output: Answer data ("Tomorrow's weather will be sunny")
[1333] Specific operation: The text data is analyzed using a natural language processing API to analyze the intent, a weather information API is called to obtain weather data, and the answer data "Tomorrow's weather will be sunny" is generated.
[1334] Step 4:
[1335] The device sends the generated answer data to a speech synthesis program (Google Text-to-Speech API), which converts the text data into audio data, which is then played back to the user through the speaker.
[1336] Input: Answer data ("Tomorrow's weather will be sunny")
[1337] Output: Voice data (voice response: "Tomorrow's weather will be sunny")
[1338] Specific operation: The answer data is converted into audio data using a speech synthesis API, and the speaker responds by saying, "Tomorrow's weather will be sunny."
[1339] Step 5:
[1340] The device sends the voice data collected during the conversation with the user to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion recognition engine identifies the user's emotions from the tone, speed, and strength of the voice.
[1341] Input: User's voice data
[1342] Output: Emotion recognition data ("User is excited")
[1343] Specific operation: Sends voice data to an emotion recognition engine to recognize the user's state of excitement.
[1344] Step 6:
[1345] The server then adjusts the answer data to be more appropriate based on the emotion recognition results obtained from the emotion recognition engine. For example, if the user is excited, the server adjusts the response to something like, "It looks like it's going to be a great day, have fun!"
[1346] Input: Emotion recognition data ("user is excited"), Answer data
[1347] Output: Adjusted answer data ("It looks like you're having a great day, have fun!")
[1348] Specific operation: Adjust answer data based on emotion recognition data to generate a response that takes into account the user's emotions.
[1349] Step 7:
[1350] The device periodically calls the calendar program (Google Calendar API) to obtain the user's schedule data. The obtained calendar data is analyzed and information about the next schedule is saved in the database. For example, if a user schedules a meeting at 10:00 AM tomorrow, that information is saved.
[1351] Input: None (or a periodic timer condition)
[1352] Output: Calendar data ("Meeting tomorrow at 10 AM")
[1353] Specific behavior: Calls the Google Calendar API to obtain new schedule data.
[1354] Step 8:
[1355] The server analyzes the calendar data and generates an alert message when it detects an approaching appointment time, which is also tailored based on emotion.
[1356] Input: Calendar data ("Meeting tomorrow at 10 AM")
[1357] Output: Alert message ("You have a meeting in 15 minutes. Good luck!")
[1358] Specific behavior: Check the appointment time and generate an alert message to notify you.
[1359] Step 9:
[1360] The device visualizes conversations with the user and schedule information on a display device. When the user speaks or responds, interactive responses such as character animations are displayed. Emotion recognition results are also displayed.
[1361] Input: User emotion recognition data, dialogue data
[1362] Output: Visually displayed animations and character facial expressions
[1363] Specific operation: Animation is displayed on the display device, and the character's facial expression changes based on the user's emotions and the content of the dialogue.
[1364] (Application example 2)
[1365] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1366] While conventional factory robot assistant systems can provide natural voice interaction with workers and schedule management, they lack the ability to recognize workers' emotions and optimize responses based on them, which hinders improvements in user experience and work efficiency. Additionally, changes to work processes and alert notifications within the factory are difficult to intuitively understand. This can increase the burden on workers and ultimately reduce production efficiency.
[1367] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for performing emotion recognition and generating emotion data, means for adjusting answer data based on the emotion data, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This not only enables natural voice dialogue with the worker, but also optimizes responses based on emotion recognition and makes it possible to provide intuitively understandable alert notifications and work processes.
[1368] "Speech recognition" is a technology that converts a user's voice into text data.
[1369] "Voice data" is digital sound information generated by voice recognition.
[1370] "Text data" is character information obtained by analyzing voice data.
[1371] "Answer data" is answer information to questions and instructions generated by analyzing text data.
[1372] "Speech synthesis" is a technology that converts text data into voice data.
[1373] "Emotion recognition" is a technology that generates emotional data from a user's voice or text.
[1374] "Emotion data" is information indicating the emotional state of the user obtained by emotion recognition.
[1375] "Schedule information" is time information for tasks and plans obtained from calendar data.
[1376] An "alert" is a notification message that is generated based on schedule information.
[1377] A "display device" is a device for providing information visually.
[1378] The present invention relates to a smart factory robot assistant system that intuitively interacts with workers in a factory to improve work efficiency. This system has speech recognition and emotion recognition functions, and provides schedule management and work instructions. A specific example of the system is shown below.
[1379] Hardware Configuration
[1380] The system of the present invention consists of smart glasses worn by a user (worker), a microphone, an internal speaker, and a display device. The server also includes multiple APIs and data storage.
[1381] Software Configuration
[1382] The software used is as follows:
[1383] Speech recognition API (e.g. Google Speech-to-Text)
[1384] Natural Language Processing API (e.g. Google NLP API)
[1385] Text-to-speech API (e.g. Google Text-to-Speech)
[1386] Emotion recognition API (e.g. Affectiva)
[1387] Calendar API (e.g. Google Calendar API)
[1388] System Operation
[1389] The system operates in the following manner.
[1390] 1. Voice Recognition
[1391] When a user speaks to the smart glasses, the built-in microphone collects the voice, which is then sent to a speech recognition API and converted into text data.
[1392] 2. Text Analysis
[1393] The text data is analyzed using a natural language processing API and used to understand the intent of the user's questions and instructions. Appropriate answer data is generated based on the analysis results.
[1394] 3. Emotional Recognition
[1395] The created text data and voice tone are sent to an emotion recognition API to generate user emotion data, which indicates the user's emotions (joy, anger, sadness, etc.).
[1396] 4. Adjusting response data
[1397] Based on the emotion data, the generated answer data is adjusted. For example, if the user expresses anger, an apology or an encouraging message is added to the response.
[1398] 5. Audio and visual representation
[1399] The adjusted answer data is converted into audio data using a speech synthesis API and played back through the smart glasses' speakers, while visual feedback such as work progress and alerts is displayed on the smart glasses' display.
[1400] 6. Schedule Management
[1401] Schedule information is retrieved using the calendar API, and alerts are generated when important events or deadlines are approaching, with audio and visual notifications.
[1402] Specific examples
[1403] For example, if a worker says, "Please tell me the next work step," their intention is analyzed using voice recognition and natural language processing technology. Based on the analysis results, answer data that provides the next work step is generated. Furthermore, emotion recognition is used to determine the worker's emotions, so if the worker is tired, an appropriate response will be given according to their emotions, such as, "The next task is △△. You seem tired, let's take a short break."
[1404] Prompt Sentence Examples
[1405] An example of a prompt sentence to input to the generative AI model is as follows:
[1406] If the question "Please tell me the next work step" is input, please output the answer text as "The next work step is ____." Also, if the questioner is feeling sad, please add a message of encouragement.
[1407] As described above, the system of the present invention integrates voice input, voice recognition, natural language processing, voice synthesis, emotion recognition, emotion-based response adjustment, schedule linkage, and visual feedback. This allows users to effectively acquire information through natural voice dialogue and improve work efficiency. In addition, responses that take the user's emotions into consideration can provide a more comfortable work environment.
[1408] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1409] Step 1:
[1410] The user speaks to the smart glasses. The built-in microphone of the smart glasses collects the user's voice and transmits it to the terminal as voice data. The input is the user's voice, and the output is voice data.
[1411] Step 2:
[1412] The device sends voice data to the voice recognition API. The API analyzes the voice data and converts it into text data. In this process, the voice data is analyzed and text data is generated as character information. The input is voice data and the output is text data.
[1413] Step 3:
[1414] The device sends the generated text data to a natural language processing API. The natural language processing API analyzes the text data and understands the intent of the user's question or instruction. This process generates appropriate answer data based on the analyzed intent. The input is text data, and the output is answer data.
[1415] Step 4:
[1416] The device sends the generated answer data to the emotion recognition API, and simultaneously analyzes the worker's voice tone. The emotion recognition API generates emotion data and identifies the user's emotion (e.g., joy, anger, sadness, etc.). The input is the answer data and voice tone, and the output is emotion data.
[1417] Step 5:
[1418] The server adjusts the answer data based on the emotion data. For example, if the user expresses anger, it adds an apology or comforting message to the answer data. The input is the answer data and emotion data, and the output is the adjusted answer data.
[1419] Step 6:
[1420] The device sends the adjusted answer data to a speech synthesis API, which converts the text data into speech data. In this process, the adjusted answer data is converted into speech. The input is the adjusted answer data, and the output is speech data.
[1421] Step 7:
[1422] The terminal plays the generated voice data through the speaker of the smart glasses, allowing the user to receive a voice response. The input is voice data, and the output is voice output.
[1423] Step 8:
[1424] The server periodically calls the calendar API to retrieve calendar data. The retrieved calendar data is parsed and schedule information is extracted. The input is the calendar data, and the output is the schedule information.
[1425] Step 9:
[1426] The server generates alerts based on schedule information and sends the alert data to the terminal. Alert messages are generated when important events or work deadlines approach. The input is schedule information and the output is alert data.
[1427] Step 10:
[1428] The device sends the generated alert data to the speech synthesis API, which converts the alert message into voice data. It also generates data to display the alert content on the smart glasses display. The input is the alert data, and the output is voice data and display data.
[1429] Step 11:
[1430] The device notifies the user of the alert through the smart glasses' speaker and displays the alert content on the display. The input is audio data and display data, and the output is audio and display notifications.
[1431] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1432] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1433] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1434] [Fourth embodiment]
[1435] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1436] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1437] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1438] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1439] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1440] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1441] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1442] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1443] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1444] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1445] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1446] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1447] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1448] The present invention aims to realize a desktop secretary system that can intuitively interact with a user, provide information, and manage schedules. The program processing of this system will be explained below.
[1449] The user speaks
[1450] The user asks questions and gives instructions to the device in natural language. For example, say, "What's the weather like tomorrow?"
[1451] Voice Recognition
[1452] The device collects the user's voice using a built-in microphone, saves it as voice data, and then sends the voice data to a voice recognition API (e.g., a publicly available voice recognition API) to convert the voice data into text data.
[1453] Text Analysis
[1454] The device sends the converted text data to a natural language processing API (e.g., a publicly available natural language processing API) to analyze the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, based on the analysis result of "What's the weather going to be like tomorrow?", it generates answer data such as "Tomorrow's weather will be sunny."
[1455] Audio response
[1456] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API) to convert the text data into voice data. The converted voice data is played back to the user through the speaker. As a result, the device responds to the user by voice, saying, "Tomorrow's weather will be sunny."
[1457] Calendar import and integration
[1458] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. The calendar data contains events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The schedule information is periodically checked, and an alert message is generated when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API and notified to the user through the speaker. This causes the device to alert the user by voice, saying, "You have a meeting in 15 minutes."
[1459] Visualization by display device
[1460] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move, blink, or other animations may be displayed when the user speaks or responds. This provides a more intuitive and enjoyable operating experience for the user.
[1461] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, calendar integration, and visual feedback. This allows the user to effectively obtain information and manage their schedule through natural voice dialogue and intuitive interaction without the need for complex operations.
[1462] The processing flow will be explained below.
[1463] Step 1:
[1464] The user speaks to the device, for example, "What's the weather like tomorrow?"
[1465] Step 2:
[1466] The device collects the user's voice using a built-in microphone and saves it as voice data.
[1467] Step 3:
[1468] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[1469] Step 4:
[1470] The server receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends it to the natural language processing API.
[1471] Step 5:
[1472] A natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[1473] Step 6:
[1474] The server receives the generated answer data and transmits it to the terminal.
[1475] Step 7:
[1476] The device sends the answer data to the speech synthesis API, which converts the answer data into voice data.
[1477] Step 8:
[1478] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[1479] Step 9:
[1480] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[1481] Step 10:
[1482] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[1483] Step 11:
[1484] The server sends the generated alert information to the terminal.
[1485] Step 12:
[1486] The device sends the alert information to a speech synthesis API, which converts it into voice data.
[1487] Step 13:
[1488] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[1489] Step 14:
[1490] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[1491] Step 15:
[1492] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[1493] Example 1
[1494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1495] In modern society, there is a demand for systems that users can operate intuitively and that allow them to acquire information and manage schedules without requiring complex operations. However, existing systems are complex to operate, and many users are unable to master them. In addition, they lack an intuitive operating experience and visual feedback through user interaction, making them insufficient to improve user satisfaction.
[1496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1497] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for analyzing the text data and generating answer data, means for outputting the generated answer data as voice data, means for acquiring calendar data and generating and outputting an alert including schedule information, means for visualizing the execution of the above means on a display device, and means for generating and displaying dynamic character animation to provide an intuitive operating experience through dialogue with the user. This allows the user to easily acquire information and manage their schedule through natural voice dialogue, and further enjoy an intuitive operating experience through visual feedback.
[1498] "User" refers to a person who uses this system to obtain information by voice and perform operations.
[1499] "Voice data" refers to data that is a digital recording of a user's speech.
[1500] "Text data" refers to data in the form of a character string converted from audio data.
[1501] "Answer data" refers to data that contains answers to questions and instructions from the user, which are obtained by analyzing text data.
[1502] The term "voice input device" refers to a device for collecting a user's voice and recording it as voice data.
[1503] "Speech conversion API" refers to an application program interface for converting voice data into text data.
[1504] "Calendar data" refers to data containing a user's schedule information.
[1505] An "alert" refers to a notification that notifies a user that a scheduled time is approaching.
[1506] "Display device" refers to a device that visually presents the execution status of the system and the contents of the dialogue.
[1507] "Dynamic character animation" refers to a moving character that is displayed on a display device to visually represent interaction with a user.
[1508] The present invention relates to a desktop secretary system that intuitively interacts with users, provides information, and manages schedules. This system integrates speech recognition, natural language processing, and speech synthesis technologies to provide users with information efficiently and effectively. Specific embodiments of this system are described below.
[1509] The user speaks
[1510] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather going to be like tomorrow?" This speech is collected through the device's microphone.
[1511] Voice Recognition
[1512] The device collects the user's voice using a built-in microphone and saves it as voice data. The voice data is then sent to an existing voice recognition API (for example, a publicly available voice recognition API) and converted into text data. Specifically, the voice recognition API used is something like the Google Speech-to-Text API.
[1513] Text Analysis
[1514] The device sends the converted text data to a natural language processing API (for example, a publicly available natural language processing API) and analyzes the intent of the user's question or instruction. Appropriate answer data is generated based on the analysis results. For example, if the question is "What's the weather like tomorrow?", the answer data generated based on the analysis results will be "Tomorrow's weather will be sunny." An example of a natural language processing API that can be used is "Google Dialogflow."
[1515] Audio response
[1516] The device sends the generated answer data to a speech synthesis API (for example, a publicly available speech synthesis API), which converts the text data into voice data. The converted voice data is played back to the user through the speaker. This causes the device to respond to the user by voice, saying, "Tomorrow's weather will be sunny." The speech synthesis API used is something like the "Google Text-to-Speech API."
[1517] Calendar import and integration
[1518] The device periodically calls a calendar API (for example, a publicly available calendar API) to obtain the user's schedule data. Calendar data includes events entered by the user. For example, a "meeting tomorrow at 10:00 AM" registered by the user is obtained as calendar data. The device periodically checks the schedule information, and generates an alert message when the scheduled time approaches. This alert message is also converted into audio data using a speech synthesis API, and notified to the user through the speaker. For example, an alert message saying "There is a meeting in 15 minutes" may be played.
[1519] Visualization by display device
[1520] The device includes a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character may be used, and the character's mouth may move or its eyes may blink when the user speaks or responds. This provides the user with a more intuitive and enjoyable operating experience.
[1521] Specific examples
[1522] For example, consider the following prompt:
[1523] "Look at my calendar, what's next?"
[1524] "Please tell me what the weather will be like tomorrow."
[1525] "I have a meeting at 3 PM. Set a reminder."
[1526] As a result, the system of the present invention not only allows users to obtain information and manage their schedules through natural voice dialogue, but also allows them to enjoy an intuitive operating experience through visual feedback.
[1527] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1528] Step 1:
[1529] The user speaks questions or commands.
[1530] Specific behavior:
[1531] The user speaks to the terminal, "Please tell me what the weather will be like tomorrow."
[1532] Input: User voice input
[1533] Output: Audio data collected by the device's microphone
[1534] Step 2:
[1535] The device stores the voice data and sends it to the voice recognition API.
[1536] Specific behavior:
[1537] The device's microphone records the user's speech as voice data, temporarily stores it in memory, and then sends the voice data to an existing speech recognition API (e.g., Google Speech-to-Text API).
[1538] Input: Audio data
[1539] Output: Audio data sent to the API
[1540] Step 3:
[1541] The server converts the voice data into text data and sends it back to the terminal.
[1542] Specific behavior:
[1543] The speech recognition API analyzes the voice data and converts it into text data such as "Please tell me what the weather will be like tomorrow." The text data is then sent from the server to the device.
[1544] Input: Audio data
[1545] Output: Text data
[1546] Step 4:
[1547] The device sends the text data to a natural language processing API.
[1548] Specific behavior:
[1549] The device sends the received text data to a natural language processing API (e.g., Google Dialogflow).
[1550] Input: Text data
[1551] Output: Text data sent to the API
[1552] Step 5:
[1553] The server analyzes the text data and generates answer data.
[1554] Specific behavior:
[1555] The natural language processing API analyzes the text data, understands the intent, and generates the answer data, "The weather will be sunny tomorrow."
[1556] Input: Text data
[1557] Output: Answer data
[1558] Step 6:
[1559] The device sends the answer data to the speech synthesis API and outputs it as voice data.
[1560] Specific behavior:
[1561] The device sends the generated answer data to a speech synthesis API (for example, Google Text-to-Speech API), which generates speech data saying, "Tomorrow's weather will be sunny."
[1562] Input: Answer data
[1563] Output: Audio data
[1564] Step 7:
[1565] The terminal plays the audio data to the user through the speaker.
[1566] Specific behavior:
[1567] The terminal receives the voice data and responds to the user through the speaker by voice, "Tomorrow's weather will be sunny."
[1568] Input: Audio data
[1569] Output: The audio response that is played to the user
[1570] Step 8:
[1571] The device periodically calls the calendar API to obtain the user's schedule data.
[1572] Specific behavior:
[1573] The device accesses a calendar API (for example, Google Calendar API) to obtain the user's schedule data. For example, it obtains information such as "Meeting tomorrow at 10:00 AM."
[1574] Input: None (periodic processing)
[1575] Output: Retrieved calendar data
[1576] Step 9:
[1577] The device checks the schedule information, and when the scheduled time approaches, it generates an alert message and converts it into voice data.
[1578] Specific behavior:
[1579] The device analyzes the calendar data it has acquired, generates an alert message 15 minutes before the scheduled meeting saying "You have a meeting in 15 minutes," and converts it into voice data using a speech synthesis API.
[1580] Input: Calendar data
[1581] Output: Audio data of the alert message
[1582] Step 10:
[1583] The terminal notifies the user of the alert audio data through the speaker.
[1584] Specific behavior:
[1585] The device plays the generated alert audio data through the speaker and notifies the user that "there is a meeting in 15 minutes."
[1586] Input: Alert audio data
[1587] Output: The alert notification that is played to the user
[1588] Step 11:
[1589] The terminal visualizes the content of the dialogue with the user on a display device and generates and displays dynamic character animations.
[1590] Specific behavior:
[1591] In order for the terminal to provide interactive responses and animations, animations such as the character's mouth movements and blinking eyes are displayed on the display device in accordance with the content of the user's dialogue.
[1592] Input: Dialogue
[1593] Output: Animation displayed on a display device
[1594] (Application example 1)
[1595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1596] Conventional voice recognition systems provide information based on user voice input, but this is limited to schedule management and general information provision. This makes it difficult to support users in more specific actions and decision-making in their daily lives. In particular, the lack of support for financial management, such as daily payment history and coupon information, is inconvenient for users and is an issue that needs to be resolved.
[1597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1598] In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to start with a user's voice input, output the answer data as voice, obtain payment history data, generate alerts, and visualize them.
[1599] "Voice data" refers to data obtained by converting a user's voice into digital form.
[1600] "Text data" refers to data obtained by converting voice data into character information.
[1601] "Answer data" is data of the answer content generated based on the analysis results of the text data.
[1602] "Payment history data" is data that records a series of information regarding payments made by a user in the past.
[1603] An "alert" is a message that conveys information that requires the user's attention or notification.
[1604] A "display device" is a device for visually presenting information to a user.
[1605] The present invention is a system for recognizing a user's voice and generating, analyzing, and displaying voice data. This system is configured using a variety of hardware and software, and provides intuitive and useful information and schedule management to the user. Detailed embodiments of the system are described below.
[1606] Hardware and Software Configuration
[1607] This system uses the following hardware and software:
[1608] 1. Voice input device
[1609] Hardware: Smartphone microphone
[1610] Description: A device that allows a user to enter verbal instructions or questions.
[1611] 2. Voice Recognition Software
[1612] Software: Google Cloud Speech-to-Text API
[1613] Description: Software for converting audio data into text data.
[1614] 3. Natural Language Processing Software
[1615] Software: Natural Language Processing APIs such as OpenAI GPT-3
[1616] Description: Software for analyzing text data and generating answer data.
[1617] 4. Text-to-Speech Software
[1618] Software: Microsoft Azure Text-to-Speech API
[1619] Description: Software that converts answer data into audio data and provides auditory feedback to the user.
[1620] 5. Payment History Acquisition Software
[1621] Software: Stripe API
[1622] Description: Software for acquiring payment history data and managing user payment history information.
[1623] 6. Display device
[1624] Hardware: Smartphone screen
[1625] Description: A device for displaying visual information to a user.
[1626] System processing procedure
[1627] We will now explain how the system of the present invention works.
[1628] 1. User voice input
[1629] The user gives voice instructions to the smartphone, for example, saying, "Tell me my latest payment history."
[1630] 2. Speech-to-text
[1631] The smartphone sends the voice data to the Google Cloud Speech-to-Text API, which converts it into text data.
[1632] 3. Text Data Analysis
[1633] The text data is analyzed using OpenAI GPT-3 to understand the user's intent. For example, the intent "Check payment history" can be recognized.
[1634] 4. Obtaining the necessary data
[1635] Call the Stripe API to get the latest payment history data.
[1636] 5. Generating Audio Data
[1637] Using the Microsoft Azure Text-to-Speech API, the answer provided to the user is converted into audio data and played back to the user.
[1638] 6. Visual Feedback
[1639] In addition to providing voice responses, payment history and related information will be displayed on the smartphone screen.
[1640] Examples of concrete examples and prompts
[1641] Examples:
[1642] When a user says to their smartphone, "Tell me my recent payment history," the system converts the speech into text, analyzes it, and recognizes the intent of "check payment history." It then calls the Stripe API to retrieve the payment history, provides that information to the user via voice, and displays the payment history on the smartphone screen.
[1643] Example prompt sentence:
[1644] Text sent to the natural language processing API: "What is your recent payment history?"
[1645] Expected intent: "Check payment history"
[1646] Required API call: Stripe API
[1647] Output data: Latest payment history information
[1648] Example of synthesized voice: "The most recent payment history is a payment of XXXX yen on March 15th."
[1649] Display example: "March 15th, XXXX yen" on a smartphone screen.
[1650] This format allows users to intuitively obtain payment history information through voice, making daily financial management more convenient.
[1651] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1652] Step 1:
[1653] The user speaks into the smartphone's microphone, for example, saying, "Tell me my recent payment history."
[1654] Input: User's voice
[1655] Output: Audio data
[1656] Specific operation: The user's voice is collected by the smartphone's microphone and saved as audio data within the application.
[1657] Step 2:
[1658] The device sends the collected voice data to the Google Cloud Speech-to-Text API, which converts the voice data into text data.
[1659] Input: Audio data
[1660] Output: Text data
[1661] Specific operation: The application on the device sends the voice data to the cloud service and receives and stores the returned text data.
[1662] Step 3:
[1663] The device then sends the converted text data to a natural language processing API such as OpenAI GPT-3 for analysis. Based on the analysis results, the device understands the user's intent and generates appropriate answer data.
[1664] Input: Text data
[1665] Output: Answer data
[1666] Specific operation: The device sends text data to a natural language processing API, understands the instruction "Tell me my payment history," and generates data to respond.
[1667] Step 4:
[1668] Based on the generated answer data, the terminal calls the Stripe API to obtain payment history data.
[1669] Input: Answer data
[1670] Output: Payment history data
[1671] Specific operation: The terminal calls the Stripe API based on the answer data and obtains the user's latest payment history information.
[1672] Step 5:
[1673] The terminal uses the acquired payment history data to generate a response and converts it into voice data using the Microsoft Azure Text-to-Speech API.
[1674] Input: Payment history data
[1675] Output: Audio data
[1676] Specific operation: The terminal prepares the payment history in text format and sends it to the speech synthesis API to obtain voice data.
[1677] Step 6:
[1678] The terminal reproduces the generated voice data to the user through a speaker.
[1679] Input: Audio data
[1680] Output: The audio the user hears
[1681] Specific operation: The device plays audio data and notifies the user by voice that "The most recent payment history is a payment of XXXX yen on March 15th."
[1682] Step 7:
[1683] At the same time, the terminal displays the acquired payment history information on the smartphone screen.
[1684] Input: Payment history data
[1685] Output: Visual display
[1686] Specific operation: The terminal application reads the payment history data and displays the latest payment history to the user through a visual interface.
[1687] This allows users to intuitively obtain payment history information through both voice input and visual display.
[1688] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1689] The present invention is a system that combines a desktop secretary system that intuitively interacts with a user, provides information, and manages schedules with an emotion engine that recognizes the user's emotions. The program processing of this system is explained below.
[1690] The user speaks
[1691] The user asks questions or gives instructions to the device in natural language, for example, "What's the weather like tomorrow?"
[1692] Voice Recognition
[1693] The device collects the user's voice using a built-in microphone and saves it as voice data, which is then sent to a voice recognition API and converted into text data.
[1694] Text Analysis
[1695] The device sends the converted text data to a natural language processing API, which analyzes the intent of the user's question or instruction. It then generates appropriate answer data based on the analysis results. For example, in response to the input "What's the weather like tomorrow?", the device generates the answer "Tomorrow's weather will be sunny."
[1696] Audio response
[1697] The device sends the generated answer data to a speech synthesis API, which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the device may respond with a voice saying, "Tomorrow's weather will be sunny."
[1698] Emotion recognition
[1699] The device sends the voice data to the emotion engine to recognize the user's emotions. The emotion engine identifies the user's emotions (e.g., joy, anger, sadness, surprise, etc.) from the tone, speed, and strength of the voice.
[1700] Regulating responses based on emotions
[1701] The server then adjusts the response data to be more appropriate based on the emotion recognition results obtained from the emotion engine. For example, if the user expresses anger, the server adjusts the response data to say, "I'm sorry, I'll try my best to help you."
[1702] Calendar import and integration
[1703] The device periodically calls the calendar API to obtain the user's schedule data. From the obtained calendar data, information about the next appointment is saved in the database. For example, a "meeting tomorrow at 10:00 AM" registered by the user is included in the calendar data.
[1704] Alert generation and output
[1705] The server analyzes the calendar data and generates an alert message when it detects an approaching scheduled time. This alert message is also tailored based on emotion. For example, "You have a meeting in 15 minutes" may be supplemented with an emotionally-based message of encouragement depending on the importance and situation, such as "You have a meeting in 15 minutes, good luck!"
[1706] Visualization by display device
[1707] The device is equipped with a display device to visualize the dialogue with the user, providing interactive responses and animations. For example, an animated character is used, and animations such as the character's mouth moving and blinking are displayed when the user speaks or responds. The results of user emotion recognition are also visualized. For example, the character's facial expression changes depending on the user's emotion.
[1708] As described above, the system of the present invention starts with the user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, by adjusting responses based on the user's emotions, a more satisfying user experience can be provided.
[1709] The processing flow will be explained below.
[1710] Step 1:
[1711] The user speaks to the device, for example, "What's the weather like tomorrow?"
[1712] Step 2:
[1713] The device collects the user's voice using a built-in microphone and saves it as voice data.
[1714] Step 3:
[1715] The device sends the saved voice data to the voice recognition API, which receives the voice data and converts it into text data.
[1716] Step 4:
[1717] The device receives text data (e.g., "What's the weather like tomorrow?") from the speech recognition API and sends this data to the natural language processing API.
[1718] Step 5:
[1719] The natural language processing API analyzes the text data, understands the user's intent, and generates appropriate answer data (e.g., "The weather will be sunny tomorrow").
[1720] Step 6:
[1721] The device receives the answer data generated by the natural language processing API and sends it to the speech synthesis API, which converts the answer data into voice data.
[1722] Step 7:
[1723] The device plays the voice data to the user through the speaker, for example, saying "Tomorrow's weather will be sunny."
[1724] Step 8:
[1725] The device sends the voice data to the emotion engine, which analyzes the voice data and recognizes the user's emotions (e.g., joy, anger, sadness, surprise, etc.).
[1726] Step 9:
[1727] The terminal transmits the emotion recognition result obtained from the emotion engine to the server.
[1728] Step 10:
[1729] The server adjusts the answer data based on the emotion recognition results. For example, if the user expresses anger, it generates an additional message saying, "I'm sorry, I'll try to help you."
[1730] Step 11:
[1731] The device then sends the adjusted answer data back to the speech synthesis API, where it is converted into voice data.
[1732] Step 12:
[1733] The device plays the adjusted voice data to the user through the speaker, for example, responding with a voice message saying, "The weather will be sunny tomorrow. Sorry, we'll do our best to help you."
[1734] Step 13:
[1735] The device periodically calls the calendar API to retrieve the user's calendar data, which includes the schedule information registered by the user.
[1736] Step 14:
[1737] The server parses the calendar data obtained from the calendar API and generates information about upcoming appointments (e.g., "I have a meeting in 15 minutes").
[1738] Step 15:
[1739] The device sends the generated alert information to a speech synthesis API, which converts it into voice data.
[1740] Step 16:
[1741] The device plays the audio data to the user through the speaker, for example, an audio alert saying "I have a meeting in 15 minutes."
[1742] Step 17:
[1743] The device invokes the Live2D engine to display animations in response to user interactions. The animated character is set to move when the user speaks and while analysis is in progress.
[1744] Step 18:
[1745] The device uses the Live2D engine to animate the animated character and provide visual feedback to the user, such as the character's mouth moving and eyes blinking.
[1746] Example 2
[1747] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1748] Conventional voice recognition systems and digital assistants provide information and schedule management based on voice input, but they do not provide responses that take the user's emotions into account, making it difficult to provide a satisfying user experience. Furthermore, schedule notifications are monotonous, requiring flexible responses that respond to the user's emotions and circumstances. The present invention aims to solve these problems by providing a system that provides integrated responses that take the user's emotions into account and schedule management.
[1749] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, and means for analyzing the text data and generating answer data. This makes it possible to provide information based on the user's voice input. Furthermore, it includes means for recognizing a user's emotion, means for adjusting the answer data based on the emotion recognition result, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This makes it possible to provide an integrated response and schedule management that takes the user's emotion into consideration.
[1750] "Voice Data" refers to information that is stored and processed digitally from a user's voice.
[1751] "Text data" refers to voice data converted into character string information, and refers to user questions and instructions that are handled as character information.
[1752] "Answer data" refers to answer information to a user's question or instruction, which is generated based on the analyzed text data.
[1753] A "voice recognition program" refers to software or algorithms for converting voice data into text data.
[1754] An "emotion engine" refers to a program or algorithm that analyzes and recognizes a user's emotional state from voice and text data.
[1755] "Calendar data" refers to digital data containing a user's schedule information.
[1756] The term "display device" refers to an output device, such as a monitor or display, for visually interacting with a user or providing information.
[1757] An "alert message" refers to a message that is sent based on the user's schedule information, and the content of the message is adjusted according to the user's emotional state and situation.
[1758] "User" refers to a person who uses this system to input voice and receive information provision and schedule management services.
[1759] "Terminal" refers to a device that receives a user's voice input, generates voice data, generates text data, plays back the voice data, recognizes emotions, and visualizes the data on a display device.
[1760] "Server" refers to a computer system that has the function of adjusting answer data based on emotion recognition results and the function of acquiring and analyzing calendar data.
[1761] This invention relates to a desktop secretary system that provides information and manages schedules through voice dialogue with a user. The system also has the function of recognizing the user's emotions and adjusting responses based on the results.
[1762] This system operates using the following hardware and software.
[1763] Hardware used
[1764] Device: Desktop or laptop computer.
[1765] Sound recording device (microphone): A device for collecting the user's voice.
[1766] Display Device: A monitor or display for visualizing user interactions.
[1767] Speaker: A device for playing back generated audio data.
[1768] Software used
[1769] Speech recognition program (API): For example, the Google Speech-to-Text API, which is used to convert voice data into text data.
[1770] Natural language processing program (API): For example, the OpenAI GPT-4 API used to analyze text data and generate answer data.
[1771] Speech synthesis program (API): For example, Google Text-to-Speech API, which is used to convert the generated answer data into audio data.
[1772] Emotion recognition engine: For example, IBM Watson Tone Analyzer, which is used to recognize user emotions from voice data.
[1773] Calendar management program (API): For example, use the Google Calendar API to retrieve and manage calendar data.
[1774] System Operation Overview
[1775] The user asks questions or gives instructions to the device in natural language. For example, they might say, "What's the weather going to be like tomorrow?" This voice input is collected as voice data through the device's built-in sound recording device (microphone).
[1776] The device sends this voice data to a voice recognition program (Google Speech-to-Text API), which converts the voice data into text data. The converted text data is then sent to a natural language processing program (OpenAI GPT-4 API), which analyzes the intent of the user's question or instruction. Based on the results of this analysis, for example, in response to an input such as "What's the weather like tomorrow?", the answer "Tomorrow's weather will be sunny" is generated.
[1777] This answer data is sent to a speech synthesis program (Google Text-to-Speech API), which converts the text data into voice data. The converted voice data is then played back to the user through a speaker. For example, the answer may be "Tomorrow's weather will be sunny."
[1778] The device sends the voice data to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. Based on the emotion recognition results, the server adjusts the response data to be more appropriate. For example, if the user expresses anger, the response may be adjusted to say, "I'm sorry, I'll try my best to help you."
[1779] The device also periodically calls the calendar management program (Google Calendar API) to obtain the user's schedule data. This data is analyzed, and information about upcoming appointments is saved in the database. For example, if the calendar data includes a "meeting tomorrow at 10:00 AM," this information is saved.
[1780] The server analyzes calendar data and generates alert messages when an appointment time approaches, which are also tailored based on emotion: for example, "You have a meeting in 15 minutes," plus a message like "Good luck!" depending on the importance and situation.
[1781] The device uses a display device to visualize the dialogue with the user. Interactive responses and animations are provided, such as the character's mouth moving or blinking when the user speaks or responds. Furthermore, the user's emotion recognition results are also visualized, and the character's facial expression changes depending on the user's emotion.
[1782] Examples of prompt statements
[1783] When a user says, "What's the weather going to be like tomorrow?", the device performs voice recognition and generates text data saying, "What's the weather going to be like tomorrow?"
[1784] Next, the terminal analyzes the text data and generates answer data such as "Tomorrow's weather will be sunny."
[1785] The answer data is converted into voice data via a voice synthesis program, and the answer "Tomorrow's weather will be sunny" is spoken.
[1786] If the emotion recognition engine determines that the user is excited, the device will respond with a tailored response: "Looks like it's going to be a great day, have fun!"
[1787] As described above, this invention starts with a user's voice input and integrates speech recognition, natural language processing, speech synthesis, emotion recognition, response adjustment based on emotion, calendar integration, and visual feedback. This allows users to effectively obtain information and manage their schedules through natural voice dialogue and intuitive interaction. Furthermore, adjusting responses based on the user's emotions can provide a more satisfying user experience.
[1788] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1789] Step 1:
[1790] The user speaks a question or instruction to the device in natural language. For example, they might say, "What's the weather like tomorrow?" This voice input becomes the initial input to the system.
[1791] Input: User's voice input ("What's the weather like tomorrow?")
[1792] Output: None (pass to next step)
[1793] Step 2:
[1794] The device collects the user's voice using a built-in audio recording device (microphone) and saves it as voice data.The device then sends the voice data to a voice recognition program (Google Speech-to-Text API) and converts it into text data.
[1795] Input: User's voice data
[1796] Output: Text data ("What's the weather like tomorrow?")
[1797] Specific operation: The user's voice is collected by the microphone and sent to the speech recognition API to obtain text data such as "What's the weather like tomorrow?"
[1798] Step 3:
[1799] The device sends the converted text data to a natural language processing program (OpenAI GPT-4 API) and analyzes the information the user is looking for. Based on the analysis results, the device calls an external weather information API (such as the OpenWeatherMap API) to obtain weather data. Answer data is generated based on the analysis results and the obtained data.
[1800] Input: Text data ("What's the weather like tomorrow?")
[1801] Output: Answer data ("Tomorrow's weather will be sunny")
[1802] Specific operation: The text data is analyzed using a natural language processing API to analyze the intent, a weather information API is called to obtain weather data, and the answer data "Tomorrow's weather will be sunny" is generated.
[1803] Step 4:
[1804] The device sends the generated answer data to a speech synthesis program (Google Text-to-Speech API), which converts the text data into audio data, which is then played back to the user through the speaker.
[1805] Input: Answer data ("Tomorrow's weather will be sunny")
[1806] Output: Voice data (voice response: "Tomorrow's weather will be sunny")
[1807] Specific operation: The answer data is converted into audio data using a speech synthesis API, and the speaker responds by saying, "Tomorrow's weather will be sunny."
[1808] Step 5:
[1809] The device sends the voice data collected during the conversation with the user to an emotion recognition engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The emotion recognition engine identifies the user's emotions from the tone, speed, and strength of the voice.
[1810] Input: User's voice data
[1811] Output: Emotion recognition data ("User is excited")
[1812] Specific operation: Sends voice data to an emotion recognition engine to recognize the user's state of excitement.
[1813] Step 6:
[1814] The server then adjusts the answer data to be more appropriate based on the emotion recognition results obtained from the emotion recognition engine. For example, if the user is excited, the server adjusts the response to something like, "It looks like it's going to be a great day, have fun!"
[1815] Input: Emotion recognition data ("user is excited"), Answer data
[1816] Output: Adjusted answer data ("It looks like you're having a great day, have fun!")
[1817] Specific operation: Adjust answer data based on emotion recognition data to generate a response that takes into account the user's emotions.
[1818] Step 7:
[1819] The device periodically calls the calendar program (Google Calendar API) to obtain the user's schedule data. The obtained calendar data is analyzed and information about the next schedule is saved in the database. For example, if a user schedules a meeting at 10:00 AM tomorrow, that information is saved.
[1820] Input: None (or a periodic timer condition)
[1821] Output: Calendar data ("Meeting tomorrow at 10 AM")
[1822] Specific behavior: Calls the Google Calendar API to obtain new schedule data.
[1823] Step 8:
[1824] The server analyzes the calendar data and generates an alert message when it detects an approaching appointment time, which is also tailored based on emotion.
[1825] Input: Calendar data ("Meeting tomorrow at 10 AM")
[1826] Output: Alert message ("You have a meeting in 15 minutes. Good luck!")
[1827] Specific behavior: Check the appointment time and generate an alert message to notify you.
[1828] Step 9:
[1829] The device visualizes conversations with the user and schedule information on a display device. When the user speaks or responds, interactive responses such as character animations are displayed. Emotion recognition results are also displayed.
[1830] Input: User emotion recognition data, dialogue data
[1831] Output: Visually displayed animations and character facial expressions
[1832] Specific operation: Animation is displayed on the display device, and the character's facial expression changes based on the user's emotions and the content of the dialogue.
[1833] (Application example 2)
[1834] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1835] While conventional factory robot assistant systems can provide natural voice interaction with workers and schedule management, they lack the ability to recognize workers' emotions and optimize responses based on them, which hinders improvements in user experience and work efficiency. Additionally, changes to work processes and alert notifications within the factory are difficult to intuitively understand. This can increase the burden on workers and ultimately reduce production efficiency.
[1836] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's voice and generating voice data, means for converting the generated voice data into text data, means for performing emotion recognition and generating emotion data, means for adjusting answer data based on the emotion data, means for acquiring calendar data and generating and outputting an alert including schedule information, and means for visualizing the execution of the above means on a display device. This not only enables natural voice dialogue with the worker, but also optimizes responses based on emotion recognition and makes it possible to provide intuitively understandable alert notifications and work processes.
[1837] "Speech recognition" is a technology that converts a user's voice into text data.
[1838] "Voice data" is digital sound information generated by voice recognition.
[1839] "Text data" is character information obtained by analyzing voice data.
[1840] "Answer data" is answer information to questions and instructions generated by analyzing text data.
[1841] "Speech synthesis" is a technology that converts text data into voice data.
[1842] "Emotion recognition" is a technology that generates emotional data from a user's voice or text.
[1843] "Emotion data" is information indicating the emotional state of the user obtained by emotion recognition.
[1844] "Schedule information" is time information for tasks and plans obtained from calendar data.
[1845] An "alert" is a notification message that is generated based on schedule information.
[1846] A "display device" is a device for providing information visually.
[1847] The present invention relates to a smart factory robot assistant system that intuitively interacts with workers in a factory to improve work efficiency. This system has speech recognition and emotion recognition functions, and provides schedule management and work instructions. A specific example of the system is shown below.
[1848] Hardware Configuration
[1849] The system of the present invention consists of smart glasses worn by a user (worker), a microphone, an internal speaker, and a display device. The server also includes multiple APIs and data storage.
[1850] Software Configuration
[1851] The software used is as follows:
[1852] Speech recognition API (e.g. Google Speech-to-Text)
[1853] Natural Language Processing API (e.g. Google NLP API)
[1854] Text-to-speech API (e.g. Google Text-to-Speech)
[1855] Emotion recognition API (e.g. Affectiva)
[1856] Calendar API (e.g. Google Calendar API)
[1857] System Operation
[1858] The system operates in the following manner.
[1859] 1. Voice Recognition
[1860] When a user speaks to the smart glasses, the built-in microphone collects the voice, which is then sent to a speech recognition API and converted into text data.
[1861] 2. Text Analysis
[1862] The text data is analyzed using a natural language processing API and used to understand the intent of the user's questions and instructions. Appropriate answer data is generated based on the analysis results.
[1863] 3. Emotional Recognition
[1864] The created text data and voice tone are sent to an emotion recognition API to generate user emotion data, which indicates the user's emotions (joy, anger, sadness, etc.).
[1865] 4. Adjusting response data
[1866] Based on the emotion data, the generated answer data is adjusted. For example, if the user expresses anger, an apology or an encouraging message is added to the response.
[1867] 5. Audio and visual representation
[1868] The adjusted answer data is converted into audio data using a speech synthesis API and played back through the smart glasses' speakers, while visual feedback such as work progress and alerts is displayed on the smart glasses' display.
[1869] 6. Schedule Management
[1870] Schedule information is retrieved using the calendar API, and alerts are generated when important events or deadlines are approaching, with audio and visual notifications.
[1871] Specific examples
[1872] For example, if a worker says, "Please tell me the next work step," their intention is analyzed using voice recognition and natural language processing technology. Based on the analysis results, answer data that provides the next work step is generated. Furthermore, emotion recognition is used to determine the worker's emotions, so if the worker is tired, an appropriate response will be given according to their emotions, such as, "The next task is △△. You seem tired, let's take a short break."
[1873] Prompt Sentence Examples
[1874] An example of a prompt sentence to input to the generative AI model is as follows:
[1875] If the question "Please tell me the next work step" is input, please output the answer text as "The next work step is ____." Also, if the questioner is feeling sad, please add a message of encouragement.
[1876] As described above, the system of the present invention integrates voice input, voice recognition, natural language processing, voice synthesis, emotion recognition, emotion-based response adjustment, schedule linkage, and visual feedback. This allows users to effectively acquire information through natural voice dialogue and improve work efficiency. In addition, responses that take the user's emotions into consideration can provide a more comfortable work environment.
[1877] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1878] Step 1:
[1879] The user speaks to the smart glasses. The built-in microphone of the smart glasses collects the user's voice and transmits it to the terminal as voice data. The input is the user's voice, and the output is voice data.
[1880] Step 2:
[1881] The device sends voice data to the voice recognition API. The API analyzes the voice data and converts it into text data. In this process, the voice data is analyzed and text data is generated as character information. The input is voice data and the output is text data.
[1882] Step 3:
[1883] The device sends the generated text data to a natural language processing API. The natural language processing API analyzes the text data and understands the intent of the user's question or instruction. This process generates appropriate answer data based on the analyzed intent. The input is text data, and the output is answer data.
[1884] Step 4:
[1885] The device sends the generated answer data to the emotion recognition API, and simultaneously analyzes the worker's voice tone. The emotion recognition API generates emotion data and identifies the user's emotion (e.g., joy, anger, sadness, etc.). The input is the answer data and voice tone, and the output is emotion data.
[1886] Step 5:
[1887] The server adjusts the answer data based on the emotion data. For example, if the user expresses anger, it adds an apology or comforting message to the answer data. The input is the answer data and emotion data, and the output is the adjusted answer data.
[1888] Step 6:
[1889] The device sends the adjusted answer data to a speech synthesis API, which converts the text data into speech data. In this process, the adjusted answer data is converted into speech. The input is the adjusted answer data, and the output is speech data.
[1890] Step 7:
[1891] The terminal plays the generated voice data through the speaker of the smart glasses, allowing the user to receive a voice response. The input is voice data, and the output is voice output.
[1892] Step 8:
[1893] The server periodically calls the calendar API to retrieve calendar data. The retrieved calendar data is parsed and schedule information is extracted. The input is the calendar data, and the output is the schedule information.
[1894] Step 9:
[1895] The server generates alerts based on schedule information and sends the alert data to the terminal. Alert messages are generated when important events or work deadlines approach. The input is schedule information and the output is alert data.
[1896] Step 10:
[1897] The device sends the generated alert data to the speech synthesis API, which converts the alert message into voice data. It also generates data to display the alert content on the smart glasses display. The input is the alert data, and the output is voice data and display data.
[1898] Step 11:
[1899] The device notifies the user of the alert through the smart glasses' speaker and displays the alert content on the display. The input is audio data and display data, and the output is audio and display notifications.
[1900] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1901] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1902] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1903] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1904] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1905] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1906] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1907] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1908] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1909] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1910] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1911] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1912] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1913] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1914] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1915] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1916] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1917] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1918] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1919] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1920] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1921] The following is further disclosed regarding the above embodiment.
[1922] (Claim 1)
[1923] means for recognizing a user's voice and generating voice data;
[1924] means for converting the generated voice data into text data;
[1925] means for analyzing text data and generating answer data;
[1926] means for outputting the generated answer data as voice data;
[1927] A means for retrieving calendar data and generating and outputting alerts containing schedule information;
[1928] means for visualizing the execution of the means on a display device;
[1929] A system including:
[1930] (Claim 2)
[1931] 10. The system of claim 1, wherein the audio data is collected using a microphone.
[1932] (Claim 3)
[1933] The system according to claim 1, wherein a speech recognition API is used for conversion to text data.
[1934] "Example 1"
[1935] Claims
[1936] (Claim 1)
[1937] means for recognizing a user's voice and generating voice data;
[1938] means for converting the generated voice data into text data;
[1939] means for analyzing text data and generating answer data;
[1940] means for outputting the generated answer data as voice data;
[1941] A means for retrieving calendar data and generating and outputting alerts containing schedule information;
[1942] means for visualizing the execution of the means on a display device;
[1943] A means for generating and displaying dynamic character animations to provide an intuitive operating experience through user interaction;
[1944] A system including:
[1945] (Claim 2)
[1946] 10. The system of claim 1, wherein the collection of the voice data uses a voice input device.
[1947] (Claim 3)
[1948] The system according to claim 1, wherein a speech conversion API is used for conversion to text data.
[1949] "Application Example 1"
[1950] (Claim 1)
[1951] means for recognizing a user's voice and generating voice data;
[1952] means for converting the generated voice data into text data;
[1953] means for analyzing text data and generating answer data;
[1954] means for outputting the generated answer data as voice data;
[1955] A means for acquiring payment history data and generating and outputting alerts containing payment information;
[1956] means for visualizing the execution of the means on a display device;
[1957] A system including:
[1958] (Claim 2)
[1959] 10. The system of claim 1, wherein the audio data is collected using a microphone.
[1960] (Claim 3)
[1961] The system according to claim 1, wherein a speech recognition API is used for conversion to text data.
[1962] "Example 2: Combining Emotion Engines"
[1963] (Claim 1)
[1964] means for recognizing a user's voice and generating voice data;
[1965] means for converting the generated voice data into text data;
[1966] means for analyzing text data and generating answer data;
[1967] means for outputting the generated answer data as voice data;
[1968] means for recognizing a user's emotion;
[1969] means for adjusting answer data based on emotion recognition results;
[1970] A means for retrieving calendar data and generating and outputting alerts containing schedule information;
[1971] means for visualizing the execution of the means on a display device;
[1972] A system including:
[1973] (Claim 2)
[1974] 10. The system of claim 1, wherein the audio data is collected using an audio recording device.
[1975] (Claim 3)
[1976] 2. The system according to claim 1, wherein a speech recognition program is used for the conversion into text data.
[1977] "Application example 2 when combining emotion engines"
[1978] (Claim 1)
[1979] means for recognizing a user's voice and generating voice data;
[1980] means for converting the generated voice data into text data;
[1981] means for analyzing text data and generating answer data;
[1982] means for outputting the generated answer data as voice data;
[1983] A means for performing emotion recognition and generating emotion data;
[1984] means for adjusting answer data based on emotion data;
[1985] A means for retrieving calendar data and generating and outputting alerts containing schedule information;
[1986] means for visualizing the execution of the means on a display device;
[1987] A system including:
[1988] (Claim 2)
[1989] 10. The system of claim 1, wherein the audio data is collected using an acoustic device.
[1990] (Claim 3)
[1991] The system according to claim 1, wherein a speech recognition API is used for conversion to text data. [Explanation of symbols]
[1992] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for recognizing a user's voice and generating voice data; means for converting the generated voice data into text data; means for analyzing text data and generating answer data; means for outputting the generated answer data as voice data; A means for acquiring calendar data and generating and outputting alerts containing schedule information; means for visualizing the execution of the means on a display device; A system including:
2. 10. The system of claim 1, wherein the audio data is collected using a microphone.
3. 2. The system according to claim 1, wherein a speech recognition API is used for conversion into text data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A