system
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-16
- Publication Date
- 2026-06-26
AI Technical Summary
Modern answering machine systems lack efficient voice recording and content confirmation, struggle with communication in different languages, and insufficient automation of schedule management.
A system that digitizes speech into text data, extracts schedule information, translates text into different languages, and notifies users efficiently, using speech recognition, natural language processing, and translation technologies.
Streamlines voice message confirmation and schedule management, enabling efficient communication across languages and ensuring users do not miss important information.
Smart Images

Figure 2026105389000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern answering machine systems only provide voice recording, and subsequent content confirmation and information extraction are often cumbersome. Also, communication in different languages and automation of schedule management are insufficient, and means for performing these operations more efficiently are required.
Means for Solving the Problems
[0005] This invention includes a speech recognition means that digitizes speech and instantly converts it into text data. Furthermore, by implementing a schedule registration means that extracts schedule information from the text data and automatically registers it in a calendar, and a translation means that translates text into different languages, it facilitates communication in different languages. In addition, it constructs a system that efficiently transmits the converted data to the user via a notification means, significantly improving the management of speech information.
[0006] "Audio data" refers to information that represents sound in a digital format.
[0007] A "voice receiving means" is a mechanism that receives voice data and converts that information into a format that can be processed within the system.
[0008] "Speech recognition means" refers to a function that analyzes speech data and converts it into corresponding text data.
[0009] "Text data" refers to information where audio data is represented in written form.
[0010] "Natural language processing" refers to technologies that analyze text data, understand its meaning and intent, and extract information.
[0011] "Schedule information" refers to data related to dates and times, used for schedule management.
[0012] A "schedule registration method" is a mechanism that automatically registers appointments in a calendar system based on extracted appointment information.
[0013] "Translation means" refers to a device or function that has the ability to convert text written in one language into another language.
[0014] "Notification means" refers to a communication method or device for informing the user of processed text data or translated data. [Brief explanation of the drawing]
[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] The AI personal assistant system of this invention is built using a mobile device and a cloud server. When a call comes in while the user is away, the device automatically records the audio and transfers the data to the server. The server receives this audio data and converts it into text data using speech recognition technology. The converted text is analyzed using natural language processing to extract information, particularly related to dates, times, and events. As a result, information related to the user's schedule is automatically registered in the calendar.
[0037] Furthermore, the server has the capability to translate text data into a specified language if the text data is in a different language. The translated text is then notified to the user. For example, if English is used in a phone call from overseas, the server will translate the content into Japanese and provide it to the user in an easily understandable format.
[0038] As a concrete example, consider a scenario where a user is away from home and someone leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday." In this case, the device records the audio and sends the data to the server. The server transcribes the audio into text, extracts the necessary information, and automatically registers the meeting in the user's calendar. Simultaneously, it translates the content into the user's preferred language and sends a summary as a notification.
[0039] These features allow the system to streamline voice message confirmation and schedule management, and it can be used as an advanced platform that also supports communication in different languages.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The device automatically starts recording audio when a call is missed. Once recording is complete, it sends the audio data to the server in digital format.
[0043] Step 2:
[0044] The server receives the voice data sent from the terminal and converts it into text data using speech recognition technology. This text data is then stored in a database.
[0045] Step 3:
[0046] The server performs natural language processing to parse text data and extracts information about dates, times, and events. This information is used to adjust schedules.
[0047] Step 4:
[0048] The server generates a schedule based on the extracted date and time information and automatically registers the appointment in conjunction with the terminal's calendar application.
[0049] Step 5:
[0050] The server determines if the text data is in a specific language and translates it into the specified language if necessary. The translated text is then saved back into the database.
[0051] Step 6:
[0052] The server summarizes the text data and translated content as needed and sends it to the terminal for notification to the user.
[0053] Step 7:
[0054] The device notifies the user via push notification or email of text data, translation results, and summaries received from the server.
[0055] This process allows the system to efficiently process phone messages and provide users with relevant information.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] In modern information and communication, it is essential to efficiently manage voice messages even when users are absent, and to facilitate smooth communication, especially between languages. However, conventional technologies have not adequately addressed this need for automated voice processing and translation, as well as centralized schedule management. This puts users at risk of overlooking important information.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes acoustic receiving means for receiving acoustic signals and converting them into information signals, acoustic recognition means for converting acoustic signals into text information, and natural language processing means for analyzing the text information and extracting time information. This enables the system to appropriately process voice messages even when the user is absent, support communication between different languages, and automatically manage schedules.
[0061] An "acoustic signal" is an electrical signal obtained by converting sound into an electrical signal, and it is the basic unit for processing sound as digital data.
[0062] An "information signal" is digital data converted from an acoustic signal, and is data in a format that can be processed by a computer.
[0063] "Acoustic receiving means" refers to a device or function for receiving acoustic signals and converting them into information signals.
[0064] "Acoustic recognition means" refers to a device or function that analyzes data input as an acoustic signal and converts its content into textual information.
[0065] "Textual information" refers to text data converted from acoustic signals, and is information represented as a string of characters.
[0066] "Natural language processing means" refers to a device or function for analyzing textual information and extracting specific patterns or meanings.
[0067] "Time information" refers to date and time information extracted by natural language processing tools, and is data used for schedule management.
[0068] A "schedule registration means" is a device or function for registering information in a calendar system or schedule management service based on extracted time information.
[0069] A "language conversion means" is a device or function that translates textual information into another language, enabling communication between different languages.
[0070] A "notification means" is a device or function that notifies the user of converted character information or time information, enabling the recipient of the information to properly understand its content.
[0071] "Communication means" refers to a device or function for sharing voice data or information signals with remote devices or services via a network.
[0072] This invention enables efficient processing of voice messages even when the user is absent, and facilitates smooth communication between multiple languages, through a system consisting of a mobile terminal and a cloud server.
[0073] Calls received while the user is unable to answer are recorded as audio by the device. This audio signal is converted into a digital signal using the built-in microphone and transmitted from the device to a cloud server. Communication is conducted via an internet connection, and data security is ensured using the HTTPS protocol.
[0074] The server converts the received acoustic signal into text information using acoustic recognition means. This process utilizes general speech recognition software, such as a speech recognition API. The converted text information is then analyzed by natural language processing means to extract important information such as date, time, and event information. This step uses natural language processing libraries such as spaCy to analyze patterns and meaning.
[0075] The extracted time information is registered in the calendar system by the server using a schedule registration method. For example, to integrate with an external scheduling service, the schedule is automatically set using a cloud calendar API.
[0076] Furthermore, the server uses language conversion means to translate text information if it differs from the language set by the user. The translation technology utilizes a machine translation API to facilitate communication between different languages.
[0077] Finally, the server notifies the user of the information that has been converted and translated by the notification system. This is done using methods such as push notifications or email so that the user can check the information immediately.
[0078] As a concrete example, if a user leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday," while they are away, the device records the audio and sends it to the server. The server processes the audio according to a prompt that says, "Translate the following audio message into text, extract important information, and add it to the calendar. Also, translate and notify the user if necessary." The server adds the necessary information to the calendar, translates it if necessary, and notifies the user of a summary. In this way, the user can manage important appointments efficiently without missing any.
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] The device automatically records calls if the user is unable to answer. The input is an analog audio signal captured via the telephone line, which is then converted into a digital audio signal. This process utilizes the device's built-in microphone and voice recording application, and the resulting digital audio is temporarily stored in the device's memory.
[0082] Step 2:
[0083] The device transfers the recorded audio data to a cloud server. The input for this step is the digital audio obtained in step 1, and the output is the audio data sent to the server. An internet connection is used to transmit the data, and the HTTPS protocol ensures data security and privacy.
[0084] Step 3:
[0085] The server converts received acoustic signals into text information using acoustic recognition means. The input is an acoustic signal sent from a terminal, and text data is output by analyzing it using speech recognition software. For example, the audio "Let's schedule a meeting for 3 o'clock next Wednesday" is accurately displayed as text.
[0086] Step 4:
[0087] The server analyzes the text information using natural language processing (NLP) and extracts date, time, and event information. The input in this step is the text information obtained in step 3, and the output is the extracted time and event information. Important information is identified from the text using a natural language processing library and registered in the database.
[0088] Step 5:
[0089] The server adds an event to the calendar using the event registration method based on the extracted date and time information. The input data is the time information analyzed in step 4, and the information is reflected in the user's calendar using the Google® Calendar API, etc. The output is the newly added calendar event.
[0090] Step 6:
[0091] The server uses a language conversion mechanism to translate text information if it differs from the language set by the user. The input data is the original text information, and the output is the translated text in the specified language. A machine translation API is used to facilitate smooth information sharing between different languages.
[0092] Step 7:
[0093] The server notifies the user of the converted and translated information through a notification system. The input for this final step is the various data obtained from step 3 onward, and the server informs the user of the schedule and translated information extracted from the voice message through the notification. The output is detailed information received by the user via push notifications or email.
[0094] (Application Example 1)
[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0096] In modern households, important phone calls and messages are often missed when family members are away, and communication in different languages makes understanding information and managing schedules particularly challenging. Furthermore, there is a need to efficiently manage the schedules of all family members and receive real-time notifications to improve their quality of life.
[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0098] In this invention, the server includes means for receiving voice data and converting it into a digital format, means for converting the voice data into text data, and means for analyzing the text data and extracting schedule information. This makes it possible to use a home robotic device to reliably manage schedules even when family members are away, quickly translate information in different languages, and provide real-time notifications in an easy-to-understand format.
[0099] "Audio data" refers to a data format in which sound is recorded digitally and used for information processing or communication.
[0100] A "digital format" is a format that can be processed by a computer by converting analog signals such as audio and video into numerical data.
[0101] "Audio receiving means" refers to a hardware or software configuration for receiving audio data and converting it into a digital format.
[0102] "Speech recognition means" refers to a technology that analyzes received speech data and converts it into corresponding text data.
[0103] "Text data" refers to data in string format obtained through speech recognition.
[0104] "Natural language processing methods" are technologies for analyzing text data to extract specific information (e.g., date and time, events).
[0105] A "schedule registration method" is a system that automatically registers extracted schedule information into a calendar or schedule management application.
[0106] "Translation methods" refer to technologies that convert text data in a specific language into a different language.
[0107] A "notification method" is a mechanism that informs the user's information terminal of the generated text data or translated information.
[0108] A "household robotic device" is a robot designed for use within the home and possesses a variety of functions, such as recording voice data.
[0109] To implement this invention, the home robot system is configured as follows: The robot device is equipped with a voice receiving means and converts voice data received within the home into a digital format. The converted voice is sent to a cloud server. The server converts the voice data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose.
[0110] The converted text data is analyzed using natural language processing technology on the cloud server to extract schedule information. Azure Cognitive Services' natural language processing capabilities are used for this process. The extracted schedule information is then automatically registered to the user's calendar by the server using the Google Calendar API.
[0111] Furthermore, if the audio content is in a specific language, the server uses Azure Translator to translate it into a different language. This translated text is then sent to the user's mobile device via a notification system. Real-time messaging services such as Pusher are used for notifications.
[0112] For example, if a family member is out and a robot records a message saying, "I have plans to meet a friend next Sunday," the server can transcribe the message into text, translate it from Japanese to English, register the event in Google Calendar, and then send the translated content to the user's device as a collaborator.
[0113] An example of a prompt message could be entered as follows: "Please wake me up at 3 AM tomorrow. If it's raining, please let me know instead of going to a cafe."
[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0115] Step 1:
[0116] The terminal uses a home robotic device to receive voice data from phone calls made while the user is away and convert it into a digital format. In this step, voice input is taken and output is produced in the form of digital voice data. The voice data is temporarily stored in local storage.
[0117] Step 2:
[0118] The device sends digital audio data to a cloud server. Here, digital audio data is passed to the server as input, and as a result, the audio data is uploaded to the cloud.
[0119] Step 3:
[0120] The server performs speech recognition processing in the cloud and uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Audio data is input, the words in the audio are analyzed, and the text data is output.
[0121] Step 4:
[0122] The server analyzes text data using natural language processing techniques and extracts schedule information using Azure Cognitive Services. It extracts keywords such as date, time, and location from the input text data and outputs them as schedule information.
[0123] Step 5:
[0124] The server extracts schedule information and registers it in the user's calendar using the Google Calendar API. In this step, the schedule information is used as input, and the output is that the events are reflected in the calendar.
[0125] Step 6:
[0126] The server uses translation tools to translate text data into a desired language using Azure Translator if the original text is in a specific language. The original text is the input, and the translated text is the output.
[0127] Step 7:
[0128] The server uses notification methods to send translated text and schedule information to the user's mobile device via services such as Pusher. The entered text is output as a message to the user, and the information is notified to the device in real time.
[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0130] The present invention provides an AI personal assistant that automatically handles incoming calls while the user is away. At the core of the system is a function that receives voice data and converts it into text data using speech recognition technology. Furthermore, it can analyze the text data using natural language processing, extract schedule information and summaries, and automatically register them in the user's calendar.
[0131] The emotion engine, in addition to this, plays the role of recognizing the user's emotions from the voice and generating emotion-based responses and suggestions. The server passes the received voice data to the emotion engine, and the server decides on a course of action based on the analysis results. For example, if the received voice contains tension or anxiety, the system is designed to notify the user of this information and provide situation-appropriate advice.
[0132] As a concrete example, if a user receives an important phone call while in a meeting, the device immediately sends the audio data to the server. The server converts the audio to text and analyzes it using an emotion engine. This analysis determines the emotional urgency of the message, and the user is notified with the message, "You have an urgent message." Furthermore, appropriate actions and countermeasures are suggested depending on the situation, allowing the user to respond quickly and accurately.
[0133] This system, equipped with such processes, is not merely a scheduling and translation tool, but an innovative platform that utilizes sentiment analysis to enable more human-like responses.
[0134] The following describes the processing flow.
[0135] Step 1:
[0136] The device automatically starts recording audio when a call is missed. The recorded audio data is converted to a digital format and sent to the server.
[0137] Step 2:
[0138] The server transmits the received audio data to the speech recognition engine, which converts the audio into text data. The speech recognition engine analyzes the data and stores the converted text in a database.
[0139] Step 3:
[0140] The server processes the text data using a natural language processing engine, extracts schedule information, and generates data for schedule registration. This information is synchronized with the device's calendar.
[0141] Step 4:
[0142] The server uses an emotion engine to perform sentiment analysis on the voice data. This evaluates the emotional nuances contained in the voice message and creates the foundational data for communicating the results to the user.
[0143] Step 5:
[0144] Based on the sentiment analysis results obtained from the analysis, the server notifies the user as needed and generates situation-appropriate responses and suggestions. This information is transmitted to the terminal by the notification system.
[0145] Step 6:
[0146] The device notifies the user of text data sent from the server, sentiment analysis results, and possible countermeasures. Notifications are sent via push notifications or email.
[0147] Step 7:
[0148] Users review notifications from their devices and select appropriate actions based on recommended responses derived from sentiment analysis.
[0149] This entire process allows the system to provide proactive support to the user and enables a deeper understanding of the content, including emotional insights.
[0150] (Example 2)
[0151] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0152] In telephone communication, properly handling messages received while the user is absent is a challenging task. Conventional technologies simply record voice messages or convert them to text, failing to consider the importance or emotional content of the received messages. Furthermore, there is a risk of important information being overlooked, and emotional responses may be delayed. This invention aims to solve these problems.
[0153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0154] In this invention, the server includes a voice receiving means for receiving voice data and converting it into a digital format, a voice recognition means for converting the voice data into text data, and an emotion analysis means for detecting emotions from the text data. This enables the server to comprehensively analyze the content and emotional elements of received messages, even when the user is absent, and to respond quickly and appropriately.
[0155] "Voice receiving means" refers to a device or software that has the function of receiving voice data and converting it into a digital format.
[0156] "Speech recognition means" refers to a technology or device for converting digital speech data into text data.
[0157] "Natural language processing means" refers to techniques or methods used to analyze text data and extract planned information or summaries.
[0158] A "schedule registration method" is a device or software that automatically registers extracted schedule information into the user's calendar or management system.
[0159] "Emotional analysis means" refers to a technology or device for detecting emotions from text data or audio data and analyzing the results.
[0160] "Sentiment notification means" refers to a device or software that provides appropriate notifications to the user based on the results of sentiment analysis.
[0161] "Notification means" refers to a device or method that transmits converted text data or analysis results to the user to inform them.
[0162] The present invention provides a voice processing platform for automatically handling incoming calls while the user is away. A terminal detects a call, records the voice data, and sends it to a server. The server uses speech recognition technology to convert the voice data into text data. General-purpose speech recognition software can be used for speech recognition. For example, open-source software or cloud-based speech recognition services can be utilized.
[0163] The server analyzes the converted text data using natural language processing. This natural language processing utilizes a natural language processing engine, which is also used as a generative AI model. This allows for the extraction of schedule information and conversation summaries from the text data. Specific software options include general API services and open-source natural language processing libraries.
[0164] Next, the server uses sentiment analysis tools to detect emotions from the text data. Existing sentiment analysis APIs and modules can be used for sentiment analysis, determining the emotional urgency and importance of the received message. This sentiment information is used for user notifications and additional responses.
[0165] As a concrete example, consider a situation where a user is in a meeting and unable to answer a phone call. The device automatically sends the phone's audio data to the server. The server performs speech recognition, generates text data, extracts schedule information using natural language processing, and applies sentiment analysis. Based on these results, the server can send a notification to the user stating, "You have an urgent message," and prompt them with suggestions or actions.
[0166] As a concrete example of a prompt message to the generating AI model, text such as "Extract schedule information from phone messages, perform sentiment analysis, and notify the results through the library" can be used. This allows the system of the present invention to go beyond simple schedule management and enable more human-like and effective responses by incorporating sentiment analysis.
[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0168] Step 1:
[0169] The device detects incoming calls while the user is away, records the audio data, and converts it to a digital format. The input is an analog signal from the telephone line, which is then output as digital audio data. Specifically, it activates the audio recording function upon receiving a call signal and saves the recorded audio as a digital file.
[0170] Step 2:
[0171] The terminal sends the recorded digital audio data to the server. The input is the digital audio file generated in step 1, which is then transferred to the server via the network. Specifically, a stable internet connection is used, and a file transfer protocol is employed to ensure the data is reliably transferred to the server.
[0172] Step 3:
[0173] The server converts received audio data into text data using speech recognition technology. The input is a digital audio file, and the output is a textual representation of that audio. Specifically, it launches speech recognition software and executes a process of sequentially generating text by analyzing the audio data.
[0174] Step 4:
[0175] The server analyzes the generated text data using natural language processing tools and extracts schedule information and summaries. The input is the text data obtained in step 3, and the output is the extracted schedule information and summaries. Specifically, it uses a natural language processing engine to analyze the sentence structure and identify important information.
[0176] Step 5:
[0177] The server detects emotions in text data using sentiment analysis tools and analyzes the results. The input is the text data obtained in step 3, and the output is the sentiment analysis results and the urgency of the information. Specifically, it starts the sentiment analysis module and calculates a sentiment score based on keywords and context in the text.
[0178] Step 6:
[0179] The server provides the user with appropriate notifications based on schedule information and sentiment analysis results, and suggests actions if necessary. The input is the schedule and sentiment information obtained in steps 4 and 5, and the output is a notification message to the user. Specifically, it generates the notification message and sends it to the user's mobile device or other device.
[0180] (Application Example 2)
[0181] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0182] In autonomous vehicles, drivers must dedicate their hands and attention to external communication, which can raise safety concerns. Furthermore, responding quickly and appropriately to audio information received while driving is difficult. Additionally, methods to identify emotional states and reduce the driver's psychological burden are needed.
[0183] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0184] In this invention, the server includes means for receiving voice information and converting it into a digital format, natural language processing means for analyzing text information and extracting planned information, and emotion recognition means for evaluating emotional states in response to stimuli. This enables the driver to safely and comfortably manage external communications even while driving, and to make quick decisions to ensure safety.
[0185] "Audio information" refers to information received as digital data through an input device.
[0186] "Converting to digital format" refers to the process of changing analog audio data into digital data.
[0187] "Voice receiving means" refers to an interface for collecting voice information.
[0188] "Speech recognition means" refers to technology for analyzing speech information and converting it into text information.
[0189] "Text information" refers to data in character format obtained through speech recognition.
[0190] "Natural language processing methods" are technologies for analyzing text information and extracting meaningful information.
[0191] "Schedule information" refers to schedule-related information extracted from text information using natural language processing tools.
[0192] The "schedule registration method" refers to a function for registering acquired schedule information into a time management system such as a calendar.
[0193] "Translation methods" refer to technologies that convert text information in a specific language into a different language.
[0194] "Notification means" refers to a method or device for informing a user of information.
[0195] "Emotional state" refers to the emotional state of the speaker as identified from their voice.
[0196] "Emotion recognition means" refers to technology for evaluating a user's emotional state from audio information.
[0197] "Urgency" refers to the degree to which the received information is assessed as being urgent.
[0198] "Driving assistance systems" are technologies that provide information to drivers to help them continue driving safely and support their decision-making.
[0199] The system realizing this invention achieves efficient information processing while ensuring driver safety in an autonomous vehicle through the reception, analysis, and evaluation of voice information. The server converts the input voice into a digital format via a device for collecting voice information. Subsequently, it converts the voice into text information using speech recognition technology, analyzes the intent of the information through natural language processing (NLP), and extracts scheduled information and other useful information.
[0200] The server also uses emotion recognition technology to evaluate the emotional state within text information and determine the urgency of the information. Based on this information, it notifies the driver in audio or visual form as a driving assistance measure. This allows for timely decision-making while maintaining the driver's safety.
[0201] The hardware used includes microphones and computer systems installed in autonomous vehicles. The software utilizes speech recognition engines (e.g., Google Speech-to-Text, IBM Watson®), natural language processing libraries (e.g., spaCy, NLTK), and sentiment analysis tools (e.g., IBM Watson Tone Analyzer).
[0202] As a concrete example, imagine a scenario where, when a user receives an important phone call while driving, the content is quickly converted to text, sentiment analysis is performed, and the driver is notified with a message such as, "You have an important message from your boss."
[0203] An example of a prompt is: "Please describe in detail the design of an AI assistant that analyzes voice messages left as missed calls while a user is traveling in an autonomous vehicle, provides notifications according to their urgency, and suggests responses based on sentiment data." By using this prompt and leveraging a generative AI model, even more advanced natural language processing capabilities and sentiment analysis can be achieved.
[0204] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0205] Step 1:
[0206] The server uses a microphone to receive audio information input from within the autonomous vehicle and converts it into a digital format. The input is the user's voice, and the output is digital audio data. This data is pre-processed for speech recognition.
[0207] Step 2:
[0208] The server uses a speech recognition engine to convert digital audio data into text information. The input is digital audio data, and the output is text data. The converted text data is then prepared for analysis for natural language processing.
[0209] Step 3:
[0210] The server uses natural language processing to analyze text information and extract important schedule information and instructions. The input is text data, and the output is the extracted schedule information and instructions. This information is then passed to the schedule registration and notification system.
[0211] Step 4:
[0212] The server uses an emotion recognition engine to assess the user's emotional state from audio and text information. Input is audio or text data, and output is data related to the emotional state. The results of the emotion analysis are used to assess urgency.
[0213] Step 5:
[0214] The server determines the urgency level based on the extracted schedule information and emotional state, and notifies the user as appropriate. The input is schedule information and emotional evaluation data, and the output is an emergency notification. For example, the driver might receive a notification saying, "You have an urgent message from your supervisor."
[0215] Step 6:
[0216] Users can receive information provided through driving assistance systems and take necessary actions in a safe manner. Inputs are notifications and information provided by the server, and outputs are the corresponding actions. This allows users to respond efficiently and safely even while driving.
[0217] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0218] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search)<url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0219] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0220] [Second Embodiment]
[0221] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0222] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0223] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0224] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0225] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0226] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0227] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0228] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0229] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0230] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0231] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0232] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0233] The AI personal assistant system of this invention is built using a mobile device and a cloud server. When a call comes in while the user is away, the device automatically records the audio and transfers the data to the server. The server receives this audio data and converts it into text data using speech recognition technology. The converted text is analyzed using natural language processing to extract information, particularly related to dates, times, and events. As a result, information related to the user's schedule is automatically registered in the calendar.
[0234] Furthermore, the server has the capability to translate text data into a specified language if the text data is in a different language. The translated text is then notified to the user. For example, if English is used in a phone call from overseas, the server will translate the content into Japanese and provide it to the user in an easily understandable format.
[0235] As a concrete example, consider a scenario where a user is away from home and someone leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday." In this case, the device records the audio and sends the data to the server. The server transcribes the audio into text, extracts the necessary information, and automatically registers the meeting in the user's calendar. Simultaneously, it translates the content into the user's preferred language and sends a summary as a notification.
[0236] These features allow the system to streamline voice message confirmation and schedule management, and it can be used as an advanced platform that also supports communication in different languages.
[0237] The following describes the processing flow.
[0238] Step 1:
[0239] The device automatically starts recording audio when a call is missed. Once recording is complete, it sends the audio data to the server in digital format.
[0240] Step 2:
[0241] The server receives the voice data sent from the terminal and converts it into text data using speech recognition technology. This text data is then stored in a database.
[0242] Step 3:
[0243] The server performs natural language processing to parse text data and extracts information about dates, times, and events. This information is used to adjust schedules.
[0244] Step 4:
[0245] The server generates a schedule based on the extracted date and time information and automatically registers the appointment in conjunction with the terminal's calendar application.
[0246] Step 5:
[0247] The server determines if the text data is in a specific language and translates it into the specified language if necessary. The translated text is then saved back into the database.
[0248] Step 6:
[0249] The server summarizes the text data and translated content as needed and sends it to the terminal for notification to the user.
[0250] Step 7:
[0251] The device notifies the user via push notification or email of text data, translation results, and summaries received from the server.
[0252] This process allows the system to efficiently process phone messages and provide users with relevant information.
[0253] (Example 1)
[0254] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0255] In modern information and communication, it is essential to efficiently manage voice messages even when users are absent, and to facilitate smooth communication, especially between languages. However, conventional technologies have not adequately addressed this need for automated voice processing and translation, as well as centralized schedule management. This puts users at risk of overlooking important information.
[0256] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0257] In this invention, the server includes acoustic receiving means for receiving acoustic signals and converting them into information signals, acoustic recognition means for converting acoustic signals into text information, and natural language processing means for analyzing the text information and extracting time information. This enables the system to appropriately process voice messages even when the user is absent, support communication between different languages, and automatically manage schedules.
[0258] An "acoustic signal" is an electrical signal obtained by converting sound into an electrical signal, and it is the basic unit for processing sound as digital data.
[0259] An "information signal" is digital data converted from an acoustic signal, and is data in a format that can be processed by a computer.
[0260] "Acoustic receiving means" refers to a device or function for receiving acoustic signals and converting them into information signals.
[0261] "Acoustic recognition means" refers to a device or function that analyzes data input as an acoustic signal and converts its content into textual information.
[0262] "Textual information" refers to text data converted from acoustic signals, and is information represented as a string of characters.
[0263] "Natural language processing means" refers to a device or function for analyzing textual information and extracting specific patterns or meanings.
[0264] "Time information" refers to date and time information extracted by natural language processing tools, and is data used for schedule management.
[0265] A "schedule registration means" is a device or function for registering information in a calendar system or schedule management service based on extracted time information.
[0266] A "language conversion means" is a device or function that translates textual information into another language, enabling communication between different languages.
[0267] A "notification means" is a device or function that notifies the user of converted character information or time information, enabling the recipient of the information to properly understand its content.
[0268] "Communication means" refers to a device or function for sharing voice data or information signals with remote devices or services via a network.
[0269] This invention enables efficient processing of voice messages even when the user is absent, and facilitates smooth communication between multiple languages, through a system consisting of a mobile terminal and a cloud server.
[0270] Calls received while the user is unable to answer are recorded as audio by the device. This audio signal is converted into a digital signal using the built-in microphone and transmitted from the device to a cloud server. Communication is conducted via an internet connection, and data security is ensured using the HTTPS protocol.
[0271] The server converts the received acoustic signal into text information using acoustic recognition means. This process utilizes general speech recognition software, such as a speech recognition API. The converted text information is then analyzed by natural language processing means to extract important information such as date, time, and event information. This step uses natural language processing libraries such as spaCy to analyze patterns and meaning.
[0272] The extracted time information is registered in the calendar system by the server using a schedule registration method. For example, to integrate with an external scheduling service, the schedule is automatically set using a cloud calendar API.
[0273] Furthermore, the server uses language conversion means to translate text information if it differs from the language set by the user. The translation technology utilizes a machine translation API to facilitate communication between different languages.
[0274] Finally, the server notifies the user of the information that has been converted and translated by the notification system. This is done using methods such as push notifications or email so that the user can check the information immediately.
[0275] As a concrete example, if a user leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday," while they are away, the device records the audio and sends it to the server. The server processes the audio according to a prompt that says, "Translate the following audio message into text, extract important information, and add it to the calendar. Also, translate and notify the user if necessary." The server adds the necessary information to the calendar, translates it if necessary, and notifies the user of a summary. In this way, the user can manage important appointments efficiently without missing any.
[0276] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0277] Step 1:
[0278] The device automatically records calls if the user is unable to answer. The input is an analog audio signal captured via the telephone line, which is then converted into a digital audio signal. This process utilizes the device's built-in microphone and voice recording application, and the resulting digital audio is temporarily stored in the device's memory.
[0279] Step 2:
[0280] The terminal transfers the recorded voice data to the cloud server. The input for this step is the digital voice obtained in step 1, and the output is the acoustic data transmitted to the server. An Internet connection is used for data transmission, and the HTTPS protocol ensures data security and privacy.
[0281] Step 3:
[0282] The server converts the received acoustic signal into character information by means of acoustic recognition. The input is the acoustic signal sent from the terminal, and character data is output by analyzing it using speech recognition software. For example, a voice such as "Set a meeting at 3 o'clock next Wednesday" is accurately displayed as text.
[0283] Step 4:
[0284] The server analyzes the character information using natural language analysis means and extracts date and time and event information. The input in this step is the character information obtained in step 3, and the output is the extracted time and event information. Important information is identified from the text using a natural language processing library and registered in the database.
[0285] Step 5:
[0286] The server adds an event to the calendar by means of schedule registration based on the extracted date and time information. The input data is the time information analyzed in step 4, and information is reflected in the user's calendar using, for example, the Google Calendar API. The output is the newly added calendar event.
[0287] Step 6:
[0288] The server uses a language conversion mechanism to translate text information if it differs from the language set by the user. The input data is the original text information, and the output is the translated text in the specified language. A machine translation API is used to facilitate smooth information sharing between different languages.
[0289] Step 7:
[0290] The server notifies the user of the converted and translated information through a notification system. The input for this final step is the various data obtained from step 3 onward, and the server informs the user of the schedule and translated information extracted from the voice message through the notification. The output is detailed information received by the user via push notifications or email.
[0291] (Application Example 1)
[0292] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0293] In modern households, important phone calls and messages are often missed when family members are away, and communication in different languages makes understanding information and managing schedules particularly challenging. Furthermore, there is a need to efficiently manage the schedules of all family members and receive real-time notifications to improve their quality of life.
[0294] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0295] In this invention, the server includes means for receiving voice data and converting it into a digital format, means for converting the voice data into text data, and means for analyzing the text data and extracting schedule information. This makes it possible to use a home robotic device to reliably manage schedules even when family members are away, quickly translate information in different languages, and provide real-time notifications in an easy-to-understand format.
[0296] "Audio data" refers to a data format in which sound is recorded digitally and used for information processing or communication.
[0297] A "digital format" is a format that can be processed by a computer by converting analog signals such as audio and video into numerical data.
[0298] "Audio receiving means" refers to a hardware or software configuration for receiving audio data and converting it into a digital format.
[0299] "Speech recognition means" refers to a technology that analyzes received speech data and converts it into corresponding text data.
[0300] "Text data" refers to data in string format obtained through speech recognition.
[0301] "Natural language processing methods" are technologies for analyzing text data to extract specific information (e.g., date and time, events).
[0302] A "schedule registration method" is a system that automatically registers extracted schedule information into a calendar or schedule management application.
[0303] "Translation methods" refer to technologies that convert text data in a specific language into a different language.
[0304] A "notification method" is a mechanism that informs the user's information terminal of the generated text data or translated information.
[0305] A "household robotic device" is a robot designed for use within the home and possesses a variety of functions, such as recording voice data.
[0306] To implement this invention, a household robot system is configured as follows. The robot device has voice receiving means and converts voice data received within the household into digital format. The converted voice is transmitted to a cloud server. The server uses voice recognition technology to convert the voice data into text data. Google Cloud Speech-to-Text API is used for this.
[0307] The converted text data is analyzed by natural language processing technology within the cloud server, and schedule information is extracted. At this time, the natural language processing function of Azure Cognitive Services is utilized. The extracted schedule information is automatically registered by the server in the user's calendar using Google Calendar API.
[0308] Furthermore, if the voice content is in a specific language, the server uses Azure Translator to translate it into a different language. This translated text is transmitted to the user's mobile information terminal through notification means and is notified. A real-time messaging service such as Pusher is used for the notification.
[0309] As a specific example, when the robot records a voice saying "There is a plan to meet a friend next Sunday" while the family is out, the server can textify the content, translate it from Japanese to English, register the schedule in Google Calendar, and transmit the translated content to the user's terminal as a collaborator.
[0310] As an example of a prompt sentence, it can be input in the form of "Please wake me up at 3 o'clock tomorrow morning. If the weather is rainy, please let me know not to go to the café."
[0311] The flow of the specific process in Application Example 1 will be described using FIG. 12. The terminal uses a home robotic device to receive voice data from phone calls made while the user is away and convert it into a digital format. In this step, voice input is taken and output is produced in the form of digital voice data. The voice data is temporarily stored in local storage.
[0314] Step 2:
[0315] The device sends digital audio data to a cloud server. Here, digital audio data is passed to the server as input, and as a result, the audio data is uploaded to the cloud.
[0316] Step 3:
[0317] The server performs speech recognition processing in the cloud and uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Audio data is input, the words in the audio are analyzed, and the text data is output.
[0318] Step 4:
[0319] The server analyzes text data using natural language processing techniques and extracts schedule information using Azure Cognitive Services. It extracts keywords such as date, time, and location from the input text data and outputs them as schedule information.
[0320] Step 5:
[0321] The server extracts schedule information and registers it in the user's calendar using the Google Calendar API. In this step, the schedule information is used as input, and the output is that the events are reflected in the calendar.
[0322] Step 6:
[0323] The server uses translation tools to translate text data into a desired language using Azure Translator if the original text is in a specific language. The original text is the input, and the translated text is the output.
[0324] Step 7:
[0325] The server uses notification methods to send translated text and schedule information to the user's mobile device via services such as Pusher. The entered text is output as a message to the user, and the information is notified to the device in real time.
[0326] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0327] The present invention provides an AI personal assistant that automatically handles incoming calls while the user is away. At the core of the system is a function that receives voice data and converts it into text data using speech recognition technology. Furthermore, it can analyze the text data using natural language processing, extract schedule information and summaries, and automatically register them in the user's calendar.
[0328] The emotion engine, in addition to this, plays the role of recognizing the user's emotions from the voice and generating emotion-based responses and suggestions. The server passes the received voice data to the emotion engine, and the server decides on a course of action based on the analysis results. For example, if the received voice contains tension or anxiety, the system is designed to notify the user of this information and provide situation-appropriate advice.
[0329] As a concrete example, if a user receives an important phone call while in a meeting, the device immediately sends the audio data to the server. The server converts the audio to text and analyzes it using an emotion engine. This analysis determines the emotional urgency of the message, and the user is notified with the message, "You have an urgent message." Furthermore, appropriate actions and countermeasures are suggested depending on the situation, allowing the user to respond quickly and accurately.
[0330] This system, equipped with such processes, is not merely a scheduling and translation tool, but an innovative platform that utilizes sentiment analysis to enable more human-like responses.
[0331] The following describes the processing flow.
[0332] Step 1:
[0333] The device automatically starts recording audio when a call is missed. The recorded audio data is converted to a digital format and sent to the server.
[0334] Step 2:
[0335] The server transmits the received audio data to the speech recognition engine, which converts the audio into text data. The speech recognition engine analyzes the data and stores the converted text in a database.
[0336] Step 3:
[0337] The server processes the text data using a natural language processing engine, extracts schedule information, and generates data for schedule registration. This information is synchronized with the device's calendar.
[0338] Step 4:
[0339] The server uses an emotion engine to perform sentiment analysis on the voice data. This evaluates the emotional nuances contained in the voice message and creates the foundational data for communicating the results to the user.
[0340] Step 5:
[0341] Based on the sentiment analysis results obtained from the analysis, the server notifies the user as needed and generates situation-appropriate responses and suggestions. This information is transmitted to the terminal by the notification system.
[0342] Step 6:
[0343] The device notifies the user of text data sent from the server, sentiment analysis results, and possible countermeasures. Notifications are sent via push notifications or email.
[0344] Step 7:
[0345] Users review notifications from their devices and select appropriate actions based on recommended responses derived from sentiment analysis.
[0346] This entire process allows the system to provide proactive support to the user and enables a deeper understanding of the content, including emotional insights.
[0347] (Example 2)
[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0349] In telephone communication, properly handling messages received while the user is absent is a challenging task. Conventional technologies simply record voice messages or convert them to text, failing to consider the importance or emotional content of the received messages. Furthermore, there is a risk of important information being overlooked, and emotional responses may be delayed. This invention aims to solve these problems.
[0350] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0351] In this invention, the server includes a voice receiving means for receiving voice data and converting it into a digital format, a voice recognition means for converting the voice data into text data, and an emotion analysis means for detecting emotions from the text data. This enables the server to comprehensively analyze the content and emotional elements of received messages, even when the user is absent, and to respond quickly and appropriately.
[0352] "Voice receiving means" refers to a device or software that has the function of receiving voice data and converting it into a digital format.
[0353] "Speech recognition means" refers to a technology or device for converting digital speech data into text data.
[0354] "Natural language processing means" refers to techniques or methods used to analyze text data and extract planned information or summaries.
[0355] A "schedule registration method" is a device or software that automatically registers extracted schedule information into the user's calendar or management system.
[0356] "Emotional analysis means" refers to a technology or device for detecting emotions from text data or audio data and analyzing the results.
[0357] "Sentiment notification means" refers to a device or software that provides appropriate notifications to the user based on the results of sentiment analysis.
[0358] "Notification means" refers to a device or method that transmits converted text data or analysis results to the user to inform them.
[0359] The present invention provides a voice processing platform for automatically handling incoming calls while the user is away. A terminal detects a call, records the voice data, and sends it to a server. The server uses speech recognition technology to convert the voice data into text data. General-purpose speech recognition software can be used for speech recognition. For example, open-source software or cloud-based speech recognition services can be utilized.
[0360] The server analyzes the converted text data using natural language processing. This natural language processing utilizes a natural language processing engine, which is also used as a generative AI model. This allows for the extraction of schedule information and conversation summaries from the text data. Specific software options include general API services and open-source natural language processing libraries.
[0361] Next, the server uses sentiment analysis tools to detect emotions from the text data. Existing sentiment analysis APIs and modules can be used for sentiment analysis, determining the emotional urgency and importance of the received message. This sentiment information is used for user notifications and additional responses.
[0362] As a concrete example, consider a situation where a user is in a meeting and unable to answer a phone call. The device automatically sends the phone's audio data to the server. The server performs speech recognition, generates text data, extracts schedule information using natural language processing, and applies sentiment analysis. Based on these results, the server can send a notification to the user stating, "You have an urgent message," and prompt them with suggestions or actions.
[0363] As a concrete example of a prompt message to the generating AI model, text such as "Extract schedule information from phone messages, perform sentiment analysis, and notify the results through the library" can be used. This allows the system of the present invention to go beyond simple schedule management and enable more human-like and effective responses by incorporating sentiment analysis.
[0364] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0365] Step 1:
[0366] The device detects incoming calls while the user is away, records the audio data, and converts it to a digital format. The input is an analog signal from the telephone line, which is then output as digital audio data. Specifically, it activates the audio recording function upon receiving a call signal and saves the recorded audio as a digital file.
[0367] Step 2:
[0368] The terminal sends the recorded digital audio data to the server. The input is the digital audio file generated in step 1, which is then transferred to the server via the network. Specifically, a stable internet connection is used, and a file transfer protocol is employed to ensure the data is reliably transferred to the server.
[0369] Step 3:
[0370] The server converts received audio data into text data using speech recognition technology. The input is a digital audio file, and the output is a textual representation of that audio. Specifically, it launches speech recognition software and executes a process of sequentially generating text by analyzing the audio data.
[0371] Step 4:
[0372] The server analyzes the generated text data using natural language processing tools and extracts schedule information and summaries. The input is the text data obtained in step 3, and the output is the extracted schedule information and summaries. Specifically, it uses a natural language processing engine to analyze the sentence structure and identify important information.
[0373] Step 5:
[0374] The server detects emotions in text data using sentiment analysis tools and analyzes the results. The input is the text data obtained in step 3, and the output is the sentiment analysis results and the urgency of the information. Specifically, it starts the sentiment analysis module and calculates a sentiment score based on keywords and context in the text.
[0375] Step 6:
[0376] The server provides the user with appropriate notifications based on schedule information and sentiment analysis results, and suggests actions if necessary. The input is the schedule and sentiment information obtained in steps 4 and 5, and the output is a notification message to the user. Specifically, it generates the notification message and sends it to the user's mobile device or other device.
[0377] (Application Example 2)
[0378] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0379] In autonomous vehicles, drivers must dedicate their hands and attention to external communication, which can raise safety concerns. Furthermore, responding quickly and appropriately to audio information received while driving is difficult. Additionally, methods to identify emotional states and reduce the driver's psychological burden are needed.
[0380] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0381] In this invention, the server includes means for receiving voice information and converting it into a digital format, natural language processing means for analyzing text information and extracting planned information, and emotion recognition means for evaluating emotional states in response to stimuli. This enables the driver to safely and comfortably manage external communications even while driving, and to make quick decisions to ensure safety.
[0382] "Audio information" refers to information received as digital data through an input device.
[0383] "Converting to digital format" refers to the process of changing analog audio data into digital data.
[0384] "Voice receiving means" refers to an interface for collecting voice information.
[0385] "Speech recognition means" refers to technology for analyzing speech information and converting it into text information.
[0386] "Text information" refers to data in character format obtained through speech recognition.
[0387] "Natural language processing methods" are technologies for analyzing text information and extracting meaningful information.
[0388] "Schedule information" refers to schedule-related information extracted from text information using natural language processing tools.
[0389] The "schedule registration method" refers to a function for registering acquired schedule information into a time management system such as a calendar.
[0390] "Translation methods" refer to technologies that convert text information in a specific language into a different language.
[0391] "Notification means" refers to a method or device for informing a user of information.
[0392] "Emotional state" refers to the emotional state of the speaker as identified from their voice.
[0393] "Emotion recognition means" refers to technology for evaluating a user's emotional state from audio information.
[0394] "Urgency" refers to the degree to which the received information is assessed as being urgent.
[0395] "Driving assistance systems" are technologies that provide information to drivers to help them continue driving safely and support their decision-making.
[0396] The system realizing this invention achieves efficient information processing while ensuring driver safety in an autonomous vehicle through the reception, analysis, and evaluation of voice information. The server converts the input voice into a digital format via a device for collecting voice information. Subsequently, it converts the voice into text information using speech recognition technology, analyzes the intent of the information through natural language processing (NLP), and extracts scheduled information and other useful information.
[0397] The server also uses emotion recognition technology to evaluate the emotional state within text information and determine the urgency of the information. Based on this information, it notifies the driver in audio or visual form as a driving assistance measure. This allows for timely decision-making while maintaining the driver's safety.
[0398] The hardware used includes microphones and computer systems installed in the autonomous vehicle. The software includes speech recognition engines (e.g., Google Speech-to-Text, IBM Watson), natural language processing libraries (e.g., spaCy, NLTK), and sentiment analysis tools (e.g., IBM Watson Tone Analyzer).
[0399] As a concrete example, imagine a scenario where, when a user receives an important phone call while driving, the content is quickly converted to text, sentiment analysis is performed, and the driver is notified with a message such as, "You have an important message from your boss."
[0400] An example of a prompt is: "Please describe in detail the design of an AI assistant that analyzes voice messages left as missed calls while a user is traveling in an autonomous vehicle, provides notifications according to their urgency, and suggests responses based on sentiment data." By using this prompt and leveraging a generative AI model, even more advanced natural language processing capabilities and sentiment analysis can be achieved.
[0401] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0402] Step 1:
[0403] The server uses a microphone to receive audio information input from within the autonomous vehicle and converts it into a digital format. The input is the user's voice, and the output is digital audio data. This data is pre-processed for speech recognition.
[0404] Step 2:
[0405] The server uses a speech recognition engine to convert digital audio data into text information. The input is digital audio data, and the output is text data. The converted text data is then prepared for analysis for natural language processing.
[0406] Step 3:
[0407] The server uses natural language processing to analyze text information and extract important schedule information and instructions. The input is text data, and the output is the extracted schedule information and instructions. This information is then passed to the schedule registration and notification system.
[0408] Step 4:
[0409] The server uses an emotion recognition engine to assess the user's emotional state from audio and text information. Input is audio or text data, and output is data related to the emotional state. The results of the emotion analysis are used to assess urgency.
[0410] Step 5:
[0411] The server determines the urgency level based on the extracted schedule information and emotional state, and notifies the user as appropriate. The input is schedule information and emotional evaluation data, and the output is an emergency notification. For example, the driver might receive a notification saying, "You have an urgent message from your supervisor."
[0412] Step 6:
[0413] Users can receive information provided through driving assistance systems and take necessary actions in a safe manner. Inputs are notifications and information provided by the server, and outputs are the corresponding actions. This allows users to respond efficiently and safely even while driving.
[0414] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0415] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0416] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0417] [Third Embodiment]
[0418] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0419] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0420] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0421] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0422] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0423] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0424] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0425] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0426] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0427] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0428] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0429] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0430] The AI personal assistant system of this invention is built using a mobile device and a cloud server. When a call comes in while the user is away, the device automatically records the audio and transfers the data to the server. The server receives this audio data and converts it into text data using speech recognition technology. The converted text is analyzed using natural language processing to extract information, particularly related to dates, times, and events. As a result, information related to the user's schedule is automatically registered in the calendar.
[0431] Furthermore, the server has the capability to translate text data into a specified language if the text data is in a different language. The translated text is then notified to the user. For example, if English is used in a phone call from overseas, the server will translate the content into Japanese and provide it to the user in an easily understandable format.
[0432] As a concrete example, consider a scenario where a user is away from home and someone leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday." In this case, the device records the audio and sends the data to the server. The server transcribes the audio into text, extracts the necessary information, and automatically registers the meeting in the user's calendar. Simultaneously, it translates the content into the user's preferred language and sends a summary as a notification.
[0433] These features allow the system to streamline voice message confirmation and schedule management, and it can be used as an advanced platform that also supports communication in different languages.
[0434] The following describes the processing flow.
[0435] Step 1:
[0436] The device automatically starts recording audio when a call is missed. Once recording is complete, it sends the audio data to the server in digital format.
[0437] Step 2:
[0438] The server receives the voice data sent from the terminal and converts it into text data using speech recognition technology. This text data is then stored in a database.
[0439] Step 3:
[0440] The server performs natural language processing to parse text data and extracts information about dates, times, and events. This information is used to adjust schedules.
[0441] Step 4:
[0442] The server generates a schedule based on the extracted date and time information and automatically registers the appointment in conjunction with the terminal's calendar application.
[0443] Step 5:
[0444] The server determines if the text data is in a specific language and translates it into the specified language if necessary. The translated text is then saved back into the database.
[0445] Step 6:
[0446] The server summarizes the text data and translated content as needed and sends it to the terminal for notification to the user.
[0447] Step 7:
[0448] The device notifies the user via push notification or email of text data, translation results, and summaries received from the server.
[0449] This process allows the system to efficiently process phone messages and provide users with relevant information.
[0450] (Example 1)
[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0452] In modern information and communication, it is essential to efficiently manage voice messages even when users are absent, and to facilitate smooth communication, especially between languages. However, conventional technologies have not adequately addressed this need for automated voice processing and translation, as well as centralized schedule management. This puts users at risk of overlooking important information.
[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0454] In this invention, the server includes acoustic receiving means for receiving acoustic signals and converting them into information signals, acoustic recognition means for converting acoustic signals into text information, and natural language processing means for analyzing the text information and extracting time information. This enables the system to appropriately process voice messages even when the user is absent, support communication between different languages, and automatically manage schedules.
[0455] An "acoustic signal" is an electrical signal obtained by converting sound into an electrical signal, and it is the basic unit for processing sound as digital data.
[0456] An "information signal" is digital data converted from an acoustic signal, and is data in a format that can be processed by a computer.
[0457] "Acoustic receiving means" refers to a device or function for receiving acoustic signals and converting them into information signals.
[0458] "Acoustic recognition means" refers to a device or function that analyzes data input as an acoustic signal and converts its content into textual information.
[0459] "Textual information" refers to text data converted from acoustic signals, and is information represented as a string of characters.
[0460] "Natural language processing means" refers to a device or function for analyzing textual information and extracting specific patterns or meanings.
[0461] "Time information" refers to date and time information extracted by natural language processing tools, and is data used for schedule management.
[0462] A "schedule registration means" is a device or function for registering information in a calendar system or schedule management service based on extracted time information.
[0463] A "language conversion means" is a device or function that translates textual information into another language, enabling communication between different languages.
[0464] A "notification means" is a device or function that notifies the user of converted character information or time information, enabling the recipient of the information to properly understand its content.
[0465] "Communication means" refers to a device or function for sharing voice data or information signals with remote devices or services via a network.
[0466] This invention enables efficient processing of voice messages even when the user is absent, and facilitates smooth communication between multiple languages, through a system consisting of a mobile terminal and a cloud server.
[0467] Calls received while the user is unable to answer are recorded as audio by the device. This audio signal is converted into a digital signal using the built-in microphone and transmitted from the device to a cloud server. Communication is conducted via an internet connection, and data security is ensured using the HTTPS protocol.
[0468] The server converts the received acoustic signal into text information using acoustic recognition means. This process utilizes general speech recognition software, such as a speech recognition API. The converted text information is then analyzed by natural language processing means to extract important information such as date, time, and event information. This step uses natural language processing libraries such as spaCy to analyze patterns and meaning.
[0469] The extracted time information is registered in the calendar system by the server using a schedule registration method. For example, to integrate with an external scheduling service, the schedule is automatically set using a cloud calendar API.
[0470] Furthermore, the server uses language conversion means to translate text information if it differs from the language set by the user. The translation technology utilizes a machine translation API to facilitate communication between different languages.
[0471] Finally, the server notifies the user of the information that has been converted and translated by the notification system. This is done using methods such as push notifications or email so that the user can check the information immediately.
[0472] As a concrete example, if a user leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday," while they are away, the device records the audio and sends it to the server. The server processes the audio according to a prompt that says, "Translate the following audio message into text, extract important information, and add it to the calendar. Also, translate and notify the user if necessary." The server adds the necessary information to the calendar, translates it if necessary, and notifies the user of a summary. In this way, the user can manage important appointments efficiently without missing any.
[0473] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0474] Step 1:
[0475] The device automatically records calls if the user is unable to answer. The input is an analog audio signal captured via the telephone line, which is then converted into a digital audio signal. This process utilizes the device's built-in microphone and voice recording application, and the resulting digital audio is temporarily stored in the device's memory.
[0476] Step 2:
[0477] The device transfers the recorded audio data to a cloud server. The input for this step is the digital audio obtained in step 1, and the output is the audio data sent to the server. An internet connection is used to transmit the data, and the HTTPS protocol ensures data security and privacy.
[0478] Step 3:
[0479] The server converts received acoustic signals into text information using acoustic recognition means. The input is an acoustic signal sent from a terminal, and text data is output by analyzing it using speech recognition software. For example, the audio "Let's schedule a meeting for 3 o'clock next Wednesday" is accurately displayed as text.
[0480] Step 4:
[0481] The server analyzes the text information using natural language processing (NLP) and extracts date, time, and event information. The input in this step is the text information obtained in step 3, and the output is the extracted time and event information. Important information is identified from the text using a natural language processing library and registered in the database.
[0482] Step 5:
[0483] The server adds an event to the calendar using the event registration method based on the extracted date and time information. The input data is the time information analyzed in step 4, and the information is reflected in the user's calendar using the Google Calendar API or similar. The output is the newly added calendar event.
[0484] Step 6:
[0485] The server uses a language conversion mechanism to translate text information if it differs from the language set by the user. The input data is the original text information, and the output is the translated text in the specified language. A machine translation API is used to facilitate smooth information sharing between different languages.
[0486] Step 7:
[0487] The server notifies the user of the converted and translated information through a notification system. The input for this final step is the various data obtained from step 3 onward, and the server informs the user of the schedule and translated information extracted from the voice message through the notification. The output is detailed information received by the user via push notifications or email.
[0488] (Application Example 1)
[0489] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0490] In modern households, important phone calls and messages are often missed when family members are away, and communication in different languages makes understanding information and managing schedules particularly challenging. Furthermore, there is a need to efficiently manage the schedules of all family members and receive real-time notifications to improve their quality of life.
[0491] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0492] In this invention, the server includes means for receiving voice data and converting it into a digital format, means for converting the voice data into text data, and means for analyzing the text data and extracting schedule information. This makes it possible to use a home robotic device to reliably manage schedules even when family members are away, quickly translate information in different languages, and provide real-time notifications in an easy-to-understand format.
[0493] "Audio data" refers to a data format in which sound is recorded digitally and used for information processing or communication.
[0494] A "digital format" is a format that can be processed by a computer by converting analog signals such as audio and video into numerical data.
[0495] "Audio receiving means" refers to a hardware or software configuration for receiving audio data and converting it into a digital format.
[0496] "Speech recognition means" refers to a technology that analyzes received speech data and converts it into corresponding text data.
[0497] "Text data" refers to data in string format obtained through speech recognition.
[0498] "Natural language processing methods" are technologies for analyzing text data to extract specific information (e.g., date and time, events).
[0499] A "schedule registration method" is a system that automatically registers extracted schedule information into a calendar or schedule management application.
[0500] "Translation methods" refer to technologies that convert text data in a specific language into a different language.
[0501] A "notification method" is a mechanism that informs the user's information terminal of the generated text data or translated information.
[0502] A "household robotic device" is a robot designed for use within the home and possesses a variety of functions, such as recording voice data.
[0503] To implement this invention, the home robot system is configured as follows: The robot device is equipped with a voice receiving means and converts voice data received within the home into a digital format. The converted voice is sent to a cloud server. The server converts the voice data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose.
[0504] The converted text data is analyzed using natural language processing technology on the cloud server to extract schedule information. Azure Cognitive Services' natural language processing capabilities are used for this process. The extracted schedule information is then automatically registered to the user's calendar by the server using the Google Calendar API.
[0505] Furthermore, if the audio content is in a specific language, the server uses Azure Translator to translate it into a different language. This translated text is then sent to the user's mobile device via a notification system. Real-time messaging services such as Pusher are used for notifications.
[0506] For example, if a family member is out and a robot records a message saying, "I have plans to meet a friend next Sunday," the server can transcribe the message into text, translate it from Japanese to English, register the event in Google Calendar, and then send the translated content to the user's device as a collaborator.
[0507] An example of a prompt message could be entered as follows: "Please wake me up at 3 AM tomorrow. If it's raining, please let me know instead of going to a cafe."
[0508] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0509] Step 1:
[0510] The terminal uses a home robotic device to receive voice data from phone calls made while the user is away and convert it into a digital format. In this step, voice input is taken and output is produced in the form of digital voice data. The voice data is temporarily stored in local storage.
[0511] Step 2:
[0512] The device sends digital audio data to a cloud server. Here, digital audio data is passed to the server as input, and as a result, the audio data is uploaded to the cloud.
[0513] Step 3:
[0514] The server performs speech recognition processing in the cloud and uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Audio data is input, the words in the audio are analyzed, and the text data is output.
[0515] Step 4:
[0516] The server analyzes text data using natural language processing techniques and extracts schedule information using Azure Cognitive Services. It extracts keywords such as date, time, and location from the input text data and outputs them as schedule information.
[0517] Step 5:
[0518] The server extracts schedule information and registers it in the user's calendar using the Google Calendar API. In this step, the schedule information is used as input, and the output is that the events are reflected in the calendar.
[0519] Step 6:
[0520] The server uses translation tools to translate text data into a desired language using Azure Translator if the original text is in a specific language. The original text is the input, and the translated text is the output.
[0521] Step 7:
[0522] The server uses notification methods to send translated text and schedule information to the user's mobile device via services such as Pusher. The entered text is output as a message to the user, and the information is notified to the device in real time.
[0523] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0524] The present invention provides an AI personal assistant that automatically handles incoming calls while the user is away. At the core of the system is a function that receives voice data and converts it into text data using speech recognition technology. Furthermore, it can analyze the text data using natural language processing, extract schedule information and summaries, and automatically register them in the user's calendar.
[0525] The emotion engine, in addition to this, plays the role of recognizing the user's emotions from the voice and generating emotion-based responses and suggestions. The server passes the received voice data to the emotion engine, and the server decides on a course of action based on the analysis results. For example, if the received voice contains tension or anxiety, the system is designed to notify the user of this information and provide situation-appropriate advice.
[0526] As a concrete example, if a user receives an important phone call while in a meeting, the device immediately sends the audio data to the server. The server converts the audio to text and analyzes it using an emotion engine. This analysis determines the emotional urgency of the message, and the user is notified with the message, "You have an urgent message." Furthermore, appropriate actions and countermeasures are suggested depending on the situation, allowing the user to respond quickly and accurately.
[0527] This system, equipped with such processes, is not merely a scheduling and translation tool, but an innovative platform that utilizes sentiment analysis to enable more human-like responses.
[0528] The following describes the processing flow.
[0529] Step 1:
[0530] The device automatically starts recording audio when a call is missed. The recorded audio data is converted to a digital format and sent to the server.
[0531] Step 2:
[0532] The server transmits the received audio data to the speech recognition engine, which converts the audio into text data. The speech recognition engine analyzes the data and stores the converted text in a database.
[0533] Step 3:
[0534] The server processes the text data using a natural language processing engine, extracts schedule information, and generates data for schedule registration. This information is synchronized with the device's calendar.
[0535] Step 4:
[0536] The server uses an emotion engine to perform sentiment analysis on the voice data. This evaluates the emotional nuances contained in the voice message and creates the foundational data for communicating the results to the user.
[0537] Step 5:
[0538] Based on the sentiment analysis results obtained from the analysis, the server notifies the user as needed and generates situation-appropriate responses and suggestions. This information is transmitted to the terminal by the notification system.
[0539] Step 6:
[0540] The device notifies the user of text data sent from the server, sentiment analysis results, and possible countermeasures. Notifications are sent via push notifications or email.
[0541] Step 7:
[0542] Users review notifications from their devices and select appropriate actions based on recommended responses derived from sentiment analysis.
[0543] This entire process allows the system to provide proactive support to the user and enables a deeper understanding of the content, including emotional insights.
[0544] (Example 2)
[0545] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0546] In telephone communication, properly handling messages received while the user is absent is a challenging task. Conventional technologies simply record voice messages or convert them to text, failing to consider the importance or emotional content of the received messages. Furthermore, there is a risk of important information being overlooked, and emotional responses may be delayed. This invention aims to solve these problems.
[0547] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0548] In this invention, the server includes a voice receiving means for receiving voice data and converting it into a digital format, a voice recognition means for converting the voice data into text data, and an emotion analysis means for detecting emotions from the text data. This enables the server to comprehensively analyze the content and emotional elements of received messages, even when the user is absent, and to respond quickly and appropriately.
[0549] "Voice receiving means" refers to a device or software that has the function of receiving voice data and converting it into a digital format.
[0550] "Speech recognition means" refers to a technology or device for converting digital speech data into text data.
[0551] "Natural language processing means" refers to techniques or methods used to analyze text data and extract planned information or summaries.
[0552] A "schedule registration method" is a device or software that automatically registers extracted schedule information into the user's calendar or management system.
[0553] "Emotional analysis means" refers to a technology or device for detecting emotions from text data or audio data and analyzing the results.
[0554] "Sentiment notification means" refers to a device or software that provides appropriate notifications to the user based on the results of sentiment analysis.
[0555] "Notification means" refers to a device or method that transmits converted text data or analysis results to the user to inform them.
[0556] The present invention provides a voice processing platform for automatically handling incoming calls while the user is away. A terminal detects a call, records the voice data, and sends it to a server. The server uses speech recognition technology to convert the voice data into text data. General-purpose speech recognition software can be used for speech recognition. For example, open-source software or cloud-based speech recognition services can be utilized.
[0557] The server analyzes the converted text data using natural language processing. This natural language processing utilizes a natural language processing engine, which is also used as a generative AI model. This allows for the extraction of schedule information and conversation summaries from the text data. Specific software options include general API services and open-source natural language processing libraries.
[0558] Next, the server uses sentiment analysis tools to detect emotions from the text data. Existing sentiment analysis APIs and modules can be used for sentiment analysis, determining the emotional urgency and importance of the received message. This sentiment information is used for user notifications and additional responses.
[0559] As a concrete example, consider a situation where a user is in a meeting and unable to answer a phone call. The device automatically sends the phone's audio data to the server. The server performs speech recognition, generates text data, extracts schedule information using natural language processing, and applies sentiment analysis. Based on these results, the server can send a notification to the user stating, "You have an urgent message," and prompt them with suggestions or actions.
[0560] As a concrete example of a prompt message to the generating AI model, text such as "Extract schedule information from phone messages, perform sentiment analysis, and notify the results through the library" can be used. This allows the system of the present invention to go beyond simple schedule management and enable more human-like and effective responses by incorporating sentiment analysis.
[0561] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0562] Step 1:
[0563] The device detects incoming calls while the user is away, records the audio data, and converts it to a digital format. The input is an analog signal from the telephone line, which is then output as digital audio data. Specifically, it activates the audio recording function upon receiving a call signal and saves the recorded audio as a digital file.
[0564] Step 2:
[0565] The terminal sends the recorded digital audio data to the server. The input is the digital audio file generated in step 1, which is then transferred to the server via the network. Specifically, a stable internet connection is used, and a file transfer protocol is employed to ensure the data is reliably transferred to the server.
[0566] Step 3:
[0567] The server converts received audio data into text data using speech recognition technology. The input is a digital audio file, and the output is a textual representation of that audio. Specifically, it launches speech recognition software and executes a process of sequentially generating text by analyzing the audio data.
[0568] Step 4:
[0569] The server analyzes the generated text data using natural language processing tools and extracts schedule information and summaries. The input is the text data obtained in step 3, and the output is the extracted schedule information and summaries. Specifically, it uses a natural language processing engine to analyze the sentence structure and identify important information.
[0570] Step 5:
[0571] The server detects emotions in text data using sentiment analysis tools and analyzes the results. The input is the text data obtained in step 3, and the output is the sentiment analysis results and the urgency of the information. Specifically, it starts the sentiment analysis module and calculates a sentiment score based on keywords and context in the text.
[0572] Step 6:
[0573] The server provides the user with appropriate notifications based on schedule information and sentiment analysis results, and suggests actions if necessary. The input is the schedule and sentiment information obtained in steps 4 and 5, and the output is a notification message to the user. Specifically, it generates the notification message and sends it to the user's mobile device or other device.
[0574] (Application Example 2)
[0575] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0576] In autonomous vehicles, drivers must dedicate their hands and attention to external communication, which can raise safety concerns. Furthermore, responding quickly and appropriately to audio information received while driving is difficult. Additionally, methods to identify emotional states and reduce the driver's psychological burden are needed.
[0577] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0578] In this invention, the server includes means for receiving voice information and converting it into a digital format, natural language processing means for analyzing text information and extracting planned information, and emotion recognition means for evaluating emotional states in response to stimuli. This enables the driver to safely and comfortably manage external communications even while driving, and to make quick decisions to ensure safety.
[0579] "Audio information" refers to information received as digital data through an input device.
[0580] "Converting to digital format" refers to the process of changing analog audio data into digital data.
[0581] "Voice receiving means" refers to an interface for collecting voice information.
[0582] "Speech recognition means" refers to technology for analyzing speech information and converting it into text information.
[0583] "Text information" refers to data in character format obtained through speech recognition.
[0584] "Natural language processing methods" are technologies for analyzing text information and extracting meaningful information.
[0585] "Schedule information" refers to schedule-related information extracted from text information using natural language processing tools.
[0586] The "schedule registration method" refers to a function for registering acquired schedule information into a time management system such as a calendar.
[0587] "Translation methods" refer to technologies that convert text information in a specific language into a different language.
[0588] "Notification means" refers to a method or device for informing a user of information.
[0589] "Emotional state" refers to the emotional state of the speaker as identified from their voice.
[0590] "Emotion recognition means" refers to technology for evaluating a user's emotional state from audio information.
[0591] "Urgency" refers to the degree to which the received information is assessed as being urgent.
[0592] "Driving assistance systems" are technologies that provide information to drivers to help them continue driving safely and support their decision-making.
[0593] The system realizing this invention achieves efficient information processing while ensuring driver safety in an autonomous vehicle through the reception, analysis, and evaluation of voice information. The server converts the input voice into a digital format via a device for collecting voice information. Subsequently, it converts the voice into text information using speech recognition technology, analyzes the intent of the information through natural language processing (NLP), and extracts scheduled information and other useful information.
[0594] The server also uses emotion recognition technology to evaluate the emotional state within text information and determine the urgency of the information. Based on this information, it notifies the driver in audio or visual form as a driving assistance measure. This allows for timely decision-making while maintaining the driver's safety.
[0595] The hardware used includes microphones and computer systems installed in the autonomous vehicle. The software includes speech recognition engines (e.g., Google Speech-to-Text, IBM Watson), natural language processing libraries (e.g., spaCy, NLTK), and sentiment analysis tools (e.g., IBM Watson Tone Analyzer).
[0596] As a concrete example, imagine a scenario where, when a user receives an important phone call while driving, the content is quickly converted to text, sentiment analysis is performed, and the driver is notified with a message such as, "You have an important message from your boss."
[0597] An example of a prompt is: "Please describe in detail the design of an AI assistant that analyzes voice messages left as missed calls while a user is traveling in an autonomous vehicle, provides notifications according to their urgency, and suggests responses based on sentiment data." By using this prompt and leveraging a generative AI model, even more advanced natural language processing capabilities and sentiment analysis can be achieved.
[0598] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0599] Step 1:
[0600] The server uses a microphone to receive audio information input from within the autonomous vehicle and converts it into a digital format. The input is the user's voice, and the output is digital audio data. This data is pre-processed for speech recognition.
[0601] Step 2:
[0602] The server uses a speech recognition engine to convert digital audio data into text information. The input is digital audio data, and the output is text data. The converted text data is then prepared for analysis for natural language processing.
[0603] Step 3:
[0604] The server uses natural language processing to analyze text information and extract important schedule information and instructions. The input is text data, and the output is the extracted schedule information and instructions. This information is then passed to the schedule registration and notification system.
[0605] Step 4:
[0606] The server uses an emotion recognition engine to assess the user's emotional state from audio and text information. Input is audio or text data, and output is data related to the emotional state. The results of the emotion analysis are used to assess urgency.
[0607] Step 5:
[0608] The server determines the urgency level based on the extracted schedule information and emotional state, and notifies the user as appropriate. The input is schedule information and emotional evaluation data, and the output is an emergency notification. For example, the driver might receive a notification saying, "You have an urgent message from your supervisor."
[0609] Step 6:
[0610] Users can receive information provided through driving assistance systems and take necessary actions in a safe manner. Inputs are notifications and information provided by the server, and outputs are the corresponding actions. This allows users to respond efficiently and safely even while driving.
[0611] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0612] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0613] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0614] [Fourth Embodiment]
[0615] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0616] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0617] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0618] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0619] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0620] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0621] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0622] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0623] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0624] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0625] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0626] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0627] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0628] The AI personal assistant system of this invention is built using a mobile device and a cloud server. When a call comes in while the user is away, the device automatically records the audio and transfers the data to the server. The server receives this audio data and converts it into text data using speech recognition technology. The converted text is analyzed using natural language processing to extract information, particularly related to dates, times, and events. As a result, information related to the user's schedule is automatically registered in the calendar.
[0629] Furthermore, the server has the capability to translate text data into a specified language if the text data is in a different language. The translated text is then notified to the user. For example, if English is used in a phone call from overseas, the server will translate the content into Japanese and provide it to the user in an easily understandable format.
[0630] As a concrete example, consider a scenario where a user is away from home and someone leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday." In this case, the device records the audio and sends the data to the server. The server transcribes the audio into text, extracts the necessary information, and automatically registers the meeting in the user's calendar. Simultaneously, it translates the content into the user's preferred language and sends a summary as a notification.
[0631] These features allow the system to streamline voice message confirmation and schedule management, and it can be used as an advanced platform that also supports communication in different languages.
[0632] The following describes the processing flow.
[0633] Step 1:
[0634] The device automatically starts recording audio when a call is missed. Once recording is complete, it sends the audio data to the server in digital format.
[0635] Step 2:
[0636] The server receives the voice data sent from the terminal and converts it into text data using speech recognition technology. This text data is then stored in a database.
[0637] Step 3:
[0638] The server performs natural language processing to parse text data and extracts information about dates, times, and events. This information is used to adjust schedules.
[0639] Step 4:
[0640] The server generates a schedule based on the extracted date and time information and automatically registers the appointment in conjunction with the terminal's calendar application.
[0641] Step 5:
[0642] The server determines if the text data is in a specific language and translates it into the specified language if necessary. The translated text is then saved back into the database.
[0643] Step 6:
[0644] The server summarizes the text data and translated content as needed and sends it to the terminal for notification to the user.
[0645] Step 7:
[0646] The device notifies the user via push notification or email of text data, translation results, and summaries received from the server.
[0647] This process allows the system to efficiently process phone messages and provide users with relevant information.
[0648] (Example 1)
[0649] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0650] In modern information and communication, it is essential to efficiently manage voice messages even when users are absent, and to facilitate smooth communication, especially between languages. However, conventional technologies have not adequately addressed this need for automated voice processing and translation, as well as centralized schedule management. This puts users at risk of overlooking important information.
[0651] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0652] In this invention, the server includes acoustic receiving means for receiving acoustic signals and converting them into information signals, acoustic recognition means for converting acoustic signals into text information, and natural language processing means for analyzing the text information and extracting time information. This enables the system to appropriately process voice messages even when the user is absent, support communication between different languages, and automatically manage schedules.
[0653] An "acoustic signal" is an electrical signal obtained by converting sound into an electrical signal, and it is the basic unit for processing sound as digital data.
[0654] An "information signal" is digital data converted from an acoustic signal, and is data in a format that can be processed by a computer.
[0655] "Acoustic receiving means" refers to a device or function for receiving acoustic signals and converting them into information signals.
[0656] "Acoustic recognition means" refers to a device or function that analyzes data input as an acoustic signal and converts its content into textual information.
[0657] "Textual information" refers to text data converted from acoustic signals, and is information represented as a string of characters.
[0658] "Natural language processing means" refers to a device or function for analyzing textual information and extracting specific patterns or meanings.
[0659] "Time information" refers to date and time information extracted by natural language processing tools, and is data used for schedule management.
[0660] A "schedule registration means" is a device or function for registering information in a calendar system or schedule management service based on extracted time information.
[0661] A "language conversion means" is a device or function that translates textual information into another language, enabling communication between different languages.
[0662] A "notification means" is a device or function that notifies the user of converted character information or time information, enabling the recipient of the information to properly understand its content.
[0663] "Communication means" refers to a device or function for sharing voice data or information signals with remote devices or services via a network.
[0664] This invention enables efficient processing of voice messages even when the user is absent, and facilitates smooth communication between multiple languages, through a system consisting of a mobile terminal and a cloud server.
[0665] Calls received while the user is unable to answer are recorded as audio by the device. This audio signal is converted into a digital signal using the built-in microphone and transmitted from the device to a cloud server. Communication is conducted via an internet connection, and data security is ensured using the HTTPS protocol.
[0666] The server converts the received acoustic signal into text information using acoustic recognition means. This process utilizes general speech recognition software, such as a speech recognition API. The converted text information is then analyzed by natural language processing means to extract important information such as date, time, and event information. This step uses natural language processing libraries such as spaCy to analyze patterns and meaning.
[0667] The extracted time information is registered in the calendar system by the server using a schedule registration method. For example, to integrate with an external scheduling service, the schedule is automatically set using a cloud calendar API.
[0668] Furthermore, the server uses language conversion means to translate text information if it differs from the language set by the user. The translation technology utilizes a machine translation API to facilitate communication between different languages.
[0669] Finally, the server notifies the user of the information that has been converted and translated by the notification system. This is done using methods such as push notifications or email so that the user can check the information immediately.
[0670] As a concrete example, if a user leaves a message saying, "Let's schedule a meeting for 3 PM next Wednesday," while they are away, the device records the audio and sends it to the server. The server processes the audio according to a prompt that says, "Translate the following audio message into text, extract important information, and add it to the calendar. Also, translate and notify the user if necessary." The server adds the necessary information to the calendar, translates it if necessary, and notifies the user of a summary. In this way, the user can manage important appointments efficiently without missing any.
[0671] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0672] Step 1:
[0673] The device automatically records calls if the user is unable to answer. The input is an analog audio signal captured via the telephone line, which is then converted into a digital audio signal. This process utilizes the device's built-in microphone and voice recording application, and the resulting digital audio is temporarily stored in the device's memory.
[0674] Step 2:
[0675] The device transfers the recorded audio data to a cloud server. The input for this step is the digital audio obtained in step 1, and the output is the audio data sent to the server. An internet connection is used to transmit the data, and the HTTPS protocol ensures data security and privacy.
[0676] Step 3:
[0677] The server converts received acoustic signals into text information using acoustic recognition means. The input is an acoustic signal sent from a terminal, and text data is output by analyzing it using speech recognition software. For example, the audio "Let's schedule a meeting for 3 o'clock next Wednesday" is accurately displayed as text.
[0678] Step 4:
[0679] The server analyzes the text information using natural language processing (NLP) and extracts date, time, and event information. The input in this step is the text information obtained in step 3, and the output is the extracted time and event information. Important information is identified from the text using a natural language processing library and registered in the database.
[0680] Step 5:
[0681] The server adds an event to the calendar using the event registration method based on the extracted date and time information. The input data is the time information analyzed in step 4, and the information is reflected in the user's calendar using the Google Calendar API or similar. The output is the newly added calendar event.
[0682] Step 6:
[0683] The server uses a language conversion mechanism to translate text information if it differs from the language set by the user. The input data is the original text information, and the output is the translated text in the specified language. A machine translation API is used to facilitate smooth information sharing between different languages.
[0684] Step 7:
[0685] The server notifies the user of the converted and translated information through a notification system. The input for this final step is the various data obtained from step 3 onward, and the server informs the user of the schedule and translated information extracted from the voice message through the notification. The output is detailed information received by the user via push notifications or email.
[0686] (Application Example 1)
[0687] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0688] In modern households, important phone calls and messages are often missed when family members are away, and communication in different languages makes understanding information and managing schedules particularly challenging. Furthermore, there is a need to efficiently manage the schedules of all family members and receive real-time notifications to improve their quality of life.
[0689] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0690] In this invention, the server includes means for receiving voice data and converting it into a digital format, means for converting the voice data into text data, and means for analyzing the text data and extracting schedule information. This makes it possible to use a home robotic device to reliably manage schedules even when family members are away, quickly translate information in different languages, and provide real-time notifications in an easy-to-understand format.
[0691] "Audio data" refers to a data format in which sound is recorded digitally and used for information processing or communication.
[0692] A "digital format" is a format that can be processed by a computer by converting analog signals such as audio and video into numerical data.
[0693] "Audio receiving means" refers to a hardware or software configuration for receiving audio data and converting it into a digital format.
[0694] "Speech recognition means" refers to a technology that analyzes received speech data and converts it into corresponding text data.
[0695] "Text data" refers to data in string format obtained through speech recognition.
[0696] "Natural language processing methods" are technologies for analyzing text data to extract specific information (e.g., date and time, events).
[0697] A "schedule registration method" is a system that automatically registers extracted schedule information into a calendar or schedule management application.
[0698] "Translation methods" refer to technologies that convert text data in a specific language into a different language.
[0699] A "notification method" is a mechanism that informs the user's information terminal of the generated text data or translated information.
[0700] A "household robotic device" is a robot designed for use within the home and possesses a variety of functions, such as recording voice data.
[0701] To implement this invention, the home robot system is configured as follows: The robot device is equipped with a voice receiving means and converts voice data received within the home into a digital format. The converted voice is sent to a cloud server. The server converts the voice data into text data using speech recognition technology. The Google Cloud Speech-to-Text API is used for this purpose.
[0702] The converted text data is analyzed using natural language processing technology on the cloud server to extract schedule information. Azure Cognitive Services' natural language processing capabilities are used for this process. The extracted schedule information is then automatically registered to the user's calendar by the server using the Google Calendar API.
[0703] Furthermore, if the audio content is in a specific language, the server uses Azure Translator to translate it into a different language. This translated text is then sent to the user's mobile device via a notification system. Real-time messaging services such as Pusher are used for notifications.
[0704] For example, if a family member is out and a robot records a message saying, "I have plans to meet a friend next Sunday," the server can transcribe the message into text, translate it from Japanese to English, register the event in Google Calendar, and then send the translated content to the user's device as a collaborator.
[0705] An example of a prompt message could be entered as follows: "Please wake me up at 3 AM tomorrow. If it's raining, please let me know instead of going to a cafe."
[0706] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0707] Step 1:
[0708] The terminal uses a home robotic device to receive voice data from phone calls made while the user is away and convert it into a digital format. In this step, voice input is taken and output is produced in the form of digital voice data. The voice data is temporarily stored in local storage.
[0709] Step 2:
[0710] The device sends digital audio data to a cloud server. Here, digital audio data is passed to the server as input, and as a result, the audio data is uploaded to the cloud.
[0711] Step 3:
[0712] The server performs speech recognition processing in the cloud and uses the Google Cloud Speech-to-Text API to convert the audio data into text data. Audio data is input, the words in the audio are analyzed, and the text data is output.
[0713] Step 4:
[0714] The server analyzes text data using natural language processing techniques and extracts schedule information using Azure Cognitive Services. It extracts keywords such as date, time, and location from the input text data and outputs them as schedule information.
[0715] Step 5:
[0716] The server extracts schedule information and registers it in the user's calendar using the Google Calendar API. In this step, the schedule information is used as input, and the output is that the events are reflected in the calendar.
[0717] Step 6:
[0718] The server uses translation tools to translate text data into a desired language using Azure Translator if the original text is in a specific language. The original text is the input, and the translated text is the output.
[0719] Step 7:
[0720] The server uses notification methods to send translated text and schedule information to the user's mobile device via services such as Pusher. The entered text is output as a message to the user, and the information is notified to the device in real time.
[0721] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0722] The present invention provides an AI personal assistant that automatically handles incoming calls while the user is away. At the core of the system is a function that receives voice data and converts it into text data using speech recognition technology. Furthermore, it can analyze the text data using natural language processing, extract schedule information and summaries, and automatically register them in the user's calendar.
[0723] The emotion engine, in addition to this, plays the role of recognizing the user's emotions from the voice and generating emotion-based responses and suggestions. The server passes the received voice data to the emotion engine, and the server decides on a course of action based on the analysis results. For example, if the received voice contains tension or anxiety, the system is designed to notify the user of this information and provide situation-appropriate advice.
[0724] As a concrete example, if a user receives an important phone call while in a meeting, the device immediately sends the audio data to the server. The server converts the audio to text and analyzes it using an emotion engine. This analysis determines the emotional urgency of the message, and the user is notified with the message, "You have an urgent message." Furthermore, appropriate actions and countermeasures are suggested depending on the situation, allowing the user to respond quickly and accurately.
[0725] This system, equipped with such processes, is not merely a scheduling and translation tool, but an innovative platform that utilizes sentiment analysis to enable more human-like responses.
[0726] The following describes the processing flow.
[0727] Step 1:
[0728] The device automatically starts recording audio when a call is missed. The recorded audio data is converted to a digital format and sent to the server.
[0729] Step 2:
[0730] The server transmits the received audio data to the speech recognition engine, which converts the audio into text data. The speech recognition engine analyzes the data and stores the converted text in a database.
[0731] Step 3:
[0732] The server processes the text data using a natural language processing engine, extracts schedule information, and generates data for schedule registration. This information is synchronized with the device's calendar.
[0733] Step 4:
[0734] The server uses an emotion engine to perform sentiment analysis on the voice data. This evaluates the emotional nuances contained in the voice message and creates the foundational data for communicating the results to the user.
[0735] Step 5:
[0736] Based on the sentiment analysis results obtained from the analysis, the server notifies the user as needed and generates situation-appropriate responses and suggestions. This information is transmitted to the terminal by the notification system.
[0737] Step 6:
[0738] The device notifies the user of text data sent from the server, sentiment analysis results, and possible countermeasures. Notifications are sent via push notifications or email.
[0739] Step 7:
[0740] Users review notifications from their devices and select appropriate actions based on recommended responses derived from sentiment analysis.
[0741] This entire process allows the system to provide proactive support to the user and enables a deeper understanding of the content, including emotional insights.
[0742] (Example 2)
[0743] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] In telephone communication, properly handling messages received while the user is absent is a challenging task. Conventional technologies simply record voice messages or convert them to text, failing to consider the importance or emotional content of the received messages. Furthermore, there is a risk of important information being overlooked, and emotional responses may be delayed. This invention aims to solve these problems.
[0745] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0746] In this invention, the server includes a voice receiving means for receiving voice data and converting it into a digital format, a voice recognition means for converting the voice data into text data, and an emotion analysis means for detecting emotions from the text data. This enables the server to comprehensively analyze the content and emotional elements of received messages, even when the user is absent, and to respond quickly and appropriately.
[0747] "Voice receiving means" refers to a device or software that has the function of receiving voice data and converting it into a digital format.
[0748] "Speech recognition means" refers to a technology or device for converting digital speech data into text data.
[0749] "Natural language processing means" refers to techniques or methods used to analyze text data and extract planned information or summaries.
[0750] A "schedule registration method" is a device or software that automatically registers extracted schedule information into the user's calendar or management system.
[0751] "Emotional analysis means" refers to a technology or device for detecting emotions from text data or audio data and analyzing the results.
[0752] "Sentiment notification means" refers to a device or software that provides appropriate notifications to the user based on the results of sentiment analysis.
[0753] "Notification means" refers to a device or method that transmits converted text data or analysis results to the user to inform them.
[0754] The present invention provides a voice processing platform for automatically handling incoming calls while the user is away. A terminal detects a call, records the voice data, and sends it to a server. The server uses speech recognition technology to convert the voice data into text data. General-purpose speech recognition software can be used for speech recognition. For example, open-source software or cloud-based speech recognition services can be utilized.
[0755] The server analyzes the converted text data using natural language processing. This natural language processing utilizes a natural language processing engine, which is also used as a generative AI model. This allows for the extraction of schedule information and conversation summaries from the text data. Specific software options include general API services and open-source natural language processing libraries.
[0756] Next, the server uses sentiment analysis tools to detect emotions from the text data. Existing sentiment analysis APIs and modules can be used for sentiment analysis, determining the emotional urgency and importance of the received message. This sentiment information is used for user notifications and additional responses.
[0757] As a concrete example, consider a situation where a user is in a meeting and unable to answer a phone call. The device automatically sends the phone's audio data to the server. The server performs speech recognition, generates text data, extracts schedule information using natural language processing, and applies sentiment analysis. Based on these results, the server can send a notification to the user stating, "You have an urgent message," and prompt them with suggestions or actions.
[0758] As a concrete example of a prompt message to the generating AI model, text such as "Extract schedule information from phone messages, perform sentiment analysis, and notify the results through the library" can be used. This allows the system of the present invention to go beyond simple schedule management and enable more human-like and effective responses by incorporating sentiment analysis.
[0759] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0760] Step 1:
[0761] The device detects incoming calls while the user is away, records the audio data, and converts it to a digital format. The input is an analog signal from the telephone line, which is then output as digital audio data. Specifically, it activates the audio recording function upon receiving a call signal and saves the recorded audio as a digital file.
[0762] Step 2:
[0763] The terminal sends the recorded digital audio data to the server. The input is the digital audio file generated in step 1, which is then transferred to the server via the network. Specifically, a stable internet connection is used, and a file transfer protocol is employed to ensure the data is reliably transferred to the server.
[0764] Step 3:
[0765] The server converts received audio data into text data using speech recognition technology. The input is a digital audio file, and the output is a textual representation of that audio. Specifically, it launches speech recognition software and executes a process of sequentially generating text by analyzing the audio data.
[0766] Step 4:
[0767] The server analyzes the generated text data using natural language processing tools and extracts schedule information and summaries. The input is the text data obtained in step 3, and the output is the extracted schedule information and summaries. Specifically, it uses a natural language processing engine to analyze the sentence structure and identify important information.
[0768] Step 5:
[0769] The server detects emotions in text data using sentiment analysis tools and analyzes the results. The input is the text data obtained in step 3, and the output is the sentiment analysis results and the urgency of the information. Specifically, it starts the sentiment analysis module and calculates a sentiment score based on keywords and context in the text.
[0770] Step 6:
[0771] The server provides the user with appropriate notifications based on schedule information and sentiment analysis results, and suggests actions if necessary. The input is the schedule and sentiment information obtained in steps 4 and 5, and the output is a notification message to the user. Specifically, it generates the notification message and sends it to the user's mobile device or other device.
[0772] (Application Example 2)
[0773] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0774] In autonomous vehicles, drivers must dedicate their hands and attention to external communication, which can raise safety concerns. Furthermore, responding quickly and appropriately to audio information received while driving is difficult. Additionally, methods to identify emotional states and reduce the driver's psychological burden are needed.
[0775] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0776] In this invention, the server includes means for receiving voice information and converting it into a digital format, natural language processing means for analyzing text information and extracting planned information, and emotion recognition means for evaluating emotional states in response to stimuli. This enables the driver to safely and comfortably manage external communications even while driving, and to make quick decisions to ensure safety.
[0777] "Audio information" refers to information received as digital data through an input device.
[0778] "Converting to digital format" refers to the process of changing analog audio data into digital data.
[0779] "Voice receiving means" refers to an interface for collecting voice information.
[0780] "Speech recognition means" refers to technology for analyzing speech information and converting it into text information.
[0781] "Text information" refers to data in character format obtained through speech recognition.
[0782] "Natural language processing methods" are technologies for analyzing text information and extracting meaningful information.
[0783] "Schedule information" refers to schedule-related information extracted from text information using natural language processing tools.
[0784] The "schedule registration method" refers to a function for registering acquired schedule information into a time management system such as a calendar.
[0785] "Translation methods" refer to technologies that convert text information in a specific language into a different language.
[0786] "Notification means" refers to a method or device for informing a user of information.
[0787] "Emotional state" refers to the emotional state of the speaker as identified from their voice.
[0788] "Emotion recognition means" refers to technology for evaluating a user's emotional state from audio information.
[0789] "Urgency" refers to the degree to which the received information is assessed as being urgent.
[0790] "Driving assistance systems" are technologies that provide information to drivers to help them continue driving safely and support their decision-making.
[0791] The system realizing this invention achieves efficient information processing while ensuring driver safety in an autonomous vehicle through the reception, analysis, and evaluation of voice information. The server converts the input voice into a digital format via a device for collecting voice information. Subsequently, it converts the voice into text information using speech recognition technology, analyzes the intent of the information through natural language processing (NLP), and extracts scheduled information and other useful information.
[0792] The server also uses emotion recognition technology to evaluate the emotional state within text information and determine the urgency of the information. Based on this information, it notifies the driver in audio or visual form as a driving assistance measure. This allows for timely decision-making while maintaining the driver's safety.
[0793] The hardware used includes microphones and computer systems installed in the autonomous vehicle. The software includes speech recognition engines (e.g., Google Speech-to-Text, IBM Watson), natural language processing libraries (e.g., spaCy, NLTK), and sentiment analysis tools (e.g., IBM Watson Tone Analyzer).
[0794] As a concrete example, imagine a scenario where, when a user receives an important phone call while driving, the content is quickly converted to text, sentiment analysis is performed, and the driver is notified with a message such as, "You have an important message from your boss."
[0795] An example of a prompt is: "Please describe in detail the design of an AI assistant that analyzes voice messages left as missed calls while a user is traveling in an autonomous vehicle, provides notifications according to their urgency, and suggests responses based on sentiment data." By using this prompt and leveraging a generative AI model, even more advanced natural language processing capabilities and sentiment analysis can be achieved.
[0796] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0797] Step 1:
[0798] The server uses a microphone to receive audio information input from within the autonomous vehicle and converts it into a digital format. The input is the user's voice, and the output is digital audio data. This data is pre-processed for speech recognition.
[0799] Step 2:
[0800] The server uses a speech recognition engine to convert digital audio data into text information. The input is digital audio data, and the output is text data. The converted text data is then prepared for analysis for natural language processing.
[0801] Step 3:
[0802] The server uses natural language processing to analyze text information and extract important schedule information and instructions. The input is text data, and the output is the extracted schedule information and instructions. This information is then passed to the schedule registration and notification system.
[0803] Step 4:
[0804] The server uses an emotion recognition engine to assess the user's emotional state from audio and text information. Input is audio or text data, and output is data related to the emotional state. The results of the emotion analysis are used to assess urgency.
[0805] Step 5:
[0806] The server determines the urgency level based on the extracted schedule information and emotional state, and notifies the user as appropriate. The input is schedule information and emotional evaluation data, and the output is an emergency notification. For example, the driver might receive a notification saying, "You have an urgent message from your supervisor."
[0807] Step 6:
[0808] Users can receive information provided through driving assistance systems and take necessary actions in a safe manner. Inputs are notifications and information provided by the server, and outputs are the corresponding actions. This allows users to respond efficiently and safely even while driving.
[0809] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0810] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0811] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0812] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0813] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0814] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0815] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0816] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0817] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0818] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0819] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0820] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0821] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0822] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0823] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0824] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0825] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0826] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0827] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0828] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0829] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0830] The following is further disclosed regarding the embodiments described above.
[0831] (Claim 1)
[0832] A means for receiving audio data and converting it into a digital format,
[0833] A speech recognition means for converting the aforementioned speech data into text data,
[0834] A natural language processing means for analyzing the aforementioned text data and extracting schedule information,
[0835] A schedule registration means that automatically registers the aforementioned schedule information in a calendar,
[0836] A translation means for translating the aforementioned text data into a different language if it is in a specific language,
[0837] A notification means for notifying the converted text data and translated data,
[0838] A system that includes this.
[0839] (Claim 2)
[0840] The system according to claim 1, wherein the natural language processing means generates and notifies a summary of the conversation content.
[0841] (Claim 3)
[0842] The system according to claim 1, wherein the schedule registration means registers appointments in cooperation with an external calendar service.
[0843] "Example 1"
[0844] (Claim 1)
[0845] An acoustic receiving means that receives an acoustic signal and converts it into an information signal,
[0846] The acoustic recognition means converts the aforementioned acoustic signal into character information,
[0847] A natural language processing means for analyzing the aforementioned character information and extracting time information,
[0848] A schedule registration means that automatically registers the aforementioned time information to the schedule,
[0849] A language conversion means for translating the aforementioned textual information into a different language when it is in a specific language,
[0850] A notification means for notifying the converted text information and translated information,
[0851] A communication means for transmitting audio data to a remote processing unit via the internet,
[0852] A system that includes this.
[0853] (Claim 2)
[0854] The system according to claim 1, wherein the natural language processing means generates and reports an outline of the dialogue content.
[0855] (Claim 3)
[0856] The system according to claim 1, wherein the schedule registration means registers a schedule in cooperation with an external scheduling service.
[0857] "Application Example 1"
[0858] (Claim 1)
[0859] A means for receiving audio data and converting it into a digital format,
[0860] A speech recognition means for converting the aforementioned speech data into text data,
[0861] A natural language processing means for analyzing the aforementioned text data and extracting schedule information,
[0862] A schedule registration means that automatically registers the aforementioned schedule information in a calendar,
[0863] A translation means for translating the aforementioned text data into a different language if it is in a specific language,
[0864] A notification means for notifying the user's mobile device of the converted text data and translated data,
[0865] A home robot device installed in a home environment that records the aforementioned voice data,
[0866] A system that includes this.
[0867] (Claim 2)
[0868] The system according to claim 1, wherein the natural language processing means generates and notifies a summary of the conversation content.
[0869] (Claim 3)
[0870] The system according to claim 1, wherein the schedule registration means registers a schedule in cooperation with an external schedule management service.
[0871] "Example 2 of combining an emotion engine"
[0872] (Claim 1)
[0873] A means for receiving audio data and converting it into a digital format,
[0874] A speech recognition means for converting the aforementioned speech data into text data,
[0875] A natural language processing means for analyzing the aforementioned text data and extracting schedule information,
[0876] A schedule registration means that automatically registers the aforementioned schedule information in a calendar,
[0877] An emotion analysis means for detecting emotions from the aforementioned audio data,
[0878] A sentiment notification means that notifies the user based on the results of the aforementioned sentiment analysis,
[0879] A notification means for transmitting the converted text data and the notified data,
[0880] A system that includes this.
[0881] (Claim 2)
[0882] The system according to claim 1, wherein the natural language processing means generates and notifies a summary of the conversation content.
[0883] (Claim 3)
[0884] The system according to claim 1, wherein the schedule registration means and emotion notification means register information in cooperation with an external information management service.
[0885] "Application example 2 when combining with an emotional engine"
[0886] (Claim 1)
[0887] A means for receiving audio information and converting it into a digital format,
[0888] A speech recognition means that converts the aforementioned speech information into text information,
[0889] A natural language processing means for analyzing the aforementioned text information and extracting schedule information,
[0890] A schedule registration means that automatically registers the aforementioned schedule information into a schedule,
[0891] A translation means for translating the aforementioned text information into a different language if it is in a specific language,
[0892] A notification means for notifying the converted text information and translated information,
[0893] An emotion recognition means for identifying an emotional state from the aforementioned audio information and evaluating the degree of urgency,
[0894] A driver assistance system that provides drivers with information based on the urgency of the situation and supports them in responding in a safe manner,
[0895] A system that includes this.
[0896] (Claim 2)
[0897] The system according to claim 1, wherein the natural language processing means has the function of generating and notifying a summary of the conversation content, and playing relaxing music based on sentiment analysis.
[0898] (Claim 3)
[0899] The system according to claim 1, wherein the schedule registration means cooperates with an external calendar service to register schedules and manage notifications while driving a vehicle. [Explanation of Symbols]
[0900] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving audio data and converting it into a digital format, A speech recognition means for converting the aforementioned speech data into text data, A natural language processing means for analyzing the aforementioned text data and extracting schedule information, A schedule registration means that automatically registers the aforementioned schedule information in a calendar, A translation means for translating the aforementioned text data into a different language if it is in a specific language, A notification means for notifying the user's mobile device of the converted text data and translated data, A home robot device installed in a home environment that records the aforementioned voice data, A system that includes this.
2. The system according to claim 1, wherein the natural language processing means generates and notifies a summary of the conversation content.
3. The system according to claim 1, wherein the schedule registration means registers a schedule in cooperation with an external schedule management service.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A