system
The system converts and summarizes voice messages to text for immediate understanding, addressing the challenge of cumbersome voice message review by providing quick and efficient text summaries.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Checking voice messages on answering machines is cumbersome, especially for busy individuals or those with hearing impairments, and there is a need for a system that can quickly and efficiently convert and summarize voice data into readable text for immediate understanding.
A system that retrieves voice data from answering machines, converts it to text, summarizes important information, and sends the summary as a short message, utilizing speech recognition and summarization algorithms to ensure quick and effective review.
Enables users to grasp the content of voice messages quickly and efficiently without listening, accommodating those with hearing difficulties and busy environments.
Smart Images

Figure 2026071052000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When checking a message recorded on an answering machine, it is troublesome to play the voice data, and it is particularly difficult to immediately grasp the content when one is busy or in an inconvenient environment. Also, for people with hearing constraints, it is difficult to check voice messages. There is a need to quickly and efficiently solve such problems and enable immediate recognition of the content of the recorded message.
Means for Solving the Problems
[0005] This system retrieves voice data recorded on an answering machine, converts it to text data, shortens the content by summarizing important information, and sends the summarized information to the user as a short message. By including means for converting voice data to text data, means for extracting and summarizing necessary information from the text data, and means for sending the summarized information as a short message, this system enables the quick and effective review of recorded voice messages.
[0006] "Audio data" refers to audio information recorded in digital format.
[0007] "Text data" refers to digital data obtained by converting audio information into text.
[0008] A "summary" is a shortened version of text data, created by extracting essential information from the original text.
[0009] "Short message" refers to a communication method that involves sending short messages via digital communication.
[0010] A "system" is a whole composed of multiple parts or elements that work together to achieve a specific function or purpose. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] This system aims to quickly and easily summarize voice messages recorded on answering machines and send them to users as short messages. The main components of the system include a server that acquires voice data, a server that converts the voice to text and summarizes it, and a server that sends the summarized information to the user's terminal.
[0033] The server first connects with the answering machine system and immediately retrieves the digital audio data as soon as a voice message is recorded. This data is securely transferred with the user's permission. Next, the server uses speech recognition technology to convert the audio data into text. This conversion utilizes advanced algorithms that analyze the context of spoken language to generate accurate text.
[0034] Next, the server processes the generated text using a summarization algorithm to extract the most important information. This summarization allows for the concise transmission of the essence of the information without compromising the overall meaning of the message. For example, if a voice message is left on the answering machine saying, "Tomorrow's meeting has been changed from 10:00 to 11:00," the server will summarize this as, "Meeting changed from 10:00 to 11:00."
[0035] Finally, the server sends the summarized text to the user's terminal in the form of a short message. A secure messaging protocol is used for the transmission process to prevent information leakage. The user can check the message on their terminal and understand the content without listening to the voicemail message. This allows users who cannot access audio data or who have hearing impairments to efficiently grasp the information.
[0036] The following describes the processing flow.
[0037] Step 1:
[0038] The server connects to the answering machine system and detects that a voice message has been recorded. The server automatically downloads and saves the voice data, preparing it for the next step.
[0039] Step 2:
[0040] The server inputs the stored audio data into the speech recognition engine, which then converts the audio into text. The speech recognition engine analyzes the characteristics of the audio and generates accurate text data.
[0041] Step 3:
[0042] The server applies a summarization algorithm to the generated text data. The summarization algorithm extracts the important parts of the text and creates a compressed summary.
[0043] Step 4:
[0044] The server converts the summarized text into a short message format for transmission via Short Message Service (SMS). The server then securely sends the converted data to the user's device.
[0045] Step 5:
[0046] The user's device receives an SMS message and displays a notification. The user can then open the SMS application on their device to view the received message and easily understand the content recorded on their voicemail.
[0047] (Example 1)
[0048] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0049] There is a need to efficiently summarize voice messages recorded on answering machines, allowing users to quickly grasp important information without having to play the audio. In particular, a system is needed that can be easily implemented while ensuring security and accuracy throughout the process from acquiring voice data to summarizing and communicating it.
[0050] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0051] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, means for summarizing the text information, and means for using a secure communication protocol for the communication. This enables the user to instantly understand important information without directly listening to the audio.
[0052] "Audio information" refers to analog or digital audio recordings expressed in audio data or audio file formats.
[0053] "Means of acquisition" refers to mechanisms and methods for collecting and acquiring audio information in cooperation with external devices or systems.
[0054] "Textual information" refers to text data converted from audio information, and is linguistic information expressed in a readable form.
[0055] "Means of conversion" refers to algorithms and technologies used to analyze audio information and replace it with text information.
[0056] "Methods of summarization" refer to methods of extracting the main points from the original text and expressing them in the fewest possible words.
[0057] "Means of communication in short message format" refers to methods for sending summarized textual information as a concise message.
[0058] A "secure communication protocol" refers to a standardized communication method designed to protect and ensure the security of communication content.
[0059] To implement this invention, a server plays a central role in configuring a system that integrates and performs the processes of acquiring, converting, summarizing, and communicating voice information.
[0060] First, the server acquires audio information through an interface with the user's audio recording system. Specific hardware options for this include digital recording devices and network-connected storage. This audio information is stored in common digital audio formats such as WAV and MP3.
[0061] Next, the server uses speech recognition software to convert the acquired audio information into text information. Specific examples of such software include general-purpose speech recognition APIs (e.g., Google® Speech-to-Text API). This converts the content of the audio data into text format.
[0062] The converted text information is processed by a server using a summarization algorithm. This process uses a generative AI model (e.g., a BERT-based model) to extract the essential information from the text and shorten it. This summary eliminates any excess or omission of information, efficiently conveying the content. For example, the audio information "Tomorrow's meeting has been changed from 10:00 to 11:00" is summarized as "Meeting changed from 10:00 to 11:00".
[0063] Finally, the server transmits the summarized text information to the user's device in short message format. During this process, the information is securely transmitted using a secure communication protocol (e.g., TLS). The user can then view the received information through a messaging app on their device.
[0064] An example of a prompt message is, "Please summarize the voice message recorded on the answering machine and send it as a text message. Example voice message: 'Tomorrow's meeting has been changed from 10:00 to 11:00.'" This message is entered into the generation AI model, and the summary is performed. This system allows users to quickly grasp the main points without having to play the audio directly.
[0065] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0066] Step 1:
[0067] The server interacts with the audio recording system to acquire audio information. It receives digital audio files (e.g., WAV or MP3 format) from the audio recording system as input. The server uses APIs and network protocols to securely store these audio files. The output is an audio data file.
[0068] Step 2:
[0069] The server converts the audio information into text information. Using the audio data acquired in Step 1 as input, the audio data is analyzed using speech recognition software (such as the Google Speech-to-Text API). This results in the spoken words being output as text. Specifically, the waveform data is converted into a string by a language model and saved in text format.
[0070] Step 3:
[0071] The server uses a generative AI model to summarize text information. Using the text data generated in Step 2 as input, a BERT-based summarization algorithm is used to extract the most important points and shorten the text. Specifically, it selects highly relevant words and phrases from the long text and reconstructs them into a concise form. The output is a summarized short text.
[0072] Step 4:
[0073] The server sends a summarized short text to the terminal. The summarized text created in step 3 is used as input and sent to the user's terminal using a secure communication protocol. Specifically, the short message is sent via an SMS gateway using TLS encryption. The output is a short message displayed on the terminal.
[0074] Step 5:
[0075] The user checks the short message received on their device. They open and read the summary information in their device's messaging app. This allows the user to understand the important content without directly listening to the audio. The text displayed on the device screen is the output.
[0076] (Application Example 1)
[0077] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0078] In environments where numerous voice alerts and warning messages are generated, there is a need for information transmission methods that allow for quick and accurate understanding and response to them. However, voice information is often difficult to use directly, and there is a challenge in that it is difficult to confirm voice messages, especially when hearing or communication conditions are limited.
[0079] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0080] In this invention, the server includes means for acquiring audio information, means for converting the audio information into document information, and means for summarizing the document information. This enables more efficient management of audio information and allows for quick acquisition of important information even in situations with hearing limitations or in emergency situations.
[0081] "Audio information" refers to the information that makes up audio data, including signals and digital data generated from audio.
[0082] "Document information" refers to text data converted from speech, in a format that can be visually displayed.
[0083] "Summarization means" refers to a technology that extracts important information from acquired document information and processes it to express it in a shortened form.
[0084] A "short text" refers to a piece of writing that expresses summarized information in a concise text format, designed to quickly convey information.
[0085] "Visual devices" are devices used by users to visually confirm information, and include displays and smart glasses.
[0086] An "information management system" refers to a system that comprises a series of processes, from acquiring audio information to generating document information, summarizing it, and transmitting it.
[0087] This system is an information management system that efficiently manages audio information, summarizes it as visually verifiable document information, and displays it in short sentences on a visual device. An embodiment of the system is shown below.
[0088] The server acquires audio information from security devices and alarm systems within the building using means to acquire audio information. The audio information is processed as digital audio data. The server converts this audio information into document information using speech recognition software such as the Google Cloud Speech-to-Text API. The document information is text data accurately transcribed from the audio. Subsequently, the server summarizes the document information using generative AI technology such as OpenAI's GPT-3 model. The summarization means extracts important information from the document information and converts it into concise short sentences.
[0089] The device receives summarized short messages and displays them on a visual device. For example, by displaying the short messages on the screen of a smartphone or smart glasses, users can quickly understand the situation without directly listening to audio information. This allows for faster responses in emergencies.
[0090] As a concrete example, consider a scenario where a fire alarm is activated inside a shopping mall building. This voice warning message is sent to a server and converted into document information as "Fire alarm activated. Location: 3rd floor. Please check immediately." Then, a generation AI model summarizes it as "Fire alarm activated on the 3rd floor. Check required," and sends it to a mobile device or visual device. An example of a prompt would be, "Please summarize the following emergency voice alert. Original text: Fire alarm activated. Location: 3rd floor. Please check immediately."
[0091] This system enables users to quickly and accurately obtain important information and take appropriate action.
[0092] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0093] Step 1:
[0094] The server acquires audio information in real time from the audio sensor. The input is audio data from microphones within the building. This data is securely transferred to the server as digital data.
[0095] Step 2:
[0096] The server uses the Google Cloud Speech-to-Text API to convert acquired audio data into document information. The input is audio data, and the output is text data as document information. By using speech recognition technology, a string of characters that accurately transcribes spoken language is generated.
[0097] Step 3:
[0098] The server uses a generative AI model to summarize document information. The input is text data, and the output is a summarized short sentence. At this stage, the AI extracts and concisely summarizes the important information. The prompt used is the command, "Summarize the following emergency audio alert."
[0099] Step 4:
[0100] The server sends summarized short texts to the terminal using a secure communication protocol. The input is summarized text data, and the output is transmitted to a visual device. At this stage, the communication is encrypted to prevent information leakage.
[0101] Step 5:
[0102] The terminal displays received short messages on a visual device. The input is a short message sent from the server, and the output is a visualization on the display. The user can instantly check the summary information on the display and take the necessary action.
[0103] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0104] This system aims to provide more personalized information by quickly summarizing voice messages recorded on answering machines and combining the summary results with the user's emotion recognition. The system mainly consists of a server that acquires voice data, a server that converts the voice to text and summarizes it, a server equipped with an emotion engine, and a server that transmits the information to the user's terminal.
[0105] The server first interacts with the answering machine system to detect voice messages and simultaneously acquire digital audio data. The acquired audio data is then converted into text data using speech recognition technology. A summarization algorithm extracts the most important information from the converted text data and generates a shortened summary.
[0106] The emotion engine analyzes this text data to recognize the user's emotional state. This emotion analysis identifies the user's stress levels and emotional tendencies, allowing for the prioritization and adjustment of message content accordingly. For example, if an emotion such as "the meeting has been postponed and stress levels are high" is detected, the server will follow up by adding a detailed explanation to the user in the notification.
[0107] Finally, the server generates a short message containing a summary and sentiment analysis results, and securely sends it to the user's device. The message is received instantly on the device, allowing the user to gain a deeper understanding of the information tailored to their situation. This makes it possible to receive personalized information even in situations where audio data cannot be viewed or when there are hearing limitations.
[0108] The following describes the processing flow.
[0109] Step 1:
[0110] The server interacts with the answering machine system to detect when voice messages have been recorded. Based on this information, the server automatically retrieves and saves the voice data. This process is carried out using secure communication to prevent the loss of voice data.
[0111] Step 2:
[0112] The server passes the acquired audio data to the speech recognition engine, which converts the audio into text data. This engine analyzes the audio and generates digital data in text format. During this process, it corrects for the speaker's accent and noise to create accurate text.
[0113] Step 3:
[0114] The server inputs the generated text data into the summarization module. The summarization module uses natural language processing techniques to extract only the key points from the text and reconstruct them as a concise summary. This process employs multiple analysis algorithms to prevent information overload or omission.
[0115] Step 4:
[0116] The server analyzes the summarized text using an emotion engine to recognize the user's emotions. The emotion engine analyzes the words and context in the text in detail to determine whether the emotion is positive or negative, or whether the user is experiencing increased stress.
[0117] Step 5:
[0118] The server prioritizes messages based on sentiment analysis results. Furthermore, it adjusts message content as needed to provide information best suited to the user's emotional state. For example, if negative emotions are detected, additional cautionary wording or follow-up information may be added.
[0119] Step 6:
[0120] The server generates a short message incorporating the summary text and sentiment analysis results, and sends it to the user's device. The user's device immediately receives the message and displays a notification.
[0121] Step 7:
[0122] The user checks the short message on their device. Through the received message, they can quickly grasp the content of the voicemail and information relevant to their mood. In this way, the user can understand the important points without having to refer to detailed information.
[0123] (Example 2)
[0124] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0125] Voice messages routinely contain a wealth of information, and it is essential to quickly grasp the most important content. However, manually reviewing and summarizing voice messages is time-consuming and laborious, and providing personalized responses, especially when including emotional information, is difficult. Furthermore, there is a need for efficient ways to acquire information in situations where hearing is impaired or in environments where voice confirmation is not possible.
[0126] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0127] In this invention, the server includes means for acquiring voice information, means for converting the voice information into document information, means for summarizing the document information, means for recognizing the emotional state and adjusting the message content, and means for sending the adjusted message to the terminal. This enables the rapid extraction of important information from voice data and the provision of personalized information according to the user's emotional state.
[0128] "Audio information" refers to a representation of sound waveforms as digital data, and specifically, recordings of human speech.
[0129] "Document information" refers to audio information converted into text data, which is text data in a format that humans can read.
[0130] "Summarization" is the process of extracting key points from document information and generating a shortened version of the text.
[0131] "Emotional state" refers to data obtained by analyzing document information that indicates the emotions and psychological tendencies of the information provider.
[0132] A "message" is the information ultimately sent to the terminal, and includes summarized document information and tailored additional information.
[0133] A "terminal" is an electronic device that ultimately receives messages and is used by the user to verify information.
[0134] To implement this invention, a server, speech recognition software, a generative AI model, an emotion analysis engine, a message delivery system, and a user terminal are used. The following describes each component and its operation.
[0135] The server first acquires audio information. Specifically, it works in conjunction with the recording device system to acquire voice messages as digital data as soon as they are recorded. This process is carried out by retrieving data directly from the voice messaging system using an API.
[0136] Subsequently, the server converts the audio information into document information using speech recognition software. This "speech recognition software" utilizes "speech recognition technology" to convert speech into text. This technology analyzes the audio waveform and generates corresponding character data.
[0137] Next, the server summarizes the document information using a generative AI model. In this step, generative AI models such as "BERTSUM" and "GPT-3" are used to extract key points and generate a concise text. For the summary, a prompt such as "Summarize this text in 50 characters or less" is entered.
[0138] Subsequently, the server uses a sentiment analysis engine to analyze the summarized document information and recognize the emotional state. The sentiment analysis engine analyzes the sender's emotions based on the text information, understanding their stress level and emotional tendencies. This makes it possible to adjust the message content.
[0139] Ultimately, the server sends the reconciled message to the terminal via a message delivery system. This "message delivery system" could be something like "Twilio" or "Amazon SNS," ensuring the message is delivered quickly and securely to the user's terminal.
[0140] The device receives messages and is used by the user to verify the information. This allows users to instantly grasp personalized information in document form, even when they cannot access audio data.
[0141] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0142] Step 1:
[0143] The server works in conjunction with the recording device to detect a voice message and immediately acquires digital audio data. The input is the digital data recorded as a voice message. At this stage, the server receives the audio file from the recording device. The output is the digital audio data that is input to the speech recognition engine.
[0144] Step 2:
[0145] The server utilizes speech recognition technology to convert acquired digital audio data into document information. The input is the digital audio data obtained in step 1. The server analyzes the audio waveform and generates a string based on it. The output is the document information used in the summarization algorithm.
[0146] Step 3:
[0147] The server uses a generative AI model to summarize the transformed document information. The input is the document information provided in step 2. The server inputs prompt sentences into the generative AI model to obtain a concise summary with key points extracted. The output is the summary provided for sentiment analysis.
[0148] Step 4:
[0149] The server uses a sentiment analysis engine to analyze the summary text and recognize the emotional state. The input is the summary text obtained in step 3. The server analyzes the emotional expressions in the document and estimates the user's emotional state. The output is sentiment data used to adjust the message content.
[0150] Step 5:
[0151] The server adjusts the message content based on the determined emotional state and generates the final message. The input is the emotional data obtained in step 4 and the summary text from step 3. The server personalizes the message and adds relevant information. The output is the final message ready to be sent to the user's terminal.
[0152] Step 6:
[0153] The server sends the coordinated message to the user's terminal via the message delivery system. The input is the final message generated in step 5. The message delivery system transfers the data through a secure communication channel. The output is the message received at the user's terminal.
[0154] Step 7:
[0155] The user's device receives and displays the message. The input is the final message sent in step 6. The device provides information to the user through its notification function, allowing them to review personalized content. The output is the displayed message, which the user can understand.
[0156] (Application Example 2)
[0157] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0158] The challenge lies in reducing the information overload and emotional burden users face when reviewing voice messages, and in achieving efficient and personalized information delivery. In particular, considering an individual's emotional state is necessary to provide more appropriate notifications and suggestions, thereby improving the user experience.
[0159] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0160] In this invention, the server includes means for acquiring voice information, means for converting the voice information into text information, means for summarizing the text information, means for analyzing the summarized text information using an emotion recognition engine, and means for generating and sending personalized notifications based on the user's emotional state. This makes it possible to generate personalized notifications and suggestions that take the user's emotions into consideration.
[0161] "Voice information" refers to digital data obtained from voice, including the content of messages and voice communications recorded by users.
[0162] "Textual information" refers to text data obtained by converting audio information, and is information written in a format that is readable by humans.
[0163] An "emotion recognition engine" refers to an algorithm or program that analyzes textual information to identify a user's emotional state.
[0164] "Emotional state" refers to a psychological or emotional state determined based on the user's voice or text information, and is expressed in categories such as stress, joy, and dissatisfaction.
[0165] "Personalized notifications" refer to information or recommendations that are customized and sent based on the emotional state or individual user characteristics.
[0166] A system for carrying out this invention includes a process of acquiring audio information from a user and converting it into text information. The server converts the audio data into text data using a speech recognition engine (e.g., Google Speech-to-Text API). The converted text information is summarized by a natural language processing library (e.g., NLTK, spaCy) to identify important content. The server analyzes the summarized text information using an emotion recognition engine (e.g., emotion analysis based on the BERT model) to determine the user's emotional state.
[0167] Based on the user's emotional state, the server generates personalized notifications and sends the information to the user's device. Secure SMS messages are sent using APIs such as the Twilio API.
[0168] For example, if a user leaves a voice message regarding a product defect, and sentiment analysis determines that the user is dissatisfied, the server will generate and send a message to the user containing a discount coupon and details on the exchange / refund process.
[0169] Examples of prompt statements to be input to a generative AI model include the following:
[0170] "The item I ordered yesterday still hasn't arrived. I'm very worried. Please check on it as soon as possible."
[0171] Emotions: Dissatisfaction, stress
[0172] Action: Coupon offer and delivery status confirmation message
[0173] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0174] Step 1:
[0175] The server receives audio information from the user's terminal. The input is a recorded audio file. The server receives this audio file and prepares for the next processing step.
[0176] Step 2:
[0177] The server converts audio information into text information. The input is the audio file received in step 1, and the output is text data. The server uses the Google Speech-to-Text API to perform speech recognition and converts the data obtained from the audio file into text information.
[0178] Step 3:
[0179] The server summarizes the converted character information. The input is the character data generated in step 2, and the output is the summarized text. The server uses a natural language processing library (NLTK or spaCy) to analyze the character information, extract important content, and summarize it.
[0180] Step 4:
[0181] The server analyzes the summarized text information using an emotion recognition engine. The input is the summarized text generated in step 3, and the output is the user's emotional state. The server executes an emotion analysis algorithm based on the BERT model to determine the user's emotional state from the text information.
[0182] Step 5:
[0183] The server generates personalized notifications based on the emotional state. The input is the emotional state information determined in step 4, and the output is the personalized message sent to the user. The server automatically generates appropriate messages and suggestions based on the emotional state and prepares them for sending.
[0184] Step 6:
[0185] The server sends the generated notification to the user's device. The input is the personalized message created in step 5, and the output is the notification received on the user's device. The server uses the Twilio API to securely send the SMS, ensuring the user receives the necessary information immediately.
[0186] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0187] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0188] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0189] [Second Embodiment]
[0190] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0191] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0192] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0193] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0194] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0195] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0196] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0197] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0198] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0199] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0200] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0201] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0202] This system aims to quickly and easily summarize voice messages recorded on answering machines and send them to users as short messages. The main components of the system include a server that acquires voice data, a server that converts the voice to text and summarizes it, and a server that sends the summarized information to the user's terminal.
[0203] The server first connects with the answering machine system and immediately retrieves the digital audio data as soon as a voice message is recorded. This data is securely transferred with the user's permission. Next, the server uses speech recognition technology to convert the audio data into text. This conversion utilizes advanced algorithms that analyze the context of spoken language to generate accurate text.
[0204] Next, the server processes the generated text using a summarization algorithm to extract the most important information. This summarization allows for the concise transmission of the essence of the information without compromising the overall meaning of the message. For example, if a voice message is left on the answering machine saying, "Tomorrow's meeting has been changed from 10:00 to 11:00," the server will summarize this as, "Meeting changed from 10:00 to 11:00."
[0205] Finally, the server sends the summarized text to the user's terminal in the form of a short message. A secure messaging protocol is used for the transmission process to prevent information leakage. The user can check the message on their terminal and understand the content without listening to the voicemail message. This allows users who cannot access audio data or who have hearing impairments to efficiently grasp the information.
[0206] The following describes the processing flow.
[0207] Step 1:
[0208] The server connects to the answering machine system and detects that a voice message has been recorded. The server automatically downloads and saves the voice data, preparing it for the next step.
[0209] Step 2:
[0210] The server inputs the stored audio data into the speech recognition engine, which then converts the audio into text. The speech recognition engine analyzes the characteristics of the audio and generates accurate text data.
[0211] Step 3:
[0212] The server applies a summarization algorithm to the generated text data. The summarization algorithm extracts the important parts of the text and creates a compressed summary.
[0213] Step 4:
[0214] The server converts the summarized text into a short message format for transmission via Short Message Service (SMS). The server then securely sends the converted data to the user's device.
[0215] Step 5:
[0216] The user's device receives an SMS message and displays a notification. The user can then open the SMS application on their device to view the received message and easily understand the content recorded on their voicemail.
[0217] (Example 1)
[0218] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0219] There is a need to efficiently summarize voice messages recorded on answering machines, allowing users to quickly grasp important information without having to play the audio. In particular, a system is needed that can be easily implemented while ensuring security and accuracy throughout the process from acquiring voice data to summarizing and communicating it.
[0220] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0221] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, means for summarizing the text information, and means for using a secure communication protocol for the communication. This enables the user to instantly understand important information without directly listening to the audio.
[0222] "Audio information" refers to analog or digital audio recordings expressed in audio data or audio file formats.
[0223] "Means of acquisition" refers to mechanisms and methods for collecting and acquiring audio information in cooperation with external devices or systems.
[0224] "Textual information" refers to text data converted from audio information, and is linguistic information expressed in a readable form.
[0225] "Means of conversion" refers to algorithms and technologies used to analyze audio information and replace it with text information.
[0226] "Methods of summarization" refer to methods of extracting the main points from the original text and expressing them in the fewest possible words.
[0227] "Means of communication in short message format" refers to methods for sending summarized textual information as a concise message.
[0228] A "secure communication protocol" refers to a standardized communication method designed to protect and ensure the security of communication content.
[0229] To implement this invention, a server plays a central role in configuring a system that integrates and performs the processes of acquiring, converting, summarizing, and communicating voice information.
[0230] First, the server acquires audio information through an interface with the user's audio recording system. Specific hardware options for this include digital recording devices and network-connected storage. This audio information is stored in common digital audio formats such as WAV and MP3.
[0231] Next, the server uses speech recognition software to convert the acquired audio information into text information. Specific examples of such software include general-purpose speech recognition APIs (e.g., Google Speech-to-Text API). This converts the content of the audio data into text format.
[0232] The converted text information is processed by a server using a summarization algorithm. This process uses a generative AI model (e.g., a BERT-based model) to extract the essential information from the text and shorten it. This summary eliminates any excess or omission of information, efficiently conveying the content. For example, the audio information "Tomorrow's meeting has been changed from 10:00 to 11:00" is summarized as "Meeting changed from 10:00 to 11:00".
[0233] Finally, the server transmits the summarized text information to the user's device in short message format. During this process, the information is securely transmitted using a secure communication protocol (e.g., TLS). The user can then view the received information through a messaging app on their device.
[0234] An example of a prompt message is, "Please summarize the voice message recorded on the answering machine and send it as a text message. Example voice message: 'Tomorrow's meeting has been changed from 10:00 to 11:00.'" This message is entered into the generation AI model, and the summary is performed. This system allows users to quickly grasp the main points without having to play the audio directly.
[0235] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0236] Step 1:
[0237] The server interacts with the audio recording system to acquire audio information. It receives digital audio files (e.g., WAV or MP3 format) from the audio recording system as input. The server uses APIs and network protocols to securely store these audio files. The output is an audio data file.
[0238] Step 2:
[0239] The server converts the audio information into text information. Using the audio data acquired in Step 1 as input, the audio data is analyzed using speech recognition software (such as the Google Speech-to-Text API). This results in the spoken words being output as text. Specifically, the waveform data is converted into a string by a language model and saved in text format.
[0240] Step 3:
[0241] The server uses a generative AI model to summarize text information. Using the text data generated in Step 2 as input, a BERT-based summarization algorithm is used to extract the most important points and shorten the text. Specifically, it selects highly relevant words and phrases from the long text and reconstructs them into a concise form. The output is a summarized short text.
[0242] Step 4:
[0243] The server sends a summarized short text to the terminal. The summarized text created in step 3 is used as input and sent to the user's terminal using a secure communication protocol. Specifically, the short message is sent via an SMS gateway using TLS encryption. The output is a short message displayed on the terminal.
[0244] Step 5:
[0245] The user checks the short message received on their device. They open and read the summary information in their device's messaging app. This allows the user to understand the important content without directly listening to the audio. The text displayed on the device screen is the output.
[0246] (Application Example 1)
[0247] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0248] In environments where numerous voice alerts and warning messages are generated, there is a need for information transmission methods that allow for quick and accurate understanding and response to them. However, voice information is often difficult to use directly, and there is a challenge in that it is difficult to confirm voice messages, especially when hearing or communication conditions are limited.
[0249] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0250] In this invention, the server includes means for acquiring audio information, means for converting the audio information into document information, and means for summarizing the document information. This enables more efficient management of audio information and allows for quick acquisition of important information even in situations with hearing limitations or in emergency situations.
[0251] "Audio information" refers to the information that makes up audio data, including signals and digital data generated from audio.
[0252] "Document information" refers to text data converted from speech, in a format that can be visually displayed.
[0253] "Summarization means" refers to a technology that extracts important information from acquired document information and processes it to express it in a shortened form.
[0254] A "short text" refers to a piece of writing that expresses summarized information in a concise text format, designed to quickly convey information.
[0255] "Visual devices" are devices used by users to visually confirm information, and include displays and smart glasses.
[0256] An "information management system" refers to a system that comprises a series of processes, from acquiring audio information to generating document information, summarizing it, and transmitting it.
[0257] This system is an information management system that efficiently manages audio information, summarizes it as visually verifiable document information, and displays it in short sentences on a visual device. An embodiment of the system is shown below.
[0258] The server acquires audio information from security devices and alarm systems within the building using means to acquire audio information. The audio information is processed as digital audio data. The server converts this audio information into document information using speech recognition software such as the Google Cloud Speech-to-Text API. The document information is text data accurately transcribed from the audio. Next, the server summarizes the document information using generative AI technology such as OpenAI's GPT-3 model. The summarization means extracts important information from the document information and converts it into concise short sentences.
[0259] The device receives summarized short messages and displays them on a visual device. For example, by displaying the short messages on the screen of a smartphone or smart glasses, users can quickly understand the situation without directly listening to audio information. This allows for faster responses in emergencies.
[0260] As a concrete example, consider a scenario where a fire alarm is activated inside a shopping mall building. This voice warning message is sent to a server and converted into document information as "Fire alarm activated. Location: 3rd floor. Please check immediately." Then, a generation AI model summarizes it as "Fire alarm activated on the 3rd floor. Check required," and sends it to a mobile device or visual device. An example of a prompt would be, "Please summarize the following emergency voice alert. Original text: Fire alarm activated. Location: 3rd floor. Please check immediately."
[0261] This system enables users to quickly and accurately obtain important information and take appropriate action.
[0262] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0263] Step 1:
[0264] The server acquires audio information in real time from the audio sensor. The input is audio data from microphones within the building. This data is securely transferred to the server as digital data.
[0265] Step 2:
[0266] The server uses the Google Cloud Speech-to-Text API to convert acquired audio data into document information. The input is audio data, and the output is text data as document information. By using speech recognition technology, a string of characters that accurately transcribes spoken language is generated.
[0267] Step 3:
[0268] The server uses a generative AI model to summarize document information. The input is text data, and the output is a summarized short sentence. At this stage, the AI extracts and concisely summarizes the important information. The prompt used is the command, "Summarize the following emergency audio alert."
[0269] Step 4:
[0270] The server sends summarized short texts to the terminal using a secure communication protocol. The input is summarized text data, and the output is transmitted to a visual device. At this stage, the communication is encrypted to prevent information leakage.
[0271] Step 5:
[0272] The terminal displays received short messages on a visual device. The input is a short message sent from the server, and the output is a visualization on the display. The user can instantly check the summary information on the display and take the necessary action.
[0273] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0274] This system aims to provide more personalized information by quickly summarizing voice messages recorded on answering machines and combining the summary results with the user's emotion recognition. The system mainly consists of a server that acquires voice data, a server that converts the voice to text and summarizes it, a server equipped with an emotion engine, and a server that transmits the information to the user's terminal.
[0275] The server first interacts with the answering machine system to detect voice messages and simultaneously acquire digital audio data. The acquired audio data is then converted into text data using speech recognition technology. A summarization algorithm extracts the most important information from the converted text data and generates a shortened summary.
[0276] The emotion engine analyzes this text data to recognize the user's emotional state. This emotion analysis identifies the user's stress levels and emotional tendencies, allowing for the prioritization and adjustment of message content accordingly. For example, if an emotion such as "the meeting has been postponed and stress levels are high" is detected, the server will follow up by adding a detailed explanation to the user in the notification.
[0277] Finally, the server generates a short message containing a summary and sentiment analysis results, and securely sends it to the user's device. The message is received instantly on the device, allowing the user to gain a deeper understanding of the information tailored to their situation. This makes it possible to receive personalized information even in situations where audio data cannot be viewed or when there are hearing limitations.
[0278] The following describes the processing flow.
[0279] Step 1:
[0280] The server interacts with the answering machine system to detect when voice messages have been recorded. Based on this information, the server automatically retrieves and saves the voice data. This process is carried out using secure communication to prevent the loss of voice data.
[0281] Step 2:
[0282] The server passes the acquired audio data to the speech recognition engine, which converts the audio into text data. This engine analyzes the audio and generates digital data in text format. During this process, it corrects for the speaker's accent and noise to create accurate text.
[0283] Step 3:
[0284] The server inputs the generated text data into the summarization module. The summarization module uses natural language processing technology to extract only the key points from the text and reconstruct them as a concise summary. This process utilizes multiple analysis algorithms to prevent information overload or shortage.
[0285] Step 4:
[0286] The server analyzes the summarized text with the sentiment engine to recognize the user's sentiment. The sentiment engine analyzes the words and context contained in the text in detail to determine whether it is positive, negative, or whether stress is increasing, etc.
[0287] Step 5:
[0288] Based on the results of the sentiment analysis, the server sets the priority of the message. Furthermore, the content of the message is adjusted as necessary to provide information most suitable for the user's emotional state. For example, when the sentiment is recognized as negative, words to prompt further attention or follow-up information may be added.
[0289] Step 6:
[0290] The server generates a short message incorporating the summarized text and the results of the sentiment analysis and sends it to the user's terminal. The user's terminal immediately receives the sent message and displays a notification.
[0291] Step 7:
[0292] The user checks the short message on the terminal. Through the received message, the user can quickly grasp the content of the voicemail and information suitable for their own sentiment. In this way, the user can understand the important points without referring to detailed information.
[0293] (Example 2)
[0294] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0295] Voice messages routinely contain a wealth of information, and it is essential to quickly grasp the most important content. However, manually reviewing and summarizing voice messages is time-consuming and laborious, and providing personalized responses, especially when including emotional information, is difficult. Furthermore, there is a need for efficient ways to acquire information in situations where hearing is impaired or in environments where voice confirmation is not possible.
[0296] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0297] In this invention, the server includes means for acquiring voice information, means for converting the voice information into document information, means for summarizing the document information, means for recognizing the emotional state and adjusting the message content, and means for sending the adjusted message to the terminal. This enables the rapid extraction of important information from voice data and the provision of personalized information according to the user's emotional state.
[0298] "Audio information" refers to a representation of sound waveforms as digital data, and specifically, recordings of human speech.
[0299] "Document information" refers to audio information converted into text data, which is text data in a format that humans can read.
[0300] "Summarization" is the process of extracting key points from document information and generating a shortened version of the text.
[0301] "Emotional state" refers to data obtained by analyzing document information that indicates the emotions and psychological tendencies of the information provider.
[0302] A "message" is information that is ultimately sent to a terminal and includes summarized document information and adjusted additional information.
[0303] A "terminal" is an electronic device that ultimately receives a message and is used by a user to view information.
[0304] To implement this invention, a server, speech recognition software, a generation AI model, a sentiment analysis engine, a message delivery system, and a user's terminal are used. Each component and its operation will be described below.
[0305] First, the server obtains voice information. Specifically, in cooperation with a recording device system, it immediately obtains this as digital data as soon as a voice message is recorded. This process is carried out by directly extracting data from the voice message system using an API.
[0306] After that, the server converts the voice information into document information using speech recognition software. As "speech recognition software", "speech recognition technology" is used to convert voice into text. This technology analyzes the voice waveform and generates corresponding character data.
[0307] Next, the server summarizes the document information using a generation AI model. In this step, "BERTSUM", "GPT-3", etc. are utilized as the "generation AI model" to extract important points and generate a concise text. Also, for the summary, a prompt sentence such as "Please summarize this text within 50 characters" is input.
[0308] After that, the server analyzes the summarized document information using a sentiment analysis engine to recognize the emotional state. The sentiment analysis engine analyzes the sender's emotion based on the text information to grasp the stress level and emotional tendency. This makes it possible to adjust the message content.
[0309] Ultimately, the server sends the reconciled message to the terminal via a message delivery system. This "message delivery system" could be something like "Twilio" or "Amazon SNS," ensuring the message is delivered quickly and securely to the user's terminal.
[0310] The device receives messages and is used by the user to verify the information. This allows users to instantly grasp personalized information in document form, even when they cannot access audio data.
[0311] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0312] Step 1:
[0313] The server works in conjunction with the recording device to detect a voice message and immediately acquires digital audio data. The input is the digital data recorded as a voice message. At this stage, the server receives the audio file from the recording device. The output is the digital audio data that is input to the speech recognition engine.
[0314] Step 2:
[0315] The server utilizes speech recognition technology to convert acquired digital audio data into document information. The input is the digital audio data obtained in step 1. The server analyzes the audio waveform and generates a string based on it. The output is the document information used in the summarization algorithm.
[0316] Step 3:
[0317] The server uses a generative AI model to summarize the transformed document information. The input is the document information provided in step 2. The server inputs prompt sentences into the generative AI model to obtain a concise summary with key points extracted. The output is the summary provided for sentiment analysis.
[0318] Step 4:
[0319] The server uses a sentiment analysis engine to analyze the summary text and recognize the emotional state. The input is the summary text obtained in step 3. The server analyzes the emotional expressions in the document and estimates the user's emotional state. The output is sentiment data used to adjust the message content.
[0320] Step 5:
[0321] The server adjusts the message content based on the determined emotional state and generates the final message. The input is the emotional data obtained in step 4 and the summary text from step 3. The server personalizes the message and adds relevant information. The output is the final message ready to be sent to the user's terminal.
[0322] Step 6:
[0323] The server sends the coordinated message to the user's terminal via the message delivery system. The input is the final message generated in step 5. The message delivery system transfers the data through a secure communication channel. The output is the message received at the user's terminal.
[0324] Step 7:
[0325] The user's device receives and displays the message. The input is the final message sent in step 6. The device provides information to the user through its notification function, allowing them to review personalized content. The output is the displayed message, which the user can understand.
[0326] (Application Example 2)
[0327] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0328] The challenge lies in reducing the information overload and emotional burden users face when reviewing voice messages, and in achieving efficient and personalized information delivery. In particular, considering an individual's emotional state is necessary to provide more appropriate notifications and suggestions, thereby improving the user experience.
[0329] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0330] In this invention, the server includes means for acquiring voice information, means for converting the voice information into text information, means for summarizing the text information, means for analyzing the summarized text information using an emotion recognition engine, and means for generating and sending personalized notifications based on the user's emotional state. This makes it possible to generate personalized notifications and suggestions that take the user's emotions into consideration.
[0331] "Voice information" refers to digital data obtained from voice, including the content of messages and voice communications recorded by users.
[0332] "Textual information" refers to text data obtained by converting audio information, and is information written in a format that is readable by humans.
[0333] An "emotion recognition engine" refers to an algorithm or program that analyzes textual information to identify a user's emotional state.
[0334] "Emotional state" refers to a psychological or emotional state determined based on the user's voice or text information, and is expressed in categories such as stress, joy, and dissatisfaction.
[0335] "Personalized notifications" refer to information or recommendations that are customized and sent based on the emotional state or individual user characteristics.
[0336] A system for carrying out this invention includes a process of acquiring audio information from a user and converting it into text information. The server converts the audio data into text data using a speech recognition engine (e.g., Google Speech-to-Text API). The converted text information is summarized by a natural language processing library (e.g., NLTK, spaCy) to identify important content. The server analyzes the summarized text information using an emotion recognition engine (e.g., emotion analysis based on the BERT model) to determine the user's emotional state.
[0337] Based on the user's emotional state, the server generates personalized notifications and sends the information to the user's device. Secure SMS messages are sent using APIs such as the Twilio API.
[0338] For example, if a user leaves a voice message regarding a product defect, and sentiment analysis determines that the user is dissatisfied, the server will generate and send a message to the user containing a discount coupon and details on the exchange / refund process.
[0339] Examples of prompt statements to be input to a generative AI model include the following:
[0340] "The item I ordered yesterday still hasn't arrived. I'm very worried. Please check on it as soon as possible."
[0341] Emotions: Dissatisfaction, stress
[0342] Action: Coupon offer and delivery status confirmation message
[0343] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0344] Step 1:
[0345] The server receives audio information from the user's terminal. The input is a recorded audio file. The server receives this audio file and prepares for the next processing step.
[0346] Step 2:
[0347] The server converts audio information into text information. The input is the audio file received in step 1, and the output is text data. The server uses the Google Speech-to-Text API to perform speech recognition and converts the data obtained from the audio file into text information.
[0348] Step 3:
[0349] The server summarizes the converted character information. The input is the character data generated in step 2, and the output is the summarized text. The server uses a natural language processing library (NLTK or spaCy) to analyze the character information, extract important content, and summarize it.
[0350] Step 4:
[0351] The server analyzes the summarized text information using an emotion recognition engine. The input is the summarized text generated in step 3, and the output is the user's emotional state. The server executes an emotion analysis algorithm based on the BERT model to determine the user's emotional state from the text information.
[0352] Step 5:
[0353] The server generates personalized notifications based on the emotional state. The input is the emotional state information determined in step 4, and the output is the personalized message sent to the user. The server automatically generates appropriate messages and suggestions based on the emotional state and prepares them for sending.
[0354] Step 6:
[0355] The server sends the generated notification to the user's device. The input is the personalized message created in step 5, and the output is the notification received on the user's device. The server uses the Twilio API to securely send the SMS, ensuring the user receives the necessary information immediately.
[0356] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0357] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0358] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0359] [Third Embodiment]
[0360] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0361] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0362] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0363] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0364] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0365] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0366] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0367] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0368] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0369] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0370] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0371] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0372] This system aims to quickly and easily summarize voice messages recorded on answering machines and send them to users as short messages. The main components of the system include a server that acquires voice data, a server that converts the voice to text and summarizes it, and a server that sends the summarized information to the user's terminal.
[0373] The server first connects with the answering machine system and immediately retrieves the digital audio data as soon as a voice message is recorded. This data is securely transferred with the user's permission. Next, the server uses speech recognition technology to convert the audio data into text. This conversion utilizes advanced algorithms that analyze the context of spoken language to generate accurate text.
[0374] Next, the server processes the generated text using a summarization algorithm to extract the most important information. This summarization allows for the concise transmission of the essence of the information without compromising the overall meaning of the message. For example, if a voice message is left on the answering machine saying, "Tomorrow's meeting has been changed from 10:00 to 11:00," the server will summarize this as, "Meeting changed from 10:00 to 11:00."
[0375] Finally, the server sends the summarized text to the user's terminal in the form of a short message. A secure messaging protocol is used for the transmission process to prevent information leakage. The user can check the message on their terminal and understand the content without listening to the voicemail message. This allows users who cannot access audio data or who have hearing impairments to efficiently grasp the information.
[0376] The following describes the processing flow.
[0377] Step 1:
[0378] The server connects to the answering machine system and detects that a voice message has been recorded. The server automatically downloads and saves the voice data, preparing it for the next step.
[0379] Step 2:
[0380] The server inputs the stored audio data into the speech recognition engine, which then converts the audio into text. The speech recognition engine analyzes the characteristics of the audio and generates accurate text data.
[0381] Step 3:
[0382] The server applies a summarization algorithm to the generated text data. The summarization algorithm extracts the important parts of the text and creates a compressed summary.
[0383] Step 4:
[0384] The server converts the summarized text into a short message format for transmission via Short Message Service (SMS). The server then securely sends the converted data to the user's device.
[0385] Step 5:
[0386] The user's device receives an SMS message and displays a notification. The user can then open the SMS application on their device to view the received message and easily understand the content recorded on their voicemail.
[0387] (Example 1)
[0388] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0389] There is a need to efficiently summarize voice messages recorded on answering machines, allowing users to quickly grasp important information without having to play the audio. In particular, a system is needed that can be easily implemented while ensuring security and accuracy throughout the process from acquiring voice data to summarizing and communicating it.
[0390] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0391] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, means for summarizing the text information, and means for using a secure communication protocol for the communication. This enables the user to instantly understand important information without directly listening to the audio.
[0392] "Audio information" refers to analog or digital audio recordings expressed in audio data or audio file formats.
[0393] "Means of acquisition" refers to mechanisms and methods for collecting and acquiring audio information in cooperation with external devices or systems.
[0394] "Textual information" refers to text data converted from audio information, and is linguistic information expressed in a readable form.
[0395] "Means of conversion" refers to algorithms and technologies used to analyze audio information and replace it with text information.
[0396] "Methods of summarization" refer to methods of extracting the main points from the original text and expressing them in the fewest possible words.
[0397] "Means of communication in short message format" refers to methods for sending summarized textual information as a concise message.
[0398] A "secure communication protocol" refers to a standardized communication method designed to protect and ensure the security of communication content.
[0399] To implement this invention, a server plays a central role in configuring a system that integrates and performs the processes of acquiring, converting, summarizing, and communicating voice information.
[0400] First, the server acquires audio information through an interface with the user's audio recording system. Specific hardware options for this include digital recording devices and network-connected storage. This audio information is stored in common digital audio formats such as WAV and MP3.
[0401] Next, the server uses speech recognition software to convert the acquired audio information into text information. Specific examples of such software include general-purpose speech recognition APIs (e.g., Google Speech-to-Text API). This converts the content of the audio data into text format.
[0402] The converted text information is processed by a server using a summarization algorithm. This process uses a generative AI model (e.g., a BERT-based model) to extract the essential information from the text and shorten it. This summary eliminates any excess or omission of information, efficiently conveying the content. For example, the audio information "Tomorrow's meeting has been changed from 10:00 to 11:00" is summarized as "Meeting changed from 10:00 to 11:00".
[0403] Finally, the server transmits the summarized text information to the user's device in short message format. During this process, the information is securely transmitted using a secure communication protocol (e.g., TLS). The user can then view the received information through a messaging app on their device.
[0404] An example of a prompt message is, "Please summarize the voice message recorded on the answering machine and send it as a text message. Example voice message: 'Tomorrow's meeting has been changed from 10:00 to 11:00.'" This message is entered into the generation AI model, and the summary is performed. This system allows users to quickly grasp the main points without having to play the audio directly.
[0405] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0406] Step 1:
[0407] The server interacts with the audio recording system to acquire audio information. It receives digital audio files (e.g., WAV or MP3 format) from the audio recording system as input. The server uses APIs and network protocols to securely store these audio files. The output is an audio data file.
[0408] Step 2:
[0409] The server converts the audio information into text information. Using the audio data acquired in Step 1 as input, the audio data is analyzed using speech recognition software (such as the Google Speech-to-Text API). This results in the spoken words being output as text. Specifically, the waveform data is converted into a string by a language model and saved in text format.
[0410] Step 3:
[0411] The server uses a generative AI model to summarize text information. Using the text data generated in Step 2 as input, a BERT-based summarization algorithm is used to extract the most important points and shorten the text. Specifically, it selects highly relevant words and phrases from the long text and reconstructs them into a concise form. The output is a summarized short text.
[0412] Step 4:
[0413] The server sends a summarized short text to the terminal. The summarized text created in step 3 is used as input and sent to the user's terminal using a secure communication protocol. Specifically, the short message is sent via an SMS gateway using TLS encryption. The output is a short message displayed on the terminal.
[0414] Step 5:
[0415] The user checks the short message received on their device. They open and read the summary information in their device's messaging app. This allows the user to understand the important content without directly listening to the audio. The text displayed on the device screen is the output.
[0416] (Application Example 1)
[0417] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0418] In environments where numerous voice alerts and warning messages are generated, there is a need for information transmission methods that allow for quick and accurate understanding and response to them. However, voice information is often difficult to use directly, and there is a challenge in that it is difficult to confirm voice messages, especially when hearing or communication conditions are limited.
[0419] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0420] In this invention, the server includes means for acquiring audio information, means for converting the audio information into document information, and means for summarizing the document information. This enables more efficient management of audio information and allows for quick acquisition of important information even in situations with hearing limitations or in emergency situations.
[0421] "Audio information" refers to the information that makes up audio data, including signals and digital data generated from audio.
[0422] "Document information" refers to text data converted from speech, in a format that can be visually displayed.
[0423] "Summarization means" refers to a technology that extracts important information from acquired document information and processes it to express it in a shortened form.
[0424] A "short text" refers to a piece of writing that expresses summarized information in a concise text format, designed to quickly convey information.
[0425] "Visual devices" are devices used by users to visually confirm information, and include displays and smart glasses.
[0426] An "information management system" refers to a system that comprises a series of processes, from acquiring audio information to generating document information, summarizing it, and transmitting it.
[0427] This system is an information management system that efficiently manages audio information, summarizes it as visually verifiable document information, and displays it in short sentences on a visual device. An embodiment of the system is shown below.
[0428] The server acquires audio information from security devices and alarm systems within the building using means to acquire audio information. The audio information is processed as digital audio data. The server converts this audio information into document information using speech recognition software such as the Google Cloud Speech-to-Text API. The document information is text data accurately transcribed from the audio. Next, the server summarizes the document information using generative AI technology such as OpenAI's GPT-3 model. The summarization means extracts important information from the document information and converts it into concise short sentences.
[0429] The device receives summarized short messages and displays them on a visual device. For example, by displaying the short messages on the screen of a smartphone or smart glasses, users can quickly understand the situation without directly listening to audio information. This allows for faster responses in emergencies.
[0430] As a concrete example, consider a scenario where a fire alarm is activated inside a shopping mall building. This voice warning message is sent to a server and converted into document information as "Fire alarm activated. Location: 3rd floor. Please check immediately." Then, a generation AI model summarizes it as "Fire alarm activated on the 3rd floor. Check required," and sends it to a mobile device or visual device. An example of a prompt would be, "Please summarize the following emergency voice alert. Original text: Fire alarm activated. Location: 3rd floor. Please check immediately."
[0431] This system enables users to quickly and accurately obtain important information and take appropriate action.
[0432] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0433] Step 1:
[0434] The server acquires audio information in real time from the audio sensor. The input is audio data from microphones within the building. This data is securely transferred to the server as digital data.
[0435] Step 2:
[0436] The server uses the Google Cloud Speech-to-Text API to convert acquired audio data into document information. The input is audio data, and the output is text data as document information. By using speech recognition technology, a string of characters that accurately transcribes spoken language is generated.
[0437] Step 3:
[0438] The server uses a generative AI model to summarize document information. The input is text data, and the output is a summarized short sentence. At this stage, the AI extracts and concisely summarizes the important information. The prompt used is the command, "Summarize the following emergency audio alert."
[0439] Step 4:
[0440] The server sends summarized short texts to the terminal using a secure communication protocol. The input is summarized text data, and the output is transmitted to a visual device. At this stage, the communication is encrypted to prevent information leakage.
[0441] Step 5:
[0442] The terminal displays received short messages on a visual device. The input is a short message sent from the server, and the output is a visualization on the display. The user can instantly check the summary information on the display and take the necessary action.
[0443] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0444] This system aims to provide more personalized information by quickly summarizing voice messages recorded on answering machines and combining the summary results with the user's emotion recognition. The system mainly consists of a server that acquires voice data, a server that converts the voice to text and summarizes it, a server equipped with an emotion engine, and a server that transmits the information to the user's terminal.
[0445] The server first interacts with the answering machine system to detect voice messages and simultaneously acquire digital audio data. The acquired audio data is then converted into text data using speech recognition technology. A summarization algorithm extracts the most important information from the converted text data and generates a shortened summary.
[0446] The emotion engine analyzes this text data to recognize the user's emotional state. This emotion analysis identifies the user's stress levels and emotional tendencies, allowing for the prioritization and adjustment of message content accordingly. For example, if an emotion such as "the meeting has been postponed and stress levels are high" is detected, the server will follow up by adding a detailed explanation to the user in the notification.
[0447] Finally, the server generates a short message containing a summary and sentiment analysis results, and securely sends it to the user's device. The message is received instantly on the device, allowing the user to gain a deeper understanding of the information tailored to their situation. This makes it possible to receive personalized information even in situations where audio data cannot be viewed or when there are hearing limitations.
[0448] The following describes the processing flow.
[0449] Step 1:
[0450] The server interacts with the answering machine system to detect when voice messages have been recorded. Based on this information, the server automatically retrieves and saves the voice data. This process is carried out using secure communication to prevent the loss of voice data.
[0451] Step 2:
[0452] The server passes the acquired audio data to the speech recognition engine, which converts the audio into text data. This engine analyzes the audio and generates digital data in text format. During this process, it corrects for the speaker's accent and noise to create accurate text.
[0453] Step 3:
[0454] The server inputs the generated text data into the summarization module. The summarization module uses natural language processing techniques to extract only the key points from the text and reconstruct them as a concise summary. This process employs multiple analysis algorithms to prevent information overload or omission.
[0455] Step 4:
[0456] The server analyzes the summarized text using an emotion engine to recognize the user's emotions. The emotion engine analyzes the words and context in the text in detail to determine whether the emotion is positive or negative, or whether the user is experiencing increased stress.
[0457] Step 5:
[0458] The server prioritizes messages based on sentiment analysis results. Furthermore, it adjusts message content as needed to provide information best suited to the user's emotional state. For example, if negative emotions are detected, additional cautionary wording or follow-up information may be added.
[0459] Step 6:
[0460] The server generates a short message incorporating the summary text and sentiment analysis results, and sends it to the user's device. The user's device immediately receives the message and displays a notification.
[0461] Step 7:
[0462] The user checks the short message on their device. Through the received message, they can quickly grasp the content of the voicemail and information relevant to their mood. In this way, the user can understand the important points without having to refer to detailed information.
[0463] (Example 2)
[0464] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] Voice messages routinely contain a wealth of information, and it is essential to quickly grasp the most important content. However, manually reviewing and summarizing voice messages is time-consuming and laborious, and providing personalized responses, especially when including emotional information, is difficult. Furthermore, there is a need for efficient ways to acquire information in situations where hearing is impaired or in environments where voice confirmation is not possible.
[0466] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0467] In this invention, the server includes means for acquiring voice information, means for converting the voice information into document information, means for summarizing the document information, means for recognizing the emotional state and adjusting the message content, and means for sending the adjusted message to the terminal. This enables the rapid extraction of important information from voice data and the provision of personalized information according to the user's emotional state.
[0468] "Audio information" refers to a representation of sound waveforms as digital data, and specifically, recordings of human speech.
[0469] "Document information" refers to audio information converted into text data, which is text data in a format that humans can read.
[0470] "Summarization" is the process of extracting key points from document information and generating a shortened version of the text.
[0471] "Emotional state" refers to data obtained by analyzing document information that indicates the emotions and psychological tendencies of the information provider.
[0472] A "message" is the information ultimately sent to the terminal, and includes summarized document information and tailored additional information.
[0473] A "terminal" is an electronic device that ultimately receives messages and is used by the user to verify information.
[0474] To implement this invention, a server, speech recognition software, a generative AI model, an emotion analysis engine, a message delivery system, and a user terminal are used. The following describes each component and its operation.
[0475] The server first acquires audio information. Specifically, it works in conjunction with the recording device system to acquire voice messages as digital data as soon as they are recorded. This process is carried out by retrieving data directly from the voice messaging system using an API.
[0476] Subsequently, the server converts the audio information into document information using speech recognition software. This "speech recognition software" utilizes "speech recognition technology" to convert speech into text. This technology analyzes the audio waveform and generates corresponding character data.
[0477] Next, the server summarizes the document information using a generative AI model. In this step, generative AI models such as "BERTSUM" and "GPT-3" are used to extract key points and generate a concise text. For the summary, a prompt such as "Summarize this text in 50 characters or less" is entered.
[0478] Subsequently, the server uses a sentiment analysis engine to analyze the summarized document information and recognize the emotional state. The sentiment analysis engine analyzes the sender's emotions based on the text information, understanding their stress level and emotional tendencies. This makes it possible to adjust the message content.
[0479] Ultimately, the server sends the reconciled message to the terminal via a message delivery system. This "message delivery system" could be something like "Twilio" or "Amazon SNS," ensuring the message is delivered quickly and securely to the user's terminal.
[0480] The device receives messages and is used by the user to verify the information. This allows users to instantly grasp personalized information in document form, even when they cannot access audio data.
[0481] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0482] Step 1:
[0483] The server works in conjunction with the recording device to detect a voice message and immediately acquires digital audio data. The input is the digital data recorded as a voice message. At this stage, the server receives the audio file from the recording device. The output is the digital audio data that is input to the speech recognition engine.
[0484] Step 2:
[0485] The server utilizes speech recognition technology to convert acquired digital audio data into document information. The input is the digital audio data obtained in step 1. The server analyzes the audio waveform and generates a string based on it. The output is the document information used in the summarization algorithm.
[0486] Step 3:
[0487] The server uses a generative AI model to summarize the transformed document information. The input is the document information provided in step 2. The server inputs prompt sentences into the generative AI model to obtain a concise summary with key points extracted. The output is the summary provided for sentiment analysis.
[0488] Step 4:
[0489] The server uses a sentiment analysis engine to analyze the summary text and recognize the emotional state. The input is the summary text obtained in step 3. The server analyzes the emotional expressions in the document and estimates the user's emotional state. The output is sentiment data used to adjust the message content.
[0490] Step 5:
[0491] The server adjusts the message content based on the determined emotional state and generates the final message. The input is the emotional data obtained in step 4 and the summary text from step 3. The server personalizes the message and adds relevant information. The output is the final message ready to be sent to the user's terminal.
[0492] Step 6:
[0493] The server sends the coordinated message to the user's terminal via the message delivery system. The input is the final message generated in step 5. The message delivery system transfers the data through a secure communication channel. The output is the message received at the user's terminal.
[0494] Step 7:
[0495] The user's device receives and displays the message. The input is the final message sent in step 6. The device provides information to the user through its notification function, allowing them to review personalized content. The output is the displayed message, which the user can understand.
[0496] (Application Example 2)
[0497] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0498] The challenge lies in reducing the information overload and emotional burden users face when reviewing voice messages, and in achieving efficient and personalized information delivery. In particular, considering an individual's emotional state is necessary to provide more appropriate notifications and suggestions, thereby improving the user experience.
[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0500] In this invention, the server includes means for acquiring voice information, means for converting the voice information into text information, means for summarizing the text information, means for analyzing the summarized text information using an emotion recognition engine, and means for generating and sending personalized notifications based on the user's emotional state. This makes it possible to generate personalized notifications and suggestions that take the user's emotions into consideration.
[0501] "Voice information" refers to digital data obtained from voice, including the content of messages and voice communications recorded by users.
[0502] "Textual information" refers to text data obtained by converting audio information, and is information written in a format that is readable by humans.
[0503] An "emotion recognition engine" refers to an algorithm or program that analyzes textual information to identify a user's emotional state.
[0504] "Emotional state" refers to a psychological or emotional state determined based on the user's voice or text information, and is expressed in categories such as stress, joy, and dissatisfaction.
[0505] "Personalized notifications" refer to information or recommendations that are customized and sent based on the emotional state or individual user characteristics.
[0506] A system for carrying out this invention includes a process of acquiring audio information from a user and converting it into text information. The server converts the audio data into text data using a speech recognition engine (e.g., Google Speech-to-Text API). The converted text information is summarized by a natural language processing library (e.g., NLTK, spaCy) to identify important content. The server analyzes the summarized text information using an emotion recognition engine (e.g., emotion analysis based on the BERT model) to determine the user's emotional state.
[0507] Based on the user's emotional state, the server generates personalized notifications and sends the information to the user's device. Secure SMS messages are sent using APIs such as the Twilio API.
[0508] For example, if a user leaves a voice message regarding a product defect, and sentiment analysis determines that the user is dissatisfied, the server will generate and send a message to the user containing a discount coupon and details on the exchange / refund process.
[0509] Examples of prompt statements to be input to a generative AI model include the following:
[0510] "The item I ordered yesterday still hasn't arrived. I'm very worried. Please check on it as soon as possible."
[0511] Emotions: Dissatisfaction, stress
[0512] Action: Coupon offer and delivery status confirmation message
[0513] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0514] Step 1:
[0515] The server receives audio information from the user's terminal. The input is a recorded audio file. The server receives this audio file and prepares for the next processing step.
[0516] Step 2:
[0517] The server converts audio information into text information. The input is the audio file received in step 1, and the output is text data. The server uses the Google Speech-to-Text API to perform speech recognition and converts the data obtained from the audio file into text information.
[0518] Step 3:
[0519] The server summarizes the converted character information. The input is the character data generated in step 2, and the output is the summarized text. The server uses a natural language processing library (NLTK or spaCy) to analyze the character information, extract important content, and summarize it.
[0520] Step 4:
[0521] The server analyzes the summarized text information using an emotion recognition engine. The input is the summarized text generated in step 3, and the output is the user's emotional state. The server executes an emotion analysis algorithm based on the BERT model to determine the user's emotional state from the text information.
[0522] Step 5:
[0523] The server generates personalized notifications based on the emotional state. The input is the emotional state information determined in step 4, and the output is the personalized message sent to the user. The server automatically generates appropriate messages and suggestions based on the emotional state and prepares them for sending.
[0524] Step 6:
[0525] The server sends the generated notification to the user's device. The input is the personalized message created in step 5, and the output is the notification received on the user's device. The server uses the Twilio API to securely send the SMS, ensuring the user receives the necessary information immediately.
[0526] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0527] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0528] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0529] [Fourth Embodiment]
[0530] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0531] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0532] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0533] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0534] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0535] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0536] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0537] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0538] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0539] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0540] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0541] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0542] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0543] This system aims to quickly and easily summarize voice messages recorded on answering machines and send them to users as short messages. The main components of the system include a server that acquires voice data, a server that converts the voice to text and summarizes it, and a server that sends the summarized information to the user's terminal.
[0544] The server first connects with the answering machine system and immediately retrieves the digital audio data as soon as a voice message is recorded. This data is securely transferred with the user's permission. Next, the server uses speech recognition technology to convert the audio data into text. This conversion utilizes advanced algorithms that analyze the context of spoken language to generate accurate text.
[0545] Next, the server processes the generated text using a summarization algorithm to extract the most important information. This summarization allows for the concise transmission of the essence of the information without compromising the overall meaning of the message. For example, if a voice message is left on the answering machine saying, "Tomorrow's meeting has been changed from 10:00 to 11:00," the server will summarize this as, "Meeting changed from 10:00 to 11:00."
[0546] Finally, the server sends the summarized text to the user's terminal in the form of a short message. A secure messaging protocol is used for the transmission process to prevent information leakage. The user can check the message on their terminal and understand the content without listening to the voicemail message. This allows users who cannot access audio data or who have hearing impairments to efficiently grasp the information.
[0547] The following describes the processing flow.
[0548] Step 1:
[0549] The server connects to the answering machine system and detects that a voice message has been recorded. The server automatically downloads and saves the voice data, preparing it for the next step.
[0550] Step 2:
[0551] The server inputs the stored audio data into the speech recognition engine, which then converts the audio into text. The speech recognition engine analyzes the characteristics of the audio and generates accurate text data.
[0552] Step 3:
[0553] The server applies a summarization algorithm to the generated text data. The summarization algorithm extracts the important parts of the text and creates a compressed summary.
[0554] Step 4:
[0555] The server converts the summarized text into a short message format for transmission via Short Message Service (SMS). The server then securely sends the converted data to the user's device.
[0556] Step 5:
[0557] The user's device receives an SMS message and displays a notification. The user can then open the SMS application on their device to view the received message and easily understand the content recorded on their voicemail.
[0558] (Example 1)
[0559] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0560] There is a need to efficiently summarize voice messages recorded on answering machines, allowing users to quickly grasp important information without having to play the audio. In particular, a system is needed that can be easily implemented while ensuring security and accuracy throughout the process from acquiring voice data to summarizing and communicating it.
[0561] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0562] In this invention, the server includes means for acquiring audio information, means for converting the audio information into text information, means for summarizing the text information, and means for using a secure communication protocol for the communication. This enables the user to instantly understand important information without directly listening to the audio.
[0563] "Audio information" refers to analog or digital audio recordings expressed in audio data or audio file formats.
[0564] "Means of acquisition" refers to mechanisms and methods for collecting and acquiring audio information in cooperation with external devices or systems.
[0565] "Textual information" refers to text data converted from audio information, and is linguistic information expressed in a readable form.
[0566] "Means of conversion" refers to algorithms and technologies used to analyze audio information and replace it with text information.
[0567] "Methods of summarization" refer to methods of extracting the main points from the original text and expressing them in the fewest possible words.
[0568] "Means of communication in short message format" refers to methods for sending summarized textual information as a concise message.
[0569] A "secure communication protocol" refers to a standardized communication method designed to protect and ensure the security of communication content.
[0570] To implement this invention, a server plays a central role in configuring a system that integrates and performs the processes of acquiring, converting, summarizing, and communicating voice information.
[0571] First, the server acquires audio information through an interface with the user's audio recording system. Specific hardware options for this include digital recording devices and network-connected storage. This audio information is stored in common digital audio formats such as WAV and MP3.
[0572] Next, the server uses speech recognition software to convert the acquired audio information into text information. Specific examples of such software include general-purpose speech recognition APIs (e.g., Google Speech-to-Text API). This converts the content of the audio data into text format.
[0573] The converted text information is processed by a server using a summarization algorithm. This process uses a generative AI model (e.g., a BERT-based model) to extract the essential information from the text and shorten it. This summary eliminates any excess or omission of information, efficiently conveying the content. For example, the audio information "Tomorrow's meeting has been changed from 10:00 to 11:00" is summarized as "Meeting changed from 10:00 to 11:00".
[0574] Finally, the server transmits the summarized text information to the user's device in short message format. During this process, the information is securely transmitted using a secure communication protocol (e.g., TLS). The user can then view the received information through a messaging app on their device.
[0575] An example of a prompt message is, "Please summarize the voice message recorded on the answering machine and send it as a text message. Example voice message: 'Tomorrow's meeting has been changed from 10:00 to 11:00.'" This message is entered into the generation AI model, and the summary is performed. This system allows users to quickly grasp the main points without having to play the audio directly.
[0576] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0577] Step 1:
[0578] The server interacts with the audio recording system to acquire audio information. It receives digital audio files (e.g., WAV or MP3 format) from the audio recording system as input. The server uses APIs and network protocols to securely store these audio files. The output is an audio data file.
[0579] Step 2:
[0580] The server converts the audio information into text information. Using the audio data acquired in Step 1 as input, the audio data is analyzed using speech recognition software (such as the Google Speech-to-Text API). This results in the spoken words being output as text. Specifically, the waveform data is converted into a string by a language model and saved in text format.
[0581] Step 3:
[0582] The server uses a generative AI model to summarize text information. Using the text data generated in Step 2 as input, a BERT-based summarization algorithm is used to extract the most important points and shorten the text. Specifically, it selects highly relevant words and phrases from the long text and reconstructs them into a concise form. The output is a summarized short text.
[0583] Step 4:
[0584] The server sends a summarized short text to the terminal. The summarized text created in step 3 is used as input and sent to the user's terminal using a secure communication protocol. Specifically, the short message is sent via an SMS gateway using TLS encryption. The output is a short message displayed on the terminal.
[0585] Step 5:
[0586] The user checks the short message received on their device. They open and read the summary information in their device's messaging app. This allows the user to understand the important content without directly listening to the audio. The text displayed on the device screen is the output.
[0587] (Application Example 1)
[0588] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0589] In environments where numerous voice alerts and warning messages are generated, there is a need for information transmission methods that allow for quick and accurate understanding and response to them. However, voice information is often difficult to use directly, and there is a challenge in that it is difficult to confirm voice messages, especially when hearing or communication conditions are limited.
[0590] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0591] In this invention, the server includes means for acquiring audio information, means for converting the audio information into document information, and means for summarizing the document information. This enables more efficient management of audio information and allows for quick acquisition of important information even in situations with hearing limitations or in emergency situations.
[0592] "Audio information" refers to the information that makes up audio data, including signals and digital data generated from audio.
[0593] "Document information" refers to text data converted from speech, in a format that can be visually displayed.
[0594] "Summarization means" refers to a technology that extracts important information from acquired document information and processes it to express it in a shortened form.
[0595] A "short text" refers to a piece of writing that expresses summarized information in a concise text format, designed to quickly convey information.
[0596] "Visual devices" are devices used by users to visually confirm information, and include displays and smart glasses.
[0597] An "information management system" refers to a system that comprises a series of processes, from acquiring audio information to generating document information, summarizing it, and transmitting it.
[0598] This system is an information management system that efficiently manages audio information, summarizes it as visually verifiable document information, and displays it in short sentences on a visual device. An embodiment of the system is shown below.
[0599] The server acquires audio information from security devices and alarm systems within the building using means to acquire audio information. The audio information is processed as digital audio data. The server converts this audio information into document information using speech recognition software such as the Google Cloud Speech-to-Text API. The document information is text data accurately transcribed from the audio. Next, the server summarizes the document information using generative AI technology such as OpenAI's GPT-3 model. The summarization means extracts important information from the document information and converts it into concise short sentences.
[0600] The device receives summarized short messages and displays them on a visual device. For example, by displaying the short messages on the screen of a smartphone or smart glasses, users can quickly understand the situation without directly listening to audio information. This allows for faster responses in emergencies.
[0601] As a concrete example, consider a scenario where a fire alarm is activated inside a shopping mall building. This voice warning message is sent to a server and converted into document information as "Fire alarm activated. Location: 3rd floor. Please check immediately." Then, a generation AI model summarizes it as "Fire alarm activated on the 3rd floor. Check required," and sends it to a mobile device or visual device. An example of a prompt would be, "Please summarize the following emergency voice alert. Original text: Fire alarm activated. Location: 3rd floor. Please check immediately."
[0602] This system enables users to quickly and accurately obtain important information and take appropriate action.
[0603] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0604] Step 1:
[0605] The server acquires audio information in real time from the audio sensor. The input is audio data from microphones within the building. This data is securely transferred to the server as digital data.
[0606] Step 2:
[0607] The server uses the Google Cloud Speech-to-Text API to convert acquired audio data into document information. The input is audio data, and the output is text data as document information. By using speech recognition technology, a string of characters that accurately transcribes spoken language is generated.
[0608] Step 3:
[0609] The server uses a generative AI model to summarize document information. The input is text data, and the output is a summarized short sentence. At this stage, the AI extracts and concisely summarizes the important information. The prompt used is the command, "Summarize the following emergency audio alert."
[0610] Step 4:
[0611] The server sends summarized short texts to the terminal using a secure communication protocol. The input is summarized text data, and the output is transmitted to a visual device. At this stage, the communication is encrypted to prevent information leakage.
[0612] Step 5:
[0613] The terminal displays received short messages on a visual device. The input is a short message sent from the server, and the output is a visualization on the display. The user can instantly check the summary information on the display and take the necessary action.
[0614] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0615] This system aims to provide more personalized information by quickly summarizing voice messages recorded on answering machines and combining the summary results with the user's emotion recognition. The system mainly consists of a server that acquires voice data, a server that converts the voice to text and summarizes it, a server equipped with an emotion engine, and a server that transmits the information to the user's terminal.
[0616] The server first interacts with the answering machine system to detect voice messages and simultaneously acquire digital audio data. The acquired audio data is then converted into text data using speech recognition technology. A summarization algorithm extracts the most important information from the converted text data and generates a shortened summary.
[0617] The emotion engine analyzes this text data to recognize the user's emotional state. This emotion analysis identifies the user's stress levels and emotional tendencies, allowing for the prioritization and adjustment of message content accordingly. For example, if an emotion such as "the meeting has been postponed and stress levels are high" is detected, the server will follow up by adding a detailed explanation to the user in the notification.
[0618] Finally, the server generates a short message containing a summary and sentiment analysis results, and securely sends it to the user's device. The message is received instantly on the device, allowing the user to gain a deeper understanding of the information tailored to their situation. This makes it possible to receive personalized information even in situations where audio data cannot be viewed or when there are hearing limitations.
[0619] The following describes the processing flow.
[0620] Step 1:
[0621] The server interacts with the answering machine system to detect when voice messages have been recorded. Based on this information, the server automatically retrieves and saves the voice data. This process is carried out using secure communication to prevent the loss of voice data.
[0622] Step 2:
[0623] The server passes the acquired audio data to the speech recognition engine, which converts the audio into text data. This engine analyzes the audio and generates digital data in text format. During this process, it corrects for the speaker's accent and noise to create accurate text.
[0624] Step 3:
[0625] The server inputs the generated text data into the summarization module. The summarization module uses natural language processing techniques to extract only the key points from the text and reconstruct them as a concise summary. This process employs multiple analysis algorithms to prevent information overload or omission.
[0626] Step 4:
[0627] The server analyzes the summarized text using an emotion engine to recognize the user's emotions. The emotion engine analyzes the words and context in the text in detail to determine whether the emotion is positive or negative, or whether the user is experiencing increased stress.
[0628] Step 5:
[0629] The server prioritizes messages based on sentiment analysis results. Furthermore, it adjusts message content as needed to provide information best suited to the user's emotional state. For example, if negative emotions are detected, additional cautionary wording or follow-up information may be added.
[0630] Step 6:
[0631] The server generates a short message incorporating the summary text and sentiment analysis results, and sends it to the user's device. The user's device immediately receives the message and displays a notification.
[0632] Step 7:
[0633] The user checks the short message on their device. Through the received message, they can quickly grasp the content of the voicemail and information relevant to their mood. In this way, the user can understand the important points without having to refer to detailed information.
[0634] (Example 2)
[0635] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0636] Voice messages routinely contain a wealth of information, and it is essential to quickly grasp the most important content. However, manually reviewing and summarizing voice messages is time-consuming and laborious, and providing personalized responses, especially when including emotional information, is difficult. Furthermore, there is a need for efficient ways to acquire information in situations where hearing is impaired or in environments where voice confirmation is not possible.
[0637] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0638] In this invention, the server includes means for acquiring voice information, means for converting the voice information into document information, means for summarizing the document information, means for recognizing the emotional state and adjusting the message content, and means for sending the adjusted message to the terminal. This enables the rapid extraction of important information from voice data and the provision of personalized information according to the user's emotional state.
[0639] "Audio information" refers to a representation of sound waveforms as digital data, and specifically, recordings of human speech.
[0640] "Document information" refers to audio information converted into text data, which is text data in a format that humans can read.
[0641] "Summarization" is the process of extracting key points from document information and generating a shortened version of the text.
[0642] "Emotional state" refers to data obtained by analyzing document information that indicates the emotions and psychological tendencies of the information provider.
[0643] A "message" is the information ultimately sent to the terminal, and includes summarized document information and tailored additional information.
[0644] A "terminal" is an electronic device that ultimately receives messages and is used by the user to verify information.
[0645] To implement this invention, a server, speech recognition software, a generative AI model, an emotion analysis engine, a message delivery system, and a user terminal are used. The following describes each component and its operation.
[0646] The server first acquires audio information. Specifically, it works in conjunction with the recording device system to acquire voice messages as digital data as soon as they are recorded. This process is carried out by retrieving data directly from the voice messaging system using an API.
[0647] Subsequently, the server converts the audio information into document information using speech recognition software. This "speech recognition software" utilizes "speech recognition technology" to convert speech into text. This technology analyzes the audio waveform and generates corresponding character data.
[0648] Next, the server summarizes the document information using a generative AI model. In this step, generative AI models such as "BERTSUM" and "GPT-3" are used to extract key points and generate a concise text. For the summary, a prompt such as "Summarize this text in 50 characters or less" is entered.
[0649] Subsequently, the server uses a sentiment analysis engine to analyze the summarized document information and recognize the emotional state. The sentiment analysis engine analyzes the sender's emotions based on the text information, understanding their stress level and emotional tendencies. This makes it possible to adjust the message content.
[0650] Ultimately, the server sends the reconciled message to the terminal via a message delivery system. This "message delivery system" could be something like "Twilio" or "Amazon SNS," ensuring the message is delivered quickly and securely to the user's terminal.
[0651] The device receives messages and is used by the user to verify the information. This allows users to instantly grasp personalized information in document form, even when they cannot access audio data.
[0652] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0653] Step 1:
[0654] The server works in conjunction with the recording device to detect a voice message and immediately acquires digital audio data. The input is the digital data recorded as a voice message. At this stage, the server receives the audio file from the recording device. The output is the digital audio data that is input to the speech recognition engine.
[0655] Step 2:
[0656] The server utilizes speech recognition technology to convert acquired digital audio data into document information. The input is the digital audio data obtained in step 1. The server analyzes the audio waveform and generates a string based on it. The output is the document information used in the summarization algorithm.
[0657] Step 3:
[0658] The server uses a generative AI model to summarize the transformed document information. The input is the document information provided in step 2. The server inputs prompt sentences into the generative AI model to obtain a concise summary with key points extracted. The output is the summary provided for sentiment analysis.
[0659] Step 4:
[0660] The server uses a sentiment analysis engine to analyze the summary text and recognize the emotional state. The input is the summary text obtained in step 3. The server analyzes the emotional expressions in the document and estimates the user's emotional state. The output is sentiment data used to adjust the message content.
[0661] Step 5:
[0662] The server adjusts the message content based on the determined emotional state and generates the final message. The input is the emotional data obtained in step 4 and the summary text from step 3. The server personalizes the message and adds relevant information. The output is the final message ready to be sent to the user's terminal.
[0663] Step 6:
[0664] The server sends the coordinated message to the user's terminal via the message delivery system. The input is the final message generated in step 5. The message delivery system transfers the data through a secure communication channel. The output is the message received at the user's terminal.
[0665] Step 7:
[0666] The user's device receives and displays the message. The input is the final message sent in step 6. The device provides information to the user through its notification function, allowing them to review personalized content. The output is the displayed message, which the user can understand.
[0667] (Application Example 2)
[0668] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0669] The challenge lies in reducing the information overload and emotional burden users face when reviewing voice messages, and in achieving efficient and personalized information delivery. In particular, considering an individual's emotional state is necessary to provide more appropriate notifications and suggestions, thereby improving the user experience.
[0670] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0671] In this invention, the server includes means for acquiring voice information, means for converting the voice information into text information, means for summarizing the text information, means for analyzing the summarized text information using an emotion recognition engine, and means for generating and sending personalized notifications based on the user's emotional state. This makes it possible to generate personalized notifications and suggestions that take the user's emotions into consideration.
[0672] "Voice information" refers to digital data obtained from voice, including the content of messages and voice communications recorded by users.
[0673] "Textual information" refers to text data obtained by converting audio information, and is information written in a format that is readable by humans.
[0674] An "emotion recognition engine" refers to an algorithm or program that analyzes textual information to identify a user's emotional state.
[0675] "Emotional state" refers to a psychological or emotional state determined based on the user's voice or text information, and is expressed in categories such as stress, joy, and dissatisfaction.
[0676] "Personalized notifications" refer to information or recommendations that are customized and sent based on the emotional state or individual user characteristics.
[0677] A system for carrying out this invention includes a process of acquiring audio information from a user and converting it into text information. The server converts the audio data into text data using a speech recognition engine (e.g., Google Speech-to-Text API). The converted text information is summarized by a natural language processing library (e.g., NLTK, spaCy) to identify important content. The server analyzes the summarized text information using an emotion recognition engine (e.g., emotion analysis based on the BERT model) to determine the user's emotional state.
[0678] Based on the user's emotional state, the server generates personalized notifications and sends the information to the user's device. Secure SMS messages are sent using APIs such as the Twilio API.
[0679] For example, if a user leaves a voice message regarding a product defect, and sentiment analysis determines that the user is dissatisfied, the server will generate and send a message to the user containing a discount coupon and details on the exchange / refund process.
[0680] Examples of prompt statements to be input to a generative AI model include the following:
[0681] "The item I ordered yesterday still hasn't arrived. I'm very worried. Please check on it as soon as possible."
[0682] Emotions: Dissatisfaction, stress
[0683] Action: Coupon offer and delivery status confirmation message
[0684] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0685] Step 1:
[0686] The server receives audio information from the user's terminal. The input is a recorded audio file. The server receives this audio file and prepares for the next processing step.
[0687] Step 2:
[0688] The server converts audio information into text information. The input is the audio file received in step 1, and the output is text data. The server uses the Google Speech-to-Text API to perform speech recognition and converts the data obtained from the audio file into text information.
[0689] Step 3:
[0690] The server summarizes the converted character information. The input is the character data generated in step 2, and the output is the summarized text. The server uses a natural language processing library (NLTK or spaCy) to analyze the character information, extract important content, and summarize it.
[0691] Step 4:
[0692] The server analyzes the summarized text information using an emotion recognition engine. The input is the summarized text generated in step 3, and the output is the user's emotional state. The server executes an emotion analysis algorithm based on the BERT model to determine the user's emotional state from the text information.
[0693] Step 5:
[0694] The server generates personalized notifications based on the emotional state. The input is the emotional state information determined in step 4, and the output is the personalized message sent to the user. The server automatically generates appropriate messages and suggestions based on the emotional state and prepares them for sending.
[0695] Step 6:
[0696] The server sends the generated notification to the user's device. The input is the personalized message created in step 5, and the output is the notification received on the user's device. The server uses the Twilio API to securely send the SMS, ensuring the user receives the necessary information immediately.
[0697] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0698] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0699] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0700] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0701] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0702] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0703] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0704] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0705] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0706] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0707] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0708] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0709] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0710] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0711] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0712] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0713] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0714] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0715] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0716] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0717] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0718] The following is further disclosed regarding the embodiments described above.
[0719] (Claim 1)
[0720] Means for acquiring audio data,
[0721] Means for converting the aforementioned audio data into text data,
[0722] Means for summarizing the aforementioned text data,
[0723] A means for sending the summarized text data as a short message,
[0724] A system that includes this.
[0725] (Claim 2)
[0726] The system according to claim 1, characterized in that the voice data acquisition means acquires voice data in cooperation with an answering machine system.
[0727] (Claim 3)
[0728] The system according to claim 1, characterized in that the text data summarization means uses natural language processing technology.
[0729] "Example 1"
[0730] (Claim 1)
[0731] Means for acquiring audio information,
[0732] means for converting the aforementioned audio information into text information,
[0733] A means for summarizing the aforementioned textual information,
[0734] A means for communicating the summarized text information in short sentence format,
[0735] The means of the aforementioned communication using a secure communication protocol,
[0736] A system that includes this.
[0737] (Claim 2)
[0738] The system according to claim 1, characterized in that the voice information acquisition means acquires voice information in cooperation with a voice recording system.
[0739] (Claim 3)
[0740] The system according to claim 1, characterized in that the text information summarization means uses natural language processing technology.
[0741] "Application Example 1"
[0742] (Claim 1)
[0743] Means for acquiring audio information,
[0744] Means for converting the aforementioned audio information into document information,
[0745] Means for summarizing the aforementioned document information,
[0746] A means for transferring the summarized document information as a short sentence,
[0747] Means for displaying the summarized document information on a visual device,
[0748] An information management system that includes this.
[0749] (Claim 2)
[0750] The information management system according to claim 1, characterized in that the voice information acquisition means acquires voice information in cooperation with an alarm device.
[0751] (Claim 3)
[0752] The information management system according to claim 1, characterized in that the document information summarization means uses generation AI technology.
[0753] "Example 2 of combining an emotion engine"
[0754] (Claim 1)
[0755] Means for acquiring audio information,
[0756] Means for converting the aforementioned audio information into document information,
[0757] Means for summarizing the aforementioned document information,
[0758] A means for analyzing the summarized document information to recognize an emotional state,
[0759] Means for adjusting the content of a message based on the aforementioned emotional state,
[0760] means for sending the adjusted message to the terminal,
[0761] A system that includes this.
[0762] (Claim 2)
[0763] The system according to claim 1, characterized in that the voice information acquisition means acquires voice information in cooperation with a recording device system.
[0764] (Claim 3)
[0765] The system according to claim 1, characterized in that the document information summarization means uses generative artificial intelligence technology.
[0766] "Application example 2 when combining with an emotional engine"
[0767] (Claim 1)
[0768] Means for acquiring audio information,
[0769] Means for converting the aforementioned audio information into text information,
[0770] Means for summarizing the aforementioned textual information,
[0771] The means for analyzing the summarized text information using an emotion recognition engine,
[0772] A means for generating and sending personalized notifications based on the user's emotional state,
[0773] A system that includes this.
[0774] (Claim 2)
[0775] The system according to claim 1, characterized in that the voice information acquisition means acquires voice information in cooperation with a voice recording system.
[0776] (Claim 3)
[0777] The system according to claim 1, characterized in that the character information summarization means utilizes natural language processing technology. [Explanation of Symbols]
[0778] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring audio data, Means for converting the aforementioned audio data into text data, Means for summarizing the aforementioned text data, A means for sending the summarized text data as a short message, A system that includes this.
2. The system according to claim 1, characterized in that the voice data acquisition means acquires voice data in cooperation with an answering machine system.
3. The system according to claim 1, characterized in that the text data summarization means uses natural language processing technology.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A