system
The system addresses the challenge of managing call information by converting audio to text, extracting key details, and automating schedule registration, enhancing efficiency and reducing manual effort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-12
- Publication Date
- 2026-06-24
AI Technical Summary
Existing systems fail to efficiently and accurately manage schedules and task lists during calls, often leading to ambiguous recordings, missed information, and time-consuming post-call organization.
A system that captures call audio as digital data, converts it into text using speech recognition, extracts important information through natural language processing, and automatically registers it in schedules and tasks, with real-time notification of the summary.
Enables immediate and accurate recording of call content, automating post-call management and reducing user burden by providing quick access to important information.
Smart Images

Figure 2026103433000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In business scenarios and daily life, it is required to easily manage schedules and task lists while accurately recording important information during a call without omission. In the prior art, since it is necessary to manually take notes during a call, the content may become ambiguous or the recording may be omitted. Also, there is a problem that it takes time to organize important information after the call and manually register it in schedules and tasks.
Means for Solving the Problems
[0005] This invention provides means for acquiring call audio as digital audio data in real time and speech recognition means for converting the acquired audio data into text, thereby enabling immediate recording of call content. Furthermore, important information is extracted from the text by natural language processing means and a summary is generated, allowing the user to easily review important information. In addition, a function is provided to automatically register the extracted information in a task list and schedule, automating post-call management tasks and reducing the user's burden. Moreover, by using means to notify the user of the summary generated and registered information after the call ends, the content can be quickly and reliably reviewed.
[0006] "Call audio" refers to the voice signals exchanged between people via telephone or voice communication systems.
[0007] "Digital audio data" refers to data obtained by converting audio into a digital format, and is stored in a format suitable for signal processing and transmission.
[0008] "Speech recognition means" refers to a program or device that analyzes speech data and converts its content into text data.
[0009] "Text" refers to string information generated from audio data through speech recognition.
[0010] "Natural language processing methods" are technologies that analyze meaning and information from text data and perform information extraction and summarization according to specific purposes.
[0011] A "summary" is a concise document that summarizes the content based on key information extracted from text data.
[0012] A "task list" is a list used to manage the work and tasks that are scheduled to be done.
[0013] A "schedule" is a table or calendar used to manage planned appointments or events at specific times.
[0014] A "notification" is a means by which a system informs a user of information, and may be presented as an alert or a message. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention relates to a system that efficiently and accurately records call content, extracts necessary information, and improves the user's work efficiency. Specifically, a terminal captures the call's audio signal as digital audio data. This data is then transmitted to a server. The server analyzes the captured audio data using speech recognition technology and converts it into text in real time. This text is then analyzed using natural language processing technology to extract important information and key phrases.
[0037] Furthermore, the server generates a summary based on the extracted information and stores it in a database. This summary is retained in the call history in a format that allows the user to easily review its contents. The server also automatically registers items in task lists and schedules based on specific key phrases. For example, if the call content recognizes "Let's have a meeting next Tuesday," it will be automatically registered in the corresponding schedule. After the call ends, the server notifies the user of the generated summary and registered schedule / task information, allowing the user to quickly review and use the information.
[0038] As a concrete example, suppose a user is on a business call where they discuss deadlines and meeting dates for a new project. Using this system, the call is analyzed in real time, and important project deadlines and upcoming meetings are automatically registered in the user's scheduling application. After the call ends, the user can review the summary and easily verify that the schedule has been accurately registered. This eliminates the need for manual note-taking and data entry, dramatically improving the efficiency of business processes.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] As soon as a call begins, the device acquires audio through the microphone and formats it as digital audio data in real time.
[0042] Step 2:
[0043] The terminal structures the captured digital audio data into packets and sends them to the server over the network. The data is transmitted in segments of a fixed buffer size.
[0044] Step 3:
[0045] The server inputs the received audio data into the speech recognition engine, which extracts linguistic information from the audio and converts it into text. This process is performed in real time, and the conversion results are temporarily stored in memory.
[0046] Step 4:
[0047] The server analyzes the text generated using natural language processing technology and extracts important information and key phrases. During this process, it prioritizes the information based on pre-configured business rules and context.
[0048] Step 5:
[0049] The server generates a short, concise summary based on the analysis results and records this summary in the call history within the database.
[0050] Step 6:
[0051] The server automatically registers information for schedules and task lists derived from data analysis into other management systems using APIs.
[0052] Step 7:
[0053] Once the call ends, the server notifies the user of the generated summary and any registered tasks or schedules. The user can then review this information to quickly understand the call content and their next steps.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] Many modern communication systems struggle to efficiently and accurately record call content and automatically organize and manage that information. Especially in business meetings and long calls, important information can be lost, or manual note-taking may be necessary. Therefore, users face the challenge of not being able to accurately refer to the information obtained during a call and use it to inform their next actions.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes means for acquiring call acoustics as digital acoustic information in real time, acoustic recognition means for converting the acquired acoustic information into text data, and natural language processing means for extracting key information from the converted text data and generating essential points. This enables efficient recording of call information, automatic organization of necessary information, and allows users to quickly review and utilize the information after a call.
[0059] "Call acoustics" refers to the acoustic signals generated during a voice call, and is acquired during conversations conducted via telephone or online platforms.
[0060] "Digital audio information" refers to data obtained by quantifying analog audio signals and converting them into a format that can be processed by computers and digital devices.
[0061] "Acoustic recognition means" refers to technologies and devices that analyze acoustic signals and automatically convert them into meaningful character data.
[0062] "Character data" refers to data in text format that is converted from speech by an acoustic recognition system.
[0063] "Key information" refers to information or key phrases extracted from text data that users should consider important.
[0064] "Key points" refer to content that efficiently summarizes and concisely describes the main information, provided in a way that users can quickly understand and use.
[0065] "Natural language processing means" refers to technologies and devices for analyzing text data and extracting or summarizing information.
[0066] "Automatic registration" refers to the process of adding items necessary for task lists and schedule management based on information extracted by the system, without manual intervention.
[0067] An "information storage device" refers to a recording medium or system for long-term storage of digital information, making it accessible at a later date.
[0068] "Information transmission network" refers to the network environment and communication technology used to deliver digital acoustic information to acoustic recognition devices.
[0069] "Users" refer to individuals or organizations that record call content or extract information through this system.
[0070] This invention is a system that acquires call acoustics as digital acoustic information in real time and accurately analyzes and utilizes that information. The implementation of this invention involves the terminal, server, and user.
[0071] The terminal captures the audio signal during a voice call as digital audio information using its microphone. For example, it converts the audio signal into PCM digital data using the built-in microphone of a smartphone or PC, or an externally connected microphone. The terminal then transmits this digital audio information to a server via an internet connection. In this process, the internet is used as the information transmission network, and the SSL / TLS protocol is used to encrypt the data.
[0072] The server converts received digital audio information into text data using acoustic recognition technology. This technology can utilize commercially available acoustic recognition software or cloud-based audio services. After conversion to text data, the server analyzes this data using natural language processing technology to extract key information. Natural language processing technology includes word segmentation, syntactic analysis, and sentiment analysis, and can utilize commercially available natural language processing software or cloud services.
[0073] Based on the extracted key information, the server generates summaries using a generative AI model. During this process, the prompt "Summarize the following text." is used to input the analyzed text into the AI model. The generated summaries and related information are stored in an information storage device, allowing users to search and review them at a later date.
[0074] Based on registered important information, the server automatically registers tasks in the work list and schedule. Specifically, this includes recognizing key phrases such as "meeting" and "deadline" and generating events in the schedule management application. This allows users to quickly review and utilize this information even after a call has ended, reducing the burden of manual note-taking and data entry.
[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0076] Step 1:
[0077] The device uses a microphone to capture the audio signal during a call as digital audio information. This can be done using the built-in microphone of a smartphone or computer, or an externally connected microphone. The input is an analog audio signal, which is converted into PCM format digital data using an ADC (analog-to-digital converter). The output is the digitized audio data.
[0078] Step 2:
[0079] The terminal transmits digital audio data to the server. During this process, the terminal encrypts the data using the SSL / TLS protocol and streams it to the server in real time over the internet. The input is digital audio data, and the output is a secure stream of audio data.
[0080] Step 3:
[0081] The server converts the received audio data into text data using acoustic recognition technology. Here, commercially available acoustic recognition software is used to analyze the audio. The input is streamed digital audio data, and the output is text data.
[0082] Step 4:
[0083] The server analyzes text data using natural language processing techniques and extracts key information. It utilizes tools such as the Google Cloud Natural Language API for word segmentation and information extraction. The input is text data, and the output is an analysis result containing key information and key phrases.
[0084] Step 5:
[0085] The server generates key points based on a generative AI model using the extracted main information. Using OpenAI® GPT-3®, the server receives parsed text via the prompt "Summarize the following text." The output is summarized data.
[0086] Step 6:
[0087] The server stores the generated summaries in an information storage device to facilitate later searching and verification. A database system (e.g., MongoDB or MySQL®) is used here. The input is summary data, and the output is the summary information recorded in the database.
[0088] Step 7:
[0089] The server automatically registers tasks and appointments based on important key phrases. It uses the Google Calendar API and other tools to add events related to the key phrases to the schedule. The input is the parsed key information, and the output is the new events added to the schedule data.
[0090] Step 8:
[0091] After the call ends, the server notifies the user of the generated summary and registered schedule information. Email and push notifications are used for this notification. The input is the summary and schedule information, and the output is the notification data sent to the user.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] In today's information-saturated world, efficiently managing information obtained through daily and work-related phone calls and ensuring that important dates and tasks are not missed is a challenging task. Especially in situations requiring real-time information processing and automation, manual management has its limitations. This can lead to the leakage of important information, potentially hindering schedule management and task completion. Therefore, a system is needed that can instantly analyze call content and accurately process and manage important information.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the acquired audio data into text, natural language processing means for extracting important information from the converted text and generating a summary, and means for updating the performance schedule using the generated summary and registered events. This makes it possible to manage information obtained from calls in real time and to automatically update important schedules and task lists, enabling users to achieve efficient schedule management without being troubled by the complexity of the information.
[0097] "Call audio" refers to the human voice transmitted through communication devices.
[0098] "Digital audio data" refers to data obtained by converting an audio signal into a digital format, which is a format that can be processed by computers and digital devices.
[0099] "Speech recognition means" refers to a technology or device that analyzes speech signals and converts them into corresponding text information.
[0100] "Natural language processing means" refers to technologies or devices that analyze text data, understand human language, and extract information.
[0101] A "summary" is a document that condenses the original information and extracts its main points.
[0102] "Events" refer to information indicating schedules and tasks introduced by events such as phone calls.
[0103] "Scheduled execution" refers to tasks or events that are scheduled to be performed and should be registered in the management system.
[0104] A "storage device" is a device used to store digital data for long-term or short-term storage, and is used for information retention.
[0105] A "user" refers to an individual or organization that utilizes this system.
[0106] A description of the embodiment for carrying out the invention will be provided.
[0107] This system supports users' lives and work by acquiring and processing voice data in real time and automatically managing important schedules and task lists. The server receives call audio transmitted from the user's terminal as digital voice data. Next, it converts this voice data into text using speech recognition. Specifically, the Python package speech_recognition can be used for speech recognition.
[0108] The converted text is analyzed using natural language processing (NLP) tools to extract important information. This important information includes schedules and task details, which are automatically managed as performance targets. Since generative AI models may be used in the natural language processing, the server sends the text data to an NLP service in the cloud for advanced analysis.
[0109] Once a call is complete, users can receive notifications containing a generated summary and updated schedule information. This notification feature ensures that important information is not missed and allows for efficient management of appointments and tasks. Specifically, schedule information can be managed using Google Calendar or the Task API.
[0110] For example, if a user makes a call saying, "The next garbage collection will be next Monday at 10 AM," this information will be automatically registered in the calendar, and the user will be notified so they don't forget. In this way, by efficiently reflecting the content of calls in the schedule, the system supports the user's life management.
[0111] Examples of prompts for a generative AI model include the following:
[0112] Please summarize the conversation as follows: "The next garbage collection will take place next Monday at 10:00 AM."
[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0114] Step 1:
[0115] The user's device acquires the call audio and records it as digital audio data. This audio data is the input, obtained by digitizing the audio signal using the device's microphone. The acquired digital audio data is then ready to be sent to the server.
[0116] Step 2:
[0117] The server converts digital audio data received from the terminal into text data using speech recognition. In this process, the input is digital audio data, and the output is text data. Specifically, the Python `speech_recognition` library is used to parse the audio signal and convert it into corresponding text. The server then passes this text data to the next processing step.
[0118] Step 3:
[0119] The server sends the converted text data to a natural language processing system for extracting important information. In this process, the input is the text data generated in step 2, and the output is the extracted key phrases and important information. Specifically, a generative AI model is used to identify keywords within the text and summarize that information.
[0120] Step 4:
[0121] Based on the extracted information, the server automatically registers tasks as planned in the task list and schedule. The input for this step is the important information obtained in step 3, and the output is the updated schedule data. Specifically, the Google Calendar API and Task API are used to reflect the appropriate date, time, and content in the user's calendar system.
[0122] Step 5:
[0123] After a call ends, the user receives a notification from the server containing a summary and updated schedule information. This notification is immediately available for the user to review, enabling them to utilize the information gathered during the call. The output is a visually displayed summary, specifically delivered as a push notification to the user's device.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] This invention provides a specific form of a system that processes a user's call audio in real time, simultaneously recording important information and analyzing emotions. The terminal captures the call audio as digital audio data and continuously transmits it to the server. The server utilizes speech recognition means to immediately convert the received audio data into text. This textualized data is analyzed using natural language processing technology to extract important information, and a summary is automatically generated.
[0126] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data and identify their emotional state. This emotional information is handled together with the extracted text information and, as needed, influences the content of the summary and the priority of the task list. It also enables improvements to interactions based on the user's emotions. For example, if the system detects that the user is feeling anxious during a call, it may provide information to alleviate that anxiety.
[0127] After the call ends, the user is notified of the generated summary, tasks, and schedule information, along with the sentiment analysis results obtained by the sentiment engine. This notification process allows the user to reflect on the call content and their own emotional state, enabling them to efficiently plan future actions.
[0128] As a concrete example, consider a business meeting where project progress is reported and scheduling is arranged. Using this system, the call content is recorded and analyzed in real time, and project deadlines and meeting dates are automatically incorporated into the schedule. Furthermore, if stress is detected in the speaker's voice, that element is detected by the emotion engine and reflected as information that helps improve communication. This makes it possible to carry out daily work more efficiently and smoothly.
[0129] The following describes the processing flow.
[0130] Step 1:
[0131] The device uses its microphone to capture audio as a digital signal as soon as a conversation begins, and sends this audio data to the server in real time.
[0132] Step 2:
[0133] The server inputs the received digital audio data into a speech recognition engine, which converts the spoken content into text. This process instantly transcribes the audio data, making it available for processing on the server.
[0134] Step 3:
[0135] The server passes the transcribed data to a natural language processing engine, which extracts important information and key phrases from the conversation. Furthermore, it generates a summary based on this information, concisely summarizing the content.
[0136] Step 4:
[0137] The emotion engine analyzes voice data in parallel, determining emotions based on the user's tone of voice, speaking speed, and other factors. This allows for real-time evaluation of emotional fluctuations during a call.
[0138] Step 5:
[0139] The server combines the extracted key information with sentiment analysis results to adjust the summaries, task lists, and schedule priorities. The sentiment analysis results are also recorded as information accompanying the call history.
[0140] Step 6:
[0141] After the call ends, the server notifies the user of the generated summary, tasks, schedule, and sentiment analysis results. This notification allows the user to review the call content and sentiment feedback.
[0142] Step 7:
[0143] Users can receive notifications, review call content, and reflect on their own emotional responses through sentiment analysis. They can also check their task lists and schedules, which are automatically managed, allowing them to plan their future work.
[0144] (Example 2)
[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0146] In today's communication environment, there is a need to efficiently manage and immediately utilize important information obtained during calls. However, manually recording and organizing call content is laborious, and understanding and appropriately processing emotional nuances is difficult. Therefore, there is a need for a system that can simultaneously analyze call content and the user's emotional state and utilize the information efficiently.
[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0148] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the audio data into text, and analysis means for analyzing the converted text and emotions to generate information. This enables simultaneous analysis of important information from the call content and the user's emotions, and automatic reflection of this information in task lists and schedules.
[0149] "Call audio" refers to voice information transmitted in real time using voice communication methods.
[0150] "Digital audio data" is a format of information obtained by digitizing analog audio signals.
[0151] "Speech recognition means" refers to a technology or device that converts speech data into text data.
[0152] "Natural language processing means" refers to technologies or processes that extract and understand structured information from text data.
[0153] "Analytical means" refer to the techniques and processes used to analyze data and extract new knowledge and information.
[0154] A "task list" is a means of organizing and displaying a user's activities and obligations in a list format.
[0155] A "time management chart" is a method for organizing and visualizing events and schedules along a timeline.
[0156] "Emotion" refers to the user's psychological state and includes technologies for recognizing or analyzing that state.
[0157] A "recording medium" is a physical or digital medium used to store and maintain data.
[0158] This invention is a system that processes a user's call audio as digital audio data in real time, simultaneously recording important information and analyzing emotions.
[0159] The terminal first acquires the user's call audio and converts it into digital audio data. This process uses an audio capture device with a microphone, and the resulting digital audio data is transmitted to the server via wireless communication.
[0160] The server uses speech recognition technology to convert the received audio data into text data. A speech recognition engine is used for this, with Google Cloud Speech-to-Text API being a specific example. The converted text is then analyzed using natural language processing technology to extract important information. SpaCy and NLTK libraries are used for this purpose.
[0161] Furthermore, the server uses a sentiment analysis engine to perform sentiment analysis based on text data. IBM Watson® Tone Analyzer is one such sentiment engine used. This identifies the user's emotional state, and this information is integrated with the text data. The analysis results can be automatically reflected in task lists and time management tables.
[0162] After the call ends, the user is notified of the generated summary, tasks, schedule information, and sentiment analysis results. The user can then use this information to plan their actions. This notification is delivered via an application installed on the user's mobile device or computer.
[0163] As a concrete example, consider a scenario where the content of business meetings is recorded and analyzed in real time, and project deadlines and meeting dates are automatically registered in the schedule. If stress levels of the speaker are detected, this information is provided as feedback to improve communication.
[0164] An example of a prompt message is, "Analyze the critical deadlines and stakeholders' emotional states in project progress reports, and create a summary to improve scheduling and communication efficiency." This allows the system to efficiently and accurately support the user's business operations.
[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0166] Step 1:
[0167] The terminal acquires the user's call audio in real time. This involves capturing the audio as an analog signal using the microphone built into the terminal and converting it into digital audio data. The input is the user's voice, and the output is the digitized audio data. This audio data is transmitted to the server via the communication module.
[0168] Step 2:
[0169] The server processes the digital audio data received from the terminal. This processing includes converting the data to text using a speech recognition engine. Specifically, the server analyzes the audio spectrum, extracts phonemes, and transcribes them based on a language model. The input is digital audio data, and the output is text data containing character information.
[0170] Step 3:
[0171] The server analyzes the transcribed data using natural language processing techniques. This involves extracting important information and generating a text summary. For example, it performs morphological and syntactic analysis to extract nouns and verbs, and uses semantic analysis to grasp the main points of the utterance. The input is text data, and the output is a condensed text containing summary information and keywords.
[0172] Step 4:
[0173] The server analyzes the user's emotions from voice and text data. Using an emotion analysis engine, it evaluates, for example, the tone of voice and word choice to identify the user's emotional state. The input is the converted text and original voice data, and the output is data containing the type of emotion identified (joy, anger, etc.).
[0174] Step 5:
[0175] The server integrates extracted key information and sentiment analysis results, automatically reflecting them in the task list and schedule. Specifically, it works in conjunction with task management applications to update schedules and add tasks. The input is summary information and sentiment data, and the output is the updated task list and schedule information.
[0176] Step 6:
[0177] After the call ends, the user receives a summary, task list, schedule information, and sentiment analysis results from the server. This information is notified to the user's device, and the user can check its contents via email or application. The output is notification information presented in a user-friendly format.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0180] In modern society, it has become increasingly difficult for families and individuals to efficiently manage important information arising from daily conversations and to recognize changes in emotions. There is a high risk of overlooking important tasks and appointments by simply ignoring conversations. Furthermore, there is often a lack of appropriate interaction with family members experiencing stress or anxiety. These challenges are raising concerns about a decline in quality of life and impacting the well-being of families and individuals.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for acquiring call audio as digital voice numbers in real time, speech recognition means for converting the acquired voice numbers into text information, natural language processing means for extracting important information from the converted text information and generating a summary, means for performing emotion analysis from the call to identify the user's emotional state, and means for making suggestions and improvements to interactions using the emotion analysis results. This makes it possible for families and individuals to efficiently manage important information obtained from conversations, and to respond immediately to changes in emotions and provide appropriate interactions.
[0183] "Call audio" refers to the sound information of a linguistic conversation conducted via a voice communication device.
[0184] "Real-time" refers to a state where processing or analysis is performed immediately at the moment an event occurs.
[0185] "Digital audio numerical data" refers to numerical information obtained by converting analog audio into numbers that can be processed by a computer.
[0186] A "speech recognition system" is a mechanism that performs technology to convert speech data into linguistic characters.
[0187] "Textual information" refers to data consisting of language and symbols, in a readable and writable format.
[0188] "Important matters" refer to information or events that deserve particular attention or should be given priority.
[0189] A "natural language processing system" is a system that analyzes language data and performs information extraction and summary generation.
[0190] "Emotional analysis" is the process of estimating emotional states from voice or text.
[0191] "User's emotional state" refers to the result of determining what emotional state a particular user is in.
[0192] "Suggesting and improving interactions" means providing methods and actions to improve the quality of communication with users.
[0193] The system for realizing this invention consists of a terminal and a server.
[0194] The terminal is equipped with a microphone for receiving voice communications and captures the voice during a call as digital voice data. This digital voice data is continuously transmitted to the server.
[0195] The server uses speech recognition software to perform speech recognition based on the received digital speech data. Specifically, the Python `speech_recognition` library is used. The recognized speech is converted into text information, and then the TextBlob natural language processing library is used to evaluate the importance of the information and generate a summary as needed.
[0196] In addition, the server runs an emotion analysis engine to extract emotional information from textual data. This uses the emotion analysis pipeline within the transformers library of the generative AI model, Hugging Face. This emotional information is used not only to visualize the user's emotional state but also to suggest and improve interactions. For example, if the server determines that the user is feeling anxious, it will make suggestions to alleviate that anxiety.
[0197] As a concrete example, a robot installed in a family's living room could record family conversations and register important matters as tasks. Furthermore, if it detects someone is stressed, it could play their favorite music to help them relax.
[0198] An example of a prompt message for the application of this system is: "Analyze family phone calls, identify the types of content that tend to cause stress if someone is stressed, and suggest ways to relax."
[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0200] Step 1:
[0201] The terminal receives call audio and converts it into digital audio numbers. The input is analog audio obtained through a microphone, which is sampled and quantized to generate digital audio numbers. This converts the audio data into a format that can be processed by the server.
[0202] Step 2:
[0203] The terminal transmits the generated digital voice numerical data to the server in real time. The input at this stage is the digital voice numerical data generated in step 1, which is sent to the server via network communication. The server then receives the data to proceed to the next processing step.
[0204] Step 3:
[0205] The server analyzes the received digital speech data using speech recognition software and converts it into text. The input is the digital speech data received by the server; speech recognition is performed using the speech_recognition library, and text is obtained as output. This text is essential for subsequent natural language processing.
[0206] Step 4:
[0207] The server analyzes the converted text information using natural language processing tools, extracts important information, and generates a summary as needed. The input is the text information obtained in step 3, and language analysis is performed using TextBlob. The output is a list of important information and the generated summary, and this information contributes to improving work efficiency.
[0208] Step 5:
[0209] The server performs sentiment analysis from textual information to identify the user's emotional state. The input is the textual information from step 3, and the Hugging Face transformers library is used for sentiment analysis. The output is the sentiment identification result, providing information that reflects the user's psychological state.
[0210] Step 6:
[0211] The server makes suggestions and improvements to interactions based on the sentiment analysis results. The input is the output from steps 4 and 5, and appropriate suggestions are made to the user according to the identified key points and emotional state. This improves the user experience and enables more comfortable communication.
[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0215] [Second Embodiment]
[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0228] This invention relates to a system that efficiently and accurately records call content, extracts necessary information, and improves the user's work efficiency. Specifically, a terminal captures the call's audio signal as digital audio data. This data is then transmitted to a server. The server analyzes the captured audio data using speech recognition technology and converts it into text in real time. This text is then analyzed using natural language processing technology to extract important information and key phrases.
[0229] Furthermore, the server generates a summary based on the extracted information and stores it in a database. This summary is retained in the call history in a format that allows the user to easily review its contents. The server also automatically registers items in task lists and schedules based on specific key phrases. For example, if the call content recognizes "Let's have a meeting next Tuesday," it will be automatically registered in the corresponding schedule. After the call ends, the server notifies the user of the generated summary and registered schedule / task information, allowing the user to quickly review and use the information.
[0230] As a concrete example, suppose a user is on a business call where they discuss deadlines and meeting dates for a new project. Using this system, the call is analyzed in real time, and important project deadlines and upcoming meetings are automatically registered in the user's scheduling application. After the call ends, the user can review the summary and easily verify that the schedule has been accurately registered. This eliminates the need for manual note-taking and data entry, dramatically improving the efficiency of business processes.
[0231] The following describes the processing flow.
[0232] Step 1:
[0233] As soon as a call begins, the device acquires audio through the microphone and formats it as digital audio data in real time.
[0234] Step 2:
[0235] The terminal structures the captured digital audio data into packets and sends them to the server over the network. The data is transmitted in segments of a fixed buffer size.
[0236] Step 3:
[0237] The server inputs the received audio data into the speech recognition engine, which extracts linguistic information from the audio and converts it into text. This process is performed in real time, and the conversion results are temporarily stored in memory.
[0238] Step 4:
[0239] The server analyzes the text generated using natural language processing technology and extracts important information and key phrases. During this process, it prioritizes the information based on pre-configured business rules and context.
[0240] Step 5:
[0241] The server generates a short, concise summary based on the analysis results and records this summary in the call history within the database.
[0242] Step 6:
[0243] The server automatically registers information for schedules and task lists derived from data analysis into other management systems using APIs.
[0244] Step 7:
[0245] Once the call ends, the server notifies the user of the generated summary and any registered tasks or schedules. The user can then review this information to quickly understand the call content and their next steps.
[0246] (Example 1)
[0247] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0248] Many modern communication systems struggle to efficiently and accurately record call content and automatically organize and manage that information. Especially in business meetings and long calls, important information can be lost, or manual note-taking may be necessary. Therefore, users face the challenge of not being able to accurately refer to the information obtained during a call and use it to inform their next actions.
[0249] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0250] In this invention, the server includes means for acquiring call acoustics as digital acoustic information in real time, acoustic recognition means for converting the acquired acoustic information into text data, and natural language processing means for extracting key information from the converted text data and generating essential points. This enables efficient recording of call information, automatic organization of necessary information, and allows users to quickly review and utilize the information after a call.
[0251] "Call acoustics" refers to the acoustic signals generated during a voice call, and is acquired during conversations conducted via telephone or online platforms.
[0252] "Digital audio information" refers to data obtained by quantifying analog audio signals and converting them into a format that can be processed by computers and digital devices.
[0253] "Acoustic recognition means" refers to technologies and devices that analyze acoustic signals and automatically convert them into meaningful character data.
[0254] "Character data" refers to data in text format that is converted from speech by an acoustic recognition system.
[0255] "Key information" refers to information or key phrases extracted from text data that users should consider important.
[0256] "Key points" refer to content that efficiently summarizes and concisely describes the main information, provided in a way that users can quickly understand and use.
[0257] "Natural language processing means" refers to technologies and devices for analyzing text data and extracting or summarizing information.
[0258] "Automatic registration" refers to the process of adding items necessary for task lists and schedule management based on information extracted by the system, without manual intervention.
[0259] An "information storage device" refers to a recording medium or system for storing digital information long-term and making it accessible later.
[0260] "Information transmission network" refers to the network environment and communication technology used to deliver digital acoustic information to acoustic recognition devices.
[0261] "Users" refer to individuals or organizations that record call content or extract information through this system.
[0262] This invention is a system that acquires call acoustics as digital acoustic information in real time and accurately analyzes and utilizes that information. The implementation of this invention involves the terminal, server, and user.
[0263] The terminal captures the audio signal during a voice call as digital audio information using its microphone. For example, it converts the audio signal into PCM digital data using the built-in microphone of a smartphone or PC, or an externally connected microphone. The terminal then transmits this digital audio information to a server via an internet connection. In this process, the internet is used as the information transmission network, and the SSL / TLS protocol is used to encrypt the data.
[0264] The server converts received digital audio information into text data using acoustic recognition technology. This technology can utilize commercially available acoustic recognition software or cloud-based audio services. After conversion to text data, the server analyzes this data using natural language processing technology to extract key information. Natural language processing technology includes word segmentation, syntactic analysis, and sentiment analysis, and can utilize commercially available natural language processing software or cloud services.
[0265] Based on the extracted key information, the server generates summaries using a generative AI model. During this process, the prompt "Summarize the following text." is used to input the analyzed text into the AI model. The generated summaries and related information are stored in an information storage device, allowing users to search and review them at a later date.
[0266] Based on registered important information, the server automatically registers tasks in the work list and schedule. Specifically, this includes recognizing key phrases such as "meeting" and "deadline" and generating events in the schedule management application. This allows users to quickly review and utilize this information even after a call has ended, reducing the burden of manual note-taking and data entry.
[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0268] Step 1:
[0269] The device uses a microphone to capture the audio signal during a call as digital audio information. This can be done using the built-in microphone of a smartphone or computer, or an externally connected microphone. The input is an analog audio signal, which is converted into PCM format digital data using an ADC (analog-to-digital converter). The output is the digitized audio data.
[0270] Step 2:
[0271] The terminal transmits digital audio data to the server. During this process, the terminal encrypts the data using the SSL / TLS protocol and streams it to the server in real time over the internet. The input is digital audio data, and the output is a secure stream of audio data.
[0272] Step 3:
[0273] The server converts the received audio data into text data using acoustic recognition technology. Here, commercially available acoustic recognition software is used to analyze the audio. The input is streamed digital audio data, and the output is text data.
[0274] Step 4:
[0275] The server analyzes text data using natural language processing techniques and extracts key information. It utilizes tools such as the Google Cloud Natural Language API for word segmentation and information extraction. The input is text data, and the output is an analysis result containing key information and key phrases.
[0276] Step 5:
[0277] The server generates key points based on a generative AI model using the extracted main information. Using OpenAI GPT-3 or similar tools, the server receives parsed text as input via the prompt "Summarize the following text." The output is summarized data.
[0278] Step 6:
[0279] The server stores the generated summaries in an information storage device to facilitate later searching and verification. A database system (e.g., MongoDB or MySQL) is used here. The input is summary data, and the output is the summary information recorded in the database.
[0280] Step 7:
[0281] The server automatically registers work lists and schedules based on important key phrases. Using an API such as the Google Calendar API, it adds events related to the key phrases to the schedule. The input is the analyzed main information, and the output is the new events added to the schedule data.
[0282] Step 8:
[0283] After the call ends, the server notifies the user of the generated summary and the registered schedule information. Email or push notifications are used for the notification. The input is the summary and the schedule information, and the output is the notification data to the user.
[0284] (Application Example 1)
[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0286] In modern times when information is becoming complex, it is a difficult task to efficiently manage the information obtained from calls in daily life and business and reflect important schedules and tasks in the schedule without missing them. Especially in situations where real-time information processing and automation are required, there are limitations to manual management. As a result, there is a possibility that important information may be missed, which may hinder schedule management and task execution. Therefore, there is a need for a system that immediately analyzes the call content and accurately processes and manages important information.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0288] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the acquired audio data into text, natural language processing means for extracting important information from the converted text and generating a summary, and means for updating the performance schedule using the generated summary and registered events. This makes it possible to manage information obtained from calls in real time and to automatically update important schedules and task lists, enabling users to achieve efficient schedule management without being troubled by the complexity of the information.
[0289] "Call audio" refers to the human voice transmitted through communication devices.
[0290] "Digital audio data" refers to data obtained by converting an audio signal into a digital format, which is a format that can be processed by computers and digital devices.
[0291] "Speech recognition means" refers to a technology or device that analyzes speech signals and converts them into corresponding text information.
[0292] "Natural language processing means" refers to technologies or devices that analyze text data, understand human language, and extract information.
[0293] A "summary" is a document that condenses the original information and extracts its main points.
[0294] "Events" refer to information indicating schedules and tasks introduced by events such as phone calls.
[0295] "Scheduled execution" refers to tasks or events that are scheduled to be performed and should be registered in the management system.
[0296] A "storage device" is a device used to store digital data for long-term or short-term storage, and is used for information retention.
[0297] A "user" refers to an individual or organization that utilizes this system.
[0298] A description of the embodiment for carrying out the invention will be provided.
[0299] This system supports users' lives and work by acquiring and processing voice data in real time and automatically managing important schedules and task lists. The server receives call audio transmitted from the user's terminal as digital voice data. Next, it converts this voice data into text using speech recognition. Specifically, the Python package speech_recognition can be used for speech recognition.
[0300] The converted text is analyzed using natural language processing (NLP) tools to extract important information. This important information includes schedules and task details, which are automatically managed as performance targets. Since generative AI models may be used in the natural language processing, the server sends the text data to an NLP service in the cloud for advanced analysis.
[0301] Once a call is complete, users can receive notifications containing a generated summary and updated schedule information. This notification feature ensures that important information is not missed and allows for efficient management of appointments and tasks. Specifically, schedule information can be managed using Google Calendar or the Task API.
[0302] For example, if a user makes a call saying, "The next garbage collection will be next Monday at 10 AM," this information will be automatically registered in the calendar, and the user will be notified so they don't forget. In this way, by efficiently reflecting the content of calls in the schedule, the system supports the user's life management.
[0303] Examples of prompts for a generative AI model include the following:
[0304] Please summarize the call content as follows: 'The next garbage collection will be carried out at 10:00 am next Monday.'
[0305] The flow of the specific process in Application Example 1 will be described using Figure 12.
[0306] Step 1:
[0307] The user's terminal acquires the call voice and records it as digital voice data. This voice data is the input and is obtained by digitizing the voice signal using the terminal's microphone. The acquired digital voice data is in a state ready to be transmitted to the server.
[0308] Step 2:
[0309] The server converts the digital voice data received from the terminal into text data using voice recognition means. At this time, the input is digital voice data and the output is text data. Specifically, the speech_recognition library in Python is used to parse the voice signal and convert it into the corresponding text. The server passes this text data to the next processing step.
[0310] Step 3:
[0311] The server sends the converted text data to natural language processing means to extract important information. In this process, the input is the text data generated in Step 2 and the output is the extracted key phrases and important information. As a specific operation, a generative AI model is used to identify the keywords in the text and summarize the information.
[0312] Step 4:
[0313] Based on the extracted information, the server automatically registers tasks as planned in the task list and schedule. The input for this step is the important information obtained in step 3, and the output is the updated schedule data. Specifically, the Google Calendar API and Task API are used to reflect the appropriate date, time, and content in the user's calendar system.
[0314] Step 5:
[0315] After a call ends, the user receives a notification from the server containing a summary and updated schedule information. This notification is immediately available for the user to review, enabling them to utilize the information gathered during the call. The output is a visually displayed summary, specifically delivered as a push notification to the user's device.
[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0317] This invention provides a specific form of a system that processes a user's call audio in real time, simultaneously recording important information and analyzing emotions. The terminal captures the call audio as digital audio data and continuously transmits it to the server. The server utilizes speech recognition means to immediately convert the received audio data into text. This textualized data is analyzed using natural language processing technology to extract important information, and a summary is automatically generated.
[0318] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data and identify their emotional state. This emotional information is handled together with the extracted text information and, as needed, influences the content of the summary and the priority of the task list. It also enables improvements to interactions based on the user's emotions. For example, if the system detects that the user is feeling anxious during a call, it may provide information to alleviate that anxiety.
[0319] After the call ends, the user is notified of the generated summary, tasks, and schedule information, along with the sentiment analysis results obtained by the sentiment engine. This notification process allows the user to reflect on the call content and their own emotional state, enabling them to efficiently plan future actions.
[0320] As a concrete example, consider a business meeting where project progress is reported and scheduling is arranged. Using this system, the call content is recorded and analyzed in real time, and project deadlines and meeting dates are automatically incorporated into the schedule. Furthermore, if stress is detected in the speaker's voice, that element is detected by the emotion engine and reflected as information that helps improve communication. This makes it possible to carry out daily work more efficiently and smoothly.
[0321] The following describes the processing flow.
[0322] Step 1:
[0323] The device uses its microphone to capture audio as a digital signal as soon as a conversation begins, and sends this audio data to the server in real time.
[0324] Step 2:
[0325] The server inputs the received digital audio data into a speech recognition engine, which converts the spoken content into text. This process instantly transcribes the audio data, making it available for processing on the server.
[0326] Step 3:
[0327] The server passes the transcribed data to a natural language processing engine, which extracts important information and key phrases from the conversation. Furthermore, it generates a summary based on this information, concisely summarizing the content.
[0328] Step 4:
[0329] The emotion engine analyzes voice data in parallel, determining emotions based on the user's tone of voice, speaking speed, and other factors. This allows for real-time evaluation of emotional fluctuations during a call.
[0330] Step 5:
[0331] The server combines the extracted key information with sentiment analysis results to adjust the summaries, task lists, and schedule priorities. The sentiment analysis results are also recorded as information accompanying the call history.
[0332] Step 6:
[0333] After the call ends, the server notifies the user of the generated summary, tasks, schedule, and sentiment analysis results. This notification allows the user to review the call content and sentiment feedback.
[0334] Step 7:
[0335] Users can receive notifications, review call content, and reflect on their own emotional responses through sentiment analysis. They can also check their task lists and schedules, which are automatically managed, allowing them to plan their future work.
[0336] (Example 2)
[0337] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0338] In today's communication environment, there is a need to efficiently manage and immediately utilize important information obtained during calls. However, manually recording and organizing call content is laborious, and understanding and appropriately processing emotional nuances is difficult. Therefore, there is a need for a system that can simultaneously analyze call content and the user's emotional state and utilize the information efficiently.
[0339] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0340] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the audio data into text, and analysis means for analyzing the converted text and emotions to generate information. This enables simultaneous analysis of important information from the call content and the user's emotions, and automatic reflection of this information in task lists and schedules.
[0341] "Call audio" refers to voice information transmitted in real time using voice communication methods.
[0342] "Digital audio data" is a format of information obtained by digitizing analog audio signals.
[0343] "Speech recognition means" refers to a technology or device that converts speech data into text data.
[0344] "Natural language processing means" refers to technologies or processes that extract and understand structured information from text data.
[0345] "Analytical means" refer to the techniques and processes used to analyze data and extract new knowledge and information.
[0346] A "task list" is a means of organizing and displaying a user's activities and obligations in a list format.
[0347] A "time management chart" is a method for organizing and visualizing events and schedules along a timeline.
[0348] "Emotion" refers to the user's psychological state and includes technologies for recognizing or analyzing that state.
[0349] A "recording medium" is a physical or digital medium used to store and maintain data.
[0350] This invention is a system that processes a user's call audio as digital audio data in real time, simultaneously recording important information and analyzing emotions.
[0351] The terminal first acquires the user's call audio and converts it into digital audio data. This process uses an audio capture device with a microphone, and the resulting digital audio data is transmitted to the server via wireless communication.
[0352] The server uses speech recognition technology to convert the received audio data into text data. A speech recognition engine is used for this, with Google Cloud Speech-to-Text API being a specific example. The converted text is then analyzed using natural language processing technology to extract important information. SpaCy and NLTK libraries are used for this purpose.
[0353] Furthermore, the server uses a sentiment analysis engine to perform sentiment analysis based on text data. IBM Watson Tone Analyzer is one such sentiment engine that can be used. This identifies the user's emotional state, and this information is integrated with the text data. The analysis results can be automatically reflected in task lists and time management tables.
[0354] After the call ends, the user is notified of the generated summary, tasks, schedule information, and sentiment analysis results. The user can then use this information to plan their actions. This notification is delivered via an application installed on the user's mobile device or computer.
[0355] As a concrete example, consider a scenario where the content of business meetings is recorded and analyzed in real time, and project deadlines and meeting dates are automatically registered in the schedule. If stress levels of the speaker are detected, this information is provided as feedback to improve communication.
[0356] An example of a prompt message is, "Analyze the critical deadlines and stakeholders' emotional states in project progress reports, and create a summary to improve scheduling and communication efficiency." This allows the system to efficiently and accurately support the user's business operations.
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] The terminal acquires the user's call audio in real time. This involves capturing the audio as an analog signal using the microphone built into the terminal and converting it into digital audio data. The input is the user's voice, and the output is the digitized audio data. This audio data is transmitted to the server via the communication module.
[0360] Step 2:
[0361] The server processes the digital audio data received from the terminal. This processing includes converting the data to text using a speech recognition engine. Specifically, the server analyzes the audio spectrum, extracts phonemes, and transcribes them based on a language model. The input is digital audio data, and the output is text data containing character information.
[0362] Step 3:
[0363] The server analyzes the transcribed data using natural language processing techniques. This involves extracting important information and generating a text summary. For example, it performs morphological and syntactic analysis to extract nouns and verbs, and uses semantic analysis to grasp the main points of the utterance. The input is text data, and the output is a condensed text containing summary information and keywords.
[0364] Step 4:
[0365] The server analyzes the user's emotions from voice and text data. Using an emotion analysis engine, it evaluates, for example, the tone of voice and word choice to identify the user's emotional state. The input is the converted text and original voice data, and the output is data containing the type of emotion identified (joy, anger, etc.).
[0366] Step 5:
[0367] The server integrates extracted key information and sentiment analysis results, automatically reflecting them in the task list and schedule. Specifically, it works in conjunction with task management applications to update schedules and add tasks. The input is summary information and sentiment data, and the output is the updated task list and schedule information.
[0368] Step 6:
[0369] After the call ends, the user receives a summary, task list, schedule information, and sentiment analysis results from the server. This information is notified to the user's device, and the user can check its contents via email or application. The output is notification information presented in a user-friendly format.
[0370] (Application Example 2)
[0371] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0372] In modern society, it has become increasingly difficult for families and individuals to efficiently manage important information arising from daily conversations and to recognize changes in emotions. There is a high risk of overlooking important tasks and appointments by simply ignoring conversations. Furthermore, there is often a lack of appropriate interaction with family members experiencing stress or anxiety. These challenges are raising concerns about a decline in quality of life and impacting the well-being of families and individuals.
[0373] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0374] In this invention, the server includes means for acquiring call audio as digital voice numbers in real time, speech recognition means for converting the acquired voice numbers into text information, natural language processing means for extracting important information from the converted text information and generating a summary, means for performing emotion analysis from the call to identify the user's emotional state, and means for making suggestions and improvements to interactions using the emotion analysis results. This makes it possible for families and individuals to efficiently manage important information obtained from conversations, and to respond immediately to changes in emotions and provide appropriate interactions.
[0375] "Call audio" refers to the sound information of a linguistic conversation conducted via a voice communication device.
[0376] "Real-time" refers to a state where processing or analysis is performed immediately at the moment an event occurs.
[0377] "Digital audio numerical data" refers to numerical information obtained by converting analog audio into numbers that can be processed by a computer.
[0378] A "speech recognition system" is a mechanism that performs technology to convert speech data into linguistic characters.
[0379] "Textual information" refers to data consisting of language and symbols, in a readable and writable format.
[0380] "Important matters" refer to information or events that deserve particular attention or should be given priority.
[0381] A "natural language processing system" is a system that analyzes language data and performs information extraction and summary generation.
[0382] "Emotional analysis" is the process of estimating emotional states from voice or text.
[0383] "User's emotional state" refers to the result of determining what emotional state a particular user is in.
[0384] "Suggesting and improving interactions" means providing methods and actions to improve the quality of communication with users.
[0385] The system for realizing this invention consists of a terminal and a server.
[0386] The terminal is equipped with a microphone for receiving voice communications and captures the voice during a call as digital voice data. This digital voice data is continuously transmitted to the server.
[0387] The server uses speech recognition software to perform speech recognition based on the received digital speech data. Specifically, the Python `speech_recognition` library is used. The recognized speech is converted into text information, and then the TextBlob natural language processing library is used to evaluate the importance of the information and generate a summary as needed.
[0388] In addition, the server runs an emotion analysis engine to extract emotional information from textual data. This uses the emotion analysis pipeline within the transformers library of the generative AI model, Hugging Face. This emotional information is used not only to visualize the user's emotional state but also to suggest and improve interactions. For example, if the server determines that the user is feeling anxious, it will make suggestions to alleviate that anxiety.
[0389] As a concrete example, a robot installed in a family's living room could record family conversations and register important matters as tasks. Furthermore, if it detects someone is stressed, it could play their favorite music to help them relax.
[0390] An example of a prompt message for the application of this system is: "Analyze family phone calls, identify the types of content that tend to cause stress if someone is stressed, and suggest ways to relax."
[0391] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0392] Step 1:
[0393] The terminal receives call audio and converts it into digital audio numbers. The input is analog audio obtained through a microphone, which is sampled and quantized to generate digital audio numbers. This converts the audio data into a format that can be processed by the server.
[0394] Step 2:
[0395] The terminal transmits the generated digital voice numerical data to the server in real time. The input at this stage is the digital voice numerical data generated in step 1, which is sent to the server via network communication. The server then receives the data to proceed to the next processing step.
[0396] Step 3:
[0397] The server analyzes the received digital speech data using speech recognition software and converts it into text. The input is the digital speech data received by the server; speech recognition is performed using the speech_recognition library, and text is obtained as output. This text is essential for subsequent natural language processing.
[0398] Step 4:
[0399] The server analyzes the converted text information using natural language processing tools, extracts important information, and generates a summary as needed. The input is the text information obtained in step 3, and language analysis is performed using TextBlob. The output is a list of important information and the generated summary, and this information contributes to improving work efficiency.
[0400] Step 5:
[0401] The server performs sentiment analysis from textual information to identify the user's emotional state. The input is the textual information from step 3, and the Hugging Face transformers library is used for sentiment analysis. The output is the sentiment identification result, providing information that reflects the user's psychological state.
[0402] Step 6:
[0403] The server makes suggestions and improvements to interactions based on the sentiment analysis results. The input is the output from steps 4 and 5, and appropriate suggestions are made to the user according to the identified key points and emotional state. This improves the user experience and enables more comfortable communication.
[0404] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0405] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0406] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0407] [Third Embodiment]
[0408] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0409] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0410] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0411] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0412] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0414] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0415] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0416] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0417] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0418] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0419] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0420] This invention relates to a system that efficiently and accurately records call content, extracts necessary information, and improves the user's work efficiency. Specifically, a terminal captures the call's audio signal as digital audio data. This data is then transmitted to a server. The server analyzes the captured audio data using speech recognition technology and converts it into text in real time. This text is then analyzed using natural language processing technology to extract important information and key phrases.
[0421] Furthermore, the server generates a summary based on the extracted information and stores it in a database. This summary is retained in the call history in a format that allows the user to easily review its contents. The server also automatically registers items in task lists and schedules based on specific key phrases. For example, if the call content recognizes "Let's have a meeting next Tuesday," it will be automatically registered in the corresponding schedule. After the call ends, the server notifies the user of the generated summary and registered schedule / task information, allowing the user to quickly review and use the information.
[0422] As a concrete example, suppose a user is on a business call where they discuss deadlines and meeting dates for a new project. Using this system, the call is analyzed in real time, and important project deadlines and upcoming meetings are automatically registered in the user's scheduling application. After the call ends, the user can review the summary and easily verify that the schedule has been accurately registered. This eliminates the need for manual note-taking and data entry, dramatically improving the efficiency of business processes.
[0423] The following describes the processing flow.
[0424] Step 1:
[0425] As soon as a call begins, the device acquires audio through the microphone and formats it as digital audio data in real time.
[0426] Step 2:
[0427] The terminal structures the captured digital audio data into packets and sends them to the server over the network. The data is transmitted in segments of a fixed buffer size.
[0428] Step 3:
[0429] The server inputs the received audio data into the speech recognition engine, which extracts linguistic information from the audio and converts it into text. This process is performed in real time, and the conversion results are temporarily stored in memory.
[0430] Step 4:
[0431] The server analyzes the text generated using natural language processing technology and extracts important information and key phrases. During this process, it prioritizes the information based on pre-configured business rules and context.
[0432] Step 5:
[0433] The server generates a short, concise summary based on the analysis results and records this summary in the call history within the database.
[0434] Step 6:
[0435] The server automatically registers information for schedules and task lists derived from data analysis into other management systems using APIs.
[0436] Step 7:
[0437] Once the call ends, the server notifies the user of the generated summary and any registered tasks or schedules. The user can then review this information to quickly understand the call content and their next steps.
[0438] (Example 1)
[0439] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0440] Many modern communication systems struggle to efficiently and accurately record call content and automatically organize and manage that information. Especially in business meetings and long calls, important information can be lost, or manual note-taking may be necessary. Therefore, users face the challenge of not being able to accurately refer to the information obtained during a call and use it to inform their next actions.
[0441] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0442] In this invention, the server includes means for acquiring call acoustics as digital acoustic information in real time, acoustic recognition means for converting the acquired acoustic information into text data, and natural language processing means for extracting key information from the converted text data and generating essential points. This enables efficient recording of call information, automatic organization of necessary information, and allows users to quickly review and utilize the information after a call.
[0443] "Call acoustics" refers to the acoustic signals generated during a voice call, and is acquired during conversations conducted via telephone or online platforms.
[0444] "Digital audio information" refers to data obtained by quantifying analog audio signals and converting them into a format that can be processed by computers and digital devices.
[0445] "Acoustic recognition means" refers to technologies and devices that analyze acoustic signals and automatically convert them into meaningful character data.
[0446] "Character data" refers to data in text format that is converted from speech by an acoustic recognition system.
[0447] "Key information" refers to information or key phrases extracted from text data that users should consider important.
[0448] "Key points" refer to content that efficiently summarizes and concisely describes the main information, provided in a way that users can quickly understand and use.
[0449] "Natural language processing means" refers to technologies and devices for analyzing text data and extracting or summarizing information.
[0450] "Automatic registration" refers to the process of adding items necessary for task lists and schedule management based on information extracted by the system, without manual intervention.
[0451] An "information storage device" refers to a recording medium or system for storing digital information long-term and making it accessible later.
[0452] "Information transmission network" refers to the network environment and communication technology used to deliver digital acoustic information to acoustic recognition devices.
[0453] "Users" refer to individuals or organizations that record call content or extract information through this system.
[0454] This invention is a system that acquires call acoustics as digital acoustic information in real time and accurately analyzes and utilizes that information. The implementation of this invention involves the terminal, server, and user.
[0455] The terminal captures the audio signal during a voice call as digital audio information using its microphone. For example, it converts the audio signal into PCM digital data using the built-in microphone of a smartphone or PC, or an externally connected microphone. The terminal then transmits this digital audio information to a server via an internet connection. In this process, the internet is used as the information transmission network, and the SSL / TLS protocol is used to encrypt the data.
[0456] The server converts received digital audio information into text data using acoustic recognition technology. This technology can utilize commercially available acoustic recognition software or cloud-based audio services. After conversion to text data, the server analyzes this data using natural language processing technology to extract key information. Natural language processing technology includes word segmentation, syntactic analysis, and sentiment analysis, and can utilize commercially available natural language processing software or cloud services.
[0457] Based on the extracted key information, the server generates summaries using a generative AI model. During this process, the prompt "Summarize the following text." is used to input the analyzed text into the AI model. The generated summaries and related information are stored in an information storage device, allowing users to search and review them at a later date.
[0458] Based on registered important information, the server automatically registers tasks in the work list and schedule. Specifically, this includes recognizing key phrases such as "meeting" and "deadline" and generating events in the schedule management application. This allows users to quickly review and utilize this information even after a call has ended, reducing the burden of manual note-taking and data entry.
[0459] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0460] Step 1:
[0461] The device uses a microphone to capture the audio signal during a call as digital audio information. This can be done using the built-in microphone of a smartphone or computer, or an externally connected microphone. The input is an analog audio signal, which is converted into PCM format digital data using an ADC (analog-to-digital converter). The output is the digitized audio data.
[0462] Step 2:
[0463] The terminal transmits digital audio data to the server. During this process, the terminal encrypts the data using the SSL / TLS protocol and streams it to the server in real time over the internet. The input is digital audio data, and the output is a secure stream of audio data.
[0464] Step 3:
[0465] The server converts the received audio data into text data using acoustic recognition technology. Here, commercially available acoustic recognition software is used to analyze the audio. The input is streamed digital audio data, and the output is text data.
[0466] Step 4:
[0467] The server analyzes text data using natural language processing techniques and extracts key information. It utilizes tools such as the Google Cloud Natural Language API for word segmentation and information extraction. The input is text data, and the output is an analysis result containing key information and key phrases.
[0468] Step 5:
[0469] The server generates key points based on a generative AI model using the extracted main information. Using OpenAI GPT-3 or similar tools, the server receives parsed text as input via the prompt "Summarize the following text." The output is summarized data.
[0470] Step 6:
[0471] The server stores the generated summaries in an information storage device to facilitate later searching and verification. A database system (e.g., MongoDB or MySQL) is used here. The input is summary data, and the output is the summary information recorded in the database.
[0472] Step 7:
[0473] The server automatically registers tasks and appointments based on important key phrases. It uses the Google Calendar API and other tools to add events related to the key phrases to the schedule. The input is the parsed key information, and the output is the new events added to the schedule data.
[0474] Step 8:
[0475] After the call ends, the server notifies the user of the generated summary and registered schedule information. Email and push notifications are used for this notification. The input is the summary and schedule information, and the output is the notification data sent to the user.
[0476] (Application Example 1)
[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0478] In today's information-saturated world, efficiently managing information obtained through daily and work-related phone calls and ensuring that important dates and tasks are not missed is a challenging task. Especially in situations requiring real-time information processing and automation, manual management has its limitations. This can lead to the leakage of important information, potentially hindering schedule management and task completion. Therefore, a system is needed that can instantly analyze call content and accurately process and manage important information.
[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0480] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the acquired audio data into text, natural language processing means for extracting important information from the converted text and generating a summary, and means for updating the performance schedule using the generated summary and registered events. This makes it possible to manage information obtained from calls in real time and to automatically update important schedules and task lists, enabling users to achieve efficient schedule management without being troubled by the complexity of the information.
[0481] "Call audio" refers to the human voice transmitted through communication devices.
[0482] "Digital audio data" refers to data obtained by converting an audio signal into a digital format, which is a format that can be processed by computers and digital devices.
[0483] "Speech recognition means" refers to a technology or device that analyzes speech signals and converts them into corresponding text information.
[0484] "Natural language processing means" refers to technologies or devices that analyze text data, understand human language, and extract information.
[0485] A "summary" is a document that condenses the original information and extracts its main points.
[0486] "Events" refer to information indicating schedules and tasks introduced by events such as phone calls.
[0487] "Scheduled execution" refers to tasks or events that are scheduled to be performed and should be registered in the management system.
[0488] A "storage device" is a device used to store digital data for long-term or short-term storage, and is used for information retention.
[0489] A "user" refers to an individual or organization that utilizes this system.
[0490] A description of the embodiment for carrying out the invention will be provided.
[0491] This system supports users' lives and work by acquiring and processing voice data in real time and automatically managing important schedules and task lists. The server receives call audio transmitted from the user's terminal as digital voice data. Next, it converts this voice data into text using speech recognition. Specifically, the Python package speech_recognition can be used for speech recognition.
[0492] The converted text is analyzed using natural language processing (NLP) tools to extract important information. This important information includes schedules and task details, which are automatically managed as performance targets. Since generative AI models may be used in the natural language processing, the server sends the text data to an NLP service in the cloud for advanced analysis.
[0493] Once a call is complete, users can receive notifications containing a generated summary and updated schedule information. This notification feature ensures that important information is not missed and allows for efficient management of appointments and tasks. Specifically, schedule information can be managed using Google Calendar or the Task API.
[0494] For example, if a user makes a call saying, "The next garbage collection will be next Monday at 10 AM," this information will be automatically registered in the calendar, and the user will be notified so they don't forget. In this way, by efficiently reflecting the content of calls in the schedule, the system supports the user's life management.
[0495] Examples of prompts for a generative AI model include the following:
[0496] Please summarize the conversation as follows: "The next garbage collection will take place next Monday at 10:00 AM."
[0497] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0498] Step 1:
[0499] The user's device acquires the call audio and records it as digital audio data. This audio data is the input, obtained by digitizing the audio signal using the device's microphone. The acquired digital audio data is then ready to be sent to the server.
[0500] Step 2:
[0501] The server converts digital audio data received from the terminal into text data using speech recognition. In this process, the input is digital audio data, and the output is text data. Specifically, the Python `speech_recognition` library is used to parse the audio signal and convert it into corresponding text. The server then passes this text data to the next processing step.
[0502] Step 3:
[0503] The server sends the converted text data to a natural language processing system for extracting important information. In this process, the input is the text data generated in step 2, and the output is the extracted key phrases and important information. Specifically, a generative AI model is used to identify keywords within the text and summarize that information.
[0504] Step 4:
[0505] Based on the extracted information, the server automatically registers tasks as planned in the task list and schedule. The input for this step is the important information obtained in step 3, and the output is the updated schedule data. Specifically, the Google Calendar API and Task API are used to reflect the appropriate date, time, and content in the user's calendar system.
[0506] Step 5:
[0507] After a call ends, the user receives a notification from the server containing a summary and updated schedule information. This notification is immediately available for the user to review, enabling them to utilize the information gathered during the call. The output is a visually displayed summary, specifically delivered as a push notification to the user's device.
[0508] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0509] This invention provides a specific form of a system that processes a user's call audio in real time, simultaneously recording important information and analyzing emotions. The terminal captures the call audio as digital audio data and continuously transmits it to the server. The server utilizes speech recognition means to immediately convert the received audio data into text. This textualized data is analyzed using natural language processing technology to extract important information, and a summary is automatically generated.
[0510] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data and identify their emotional state. This emotional information is handled together with the extracted text information and, as needed, influences the content of the summary and the priority of the task list. It also enables improvements to interactions based on the user's emotions. For example, if the system detects that the user is feeling anxious during a call, it may provide information to alleviate that anxiety.
[0511] After the call ends, the user is notified of the generated summary, tasks, and schedule information, along with the sentiment analysis results obtained by the sentiment engine. This notification process allows the user to reflect on the call content and their own emotional state, enabling them to efficiently plan future actions.
[0512] As a concrete example, consider a business meeting where project progress is reported and scheduling is arranged. Using this system, the call content is recorded and analyzed in real time, and project deadlines and meeting dates are automatically incorporated into the schedule. Furthermore, if stress is detected in the speaker's voice, that element is detected by the emotion engine and reflected as information that helps improve communication. This makes it possible to carry out daily work more efficiently and smoothly.
[0513] The following describes the processing flow.
[0514] Step 1:
[0515] The device uses its microphone to capture audio as a digital signal as soon as a conversation begins, and sends this audio data to the server in real time.
[0516] Step 2:
[0517] The server inputs the received digital audio data into a speech recognition engine, which converts the spoken content into text. This process instantly transcribes the audio data, making it available for processing on the server.
[0518] Step 3:
[0519] The server passes the transcribed data to a natural language processing engine, which extracts important information and key phrases from the conversation. Furthermore, it generates a summary based on this information, concisely summarizing the content.
[0520] Step 4:
[0521] The emotion engine analyzes voice data in parallel, determining emotions based on the user's tone of voice, speaking speed, and other factors. This allows for real-time evaluation of emotional fluctuations during a call.
[0522] Step 5:
[0523] The server combines the extracted key information with sentiment analysis results to adjust the summaries, task lists, and schedule priorities. The sentiment analysis results are also recorded as information accompanying the call history.
[0524] Step 6:
[0525] After the call ends, the server notifies the user of the generated summary, tasks, schedule, and sentiment analysis results. This notification allows the user to review the call content and sentiment feedback.
[0526] Step 7:
[0527] Users can receive notifications, review call content, and reflect on their own emotional responses through sentiment analysis. They can also check their task lists and schedules, which are automatically managed, allowing them to plan their future work.
[0528] (Example 2)
[0529] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0530] In today's communication environment, there is a need to efficiently manage and immediately utilize important information obtained during calls. However, manually recording and organizing call content is laborious, and understanding and appropriately processing emotional nuances is difficult. Therefore, there is a need for a system that can simultaneously analyze call content and the user's emotional state and utilize the information efficiently.
[0531] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0532] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the audio data into text, and analysis means for analyzing the converted text and emotions to generate information. This enables simultaneous analysis of important information from the call content and the user's emotions, and automatic reflection of this information in task lists and schedules.
[0533] "Call audio" refers to voice information transmitted in real time using voice communication methods.
[0534] "Digital audio data" is a format of information obtained by digitizing analog audio signals.
[0535] "Speech recognition means" refers to a technology or device that converts speech data into text data.
[0536] "Natural language processing means" refers to technologies or processes that extract and understand structured information from text data.
[0537] "Analytical means" refer to the techniques and processes used to analyze data and extract new knowledge and information.
[0538] A "task list" is a means of organizing and displaying a user's activities and obligations in a list format.
[0539] A "time management chart" is a method for organizing and visualizing events and schedules along a timeline.
[0540] "Emotion" refers to the user's psychological state and includes technologies for recognizing or analyzing that state.
[0541] A "recording medium" is a physical or digital medium used to store and maintain data.
[0542] This invention is a system that processes a user's call audio as digital audio data in real time, simultaneously recording important information and analyzing emotions.
[0543] The terminal first acquires the user's call audio and converts it into digital audio data. This process uses an audio capture device with a microphone, and the resulting digital audio data is transmitted to the server via wireless communication.
[0544] The server uses speech recognition technology to convert the received audio data into text data. A speech recognition engine is used for this, with Google Cloud Speech-to-Text API being a specific example. The converted text is then analyzed using natural language processing technology to extract important information. SpaCy and NLTK libraries are used for this purpose.
[0545] Furthermore, the server uses a sentiment analysis engine to perform sentiment analysis based on text data. IBM Watson Tone Analyzer is one such sentiment engine that can be used. This identifies the user's emotional state, and this information is integrated with the text data. The analysis results can be automatically reflected in task lists and time management tables.
[0546] After the call ends, the user is notified of the generated summary, tasks, schedule information, and sentiment analysis results. The user can then use this information to plan their actions. This notification is delivered via an application installed on the user's mobile device or computer.
[0547] As a concrete example, consider a scenario where the content of business meetings is recorded and analyzed in real time, and project deadlines and meeting dates are automatically registered in the schedule. If stress levels of the speaker are detected, this information is provided as feedback to improve communication.
[0548] An example of a prompt message is, "Analyze the critical deadlines and stakeholders' emotional states in project progress reports, and create a summary to improve scheduling and communication efficiency." This allows the system to efficiently and accurately support the user's business operations.
[0549] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0550] Step 1:
[0551] The terminal acquires the user's call audio in real time. This involves capturing the audio as an analog signal using the microphone built into the terminal and converting it into digital audio data. The input is the user's voice, and the output is the digitized audio data. This audio data is transmitted to the server via the communication module.
[0552] Step 2:
[0553] The server processes the digital audio data received from the terminal. This processing includes converting the data to text using a speech recognition engine. Specifically, the server analyzes the audio spectrum, extracts phonemes, and transcribes them based on a language model. The input is digital audio data, and the output is text data containing character information.
[0554] Step 3:
[0555] The server analyzes the transcribed data using natural language processing techniques. This involves extracting important information and generating a text summary. For example, it performs morphological and syntactic analysis to extract nouns and verbs, and uses semantic analysis to grasp the main points of the utterance. The input is text data, and the output is a condensed text containing summary information and keywords.
[0556] Step 4:
[0557] The server analyzes the user's emotions from voice and text data. Using an emotion analysis engine, it evaluates, for example, the tone of voice and word choice to identify the user's emotional state. The input is the converted text and original voice data, and the output is data containing the type of emotion identified (joy, anger, etc.).
[0558] Step 5:
[0559] The server integrates extracted key information and sentiment analysis results, automatically reflecting them in the task list and schedule. Specifically, it works in conjunction with task management applications to update schedules and add tasks. The input is summary information and sentiment data, and the output is the updated task list and schedule information.
[0560] Step 6:
[0561] After the call ends, the user receives a summary, task list, schedule information, and sentiment analysis results from the server. This information is notified to the user's device, and the user can check its contents via email or application. The output is notification information presented in a user-friendly format.
[0562] (Application Example 2)
[0563] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0564] In modern society, it has become increasingly difficult for families and individuals to efficiently manage important information arising from daily conversations and to recognize changes in emotions. There is a high risk of overlooking important tasks and appointments by simply ignoring conversations. Furthermore, there is often a lack of appropriate interaction with family members experiencing stress or anxiety. These challenges are raising concerns about a decline in quality of life and impacting the well-being of families and individuals.
[0565] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0566] In this invention, the server includes means for acquiring call audio as digital voice numbers in real time, speech recognition means for converting the acquired voice numbers into text information, natural language processing means for extracting important information from the converted text information and generating a summary, means for performing emotion analysis from the call to identify the user's emotional state, and means for making suggestions and improvements to interactions using the emotion analysis results. This makes it possible for families and individuals to efficiently manage important information obtained from conversations, and to respond immediately to changes in emotions and provide appropriate interactions.
[0567] "Call audio" refers to the sound information of a linguistic conversation conducted via a voice communication device.
[0568] "Real-time" refers to a state where processing or analysis is performed immediately at the moment an event occurs.
[0569] "Digital audio numerical data" refers to numerical information obtained by converting analog audio into numbers that can be processed by a computer.
[0570] A "speech recognition system" is a mechanism that performs technology to convert speech data into linguistic characters.
[0571] "Textual information" refers to data consisting of language and symbols, in a readable and writable format.
[0572] "Important matters" refer to information or events that deserve particular attention or should be given priority.
[0573] A "natural language processing system" is a system that analyzes language data and performs information extraction and summary generation.
[0574] "Emotional analysis" is the process of estimating emotional states from voice or text.
[0575] "User's emotional state" refers to the result of determining what emotional state a particular user is in.
[0576] "Suggesting and improving interactions" means providing methods and actions to improve the quality of communication with users.
[0577] The system for realizing this invention consists of a terminal and a server.
[0578] The terminal is equipped with a microphone for receiving voice communications and captures the voice during a call as digital voice data. This digital voice data is continuously transmitted to the server.
[0579] The server uses speech recognition software to perform speech recognition based on the received digital speech data. Specifically, the Python `speech_recognition` library is used. The recognized speech is converted into text information, and then the TextBlob natural language processing library is used to evaluate the importance of the information and generate a summary as needed.
[0580] In addition, the server runs an emotion analysis engine to extract emotional information from textual data. This uses the emotion analysis pipeline within the transformers library of the generative AI model, Hugging Face. This emotional information is used not only to visualize the user's emotional state but also to suggest and improve interactions. For example, if the server determines that the user is feeling anxious, it will make suggestions to alleviate that anxiety.
[0581] As a concrete example, a robot installed in a family's living room could record family conversations and register important matters as tasks. Furthermore, if it detects someone is stressed, it could play their favorite music to help them relax.
[0582] An example of a prompt message for the application of this system is: "Analyze family phone calls, identify the types of content that tend to cause stress if someone is stressed, and suggest ways to relax."
[0583] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0584] Step 1:
[0585] The terminal receives call audio and converts it into digital audio numbers. The input is analog audio obtained through a microphone, which is sampled and quantized to generate digital audio numbers. This converts the audio data into a format that can be processed by the server.
[0586] Step 2:
[0587] The terminal transmits the generated digital voice numerical data to the server in real time. The input at this stage is the digital voice numerical data generated in step 1, which is sent to the server via network communication. The server then receives the data to proceed to the next processing step.
[0588] Step 3:
[0589] The server analyzes the received digital speech data using speech recognition software and converts it into text. The input is the digital speech data received by the server; speech recognition is performed using the speech_recognition library, and text is obtained as output. This text is essential for subsequent natural language processing.
[0590] Step 4:
[0591] The server analyzes the converted text information using natural language processing tools, extracts important information, and generates a summary as needed. The input is the text information obtained in step 3, and language analysis is performed using TextBlob. The output is a list of important information and the generated summary, and this information contributes to improving work efficiency.
[0592] Step 5:
[0593] The server performs sentiment analysis from textual information to identify the user's emotional state. The input is the textual information from step 3, and the Hugging Face transformers library is used for sentiment analysis. The output is the sentiment identification result, providing information that reflects the user's psychological state.
[0594] Step 6:
[0595] The server makes suggestions and improvements to interactions based on the sentiment analysis results. The input is the output from steps 4 and 5, and appropriate suggestions are made to the user according to the identified key points and emotional state. This improves the user experience and enables more comfortable communication.
[0596] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0597] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0598] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0599] [Fourth Embodiment]
[0600] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0601] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0602] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0603] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0604] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0605] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0606] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0607] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0608] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0609] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0610] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0611] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0612] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0613] This invention relates to a system that efficiently and accurately records call content, extracts necessary information, and improves the user's work efficiency. Specifically, a terminal captures the call's audio signal as digital audio data. This data is then transmitted to a server. The server analyzes the captured audio data using speech recognition technology and converts it into text in real time. This text is then analyzed using natural language processing technology to extract important information and key phrases.
[0614] Furthermore, the server generates a summary based on the extracted information and stores it in a database. This summary is retained in the call history in a format that allows the user to easily review its contents. The server also automatically registers items in task lists and schedules based on specific key phrases. For example, if the call content recognizes "Let's have a meeting next Tuesday," it will be automatically registered in the corresponding schedule. After the call ends, the server notifies the user of the generated summary and registered schedule / task information, allowing the user to quickly review and use the information.
[0615] As a concrete example, suppose a user is on a business call where they discuss deadlines and meeting dates for a new project. Using this system, the call is analyzed in real time, and important project deadlines and upcoming meetings are automatically registered in the user's scheduling application. After the call ends, the user can review the summary and easily verify that the schedule has been accurately registered. This eliminates the need for manual note-taking and data entry, dramatically improving the efficiency of business processes.
[0616] The following describes the processing flow.
[0617] Step 1:
[0618] As soon as a call begins, the device acquires audio through the microphone and formats it as digital audio data in real time.
[0619] Step 2:
[0620] The terminal structures the captured digital audio data into packets and sends them to the server over the network. The data is transmitted in segments of a fixed buffer size.
[0621] Step 3:
[0622] The server inputs the received audio data into the speech recognition engine, which extracts linguistic information from the audio and converts it into text. This process is performed in real time, and the conversion results are temporarily stored in memory.
[0623] Step 4:
[0624] The server analyzes the text generated using natural language processing technology and extracts important information and key phrases. During this process, it prioritizes the information based on pre-configured business rules and context.
[0625] Step 5:
[0626] The server generates a short, concise summary based on the analysis results and records this summary in the call history within the database.
[0627] Step 6:
[0628] The server automatically registers information for schedules and task lists derived from data analysis into other management systems using APIs.
[0629] Step 7:
[0630] Once the call ends, the server notifies the user of the generated summary and any registered tasks or schedules. The user can then review this information to quickly understand the call content and their next steps.
[0631] (Example 1)
[0632] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0633] Many modern communication systems struggle to efficiently and accurately record call content and automatically organize and manage that information. Especially in business meetings and long calls, important information can be lost, or manual note-taking may be necessary. Therefore, users face the challenge of not being able to accurately refer to the information obtained during a call and use it to inform their next actions.
[0634] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0635] In this invention, the server includes means for acquiring call acoustics as digital acoustic information in real time, acoustic recognition means for converting the acquired acoustic information into text data, and natural language processing means for extracting key information from the converted text data and generating essential points. This enables efficient recording of call information, automatic organization of necessary information, and allows users to quickly review and utilize the information after a call.
[0636] "Call acoustics" refers to the acoustic signals generated during a voice call, and is acquired during conversations conducted via telephone or online platforms.
[0637] "Digital audio information" refers to data obtained by quantifying analog audio signals and converting them into a format that can be processed by computers and digital devices.
[0638] "Acoustic recognition means" refers to technologies and devices that analyze acoustic signals and automatically convert them into meaningful character data.
[0639] "Character data" refers to data in text format that is converted from speech by an acoustic recognition system.
[0640] "Key information" refers to information or key phrases extracted from text data that users should consider important.
[0641] "Key points" refer to content that efficiently summarizes and concisely describes the main information, provided in a way that users can quickly understand and use.
[0642] "Natural language processing means" refers to technologies and devices for analyzing text data and extracting or summarizing information.
[0643] "Automatic registration" refers to the process of adding items necessary for task lists and schedule management based on information extracted by the system, without manual intervention.
[0644] An "information storage device" refers to a recording medium or system for storing digital information long-term and making it accessible later.
[0645] "Information transmission network" refers to the network environment and communication technology used to deliver digital acoustic information to acoustic recognition devices.
[0646] "Users" refer to individuals or organizations that record call content or extract information through this system.
[0647] This invention is a system that acquires call acoustics as digital acoustic information in real time and accurately analyzes and utilizes that information. The implementation of this invention involves the terminal, server, and user.
[0648] The terminal captures the audio signal during a voice call as digital audio information using its microphone. For example, it converts the audio signal into PCM digital data using the built-in microphone of a smartphone or PC, or an externally connected microphone. The terminal then transmits this digital audio information to a server via an internet connection. In this process, the internet is used as the information transmission network, and the SSL / TLS protocol is used to encrypt the data.
[0649] The server converts received digital audio information into text data using acoustic recognition technology. This technology can utilize commercially available acoustic recognition software or cloud-based audio services. After conversion to text data, the server analyzes this data using natural language processing technology to extract key information. Natural language processing technology includes word segmentation, syntactic analysis, and sentiment analysis, and can utilize commercially available natural language processing software or cloud services.
[0650] Based on the extracted key information, the server generates summaries using a generative AI model. During this process, the prompt "Summarize the following text." is used to input the analyzed text into the AI model. The generated summaries and related information are stored in an information storage device, allowing users to search and review them at a later date.
[0651] Based on registered important information, the server automatically registers tasks in the work list and schedule. Specifically, this includes recognizing key phrases such as "meeting" and "deadline" and generating events in the schedule management application. This allows users to quickly review and utilize this information even after a call has ended, reducing the burden of manual note-taking and data entry.
[0652] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0653] Step 1:
[0654] The device uses a microphone to capture the audio signal during a call as digital audio information. This can be done using the built-in microphone of a smartphone or computer, or an externally connected microphone. The input is an analog audio signal, which is converted into PCM format digital data using an ADC (analog-to-digital converter). The output is the digitized audio data.
[0655] Step 2:
[0656] The terminal transmits digital audio data to the server. During this process, the terminal encrypts the data using the SSL / TLS protocol and streams it to the server in real time over the internet. The input is digital audio data, and the output is a secure stream of audio data.
[0657] Step 3:
[0658] The server converts the received audio data into text data using acoustic recognition technology. Here, commercially available acoustic recognition software is used to analyze the audio. The input is streamed digital audio data, and the output is text data.
[0659] Step 4:
[0660] The server analyzes text data using natural language processing techniques and extracts key information. It utilizes tools such as the Google Cloud Natural Language API for word segmentation and information extraction. The input is text data, and the output is an analysis result containing key information and key phrases.
[0661] Step 5:
[0662] The server generates key points based on a generative AI model using the extracted main information. Using OpenAI GPT-3 or similar tools, the server receives parsed text as input via the prompt "Summarize the following text." The output is summarized data.
[0663] Step 6:
[0664] The server stores the generated summaries in an information storage device to facilitate later searching and verification. A database system (e.g., MongoDB or MySQL) is used here. The input is summary data, and the output is the summary information recorded in the database.
[0665] Step 7:
[0666] The server automatically registers tasks and appointments based on important key phrases. It uses the Google Calendar API and other tools to add events related to the key phrases to the schedule. The input is the parsed key information, and the output is the new events added to the schedule data.
[0667] Step 8:
[0668] After the call ends, the server notifies the user of the generated summary and registered schedule information. Email and push notifications are used for this notification. The input is the summary and schedule information, and the output is the notification data sent to the user.
[0669] (Application Example 1)
[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] In today's information-saturated world, efficiently managing information obtained through daily and work-related phone calls and ensuring that important dates and tasks are not missed is a challenging task. Especially in situations requiring real-time information processing and automation, manual management has its limitations. This can lead to the leakage of important information, potentially hindering schedule management and task completion. Therefore, a system is needed that can instantly analyze call content and accurately process and manage important information.
[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0673] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the acquired audio data into text, natural language processing means for extracting important information from the converted text and generating a summary, and means for updating the performance schedule using the generated summary and registered events. This makes it possible to manage information obtained from calls in real time and to automatically update important schedules and task lists, enabling users to achieve efficient schedule management without being troubled by the complexity of the information.
[0674] "Call audio" refers to the human voice transmitted through communication devices.
[0675] "Digital audio data" refers to data obtained by converting an audio signal into a digital format, which is a format that can be processed by computers and digital devices.
[0676] "Speech recognition means" refers to a technology or device that analyzes speech signals and converts them into corresponding text information.
[0677] "Natural language processing means" refers to technologies or devices that analyze text data, understand human language, and extract information.
[0678] A "summary" is a document that condenses the original information and extracts its main points.
[0679] "Events" refer to information indicating schedules and tasks introduced by events such as phone calls.
[0680] "Scheduled execution" refers to tasks or events that are scheduled to be performed and should be registered in the management system.
[0681] A "storage device" is a device used to store digital data for long-term or short-term storage, and is used for information retention.
[0682] A "user" refers to an individual or organization that utilizes this system.
[0683] A description of the embodiment for carrying out the invention will be provided.
[0684] This system supports users' lives and work by acquiring and processing voice data in real time and automatically managing important schedules and task lists. The server receives call audio transmitted from the user's terminal as digital voice data. Next, it converts this voice data into text using speech recognition. Specifically, the Python package speech_recognition can be used for speech recognition.
[0685] The converted text is analyzed using natural language processing (NLP) tools to extract important information. This important information includes schedules and task details, which are automatically managed as performance targets. Since generative AI models may be used in the natural language processing, the server sends the text data to an NLP service in the cloud for advanced analysis.
[0686] Once a call is complete, users can receive notifications containing a generated summary and updated schedule information. This notification feature ensures that important information is not missed and allows for efficient management of appointments and tasks. Specifically, schedule information can be managed using Google Calendar or the Task API.
[0687] For example, if a user makes a call saying, "The next garbage collection will be next Monday at 10 AM," this information will be automatically registered in the calendar, and the user will be notified so they don't forget. In this way, by efficiently reflecting the content of calls in the schedule, the system supports the user's life management.
[0688] Examples of prompts for a generative AI model include the following:
[0689] Please summarize the conversation as follows: "The next garbage collection will take place next Monday at 10:00 AM."
[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0691] Step 1:
[0692] The user's device acquires the call audio and records it as digital audio data. This audio data is the input, obtained by digitizing the audio signal using the device's microphone. The acquired digital audio data is then ready to be sent to the server.
[0693] Step 2:
[0694] The server converts digital audio data received from the terminal into text data using speech recognition. In this process, the input is digital audio data, and the output is text data. Specifically, the Python `speech_recognition` library is used to parse the audio signal and convert it into corresponding text. The server then passes this text data to the next processing step.
[0695] Step 3:
[0696] The server sends the converted text data to a natural language processing system for extracting important information. In this process, the input is the text data generated in step 2, and the output is the extracted key phrases and important information. Specifically, a generative AI model is used to identify keywords within the text and summarize that information.
[0697] Step 4:
[0698] Based on the extracted information, the server automatically registers tasks as planned in the task list and schedule. The input for this step is the important information obtained in step 3, and the output is the updated schedule data. Specifically, the Google Calendar API and Task API are used to reflect the appropriate date, time, and content in the user's calendar system.
[0699] Step 5:
[0700] After a call ends, the user receives a notification from the server containing a summary and updated schedule information. This notification is immediately available for the user to review, enabling them to utilize the information gathered during the call. The output is a visually displayed summary, specifically delivered as a push notification to the user's device.
[0701] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0702] This invention provides a specific form of a system that processes a user's call audio in real time, simultaneously recording important information and analyzing emotions. The terminal captures the call audio as digital audio data and continuously transmits it to the server. The server utilizes speech recognition means to immediately convert the received audio data into text. This textualized data is analyzed using natural language processing technology to extract important information, and a summary is automatically generated.
[0703] Furthermore, the server uses an emotion engine to analyze the user's emotions from the voice data and identify their emotional state. This emotional information is handled together with the extracted text information and, as needed, influences the content of the summary and the priority of the task list. It also enables improvements to interactions based on the user's emotions. For example, if the system detects that the user is feeling anxious during a call, it may provide information to alleviate that anxiety.
[0704] After the call ends, the user is notified of the generated summary, tasks, and schedule information, along with the sentiment analysis results obtained by the sentiment engine. This notification process allows the user to reflect on the call content and their own emotional state, enabling them to efficiently plan future actions.
[0705] As a concrete example, consider a business meeting where project progress is reported and scheduling is arranged. Using this system, the call content is recorded and analyzed in real time, and project deadlines and meeting dates are automatically incorporated into the schedule. Furthermore, if stress is detected in the speaker's voice, that element is detected by the emotion engine and reflected as information that helps improve communication. This makes it possible to carry out daily work more efficiently and smoothly.
[0706] The following describes the processing flow.
[0707] Step 1:
[0708] The device uses its microphone to capture audio as a digital signal as soon as a conversation begins, and sends this audio data to the server in real time.
[0709] Step 2:
[0710] The server inputs the received digital audio data into a speech recognition engine, which converts the spoken content into text. This process instantly transcribes the audio data, making it available for processing on the server.
[0711] Step 3:
[0712] The server passes the transcribed data to a natural language processing engine, which extracts important information and key phrases from the conversation. Furthermore, it generates a summary based on this information, concisely summarizing the content.
[0713] Step 4:
[0714] The emotion engine analyzes voice data in parallel, determining emotions based on the user's tone of voice, speaking speed, and other factors. This allows for real-time evaluation of emotional fluctuations during a call.
[0715] Step 5:
[0716] The server combines the extracted key information with sentiment analysis results to adjust the summaries, task lists, and schedule priorities. The sentiment analysis results are also recorded as information accompanying the call history.
[0717] Step 6:
[0718] After the call ends, the server notifies the user of the generated summary, tasks, schedule, and sentiment analysis results. This notification allows the user to review the call content and sentiment feedback.
[0719] Step 7:
[0720] Users can receive notifications, review call content, and reflect on their own emotional responses through sentiment analysis. They can also check their task lists and schedules, which are automatically managed, allowing them to plan their future work.
[0721] (Example 2)
[0722] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0723] In today's communication environment, there is a need to efficiently manage and immediately utilize important information obtained during calls. However, manually recording and organizing call content is laborious, and understanding and appropriately processing emotional nuances is difficult. Therefore, there is a need for a system that can simultaneously analyze call content and the user's emotional state and utilize the information efficiently.
[0724] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0725] In this invention, the server includes means for acquiring call audio as digital audio data in real time, speech recognition means for converting the audio data into text, and analysis means for analyzing the converted text and emotions to generate information. This enables simultaneous analysis of important information from the call content and the user's emotions, and automatic reflection of this information in task lists and schedules.
[0726] "Call audio" refers to voice information transmitted in real time using voice communication methods.
[0727] "Digital audio data" is a format of information obtained by digitizing analog audio signals.
[0728] "Speech recognition means" refers to a technology or device that converts speech data into text data.
[0729] "Natural language processing means" refers to technologies or processes that extract and understand structured information from text data.
[0730] "Analytical means" refer to the techniques and processes used to analyze data and extract new knowledge and information.
[0731] A "task list" is a means of organizing and displaying a user's activities and obligations in a list format.
[0732] A "time management chart" is a method for organizing and visualizing events and schedules along a timeline.
[0733] "Emotion" refers to the user's psychological state and includes technologies for recognizing or analyzing that state.
[0734] A "recording medium" is a physical or digital medium used to store and maintain data.
[0735] This invention is a system that processes a user's call audio as digital audio data in real time, simultaneously recording important information and analyzing emotions.
[0736] The terminal first acquires the user's call audio and converts it into digital audio data. This process uses an audio capture device with a microphone, and the resulting digital audio data is transmitted to the server via wireless communication.
[0737] The server uses speech recognition technology to convert the received audio data into text data. A speech recognition engine is used for this, with Google Cloud Speech-to-Text API being a specific example. The converted text is then analyzed using natural language processing technology to extract important information. SpaCy and NLTK libraries are used for this purpose.
[0738] Furthermore, the server uses a sentiment analysis engine to perform sentiment analysis based on text data. IBM Watson Tone Analyzer is one such sentiment engine that can be used. This identifies the user's emotional state, and this information is integrated with the text data. The analysis results can be automatically reflected in task lists and time management tables.
[0739] After the call ends, the user is notified of the generated summary, tasks, schedule information, and sentiment analysis results. The user can then use this information to plan their actions. This notification is delivered via an application installed on the user's mobile device or computer.
[0740] As a concrete example, consider a scenario where the content of business meetings is recorded and analyzed in real time, and project deadlines and meeting dates are automatically registered in the schedule. If stress levels of the speaker are detected, this information is provided as feedback to improve communication.
[0741] An example of a prompt message is, "Analyze the critical deadlines and stakeholders' emotional states in project progress reports, and create a summary to improve scheduling and communication efficiency." This allows the system to efficiently and accurately support the user's business operations.
[0742] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0743] Step 1:
[0744] The terminal acquires the user's call audio in real time. This involves capturing the audio as an analog signal using the microphone built into the terminal and converting it into digital audio data. The input is the user's voice, and the output is the digitized audio data. This audio data is transmitted to the server via the communication module.
[0745] Step 2:
[0746] The server processes the digital audio data received from the terminal. This processing includes converting the data to text using a speech recognition engine. Specifically, the server analyzes the audio spectrum, extracts phonemes, and transcribes them based on a language model. The input is digital audio data, and the output is text data containing character information.
[0747] Step 3:
[0748] The server analyzes the transcribed data using natural language processing techniques. This involves extracting important information and generating a text summary. For example, it performs morphological and syntactic analysis to extract nouns and verbs, and uses semantic analysis to grasp the main points of the utterance. The input is text data, and the output is a condensed text containing summary information and keywords.
[0749] Step 4:
[0750] The server analyzes the user's emotions from voice and text data. Using an emotion analysis engine, it evaluates, for example, the tone of voice and word choice to identify the user's emotional state. The input is the converted text and original voice data, and the output is data containing the type of emotion identified (joy, anger, etc.).
[0751] Step 5:
[0752] The server integrates extracted key information and sentiment analysis results, automatically reflecting them in the task list and schedule. Specifically, it works in conjunction with task management applications to update schedules and add tasks. The input is summary information and sentiment data, and the output is the updated task list and schedule information.
[0753] Step 6:
[0754] After the call ends, the user receives a summary, task list, schedule information, and sentiment analysis results from the server. This information is notified to the user's device, and the user can check its contents via email or application. The output is notification information presented in a user-friendly format.
[0755] (Application Example 2)
[0756] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0757] In modern society, it has become increasingly difficult for families and individuals to efficiently manage important information arising from daily conversations and to recognize changes in emotions. There is a high risk of overlooking important tasks and appointments by simply ignoring conversations. Furthermore, there is often a lack of appropriate interaction with family members experiencing stress or anxiety. These challenges are raising concerns about a decline in quality of life and impacting the well-being of families and individuals.
[0758] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0759] In this invention, the server includes means for acquiring call audio as digital voice numbers in real time, speech recognition means for converting the acquired voice numbers into text information, natural language processing means for extracting important information from the converted text information and generating a summary, means for performing emotion analysis from the call to identify the user's emotional state, and means for making suggestions and improvements to interactions using the emotion analysis results. This makes it possible for families and individuals to efficiently manage important information obtained from conversations, and to respond immediately to changes in emotions and provide appropriate interactions.
[0760] "Call audio" refers to the sound information of a linguistic conversation conducted via a voice communication device.
[0761] "Real-time" refers to a state where processing or analysis is performed immediately at the moment an event occurs.
[0762] "Digital audio numerical data" refers to numerical information obtained by converting analog audio into numbers that can be processed by a computer.
[0763] A "speech recognition system" is a mechanism that performs technology to convert speech data into linguistic characters.
[0764] "Textual information" refers to data consisting of language and symbols, in a readable and writable format.
[0765] "Important matters" refer to information or events that deserve particular attention or should be given priority.
[0766] A "natural language processing system" is a system that analyzes language data and performs information extraction and summary generation.
[0767] "Emotional analysis" is the process of estimating emotional states from voice or text.
[0768] "User's emotional state" refers to the result of determining what emotional state a particular user is in.
[0769] "Suggesting and improving interactions" means providing methods and actions to improve the quality of communication with users.
[0770] The system for realizing this invention consists of a terminal and a server.
[0771] The terminal is equipped with a microphone for receiving voice communications and captures the voice during a call as digital voice data. This digital voice data is continuously transmitted to the server.
[0772] The server uses speech recognition software to perform speech recognition based on the received digital speech data. Specifically, the Python `speech_recognition` library is used. The recognized speech is converted into text information, and then the TextBlob natural language processing library is used to evaluate the importance of the information and generate a summary as needed.
[0773] In addition, the server runs an emotion analysis engine to extract emotional information from textual data. This uses the emotion analysis pipeline within the transformers library of the generative AI model, Hugging Face. This emotional information is used not only to visualize the user's emotional state but also to suggest and improve interactions. For example, if the server determines that the user is feeling anxious, it will make suggestions to alleviate that anxiety.
[0774] As a concrete example, a robot installed in a family's living room could record family conversations and register important matters as tasks. Furthermore, if it detects someone is stressed, it could play their favorite music to help them relax.
[0775] An example of a prompt message for the application of this system is: "Analyze family phone calls, identify the types of content that tend to cause stress if someone is stressed, and suggest ways to relax."
[0776] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0777] Step 1:
[0778] The terminal receives call audio and converts it into digital audio numbers. The input is analog audio obtained through a microphone, which is sampled and quantized to generate digital audio numbers. This converts the audio data into a format that can be processed by the server.
[0779] Step 2:
[0780] The terminal transmits the generated digital voice numerical data to the server in real time. The input at this stage is the digital voice numerical data generated in step 1, which is sent to the server via network communication. The server then receives the data to proceed to the next processing step.
[0781] Step 3:
[0782] The server analyzes the received digital speech data using speech recognition software and converts it into text. The input is the digital speech data received by the server; speech recognition is performed using the speech_recognition library, and text is obtained as output. This text is essential for subsequent natural language processing.
[0783] Step 4:
[0784] The server analyzes the converted text information using natural language processing tools, extracts important information, and generates a summary as needed. The input is the text information obtained in step 3, and language analysis is performed using TextBlob. The output is a list of important information and the generated summary, and this information contributes to improving work efficiency.
[0785] Step 5:
[0786] The server performs sentiment analysis from textual information to identify the user's emotional state. The input is the textual information from step 3, and the Hugging Face transformers library is used for sentiment analysis. The output is the sentiment identification result, providing information that reflects the user's psychological state.
[0787] Step 6:
[0788] The server makes suggestions and improvements to interactions based on the sentiment analysis results. The input is the output from steps 4 and 5, and appropriate suggestions are made to the user according to the identified key points and emotional state. This improves the user experience and enables more comfortable communication.
[0789] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0790] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0791] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0792] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0793] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0794] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0795] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0796] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0797] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0798] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0799] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0800] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0801] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0802] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0803] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0804] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0805] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0806] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0807] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0808] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0809] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0810] The following is further disclosed regarding the embodiments described above.
[0811] (Claim 1)
[0812] A means of acquiring call audio as digital audio data in real time,
[0813] A speech recognition means that converts acquired audio data into text,
[0814] A natural language processing method that extracts important information from converted text and generates a summary,
[0815] A means of automatically registering tasks to a task list and schedule based on the extracted information,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, further comprising means for notifying the user of a summary generated after the end of a call and information on registered tasks and schedules.
[0819] (Claim 3)
[0820] The system according to claim 1, comprising means for storing the generated summary in a database.
[0821] "Example 1"
[0822] (Claim 1)
[0823] A means of acquiring call audio as digital audio information in real time,
[0824] Acoustic recognition means for converting acquired acoustic information into text data,
[0825] A natural language processing method that extracts key information from converted character data and generates essential points,
[0826] A means of automatically registering tasks in a work list and schedule based on the extracted information,
[0827] Means for storing the generated key points in an information storage device,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, further comprising means for notifying the user of the key points generated after the end of a call and information on registered tasks and schedules.
[0831] (Claim 3)
[0832] The system according to claim 1, comprising means for delivering acoustic information to an acoustic recognition means using an information transmission network.
[0833] "Application Example 1"
[0834] (Claim 1)
[0835] A means of acquiring call audio as digital audio data in real time,
[0836] A speech recognition means that converts acquired audio data into text,
[0837] A natural language processing method that extracts important information from converted text and generates a summary,
[0838] A means of automatically registering tasks to a task list and schedule based on the extracted information,
[0839] A means of updating the performance schedule using the generated summary and registered events,
[0840] A system that includes this.
[0841] (Claim 2)
[0842] The system according to claim 1, further comprising means for notifying the user of a summary generated after the end of a call and information on registered scheduled performance.
[0843] (Claim 3)
[0844] The system according to claim 1, comprising means for storing the generated summary and updated performance schedule in a storage device.
[0845] "Example 2 of combining an emotion engine"
[0846] (Claim 1)
[0847] A means of acquiring call audio as digital audio data in real time,
[0848] A speech recognition means that converts acquired audio data into text,
[0849] A natural language processing method that extracts important information from converted text and generates a summary,
[0850] An analytical means that analyzes extracted information and emotions to generate information,
[0851] A means for automatically registering tasks in a task list and time management sheet based on the analyzed information,
[0852] A system that includes this.
[0853] (Claim 2)
[0854] The system according to claim 1, comprising means for notifying the user of a summary generated after the end of a call, information on registered tasks and time management tables, and analyzed sentiment data.
[0855] (Claim 3)
[0856] The system according to claim 1, comprising means for storing the generated summary and analyzed sentiment data on a recording medium.
[0857] "Application example 2 when combining with an emotional engine"
[0858] (Claim 1)
[0859] A means of acquiring call audio as digital audio numerical values in real time,
[0860] A speech recognition means that converts acquired speech numerical data into text information,
[0861] A natural language processing means for extracting important information from converted text information and generating a summary,
[0862] A means of automatically registering the extracted information into a work list and schedule,
[0863] A means of performing emotion analysis from phone calls to identify the user's emotional state,
[0864] A means of using emotion analysis results to propose and improve interactions,
[0865] A system that includes this.
[0866] (Claim 2)
[0867] The system according to claim 1, further comprising means for notifying the user of a summary generated after the end of a call and information on registered tasks and schedules.
[0868] (Claim 3)
[0869] The system according to claim 1, further comprising means for storing the generated summary in a storage device. [Explanation of Symbols]
[0870] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring call audio as digital audio data in real time, A speech recognition means that converts acquired audio data into text, A natural language processing method that extracts important information from converted text and generates a summary, A means of automatically registering tasks to a task list and schedule based on the extracted information, A means of updating the performance schedule using the generated summary and registered events, A system that includes this.
2. The system according to claim 1, further comprising means for notifying the user of a summary generated after the end of a call and registered information on scheduled performance.
3. The system according to claim 1, comprising means for storing the generated summary and updated performance schedule in a storage device.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A