system
The system automates interview recording and analysis by converting audio to text, extracting key points, and suggesting methods, addressing inefficiencies in current systems and reducing overtime.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Current interview recording systems in shops and call centers are time-consuming, inconsistent, and lack effective data analysis, leading to reduced business efficiency and overtime work.
A system that automates the recording and analysis of interviews by converting audio data to text, extracting key points, generating summaries, and analyzing past records to suggest effective interview methods, using speech recognition engines and generative AI models.
Improves work efficiency by automating interview recording and analysis, reducing overtime, and providing effective interview strategies.
Smart Images

Figure 2026063854000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In current shops and call centers, interview records are manually input, which is very time-consuming. Also, there are problems that the recorded content and format vary depending on the person in charge, and the consistency of information is not maintained. Furthermore, data collection for analysis is not properly conducted, and there is a lack of reference materials for leading to an effective interview method. As a result, business efficiency is reduced, which is a factor causing overtime work. The purpose of this invention is to solve these problems and automate the recording and analysis of interviews to improve business efficiency and reduce overtime work.
Means for Solving the Problems
[0005] The present invention provides a system that includes means for inputting audio data, means for converting the input audio data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, and means for analyzing past records and generating effective interview methods. Furthermore, the system includes means for transmitting the input audio data to a server, means for the server to convert the audio data into text data, means for the server to save the generated summary to a database, and means for retrieving and analyzing past records from the database. This automates the recording of interviews and the analysis of effective interview methods, thereby improving work efficiency and reducing overtime.
[0006] "Audio data" refers to data that records human speech in digital format.
[0007] "Means of input" refers to a device or software for acquiring audio data into the system.
[0008] "Text data" refers to data obtained by analyzing audio data and representing it as text.
[0009] "Means of conversion" refer to algorithms or software used to generate text data from audio data.
[0010] "Key points" refer to the matters or keywords in the interview that deserve particular attention.
[0011] "Means of extraction" refer to software or algorithms used to extract important points from text data.
[0012] A "summary" is a short text file that concisely compiles the necessary information from a longer text file.
[0013] "Generative means" refer to algorithms and software used to create new information or structures from multiple data sets.
[0014] The "means for recording and storing" refers to a device or software that stores the generated data in digital form for future reference and analysis.
[0015] The "past records" refer to a series of interview contents and summaries that were previously recorded and stored in a database or the like.
[0016] The "means for analyzing" refers to software or algorithms for analyzing past records to find useful information and patterns.
[0017] The "effective interview method" refers to techniques and ways of advancing topics that are optimized to achieve the purpose of the interview.
[0018] The "server" refers to a central computing device for exchanging data with multiple terminals through a network.
[0019] The "database" refers to a system for systematically storing a large amount of data and performing efficient searches and analyses.
Brief Description of the Drawings
[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7]It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0022] First, the language used in the following description will be explained.
[0023] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0028] [First Embodiment]
[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0041] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. The interview recording system based on this invention is constructed as follows.
[0042] System Overview
[0043] 1. User (device)
[0044] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[0045] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application on the device.
[0046] 2. Server
[0047] The server receives the audio file sent from the terminal and saves it to storage.
[0048] The saved audio files are converted into text data using a speech recognition engine (for example, Google® Speech-to-Text API).
[0049] The system uses generative AI to extract key points from the converted text data and generate a summary.
[0050] The generated summary is saved to the database.
[0051] By retrieving past records from a database and analyzing them using natural language processing techniques, we can identify effective interview methods.
[0052] Program processing
[0053] The user records audio and sends it to the server from their device.
[0054] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[0055] The server receives the audio file and saves it locally.
[0056] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[0057] The server converts the audio to text and generates a summary.
[0058] For text conversion, the server sends the audio file to a speech recognition engine, which then passes the acquired text data to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0059] The server saves the interview records to the database.
[0060] The generated summary and audio file paths are stored in a database on the server. For example, they might be stored in the "records" table of an SQLite database.
[0061] The server analyzes past records and suggests effective interview methods.
[0062] The server retrieves past interview records from the database and uses natural language processing technology to identify common keywords and patterns. This allows it to generate effective interview methods to improve interview efficiency and provide feedback to the user.
[0063] Specific example
[0064] For example, suppose a user has a meeting with a customer one day. They record this meeting and send it to a server via their device. The server converts the audio file into text, automatically extracts the key points of the conversation, generates a summary, and saves it to a database. A few weeks later, the user analyzes multiple meeting records and identifies frequently occurring keywords, revealing that "customer satisfaction," "product quality," and "service improvement" are important points. This allows them to focus on these keywords in future meetings, leading to more effective communication.
[0065] Thus, the system of the present invention automates the recording and analysis of interviews, thereby improving work efficiency and reducing overtime.
[0066] The following describes the processing flow.
[0067] Step 1:
[0068] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[0069] Step 2:
[0070] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[0071] Step 3:
[0072] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[0073] Step 4:
[0074] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary directory. For example, the file is saved to " / path / to / save / meeting_audio.wav".
[0075] Step 5:
[0076] The server sends the stored audio file to a speech recognition engine, which converts it to text. The server reads the audio file and uses a speech recognition API (such as the Google Speech-to-Text API) to transcribe the audio. As a result, the content of the interview is obtained as text data.
[0077] Step 6:
[0078] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. The generative AI model uses tools such as Hugging Face's Transformers to perform the summarization. As a result, a concise summary is obtained that summarizes the main points of the interview discussion.
[0079] Step 7:
[0080] The server saves the generated summary and the path to the original audio file to a database. For example, SQLite is used for the database, and the summary and the path to the corresponding audio file are inserted into a table called "records".
[0081] Step 8:
[0082] The server retrieves past interview records from the database and analyzes them using natural language processing technology. This identifies common keywords and patterns in the interview content. The analysis results serve as reference material for developing more effective interview methods.
[0083] Step 9:
[0084] Based on the analysis results, the server suggests effective interview methods to the user. Specifically, it suggests incorporating questions and topics that include the top keywords. This suggestion is provided to the user as a strategic guideline for future interviews.
[0085] (Example 1)
[0086] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] Traditional interview recording systems have the problem of requiring significant time and effort to manually record interview content and summarize key points. Furthermore, there was a lack of concrete methods for effectively analyzing past interview records and using that information to improve future interviews. This led to concerns that interview efficiency would decrease, resulting in a decline in the quality of customer service and overall work.
[0088] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0089] In this invention, the server includes means for converting audio data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating effective interview methods, means for the user to record audio data and send it to the server from a terminal, means for the server to temporarily store the received audio data, means for the server to generate a summary using natural language processing technology, and means for the server to save the summary and analyze past records. As a result, the recording and summary generation of interviews and the analysis of past records are automated, making it possible to quickly propose effective interview methods.
[0090] "Audio data" refers to digital audio recordings of conversations and statements made by the user during an interview.
[0091] "Text data" refers to audio data converted into a string of characters, representing the content of the interview in written form.
[0092] A "summary" is a concise compilation of key points extracted from text data, including the main discussion points and conclusions of the interview.
[0093] "User" refers to the person who conducts the interview, starts and stops recording, and sends the data to the server.
[0094] A "terminal" refers to a recording device or computer used by a user, which has the function of recording audio data and sending it to a server.
[0095] A "server" is a central processing unit that receives audio data, stores it, converts it to text data, generates summaries, and analyzes stored historical records.
[0096] "Natural language processing technology" is a technology that enables computers to understand and process human language, and is used for generating summaries of text data and analyzing records.
[0097] A "database" is a system for managing and storing data such as generated summaries and historical records.
[0098] "Temporary storage" refers to the process of temporarily saving audio data between the time it is received and the start of processing.
[0099] A "speech recognition engine" refers to software or services that convert speech data into text data.
[0100] This invention relates to a system for effectively recording and analyzing interviews in shops and call centers. Specific embodiments are described in detail below.
[0101] Hardware and software to be used
[0102] In this invention, a system is built through the collaboration of three parties: the user, the terminal, and the server. The hardware and software used are shown below.
[0103] User device: A device with recording capabilities, such as a smartphone, tablet, or personal computer.
[0104] Server: A device with high-performance computing power and storage, such as a cloud server or on-premises server.
[0105] Speech recognition engine: Software used to convert speech data into text data (e.g., Google Speech-to-Text API).
[0106] Generative AI models: Artificial intelligence models for generating summaries of text data (e.g., Transformers in Hugging Face).
[0107] Database: A database system (e.g., SQLite) for storing generated summaries and interview records.
[0108] Overview of program processing
[0109] The user records the interview and sends it to the server from their device.
[0110] The user records audio during the interview using an application on their device. Once recording is complete, they press the "Stop Recording" button to end the recording and save the audio file to the device's local storage. Then, the user clicks the "Send" button to send the audio file to the server via an HTTP POST request. For example, the audio file might be saved at " / local / path / to / record.wav" and sent to the server.
[0111] The server receives the audio file and saves it locally.
[0112] The server saves the received audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav"). In this step, the server verifies the format and size of the received file and saves only valid audio data.
[0113] Convert audio data to text data
[0114] The server sends the stored audio files to the speech recognition engine. For example, it uses the Google Speech-to-Text API to convert the audio files into text data. Then, it retrieves the converted text data and passes it to a generative AI model.
[0115] Generate a summary from text data
[0116] The server uses a generative AI model (e.g., Hugging Face's Transformers) to extract key points from text data and generate a summary. This summary generation is instructed using prompts. An example prompt is, "Summarize the main points of this text data." The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0117] Save the summary and audio file path to the database.
[0118] The server records the generated summary and the path to the audio file in a database. For example, it might use a "records" table in an SQLite database to store the summary and path in association. This allows for easy retrieval from the database later when needed.
[0119] We analyze past records and propose effective interview methods.
[0120] The server retrieves past interview records from the database and analyzes them using natural language processing technology. By analyzing multiple summary data and identifying frequently occurring keywords and patterns, it generates specific suggestions for the next interview. These suggestions allow the user to conduct the interview more effectively.
[0121] Specific example
[0122] For example, suppose a user conducts a meeting with a customer one day and records the conversation. After recording, they click the "Send" button to send the audio file to the server. The server receives the audio file and saves it to a path like " / path / to / save / meeting_audio.wav". Then, a speech recognition engine is used to convert it into text data, and a generative AI model is used to generate a summary. The generated summary might be text like, "This meeting mainly focused on customer satisfaction." This summary and file path are stored in a database and later analyzed along with other meeting records to provide material for suggesting more effective meeting methods.
[0123] Examples of prompt messages include, "Please summarize the main points of this text data."
[0124] The above describes a specific embodiment for carrying out the present invention. This system automates the recording, summarizing, and analysis of interviews, thereby significantly improving work efficiency.
[0125] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0126] Step 1:
[0127] The user records the interview audio.
[0128] Input: Interview audio
[0129] Output: Audio files saved to local storage
[0130] Specific operation: Recording begins when the user launches the application on their device and presses the "Start Recording" button. After the interview ends, the user presses the "Stop Recording" button to end the recording and saves the audio file to the device's local storage. For example, the audio data will be saved to the path " / local / path / to / record.wav".
[0131] Step 2:
[0132] The device sends the audio file to the server.
[0133] Input: Audio files stored in local storage
[0134] Output: Audio file sent to the server
[0135] Specific action: The user clicks the "Send" button on the application. The device reads the audio file from local storage and sends it to the server via an HTTP POST request. For example, an audio file saved at " / local / path / to / record.wav" is sent to the server.
[0136] Step 3:
[0137] The server receives the audio file and saves it to a temporary directory.
[0138] Input: Audio file sent via HTTP POST request
[0139] Output: Audio files saved in a temporary directory
[0140] Specific operation: The server receives an HTTP POST request and verifies the file format and size. After confirming that it is valid audio data, it saves the audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav").
[0141] Step 4:
[0142] The server converts the audio data into text data.
[0143] Input: Audio file saved in a temporary directory
[0144] Output: Text data converted from audio data
[0145] Specific operation: The server reads the audio file stored in a temporary directory and sends it to a speech recognition engine (e.g., Google Speech-to-Text API). It receives the text data returned by the speech recognition engine and passes it on to the next process. An example of the converted text would be in the format of "To obtain customer feedback, the following questions were asked."
[0146] Step 5:
[0147] The server generates a summary from the text data it has acquired.
[0148] Input: Text data
[0149] Output: Summary generated from text data
[0150] Specific operation: The server inputs text data into an AI model that generates data (e.g., Hugging Face's Transformers). The prompt is "Summarize the main points of this text data," and the model extracts the key points and generates a summary. An example of a generated summary would be, "This meeting focused on customer satisfaction."
[0151] Step 6:
[0152] The server saves the generated summary and the path to the audio file to the database.
[0153] Input: Generated summary, path to audio file
[0154] Output: Summary and audio file path stored in the database
[0155] Specific operation: The server adds the generated summary and the audio file path to a database (e.g., the "records" table in an SQLite database). Specifically, it uses SQL INSERT statements to record the summary and the corresponding audio file path.
[0156] Step 7:
[0157] The server analyzes past records and generates effective interview methods.
[0158] Input: Past interview records retrieved from the database
[0159] Output: Proposals for effective interview methods
[0160] Specific operation: The server retrieves past interview records from the database and analyzes them using natural language processing technology. It analyzes multiple summary data to identify frequently occurring keywords and patterns. Based on this, it proposes effective interview methods and provides feedback to the user. If the extracted keywords are "customer satisfaction" or "service improvement," it suggests focusing on these points in the next interview.
[0161] (Application Example 1)
[0162] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0163] In autonomous vehicles, there is a lack of effective means to collect and analyze important information regarding communication with passengers and driving conditions. As a result, improvements in driving efficiency and passenger experience are not being fully realized. Furthermore, there is no system that automatically analyzes passenger opinions and feedback to provide appropriate driving advice and feedback for service improvement.
[0164] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0165] In this invention, the server includes means for transmitting voice data to a cloud server and generating suggestions to improve driving efficiency and user experience based on the analysis results, means for providing feedback to the interface of an autonomous vehicle, and means for the cloud server to store the generated summaries in a database. This enables the effective collection, analysis, and summarization of important information regarding communication with passengers and driving conditions, thereby improving driving efficiency and passenger experience.
[0166] "Audio data" refers to data used to record and play back audio information in digital format.
[0167] "Text data" refers to data obtained by analyzing audio data and converting it into textual information.
[0168] A "summary" is a concise compilation of important information extracted from text data.
[0169] A "cloud server" is a remote server accessible via the internet that processes and stores large amounts of data.
[0170] "Driving efficiency" refers to the efficiency of driving a car in terms of fuel consumption, time, and energy consumption.
[0171] "User experience" refers to the overall experience of passengers using autonomous vehicles, including their satisfaction and emotional reactions.
[0172] An "interface" is a device or program that allows humans and machines to exchange information with each other.
[0173] A "database" is a system that systematically stores large amounts of data and allows for efficient searching and updating.
[0174] This invention relates to a system for automatically analyzing conversation recordings collected by autonomous vehicles during operation, extracting important information, and improving driving efficiency and user experience. Details are provided below.
[0175] System configuration and the hardware and software used.
[0176] Hardware to use
[0177] 1. Onboard computer of autonomous vehicle:
[0178] The system collects and partially processes audio data.
[0179] 2. Cloud Server:
[0180] It performs speech analysis, text conversion, database storage, and analysis.
[0181] Software to use
[0182] 1. Speech recognition engine:
[0183] The Google Speech-to-Text API is used to convert speech to text.
[0184] 2. Generative AI models:
[0185] Hugging Face's Transformers are used to extract important information and generate a summary.
[0186] 3. Database system:
[0187] Use MySQL® or SQLite to store, search, and analyze large amounts of data.
[0188] System operation
[0189] 1. Collection of audio data:
[0190] Microphones inside the autonomous vehicle record audio, and the data is saved to the onboard computer.
[0191] Send the audio file to the cloud server using an HTTP POST request.
[0192] 2. Text conversion and summary generation of audio data:
[0193] The server saves the received audio files to a temporary directory.
[0194] The speech recognition engine (Google Speech-to-Text API) converts speech to text.
[0195] A generative AI model (Hugging Face's Transformers) is used to generate a summary, which is then saved to a database.
[0196] 3. Analysis of past records and generation of feedback:
[0197] Historical records are retrieved from the database, and common keywords and patterns are identified using NLP (Neuro-Linguistic Programming) techniques.
[0198] It generates effective communication methods and driving advice, and displays them on passenger interfaces such as tablets.
[0199] Specific example of processing
[0200] For example, suppose a passenger is in an autonomous vehicle and is talking about their destination while the vehicle is in motion. Below are some examples of prompts to input into the generating AI model.
[0201] Example of a prompt
[0202] Passenger: How long will it take for this car to reach its destination?
[0203] Vehicle AI: Considering the current traffic conditions, it will take approximately 30 minutes to arrive. Do you have any other questions?
[0204] Passenger: Is it okay if there's traffic?
[0205] Extract and summarize the key points and conclusions of the conversation.
[0206] Based on this prompt, the generative AI model generates the following summary.
[0207] Generated summary
[0208] Key discussion points: Arrival time at destination, traffic conditions, and how to deal with congestion.
[0209] Conclusion: Estimated travel time is approximately 30 minutes, taking traffic conditions into consideration.
[0210] conclusion
[0211] In this way, by applying the invention to autonomous vehicles, it is possible to improve the passenger experience while further increasing driving efficiency. The generated feedback is provided to the driver and passengers through the autonomous vehicle's interface, enabling appropriate advice and service improvements tailored to the driving situation.
[0212] This embodiment of the invention enables effective information collection and analysis in autonomous vehicles, thereby improving both driving efficiency and user experience.
[0213] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0214] Step 1:
[0215] The onboard computer of the autonomous vehicle collects audio data. The onboard computer records audio data using the microphone inside the vehicle, and this data is stored in local storage. The input is conversations between passengers and the driver, and the output is a recorded audio file.
[0216] Step 2:
[0217] The device sends the recorded audio file to the cloud server. The audio file is uploaded to the cloud server using an HTTP POST request. The input is the audio file on the onboard computer, and the output is the audio file stored on the cloud server.
[0218] Step 3:
[0219] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" is used. The input is the uploaded audio file, and the output is the temporary file stored on the server.
[0220] Step 4:
[0221] The server converts the audio file into text data. It uses a speech recognition engine (Google Speech-to-Text API) to convert the audio data into text data. The input is an audio file, and the output is the converted text data.
[0222] Step 5:
[0223] The server extracts key points from text data and generates a summary. It uses a generative AI model (Hugging Face's Transformers) to analyze text data and extract key points. The input is text data, and the output is the generated summary.
[0224] Step 6:
[0225] The server saves the generated summary to a database. The summary is stored in a database (e.g., MySQL or SQLite). The input is the generated summary, and the output is the summary stored in the database.
[0226] Step 7:
[0227] The server retrieves historical records from the database and performs analysis. Natural language processing techniques are used to identify common keywords and patterns. The input is historical records from the database, and the output is the analysis results.
[0228] Step 8:
[0229] The server generates suggestions to improve driving efficiency and user experience based on the analysis results. It generates appropriate driving advice and feedback for service improvement and displays them on the autonomous vehicle's interface. The input is the analysis results, and the output is the generated suggestions and feedback.
[0230] In this way, each processing step is executed sequentially, creating a system that improves the driving efficiency and user experience of autonomous vehicles.
[0231] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0232] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. In particular, it aims to further improve the quality of interviews by incorporating an emotion engine that recognizes user emotions. The interview recording system based on this invention is constructed as follows.
[0233] System Overview
[0234] 1. User (device)
[0235] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[0236] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application that provides the recording function.
[0237] 2. Server
[0238] The server receives the audio file sent from the terminal and saves it to storage.
[0239] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[0240] The system uses generative AI to extract key points from the converted text data and generate a summary.
[0241] In addition to the generated summary, an emotion engine is used to identify the user's emotions from the audio data, and this information is added to the text data and the generated summary.
[0242] The generated summary and sentiment information are stored in a database.
[0243] We retrieve emotional data from a database along with past records, analyze it using natural language processing techniques, and identify effective interview methods.
[0244] Program processing
[0245] The user records audio and sends it to the server from their device.
[0246] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[0247] The server receives the audio file and saves it locally.
[0248] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[0249] The server converts the audio to text.
[0250] For text conversion, the server sends the audio file to the speech recognition engine, which then generates the acquired text data. This process transcribes the contents of the audio file into text.
[0251] The server generates a summary.
[0252] The generated text data is passed to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. This allows the main discussion points of the interview to be concisely summarized.
[0253] The server recognizes emotions.
[0254] In addition to the generated summary, the user's emotions are analyzed from the audio data using an emotion engine (e.g., emotion recognition AI). The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds the results to the text data and summary.
[0255] The server saves the interview records to the database.
[0256] The generated summary and sentiment information are stored in a database. For example, the format in which it is stored in the "records" table of an SQLite database is as follows:
[0257] sql
[0258] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[0259] The server analyzes past records and emotions.
[0260] The server retrieves past interview records and sentiment data from the database and performs analysis using natural language processing techniques. A comprehensive analysis, including sentiment data, is conducted to identify common keywords and patterns.
[0261] The server suggests effective interview methods.
[0262] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data.
[0263] Specific example
[0264] For example, suppose a user conducts a meeting with a customer. This meeting is recorded and sent to a server via the device. The server converts the audio file into text, extracts the key points of the conversation to generate a summary, and uses an emotion engine to identify emotional information. This information may include, for example, "Many positive emotions of customer satisfaction were recognized." This information and the summary are stored in a database and used for future analysis. In the next meeting, the user can use the emotional information from the previous meeting to communicate more effectively.
[0265] Thus, the system of the present invention automates the recording and analysis of interviews, and by further incorporating emotional information, it is possible to improve work efficiency and reduce overtime.
[0266] The following describes the processing flow.
[0267] Step 1:
[0268] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[0269] Step 2:
[0270] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[0271] Step 3:
[0272] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[0273] Step 4:
[0274] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary storage directory. The save location is, for example, a path like " / path / to / save / meeting_audio.wav".
[0275] Step 5:
[0276] The server sends the stored audio files to a speech recognition engine, which converts them into text. For example, the Google Speech-to-Text API is used as the speech recognition engine. The converted text data is a written record of the interview content.
[0277] Step 6:
[0278] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. Hugging Face's Transformers is used as the generative AI model. As a result, a summary is obtained that concisely summarizes the main discussion points of the interview.
[0279] Step 7:
[0280] Next, the server passes the audio data to the emotion engine to identify the user's emotions. An emotion recognition AI (for example, IBM Watson®'s Empathy API) is used as the emotion engine. The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds this information to the text data and summary.
[0281] Step 8:
[0282] The server stores the generated summary and sentiment information in a database. For example, SQLite is used for the database, and the summary, corresponding sentiment information, and audio file paths are inserted into a table called "records".
[0283] Step 9:
[0284] The server retrieves past interview records and sentiment data from the database and analyzes them using natural language processing techniques. A comprehensive analysis is performed, including the sentiment data, to identify common keywords and patterns.
[0285] Step 10:
[0286] Based on the analysis results, the server proposes an effective interview method to the user. Specifically, questions and topics containing the top keywords are provided, and further proposals are made to optimize the progress of the interview by referring to past sentiment data. This proposal is provided to the user as a strategic guideline for future interviews.
[0287] (Example 2)
[0288] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0289] In a conventional interview recording system, it was necessary to manually record the interview content, and it was difficult to efficiently organize and store information. In addition, it did not have a function to analyze emotions from voice data and improve the quality of interviews. As a result, there was a lack of information for finding an effective interview method, and as a result, the business efficiency decreased and the communication effect was limited. There is a demand for a new system to solve these problems.
[0290] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving and temporarily storing voice data, means for converting voice data into text data, and means for extracting important points from the converted text data using a generated AI model and generating a summary. This enables automatic recording and analysis of interview content, and further enables advanced analysis including user sentiment information. This makes it possible to identify effective interview methods and improve work efficiency.
[0291] "Interview audio data" refers to digital audio files that record the content of conversations conducted in person or remotely.
[0292] A "memory device" refers to hardware or databases used to store digital information such as audio data, text data, and sentiment analysis results.
[0293] "Communication function" refers to the function of transmitting digital data to other devices or servers via a network.
[0294] A "server" is a computer system that processes, stores, and analyzes audio data received from terminals via a network.
[0295] "Text data" refers to text information obtained by transcribing audio data.
[0296] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates summaries and responses based on the input data.
[0297] An "emotion engine" is a system that analyzes and recognizes emotions from voice data.
[0298] "Recording" refers to saving information such as generated summaries and sentiment analysis results so that they can be referenced later.
[0299] A "database" refers to a software system designed to systematically store and manage information and enable retrieval and updates as needed.
[0300] "Natural language processing technology" refers to artificial intelligence technology for analyzing, understanding, and manipulating human language.
[0301] This invention relates to a system that automatically records, analyzes voice data of interviews, and proposes effective interview methods. Specific embodiments for implementing this system will be described below.
[0302] Overview of the System
[0303] This system is composed of a user (terminal), a server, and a database. The user records interviews at a shop or call center using a terminal and transmits the recorded data to the server.
[0304] Hardware and Software Used
[0305] 1. Hardware
[0306] Terminal: A recording device used by the user. This includes smartphones, tablets, personal computers, etc.
[0307] Server: A computer system for receiving and processing voice data.
[0308] 2. Software
[0309] Voice recording application: An application for recording voice on the terminal and saving it to local storage.
[0310] Communication protocol: An HTTP POST request for transmitting voice data to the server.
[0311] Speech recognition engine: Converts speech data into text data (e.g., Google Speech-to-Text API).
[0312] Generative AI model: Generates summaries from generated text data (e.g., Transformers for Hugging Face).
[0313] Emotion engine: Software for analyzing emotions from voice data (e.g., emotion recognition AI).
[0314] Database: Stores the generated summary and sentiment information (e.g., an SQLite database).
[0315] Natural language processing technology: A technology for analyzing historical records and sentiment data.
[0316] System operation
[0317] The user uses a voice recording application on their device during the interview. Once recording is complete, the audio data is saved locally on the device and uploaded to the server using the application's transmission function. This transmission is done via an HTTP POST request. The server temporarily stores the received audio data and converts it into text data using a speech recognition engine. The converted text data is then passed to a generation AI model, which extracts key points and generates a summary.
[0318] Simultaneously, the server uses an emotion engine to analyze the user's emotions from the audio data. This emotion information is added to the summary, and the final generated summary and emotion information are stored in a database. The server then uses natural language processing techniques to analyze past interview records and emotion information stored in the database, identifying common keywords and patterns to suggest effective interview methods.
[0319] Specific example
[0320] For example, consider a scenario where a user records a meeting with a customer and sends the audio data from their device to a server. The server converts this audio data into text data, extracts key points, and creates a summary. It also analyzes the customer's emotions from the audio data and adds emotional information to the summary, such as "the customer showed many positive emotions of satisfaction." Finally, this data is stored in a database and used as a reference for future meetings.
[0321] Users can use information from previous meetings to select appropriate questions and topics for subsequent meetings. This improves the quality of meetings, leading to increased work efficiency and reduced overtime.
[0322] Example of a prompt
[0323] For example, you can generate a summary by giving the following prompt to the AI model:
[0324] "Please summarize what we discussed in today's meeting."
[0325] Thus, the present invention is a system that automatically records and analyzes voice data, enabling the proposal of interview methods that also take emotional information into account.
[0326] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0327] System program processing flow
[0328] Step 1:
[0329] Audio recording
[0330] The user records audio data during the interview using the device's audio recording application. Recording starts when the record button is pressed and ends when it is pressed again. The completed audio data (input) is saved to the device's local storage (output). For example, it is saved to the path " / local / path / to / meeting_audio.wav".
[0331] Step 2:
[0332] Sending audio files
[0333] The user clicks the upload button within the voice recording application to send the recorded audio file to the server. The sending process is performed using an HTTP POST request. The input is the audio file in local storage, and the output is the audio file uploaded to the server.
[0334] Step 3:
[0335] Receiving and saving audio files
[0336] The server saves audio files received from the user's terminal to a temporary directory. The input is the audio file sent via an HTTP POST request, and the output is the audio file stored in the server's temporary storage area. For example, the audio data is saved to " / path / to / save / meeting_audio.wav".
[0337] Step 4:
[0338] Speech-to-text conversion
[0339] The server uses the stored audio file to call a speech recognition engine (e.g., Google Speech-to-Text API) and converts the audio data into text data. The input is an audio file, and the output is the converted text data. As a result of the transcription, all statements made during the interview are obtained in text format.
[0340] Step 5:
[0341] Summary generation
[0342] The server passes the acquired text data to an AI model (e.g., Hugging Face's Transformers) to extract key points and generate a summary. The input is the acquired text data, and the output is the generated summary text. This provides a concise overview of the interview content.
[0343] Step 6:
[0344] Recognition of emotions
[0345] The server passes the audio data to an emotion engine (e.g., emotion recognition AI) to analyze the user's emotions. The input is audio data, and the output is the result of the emotion analysis. The analyzed emotion information is added to the summary text of the interview.
[0346] Step 7:
[0347] Preservation of products
[0348] The server stores the generated summary and sentiment information in a database. The input is the summary text and sentiment analysis results, and the stored data is the output. For example, it is stored in the "records" table in an SQLite database in the following format:
[0349] sql
[0350] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[0351] Step 8:
[0352] Analysis of past records and emotions
[0353] The server retrieves past interview records and sentiment information from the database and analyzes them using natural language processing techniques. The input is past interview records and sentiment data from the database, and the output is effective interview methods derived from the analysis. This process identifies common keywords and patterns.
[0354] Step 9:
[0355] Suggestions for effective interview methods
[0356] The server suggests effective interview methods to the user based on the analysis results. The input is the analysis results, and the output is the suggested content. Specifically, it suggests questions and topics that include top keywords, as well as interview procedures that take into account past sentiment data.
[0357] (Application Example 2)
[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0359] While conventional interview recording systems could transcribe audio data and analyze past records, they lacked the functionality to improve the quality of customer service in real time. Furthermore, there was no mechanism to accurately capture customer emotions and provide appropriate advice in real time based on those emotions. Therefore, improving customer satisfaction and response efficiency were identified as challenges.
[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data, means for converting the input voice data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating an effective interview method, means for analyzing the user's emotions and adding the emotional information to the summary data, and means for displaying emotional information in real time during the conversation with the customer and suggesting an appropriate conversation method. This makes it possible to improve responses in real time based on the customer's emotions.
[0361] "Audio data" refers to information that can be recorded, stored, or transmitted in digital format.
[0362] "Means of input" refers to devices or functions that allow users to supply voice data to the system.
[0363] "Text data" refers to information obtained by converting audio data into written text.
[0364] "Means of conversion" refers to the functions of software or hardware used to convert audio data into text data.
[0365] "Key points" refer to the most important information or topics extracted from the interview content.
[0366] A "means for generating summaries" refers to a function that extracts important points from text data and summarizes them concisely.
[0367] "Means of recording and preserving" refers to the functions of storage devices and databases for storing generated summaries and other data.
[0368] "Past records" refer to interview data and summary data that have been saved in the past.
[0369] "Means of analysis" refer to software and algorithms used to analyze past records and identify patterns and trends.
[0370] "Effective interview methods" refer to methods and techniques for optimizing communication, which are proposed based on the analysis results.
[0371] "User emotions" refers to the emotional state a user exhibits during an interview (for example, joy, anger, sadness, etc.).
[0372] "Means of analyzing emotions" refers to software or algorithms that identify a user's emotions from voice data and analyze that information.
[0373] "Emotional information" refers to information about a user's emotional state obtained through means of analyzing emotions.
[0374] "Means of displaying in real time" refers to the functions of displays and software that immediately show acquired emotional information to the user.
[0375] "Means of suggesting dialogue methods" refer to software or algorithms that suggest appropriate communication methods to users based on emotional information and other factors.
[0376] This invention relates to a system that automatically records interviews in shops and call centers, recognizes user emotions, and suggests effective interview methods. The system of this invention includes the following components.
[0377] System Overview
[0378] 1. Input of audio data
[0379] The user records conversations with customers using smart glasses. The recorded audio data is temporarily stored in the smart glasses. After recording is complete, the audio data is automatically sent to a server. This transmission requires an internet connection.
[0380] 2. Converting audio data
[0381] The server converts the received audio data into text data using a speech recognition engine such as the Google Speech-to-Text API. The converted text data is stored in a temporary directory.
[0382] 3. Summarization from text data
[0383] Key points are extracted from text data, and a summary is generated using generative AI models such as Hugging Face's Transformers. The generated summary data is temporarily stored for further analysis.
[0384] 4. Recognition of emotions
[0385] The server identifies the user's emotions from the voice data using an emotion recognition engine such as IBM Watson Tone Analyzer. The identified emotion information is added to the summary data.
[0386] 5. Proposal for a real-time dialogue method
[0387] Based on the generated summary and sentiment information, the server formulates a response plan in real time and displays it on the smart glasses' display. This allows staff to get the appropriate response plan on the spot.
[0388] 6. Recording and Preservation
[0389] The generated summaries and sentiment information are stored in a storage device such as an SQLite database. This makes the data reusable at any time.
[0390] 7. Analysis of past records
[0391] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. By analyzing past conversation content and emotional data, and extracting common patterns and trends, it generates and proposes effective interview methods.
[0392] Specific example
[0393] For example, if a customer asks for a product description, the smart glasses record the conversation and send it to a server. The server converts the recorded audio data into text, generates a summary using a generative AI model, and adds customer sentiment information to the summary. Next, real-time advice such as "The customer is satisfied, so we will suggest further related products" is displayed on the smart glasses' screen, allowing staff to take appropriate action immediately.
[0394] Example of a prompt
[0395] The following prompt messages can be used to efficiently analyze customer interactions and suggest appropriate responses.
[0396] "Could you please explain the latest model in detail?"
[0397] "Are there any other products with similar functions?"
[0398] "Do you have a demo that clearly explains how to use it?"
[0399] As a result, the system of the present invention can improve the quality of communication with customers and enable effective responses.
[0400] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0401] Step 1:
[0402] The user records conversations with customers using smart glasses. The input is the audio of the conversation with the customer, which the smart glasses record as audio data. The output is a temporarily stored audio file. Recording is possible through the recording function within the smart glasses.
[0403] Step 2:
[0404] After recording is complete, the audio data is automatically sent to the server. The input here is the audio file stored in the smart glasses, which is transferred to the server via an internet connection. The output is the audio file stored on the server. This transmission is performed via an HTTP POST request.
[0405] Step 3:
[0406] The server converts the received audio data into text data using the Google Speech-to-Text API. An audio file is used as input, and the output is text data. The server sends the audio data to the API and saves the returned text data to a temporary directory.
[0407] Step 4:
[0408] The server extracts key points from the transformed text data and generates a summary using a generative AI model (e.g., Hugging Face's Transformers). The input is text data, and the output is summary data. The server invokes the generative AI model, generates the summary data, and temporarily stores it.
[0409] Step 5:
[0410] The server identifies the user's emotions from audio data using IBM Watson Tone Analyzer. The input is an audio file, and the output is emotion information. The server sends the audio data to the emotion recognition engine and adds the received emotion information to the summary data.
[0411] Step 6:
[0412] The server proposes a dialogue method to the user in real time based on the generated summary and sentiment information. The input is the summary and sentiment information, and the output is the proposed dialogue method. Based on this data, the server formulates an appropriate response and displays it on the smart glasses' display.
[0413] Step 7:
[0414] The server saves the generated summary and sentiment information to an SQLite database. The input is the summary and sentiment information, and the output is the records stored in the database. The server performs the save process and persists the data.
[0415] Step 8:
[0416] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. The input is historical record data, and the output is the analysis results and effective interview methods. The server reads data from the database and applies algorithms to analyze patterns and trends.
[0417] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0418] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0419] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0420] [Second Embodiment]
[0421] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0422] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0423] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0424] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0425] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0426] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0427] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0428] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0429] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0430] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0431] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0432] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0433] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. The interview recording system based on this invention is constructed as follows.
[0434] System Overview
[0435] 1. User (device)
[0436] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[0437] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application on the device.
[0438] 2. Server
[0439] The server receives the audio file sent from the terminal and saves it to storage.
[0440] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[0441] The system uses generative AI to extract key points from the converted text data and generate a summary.
[0442] The generated summary is saved to the database.
[0443] By retrieving past records from a database and analyzing them using natural language processing techniques, we can identify effective interview methods.
[0444] Program processing
[0445] The user records audio and sends it to the server from their device.
[0446] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[0447] The server receives the audio file and saves it locally.
[0448] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[0449] The server converts the audio to text and generates a summary.
[0450] For text conversion, the server sends the audio file to a speech recognition engine, which then passes the acquired text data to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0451] The server saves the interview records to the database.
[0452] The generated summary and audio file paths are stored in a database on the server. For example, they might be stored in the "records" table of an SQLite database.
[0453] The server analyzes past records and suggests effective interview methods.
[0454] The server retrieves past interview records from the database and uses natural language processing technology to identify common keywords and patterns. This allows it to generate effective interview methods to improve interview efficiency and provide feedback to the user.
[0455] Specific example
[0456] For example, suppose a user has a meeting with a customer one day. They record this meeting and send it to a server via their device. The server converts the audio file into text, automatically extracts the key points of the conversation, generates a summary, and saves it to a database. A few weeks later, the user analyzes multiple meeting records and identifies frequently occurring keywords, revealing that "customer satisfaction," "product quality," and "service improvement" are important points. This allows them to focus on these keywords in future meetings, leading to more effective communication.
[0457] Thus, the system of the present invention automates the recording and analysis of interviews, thereby improving work efficiency and reducing overtime.
[0458] The following describes the processing flow.
[0459] Step 1:
[0460] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[0461] Step 2:
[0462] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[0463] Step 3:
[0464] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[0465] Step 4:
[0466] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary directory. For example, the file is saved to " / path / to / save / meeting_audio.wav".
[0467] Step 5:
[0468] The server sends the stored audio file to a speech recognition engine, which converts it to text. The server reads the audio file and uses a speech recognition API (such as the Google Speech-to-Text API) to transcribe the audio. As a result, the content of the interview is obtained as text data.
[0469] Step 6:
[0470] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. The generative AI model uses tools such as Hugging Face's Transformers to perform the summarization. As a result, a concise summary is obtained that summarizes the main points of the interview discussion.
[0471] Step 7:
[0472] The server saves the generated summary and the path to the original audio file to a database. For example, SQLite is used for the database, and the summary and the path to the corresponding audio file are inserted into a table called "records".
[0473] Step 8:
[0474] The server retrieves past interview records from the database and analyzes them using natural language processing technology. This identifies common keywords and patterns in the interview content. The analysis results serve as reference material for developing more effective interview methods.
[0475] Step 9:
[0476] Based on the analysis results, the server suggests effective interview methods to the user. Specifically, it suggests incorporating questions and topics that include the top keywords. This suggestion is provided to the user as a strategic guideline for future interviews.
[0477] (Example 1)
[0478] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0479] Traditional interview recording systems have the problem of requiring significant time and effort to manually record interview content and summarize key points. Furthermore, there was a lack of concrete methods for effectively analyzing past interview records and using that information to improve future interviews. This led to concerns that interview efficiency would decrease, resulting in a decline in the quality of customer service and overall work.
[0480] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0481] In this invention, the server includes means for converting audio data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating effective interview methods, means for the user to record audio data and send it to the server from a terminal, means for the server to temporarily store the received audio data, means for the server to generate a summary using natural language processing technology, and means for the server to save the summary and analyze past records. As a result, the recording and summary generation of interviews and the analysis of past records are automated, making it possible to quickly propose effective interview methods.
[0482] "Audio data" refers to digital audio recordings of conversations and statements made by the user during an interview.
[0483] "Text data" refers to audio data converted into a string of characters, representing the content of the interview in written form.
[0484] A "summary" is a concise compilation of key points extracted from text data, including the main discussion points and conclusions of the interview.
[0485] "User" refers to the person who conducts the interview, starts and stops recording, and sends the data to the server.
[0486] A "terminal" refers to a recording device or computer used by a user, which has the function of recording audio data and sending it to a server.
[0487] A "server" is a central processing unit that receives audio data, stores it, converts it to text data, generates summaries, and analyzes stored historical records.
[0488] "Natural language processing technology" is a technology that enables computers to understand and process human language, and is used for generating summaries of text data and analyzing records.
[0489] A "database" is a system for managing and storing data such as generated summaries and historical records.
[0490] "Temporary storage" refers to the process of temporarily saving audio data between the time it is received and the start of processing.
[0491] A "speech recognition engine" refers to software or services that convert speech data into text data.
[0492] This invention relates to a system for effectively recording and analyzing interviews in shops and call centers. Specific embodiments are described in detail below.
[0493] Hardware and software to be used
[0494] In this invention, a system is built through the collaboration of three parties: the user, the terminal, and the server. The hardware and software used are shown below.
[0495] User device: A device with recording capabilities, such as a smartphone, tablet, or personal computer.
[0496] Server: A device with high-performance computing power and storage, such as a cloud server or on-premises server.
[0497] Speech recognition engine: Software used to convert speech data into text data (e.g., Google Speech-to-Text API).
[0498] Generative AI models: Artificial intelligence models for generating summaries of text data (e.g., Transformers in Hugging Face).
[0499] Database: A database system (e.g., SQLite) for storing generated summaries and interview records.
[0500] Overview of program processing
[0501] The user records the interview and sends it to the server from their device.
[0502] The user records audio during the interview using an application on their device. Once recording is complete, they press the "Stop Recording" button to end the recording and save the audio file to the device's local storage. Then, the user clicks the "Send" button to send the audio file to the server via an HTTP POST request. For example, the audio file might be saved at " / local / path / to / record.wav" and sent to the server.
[0503] The server receives the audio file and saves it locally.
[0504] The server saves the received audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav"). In this step, the server verifies the format and size of the received file and saves only valid audio data.
[0505] Convert audio data to text data
[0506] The server sends the stored audio files to the speech recognition engine. For example, it uses the Google Speech-to-Text API to convert the audio files into text data. Then, it retrieves the converted text data and passes it to a generative AI model.
[0507] Generate a summary from text data
[0508] The server uses a generative AI model (e.g., Hugging Face's Transformers) to extract key points from text data and generate a summary. This summary generation is instructed using prompts. An example prompt is, "Summarize the main points of this text data." The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0509] Save the summary and audio file path to the database.
[0510] The server records the generated summary and the path to the audio file in a database. For example, it might use a "records" table in an SQLite database to store the summary and path in association. This allows for easy retrieval from the database later when needed.
[0511] We analyze past records and propose effective interview methods.
[0512] The server retrieves past interview records from the database and analyzes them using natural language processing technology. By analyzing multiple summary data and identifying frequently occurring keywords and patterns, it generates specific suggestions for the next interview. These suggestions allow the user to conduct the interview more effectively.
[0513] Specific example
[0514] For example, suppose a user conducts a meeting with a customer one day and records the conversation. After recording, they click the "Send" button to send the audio file to the server. The server receives the audio file and saves it to a path like " / path / to / save / meeting_audio.wav". Then, a speech recognition engine is used to convert it into text data, and a generative AI model is used to generate a summary. The generated summary might be text like, "This meeting mainly focused on customer satisfaction." This summary and file path are stored in a database and later analyzed along with other meeting records to provide material for suggesting more effective meeting methods.
[0515] Examples of prompt messages include, "Please summarize the main points of this text data."
[0516] The above describes a specific embodiment for carrying out the present invention. This system automates the recording, summarizing, and analysis of interviews, thereby significantly improving work efficiency.
[0517] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0518] Step 1:
[0519] The user records the interview audio.
[0520] Input: Interview audio
[0521] Output: Audio files saved to local storage
[0522] Specific operation: Recording begins when the user launches the application on their device and presses the "Start Recording" button. After the interview ends, the user presses the "Stop Recording" button to end the recording and saves the audio file to the device's local storage. For example, the audio data will be saved to the path " / local / path / to / record.wav".
[0523] Step 2:
[0524] The device sends the audio file to the server.
[0525] Input: Audio files stored in local storage
[0526] Output: Audio file sent to the server
[0527] Specific action: The user clicks the "Send" button on the application. The device reads the audio file from local storage and sends it to the server via an HTTP POST request. For example, an audio file saved at " / local / path / to / record.wav" is sent to the server.
[0528] Step 3:
[0529] The server receives the audio file and saves it to a temporary directory.
[0530] Input: Audio file sent via HTTP POST request
[0531] Output: Audio files saved in a temporary directory
[0532] Specific operation: The server receives an HTTP POST request and verifies the file format and size. After confirming that it is valid audio data, it saves the audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav").
[0533] Step 4:
[0534] The server converts the audio data into text data.
[0535] Input: Audio file saved in a temporary directory
[0536] Output: Text data converted from audio data
[0537] Specific operation: The server reads the audio file stored in a temporary directory and sends it to a speech recognition engine (e.g., Google Speech-to-Text API). It receives the text data returned by the speech recognition engine and passes it on to the next process. An example of the converted text would be in the format of "To obtain customer feedback, the following questions were asked."
[0538] Step 5:
[0539] The server generates a summary from the text data it has acquired.
[0540] Input: Text data
[0541] Output: Summary generated from text data
[0542] Specific operation: The server inputs text data into an AI model that generates data (e.g., Hugging Face's Transformers). The prompt is "Summarize the main points of this text data," and the model extracts the key points and generates a summary. An example of a generated summary would be, "This meeting focused on customer satisfaction."
[0543] Step 6:
[0544] The server saves the generated summary and the path to the audio file to the database.
[0545] Input: Generated summary, path to audio file
[0546] Output: Summary and audio file path stored in the database
[0547] Specific operation: The server adds the generated summary and the audio file path to a database (e.g., the "records" table in an SQLite database). Specifically, it uses SQL INSERT statements to record the summary and the corresponding audio file path.
[0548] Step 7:
[0549] The server analyzes past records and generates effective interview methods.
[0550] Input: Past interview records retrieved from the database
[0551] Output: Proposals for effective interview methods
[0552] Specific operation: The server retrieves past interview records from the database and analyzes them using natural language processing technology. It analyzes multiple summary data to identify frequently occurring keywords and patterns. Based on this, it proposes effective interview methods and provides feedback to the user. If the extracted keywords are "customer satisfaction" or "service improvement," it suggests focusing on these points in the next interview.
[0553] (Application Example 1)
[0554] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0555] In autonomous vehicles, there is a lack of effective means to collect and analyze important information regarding communication with passengers and driving conditions. As a result, improvements in driving efficiency and passenger experience are not being fully realized. Furthermore, there is no system that automatically analyzes passenger opinions and feedback to provide appropriate driving advice and feedback for service improvement.
[0556] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0557] In this invention, the server includes means for transmitting voice data to a cloud server and generating suggestions to improve driving efficiency and user experience based on the analysis results, means for providing feedback to the interface of an autonomous vehicle, and means for the cloud server to store the generated summaries in a database. This enables the effective collection, analysis, and summarization of important information regarding communication with passengers and driving conditions, thereby improving driving efficiency and passenger experience.
[0558] "Audio data" refers to data used to record and play back audio information in digital format.
[0559] "Text data" refers to data obtained by analyzing audio data and converting it into textual information.
[0560] A "summary" is a concise compilation of important information extracted from text data.
[0561] A "cloud server" is a remote server accessible via the internet that processes and stores large amounts of data.
[0562] "Driving efficiency" refers to the efficiency of driving a car in terms of fuel consumption, time, and energy consumption.
[0563] "User experience" refers to the overall experience of passengers using autonomous vehicles, including their satisfaction and emotional reactions.
[0564] An "interface" is a device or program that allows humans and machines to exchange information with each other.
[0565] A "database" is a system that systematically stores large amounts of data and allows for efficient searching and updating.
[0566] This invention relates to a system for automatically analyzing conversation recordings collected by autonomous vehicles during operation, extracting important information, and improving driving efficiency and user experience. Details are provided below.
[0567] System configuration and the hardware and software used.
[0568] Hardware to use
[0569] 1. Onboard computer of autonomous vehicle:
[0570] The system collects and partially processes audio data.
[0571] 2. Cloud Server:
[0572] It performs speech analysis, text conversion, database storage, and analysis.
[0573] Software to use
[0574] 1. Speech recognition engine:
[0575] The Google Speech-to-Text API is used to convert speech to text.
[0576] 2. Generative AI models:
[0577] Hugging Face's Transformers are used to extract important information and generate a summary.
[0578] 3. Database system:
[0579] Use MySQL or SQLite to store, search, and analyze large amounts of data.
[0580] System operation
[0581] 1. Collection of audio data:
[0582] Microphones inside the autonomous vehicle record audio, and the data is saved to the onboard computer.
[0583] Send the audio file to the cloud server using an HTTP POST request.
[0584] 2. Text conversion and summary generation of audio data:
[0585] The server saves the received audio files to a temporary directory.
[0586] The speech recognition engine (Google Speech-to-Text API) converts speech to text.
[0587] A generative AI model (Hugging Face's Transformers) is used to generate a summary, which is then saved to a database.
[0588] 3. Analysis of past records and generation of feedback:
[0589] Historical records are retrieved from the database, and common keywords and patterns are identified using NLP (Neuro-Linguistic Programming) techniques.
[0590] It generates effective communication methods and driving advice, and displays them on passenger interfaces such as tablets.
[0591] Specific example of processing
[0592] For example, suppose a passenger is in an autonomous vehicle and is talking about their destination while the vehicle is in motion. Below are some examples of prompts to input into the generating AI model.
[0593] Example of a prompt
[0594] Passenger: How long will it take for this car to reach its destination?
[0595] Vehicle AI: Considering the current traffic conditions, it will take approximately 30 minutes to arrive. Do you have any other questions?
[0596] Passenger: Is it okay if there's traffic?
[0597] Extract and summarize the key points and conclusions of the conversation.
[0598] Based on this prompt, the generative AI model generates the following summary.
[0599] Generated summary
[0600] Key discussion points: Arrival time at destination, traffic conditions, and how to deal with congestion.
[0601] Conclusion: Estimated travel time is approximately 30 minutes, taking traffic conditions into consideration.
[0602] conclusion
[0603] In this way, by applying the invention to autonomous vehicles, it is possible to improve the passenger experience while further increasing driving efficiency. The generated feedback is provided to the driver and passengers through the autonomous vehicle's interface, enabling appropriate advice and service improvements tailored to the driving situation.
[0604] This embodiment of the invention enables effective information collection and analysis in autonomous vehicles, thereby improving both driving efficiency and user experience.
[0605] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0606] Step 1:
[0607] The onboard computer of the autonomous vehicle collects audio data. The onboard computer records audio data using the microphone inside the vehicle, and this data is stored in local storage. The input is conversations between passengers and the driver, and the output is a recorded audio file.
[0608] Step 2:
[0609] The device sends the recorded audio file to the cloud server. The audio file is uploaded to the cloud server using an HTTP POST request. The input is the audio file on the onboard computer, and the output is the audio file stored on the cloud server.
[0610] Step 3:
[0611] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" is used. The input is the uploaded audio file, and the output is the temporary file stored on the server.
[0612] Step 4:
[0613] The server converts the audio file into text data. It uses a speech recognition engine (Google Speech-to-Text API) to convert the audio data into text data. The input is an audio file, and the output is the converted text data.
[0614] Step 5:
[0615] The server extracts key points from text data and generates a summary. It uses a generative AI model (Hugging Face's Transformers) to analyze text data and extract key points. The input is text data, and the output is the generated summary.
[0616] Step 6:
[0617] The server saves the generated summary to a database. The summary is stored in a database (e.g., MySQL or SQLite). The input is the generated summary, and the output is the summary stored in the database.
[0618] Step 7:
[0619] The server retrieves historical records from the database and performs analysis. Natural language processing techniques are used to identify common keywords and patterns. The input is historical records from the database, and the output is the analysis results.
[0620] Step 8:
[0621] The server generates suggestions to improve driving efficiency and user experience based on the analysis results. It generates appropriate driving advice and feedback for service improvement and displays them on the autonomous vehicle's interface. The input is the analysis results, and the output is the generated suggestions and feedback.
[0622] In this way, each processing step is executed sequentially, creating a system that improves the driving efficiency and user experience of autonomous vehicles.
[0623] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0624] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. In particular, it aims to further improve the quality of interviews by incorporating an emotion engine that recognizes user emotions. The interview recording system based on this invention is constructed as follows.
[0625] System Overview
[0626] 1. User (device)
[0627] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[0628] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application that provides the recording function.
[0629] 2. Server
[0630] The server receives the audio file sent from the terminal and saves it to storage.
[0631] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[0632] The system uses generative AI to extract key points from the converted text data and generate a summary.
[0633] In addition to the generated summary, an emotion engine is used to identify the user's emotions from the audio data, and this information is added to the text data and the generated summary.
[0634] The generated summary and sentiment information are stored in a database.
[0635] We retrieve emotional data from a database along with past records, analyze it using natural language processing techniques, and identify effective interview methods.
[0636] Program processing
[0637] The user records audio and sends it to the server from their device.
[0638] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[0639] The server receives the audio file and saves it locally.
[0640] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[0641] The server converts the audio to text.
[0642] For text conversion, the server sends the audio file to the speech recognition engine, which then generates the acquired text data. This process transcribes the contents of the audio file into text.
[0643] The server generates a summary.
[0644] The generated text data is passed to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. This allows the main discussion points of the interview to be concisely summarized.
[0645] The server recognizes emotions.
[0646] In addition to the generated summary, the user's emotions are analyzed from the audio data using an emotion engine (e.g., emotion recognition AI). The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds the results to the text data and summary.
[0647] The server saves the interview records to the database.
[0648] The generated summary and sentiment information are stored in a database. For example, the format in which it is stored in the "records" table of an SQLite database is as follows:
[0649] sql
[0650] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[0651] The server analyzes past records and emotions.
[0652] The server retrieves past interview records and sentiment data from the database and performs analysis using natural language processing techniques. A comprehensive analysis, including sentiment data, is conducted to identify common keywords and patterns.
[0653] The server suggests effective interview methods.
[0654] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data.
[0655] Specific example
[0656] For example, suppose a user conducts a meeting with a customer. This meeting is recorded and sent to a server via the device. The server converts the audio file into text, extracts the key points of the conversation to generate a summary, and uses an emotion engine to identify emotional information. This information may include, for example, "Many positive emotions of customer satisfaction were recognized." This information and the summary are stored in a database and used for future analysis. In the next meeting, the user can use the emotional information from the previous meeting to communicate more effectively.
[0657] Thus, the system of the present invention automates the recording and analysis of interviews, and by further incorporating emotional information, it is possible to improve work efficiency and reduce overtime.
[0658] The following describes the processing flow.
[0659] Step 1:
[0660] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[0661] Step 2:
[0662] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[0663] Step 3:
[0664] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[0665] Step 4:
[0666] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary storage directory. The save location is, for example, a path like " / path / to / save / meeting_audio.wav".
[0667] Step 5:
[0668] The server sends the stored audio files to a speech recognition engine, which converts them into text. For example, the Google Speech-to-Text API is used as the speech recognition engine. The converted text data is a written record of the interview content.
[0669] Step 6:
[0670] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. Hugging Face's Transformers is used as the generative AI model. As a result, a summary is obtained that concisely summarizes the main discussion points of the interview.
[0671] Step 7:
[0672] Next, the server passes the audio data to the emotion engine to identify the user's emotions. An emotion recognition AI (such as IBM Watson's Empathy API) is used as the emotion engine. The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds this information to the text data and summary.
[0673] Step 8:
[0674] The server stores the generated summary and sentiment information in a database. For example, SQLite is used for the database, and the summary, corresponding sentiment information, and audio file paths are inserted into a table called "records".
[0675] Step 9:
[0676] The server retrieves past interview records and emotional data from the database and analyzes them using natural language processing technology. A comprehensive analysis, including emotional data, is performed to identify common keywords and patterns.
[0677] Step 10:
[0678] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data. This suggestion is provided to the user as a strategic guideline for future interviews.
[0679] (Example 2)
[0680] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0681] Conventional interview recording systems required manual recording of interview content, making it difficult to efficiently organize and store information. Furthermore, they lacked the functionality to analyze emotions from audio data to improve interview quality. This resulted in a lack of information needed to identify effective interview methods, leading to decreased work efficiency and limited communication effectiveness. A new system is needed to address these problems.
[0682] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving and temporarily storing voice data, means for converting voice data into text data, and means for extracting important points from the converted text data using a generated AI model and generating a summary. This enables automatic recording and analysis of interview content, and further enables advanced analysis including user sentiment information. This makes it possible to identify effective interview methods and improve work efficiency.
[0683] "Interview audio data" refers to digital audio files that record the content of conversations conducted in person or remotely.
[0684] A "memory device" refers to hardware or databases used to store digital information such as audio data, text data, and sentiment analysis results.
[0685] "Communication function" refers to the function of transmitting digital data to other devices or servers via a network.
[0686] A "server" is a computer system that processes, stores, and analyzes audio data received from terminals via a network.
[0687] "Text data" refers to text information obtained by transcribing audio data.
[0688] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates summaries and responses based on the input data.
[0689] An "emotion engine" is a system that analyzes and recognizes emotions from voice data.
[0690] "Recording" refers to saving information such as generated summaries and sentiment analysis results so that they can be referenced later.
[0691] A "database" is a software system designed to systematically store and manage information, and to allow for searching and updating as needed.
[0692] "Natural language processing technology" refers to artificial intelligence technology used to analyze, understand, and manipulate human language.
[0693] This invention relates to a system that automatically records and analyzes audio data from interviews and proposes effective interview methods. The following describes a specific implementation of this system.
[0694] System Overview
[0695] This system consists of users (terminals), a server, and a database. Users record interviews at shops or call centers using their terminals and send the recorded data to the server.
[0696] Hardware and software to be used
[0697] 1. Hardware
[0698] Device: The recording device used by the user. This includes smartphones, tablets, and personal computers.
[0699] Server: A computer system for receiving and processing audio data.
[0700] 2. Software
[0701] Voice recording application: An application for recording audio on a device and saving it to local storage.
[0702] Communication protocol: HTTP POST request for sending voice data to a server.
[0703] Speech recognition engine: Converts speech data into text data (e.g., Google Speech-to-Text API).
[0704] Generative AI model: Generates summaries from generated text data (e.g., Transformers for Hugging Face).
[0705] Emotion engine: Software for analyzing emotions from voice data (e.g., emotion recognition AI).
[0706] Database: Stores the generated summary and sentiment information (e.g., an SQLite database).
[0707] Natural language processing technology: A technology for analyzing historical records and sentiment data.
[0708] System operation
[0709] The user uses a voice recording application on their device during the interview. Once recording is complete, the audio data is saved locally on the device and uploaded to the server using the application's transmission function. This transmission is done via an HTTP POST request. The server temporarily stores the received audio data and converts it into text data using a speech recognition engine. The converted text data is then passed to a generation AI model, which extracts key points and generates a summary.
[0710] Simultaneously, the server uses an emotion engine to analyze the user's emotions from the audio data. This emotion information is added to the summary, and the final generated summary and emotion information are stored in a database. The server then uses natural language processing techniques to analyze past interview records and emotion information stored in the database, identifying common keywords and patterns to suggest effective interview methods.
[0711] Specific example
[0712] For example, consider a scenario where a user records a meeting with a customer and sends the audio data from their device to a server. The server converts this audio data into text data, extracts key points, and creates a summary. It also analyzes the customer's emotions from the audio data and adds emotional information to the summary, such as "the customer showed many positive emotions of satisfaction." Finally, this data is stored in a database and used as a reference for future meetings.
[0713] Users can use information from previous meetings to select appropriate questions and topics for subsequent meetings. This improves the quality of meetings, leading to increased work efficiency and reduced overtime.
[0714] Example of a prompt
[0715] For example, you can generate a summary by giving the following prompt to the AI model:
[0716] "Please summarize what we discussed in today's meeting."
[0717] Thus, the present invention is a system that automatically records and analyzes voice data, enabling the proposal of interview methods that also take emotional information into account.
[0718] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0719] System program processing flow
[0720] Step 1:
[0721] Audio recording
[0722] The user records audio data during the interview using the device's audio recording application. Recording starts when the record button is pressed and ends when it is pressed again. The completed audio data (input) is saved to the device's local storage (output). For example, it is saved to the path " / local / path / to / meeting_audio.wav".
[0723] Step 2:
[0724] Sending audio files
[0725] The user clicks the upload button within the voice recording application to send the recorded audio file to the server. The sending process is performed using an HTTP POST request. The input is the audio file in local storage, and the output is the audio file uploaded to the server.
[0726] Step 3:
[0727] Receiving and saving audio files
[0728] The server saves audio files received from the user's terminal to a temporary directory. The input is the audio file sent via an HTTP POST request, and the output is the audio file stored in the server's temporary storage area. For example, the audio data is saved to " / path / to / save / meeting_audio.wav".
[0729] Step 4:
[0730] Speech-to-text conversion
[0731] The server uses the stored audio file to call a speech recognition engine (e.g., Google Speech-to-Text API) and converts the audio data into text data. The input is an audio file, and the output is the converted text data. As a result of the transcription, all statements made during the interview are obtained in text format.
[0732] Step 5:
[0733] Summary generation
[0734] The server passes the acquired text data to an AI model (e.g., Hugging Face's Transformers) to extract key points and generate a summary. The input is the acquired text data, and the output is the generated summary text. This provides a concise overview of the interview content.
[0735] Step 6:
[0736] Recognition of emotions
[0737] The server passes the audio data to an emotion engine (e.g., emotion recognition AI) to analyze the user's emotions. The input is audio data, and the output is the result of the emotion analysis. The analyzed emotion information is added to the summary text of the interview.
[0738] Step 7:
[0739] Preservation of products
[0740] The server stores the generated summary and sentiment information in a database. The input is the summary text and sentiment analysis results, and the stored data is the output. For example, it is stored in the "records" table in an SQLite database in the following format:
[0741] sql
[0742] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[0743] Step 8:
[0744] Analysis of past records and emotions
[0745] The server retrieves past interview records and sentiment information from the database and analyzes them using natural language processing techniques. The input is past interview records and sentiment data from the database, and the output is effective interview methods derived from the analysis. This process identifies common keywords and patterns.
[0746] Step 9:
[0747] Suggestions for effective interview methods
[0748] The server suggests effective interview methods to the user based on the analysis results. The input is the analysis results, and the output is the suggested content. Specifically, it suggests questions and topics that include top keywords, as well as interview procedures that take into account past sentiment data.
[0749] (Application Example 2)
[0750] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0751] While conventional interview recording systems could transcribe audio data and analyze past records, they lacked the functionality to improve the quality of customer service in real time. Furthermore, there was no mechanism to accurately capture customer emotions and provide appropriate advice in real time based on those emotions. Therefore, improving customer satisfaction and response efficiency were identified as challenges.
[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data, means for converting the input voice data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating an effective interview method, means for analyzing the user's emotions and adding the emotional information to the summary data, and means for displaying emotional information in real time during the conversation with the customer and suggesting an appropriate conversation method. This makes it possible to improve responses in real time based on the customer's emotions.
[0753] "Audio data" refers to information that can be recorded, stored, or transmitted in digital format.
[0754] "Means of input" refers to devices or functions that allow users to supply voice data to the system.
[0755] "Text data" refers to information obtained by converting audio data into written text.
[0756] "Means of conversion" refers to the functions of software or hardware used to convert audio data into text data.
[0757] "Key points" refer to the most important information or topics extracted from the interview content.
[0758] A "means for generating summaries" refers to a function that extracts important points from text data and summarizes them concisely.
[0759] "Means of recording and preserving" refers to the functions of storage devices and databases for storing generated summaries and other data.
[0760] "Past records" refer to interview data and summary data that have been saved in the past.
[0761] "Means of analysis" refer to software and algorithms used to analyze past records and identify patterns and trends.
[0762] "Effective interview methods" refer to methods and techniques for optimizing communication, which are proposed based on the analysis results.
[0763] "User emotions" refers to the emotional state a user exhibits during an interview (for example, joy, anger, sadness, etc.).
[0764] "Means of analyzing emotions" refers to software or algorithms that identify a user's emotions from voice data and analyze that information.
[0765] "Emotional information" refers to information about a user's emotional state obtained through means of analyzing emotions.
[0766] "Means of displaying in real time" refers to the functions of displays and software that immediately show acquired emotional information to the user.
[0767] "Means of suggesting dialogue methods" refer to software or algorithms that suggest appropriate communication methods to users based on emotional information and other factors.
[0768] This invention relates to a system that automatically records interviews in shops and call centers, recognizes user emotions, and suggests effective interview methods. The system of this invention includes the following components.
[0769] System Overview
[0770] 1. Input of audio data
[0771] The user records conversations with customers using smart glasses. The recorded audio data is temporarily stored in the smart glasses. After recording is complete, the audio data is automatically sent to a server. This transmission requires an internet connection.
[0772] 2. Converting audio data
[0773] The server converts the received audio data into text data using a speech recognition engine such as the Google Speech-to-Text API. The converted text data is stored in a temporary directory.
[0774] 3. Summarization from text data
[0775] Key points are extracted from text data, and a summary is generated using generative AI models such as Hugging Face's Transformers. The generated summary data is temporarily stored for further analysis.
[0776] 4. Recognition of emotions
[0777] The server identifies the user's emotions from the voice data using an emotion recognition engine such as IBM Watson Tone Analyzer. The identified emotion information is added to the summary data.
[0778] 5. Proposal for a real-time dialogue method
[0779] Based on the generated summary and sentiment information, the server formulates a response plan in real time and displays it on the smart glasses' display. This allows staff to get the appropriate response plan on the spot.
[0780] 6. Recording and Preservation
[0781] The generated summaries and sentiment information are stored in a storage device such as an SQLite database. This makes the data reusable at any time.
[0782] 7. Analysis of past records
[0783] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. By analyzing past conversation content and emotional data, and extracting common patterns and trends, it generates and proposes effective interview methods.
[0784] Specific example
[0785] For example, if a customer asks for a product description, the smart glasses record the conversation and send it to a server. The server converts the recorded audio data into text, generates a summary using a generative AI model, and adds customer sentiment information to the summary. Next, real-time advice such as "The customer is satisfied, so we will suggest further related products" is displayed on the smart glasses' screen, allowing staff to take appropriate action immediately.
[0786] Example of a prompt
[0787] The following prompt messages can be used to efficiently analyze customer interactions and suggest appropriate responses.
[0788] "Could you please explain the latest model in detail?"
[0789] "Are there any other products with similar functions?"
[0790] "Do you have a demo that clearly explains how to use it?"
[0791] As a result, the system of the present invention can improve the quality of communication with customers and enable effective responses.
[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0793] Step 1:
[0794] The user records conversations with customers using smart glasses. The input is the audio of the conversation with the customer, which the smart glasses record as audio data. The output is a temporarily stored audio file. Recording is possible through the recording function within the smart glasses.
[0795] Step 2:
[0796] After recording is complete, the audio data is automatically sent to the server. The input here is the audio file stored in the smart glasses, which is transferred to the server via an internet connection. The output is the audio file stored on the server. This transmission is performed via an HTTP POST request.
[0797] Step 3:
[0798] The server converts the received audio data into text data using the Google Speech-to-Text API. An audio file is used as input, and the output is text data. The server sends the audio data to the API and saves the returned text data to a temporary directory.
[0799] Step 4:
[0800] The server extracts key points from the transformed text data and generates a summary using a generative AI model (e.g., Hugging Face's Transformers). The input is text data, and the output is summary data. The server invokes the generative AI model, generates the summary data, and temporarily stores it.
[0801] Step 5:
[0802] The server identifies the user's emotions from audio data using IBM Watson Tone Analyzer. The input is an audio file, and the output is emotion information. The server sends the audio data to the emotion recognition engine and adds the received emotion information to the summary data.
[0803] Step 6:
[0804] The server proposes a dialogue method to the user in real time based on the generated summary and sentiment information. The input is the summary and sentiment information, and the output is the proposed dialogue method. Based on this data, the server formulates an appropriate response and displays it on the smart glasses' display.
[0805] Step 7:
[0806] The server saves the generated summary and sentiment information to an SQLite database. The input is the summary and sentiment information, and the output is the records stored in the database. The server performs the save process and persists the data.
[0807] Step 8:
[0808] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. The input is historical record data, and the output is the analysis results and effective interview methods. The server reads data from the database and applies algorithms to analyze patterns and trends.
[0809] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0810] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0811] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0812] [Third Embodiment]
[0813] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0814] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0815] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0816] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0817] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0818] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0819] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0820] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0821] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0822] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0823] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0824] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0825] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. The interview recording system based on this invention is constructed as follows.
[0826] System Overview
[0827] 1. User (device)
[0828] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[0829] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application on the device.
[0830] 2. Server
[0831] The server receives the audio file sent from the terminal and saves it to storage.
[0832] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[0833] The system uses generative AI to extract key points from the converted text data and generate a summary.
[0834] The generated summary is saved to the database.
[0835] By retrieving past records from a database and analyzing them using natural language processing techniques, we can identify effective interview methods.
[0836] Program processing
[0837] The user records audio and sends it to the server from their device.
[0838] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[0839] The server receives the audio file and saves it locally.
[0840] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[0841] The server converts the audio to text and generates a summary.
[0842] For text conversion, the server sends the audio file to a speech recognition engine, which then passes the acquired text data to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0843] The server saves the interview records to the database.
[0844] The generated summary and audio file paths are stored in a database on the server. For example, they might be stored in the "records" table of an SQLite database.
[0845] The server analyzes past records and suggests effective interview methods.
[0846] The server retrieves past interview records from the database and uses natural language processing technology to identify common keywords and patterns. This allows it to generate effective interview methods to improve interview efficiency and provide feedback to the user.
[0847] Specific example
[0848] For example, suppose a user has a meeting with a customer one day. They record this meeting and send it to a server via their device. The server converts the audio file into text, automatically extracts the key points of the conversation, generates a summary, and saves it to a database. A few weeks later, the user analyzes multiple meeting records and identifies frequently occurring keywords, revealing that "customer satisfaction," "product quality," and "service improvement" are important points. This allows them to focus on these keywords in future meetings, leading to more effective communication.
[0849] Thus, the system of the present invention automates the recording and analysis of interviews, thereby improving work efficiency and reducing overtime.
[0850] The following describes the processing flow.
[0851] Step 1:
[0852] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[0853] Step 2:
[0854] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[0855] Step 3:
[0856] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[0857] Step 4:
[0858] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary directory. For example, the file is saved to " / path / to / save / meeting_audio.wav".
[0859] Step 5:
[0860] The server sends the stored audio file to a speech recognition engine, which converts it to text. The server reads the audio file and uses a speech recognition API (such as the Google Speech-to-Text API) to transcribe the audio. As a result, the content of the interview is obtained as text data.
[0861] Step 6:
[0862] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. The generative AI model uses tools such as Hugging Face's Transformers to perform the summarization. As a result, a concise summary is obtained that summarizes the main points of the interview discussion.
[0863] Step 7:
[0864] The server saves the generated summary and the path to the original audio file to a database. For example, SQLite is used for the database, and the summary and the path to the corresponding audio file are inserted into a table called "records".
[0865] Step 8:
[0866] The server retrieves past interview records from the database and analyzes them using natural language processing technology. This identifies common keywords and patterns in the interview content. The analysis results serve as reference material for developing more effective interview methods.
[0867] Step 9:
[0868] Based on the analysis results, the server suggests effective interview methods to the user. Specifically, it suggests incorporating questions and topics that include the top keywords. This suggestion is provided to the user as a strategic guideline for future interviews.
[0869] (Example 1)
[0870] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0871] Traditional interview recording systems have the problem of requiring significant time and effort to manually record interview content and summarize key points. Furthermore, there was a lack of concrete methods for effectively analyzing past interview records and using that information to improve future interviews. This led to concerns that interview efficiency would decrease, resulting in a decline in the quality of customer service and overall work.
[0872] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0873] In this invention, the server includes means for converting audio data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating effective interview methods, means for the user to record audio data and send it to the server from a terminal, means for the server to temporarily store the received audio data, means for the server to generate a summary using natural language processing technology, and means for the server to save the summary and analyze past records. As a result, the recording and summary generation of interviews and the analysis of past records are automated, making it possible to quickly propose effective interview methods.
[0874] "Audio data" refers to digital audio recordings of conversations and statements made by the user during an interview.
[0875] "Text data" refers to audio data converted into a string of characters, representing the content of the interview in written form.
[0876] A "summary" is a concise compilation of key points extracted from text data, including the main discussion points and conclusions of the interview.
[0877] "User" refers to the person who conducts the interview, starts and stops recording, and sends the data to the server.
[0878] A "terminal" refers to a recording device or computer used by a user, which has the function of recording audio data and sending it to a server.
[0879] A "server" is a central processing unit that receives audio data, stores it, converts it to text data, generates summaries, and analyzes stored historical records.
[0880] "Natural language processing technology" is a technology that enables computers to understand and process human language, and is used for generating summaries of text data and analyzing records.
[0881] A "database" is a system for managing and storing data such as generated summaries and historical records.
[0882] "Temporary storage" refers to the process of temporarily saving audio data between the time it is received and the start of processing.
[0883] A "speech recognition engine" refers to software or services that convert speech data into text data.
[0884] This invention relates to a system for effectively recording and analyzing interviews in shops and call centers. Specific embodiments are described in detail below.
[0885] Hardware and software to be used
[0886] In this invention, a system is built through the collaboration of three parties: the user, the terminal, and the server. The hardware and software used are shown below.
[0887] User device: A device with recording capabilities, such as a smartphone, tablet, or personal computer.
[0888] Server: A device with high-performance computing power and storage, such as a cloud server or on-premises server.
[0889] Speech recognition engine: Software used to convert speech data into text data (e.g., Google Speech-to-Text API).
[0890] Generative AI models: Artificial intelligence models for generating summaries of text data (e.g., Transformers in Hugging Face).
[0891] Database: A database system (e.g., SQLite) for storing generated summaries and interview records.
[0892] Overview of program processing
[0893] The user records the interview and sends it to the server from their device.
[0894] The user records audio during the interview using an application on their device. Once recording is complete, they press the "Stop Recording" button to end the recording and save the audio file to the device's local storage. Then, the user clicks the "Send" button to send the audio file to the server via an HTTP POST request. For example, the audio file might be saved at " / local / path / to / record.wav" and sent to the server.
[0895] The server receives the audio file and saves it locally.
[0896] The server saves the received audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav"). In this step, the server verifies the format and size of the received file and saves only valid audio data.
[0897] Convert audio data to text data
[0898] The server sends the stored audio files to the speech recognition engine. For example, it uses the Google Speech-to-Text API to convert the audio files into text data. Then, it retrieves the converted text data and passes it to a generative AI model.
[0899] Generate a summary from text data
[0900] The server uses a generative AI model (e.g., Hugging Face's Transformers) to extract key points from text data and generate a summary. This summary generation is instructed using prompts. An example prompt is, "Summarize the main points of this text data." The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[0901] Save the summary and audio file path to the database.
[0902] The server records the generated summary and the path to the audio file in a database. For example, it might use a "records" table in an SQLite database to store the summary and path in association. This allows for easy retrieval from the database later when needed.
[0903] We analyze past records and propose effective interview methods.
[0904] The server retrieves past interview records from the database and analyzes them using natural language processing technology. By analyzing multiple summary data and identifying frequently occurring keywords and patterns, it generates specific suggestions for the next interview. These suggestions allow the user to conduct the interview more effectively.
[0905] Specific example
[0906] For example, suppose a user conducts a meeting with a customer one day and records the conversation. After recording, they click the "Send" button to send the audio file to the server. The server receives the audio file and saves it to a path like " / path / to / save / meeting_audio.wav". Then, a speech recognition engine is used to convert it into text data, and a generative AI model is used to generate a summary. The generated summary might be text like, "This meeting mainly focused on customer satisfaction." This summary and file path are stored in a database and later analyzed along with other meeting records to provide material for suggesting more effective meeting methods.
[0907] Examples of prompt messages include, "Please summarize the main points of this text data."
[0908] The above describes a specific embodiment for carrying out the present invention. This system automates the recording, summarizing, and analysis of interviews, thereby significantly improving work efficiency.
[0909] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0910] Step 1:
[0911] The user records the interview audio.
[0912] Input: Interview audio
[0913] Output: Audio files saved to local storage
[0914] Specific operation: Recording begins when the user launches the application on their device and presses the "Start Recording" button. After the interview ends, the user presses the "Stop Recording" button to end the recording and saves the audio file to the device's local storage. For example, the audio data will be saved to the path " / local / path / to / record.wav".
[0915] Step 2:
[0916] The device sends the audio file to the server.
[0917] Input: Audio files stored in local storage
[0918] Output: Audio file sent to the server
[0919] Specific action: The user clicks the "Send" button on the application. The device reads the audio file from local storage and sends it to the server via an HTTP POST request. For example, an audio file saved at " / local / path / to / record.wav" is sent to the server.
[0920] Step 3:
[0921] The server receives the audio file and saves it to a temporary directory.
[0922] Input: Audio file sent via HTTP POST request
[0923] Output: Audio files saved in a temporary directory
[0924] Specific operation: The server receives an HTTP POST request and verifies the file format and size. After confirming that it is valid audio data, it saves the audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav").
[0925] Step 4:
[0926] The server converts the audio data into text data.
[0927] Input: Audio file saved in a temporary directory
[0928] Output: Text data converted from audio data
[0929] Specific operation: The server reads the audio file stored in a temporary directory and sends it to a speech recognition engine (e.g., Google Speech-to-Text API). It receives the text data returned by the speech recognition engine and passes it on to the next process. An example of the converted text would be in the format of "To obtain customer feedback, the following questions were asked."
[0930] Step 5:
[0931] The server generates a summary from the text data it has acquired.
[0932] Input: Text data
[0933] Output: Summary generated from text data
[0934] Specific operation: The server inputs text data into an AI model that generates data (e.g., Hugging Face's Transformers). The prompt is "Summarize the main points of this text data," and the model extracts the key points and generates a summary. An example of a generated summary would be, "This meeting focused on customer satisfaction."
[0935] Step 6:
[0936] The server saves the generated summary and the path to the audio file to the database.
[0937] Input: Generated summary, path to audio file
[0938] Output: Summary and audio file path stored in the database
[0939] Specific operation: The server adds the generated summary and the audio file path to a database (e.g., the "records" table in an SQLite database). Specifically, it uses SQL INSERT statements to record the summary and the corresponding audio file path.
[0940] Step 7:
[0941] The server analyzes past records and generates effective interview methods.
[0942] Input: Past interview records retrieved from the database
[0943] Output: Proposals for effective interview methods
[0944] Specific operation: The server retrieves past interview records from the database and analyzes them using natural language processing technology. It analyzes multiple summary data to identify frequently occurring keywords and patterns. Based on this, it proposes effective interview methods and provides feedback to the user. If the extracted keywords are "customer satisfaction" or "service improvement," it suggests focusing on these points in the next interview.
[0945] (Application Example 1)
[0946] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0947] In autonomous vehicles, there is a lack of effective means to collect and analyze important information regarding communication with passengers and driving conditions. As a result, improvements in driving efficiency and passenger experience are not being fully realized. Furthermore, there is no system that automatically analyzes passenger opinions and feedback to provide appropriate driving advice and feedback for service improvement.
[0948] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0949] In this invention, the server includes means for transmitting voice data to a cloud server and generating suggestions to improve driving efficiency and user experience based on the analysis results, means for providing feedback to the interface of an autonomous vehicle, and means for the cloud server to store the generated summaries in a database. This enables the effective collection, analysis, and summarization of important information regarding communication with passengers and driving conditions, thereby improving driving efficiency and passenger experience.
[0950] "Audio data" refers to data used to record and play back audio information in digital format.
[0951] "Text data" refers to data obtained by analyzing audio data and converting it into textual information.
[0952] A "summary" is a concise compilation of important information extracted from text data.
[0953] A "cloud server" is a remote server accessible via the internet that processes and stores large amounts of data.
[0954] "Driving efficiency" refers to the efficiency of driving a car in terms of fuel consumption, time, and energy consumption.
[0955] "User experience" refers to the overall experience of passengers using autonomous vehicles, including their satisfaction and emotional reactions.
[0956] An "interface" is a device or program that allows humans and machines to exchange information with each other.
[0957] A "database" is a system that systematically stores large amounts of data and allows for efficient searching and updating.
[0958] This invention relates to a system for automatically analyzing conversation recordings collected by autonomous vehicles during operation, extracting important information, and improving driving efficiency and user experience. Details are provided below.
[0959] System configuration and the hardware and software used.
[0960] Hardware to use
[0961] 1. Onboard computer of autonomous vehicle:
[0962] The system collects and partially processes audio data.
[0963] 2. Cloud Server:
[0964] It performs speech analysis, text conversion, database storage, and analysis.
[0965] Software to use
[0966] 1. Speech recognition engine:
[0967] The Google Speech-to-Text API is used to convert speech to text.
[0968] 2. Generative AI models:
[0969] Hugging Face's Transformers are used to extract important information and generate a summary.
[0970] 3. Database system:
[0971] Use MySQL or SQLite to store, search, and analyze large amounts of data.
[0972] System operation
[0973] 1. Collection of audio data:
[0974] Microphones inside the autonomous vehicle record audio, and the data is saved to the onboard computer.
[0975] Send the audio file to the cloud server using an HTTP POST request.
[0976] 2. Text conversion and summary generation of audio data:
[0977] The server saves the received audio files to a temporary directory.
[0978] The speech recognition engine (Google Speech-to-Text API) converts speech to text.
[0979] A generative AI model (Hugging Face's Transformers) is used to generate a summary, which is then saved to a database.
[0980] 3. Analysis of past records and generation of feedback:
[0981] Historical records are retrieved from the database, and common keywords and patterns are identified using NLP (Neuro-Linguistic Programming) techniques.
[0982] It generates effective communication methods and driving advice, and displays them on passenger interfaces such as tablets.
[0983] Specific example of processing
[0984] For example, suppose a passenger is in an autonomous vehicle and is talking about their destination while the vehicle is in motion. Below are some examples of prompts to input into the generating AI model.
[0985] Example of a prompt
[0986] Passenger: How long will it take for this car to reach its destination?
[0987] Vehicle AI: Considering the current traffic conditions, it will take approximately 30 minutes to arrive. Do you have any other questions?
[0988] Passenger: Is it okay if there's traffic?
[0989] Extract and summarize the key points and conclusions of the conversation.
[0990] Based on this prompt, the generative AI model generates the following summary.
[0991] Generated summary
[0992] Key discussion points: Arrival time at destination, traffic conditions, and how to deal with congestion.
[0993] Conclusion: Estimated travel time is approximately 30 minutes, taking traffic conditions into consideration.
[0994] conclusion
[0995] In this way, by applying the invention to autonomous vehicles, it is possible to improve the passenger experience while further increasing driving efficiency. The generated feedback is provided to the driver and passengers through the autonomous vehicle's interface, enabling appropriate advice and service improvements tailored to the driving situation.
[0996] This embodiment of the invention enables effective information collection and analysis in autonomous vehicles, thereby improving both driving efficiency and user experience.
[0997] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0998] Step 1:
[0999] The onboard computer of the autonomous vehicle collects audio data. The onboard computer records audio data using the microphone inside the vehicle, and this data is stored in local storage. The input is conversations between passengers and the driver, and the output is a recorded audio file.
[1000] Step 2:
[1001] The device sends the recorded audio file to the cloud server. The audio file is uploaded to the cloud server using an HTTP POST request. The input is the audio file on the onboard computer, and the output is the audio file stored on the cloud server.
[1002] Step 3:
[1003] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" is used. The input is the uploaded audio file, and the output is the temporary file stored on the server.
[1004] Step 4:
[1005] The server converts the audio file into text data. It uses a speech recognition engine (Google Speech-to-Text API) to convert the audio data into text data. The input is an audio file, and the output is the converted text data.
[1006] Step 5:
[1007] The server extracts key points from text data and generates a summary. It uses a generative AI model (Hugging Face's Transformers) to analyze text data and extract key points. The input is text data, and the output is the generated summary.
[1008] Step 6:
[1009] The server saves the generated summary to a database. The summary is stored in a database (e.g., MySQL or SQLite). The input is the generated summary, and the output is the summary stored in the database.
[1010] Step 7:
[1011] The server retrieves historical records from the database and performs analysis. Natural language processing techniques are used to identify common keywords and patterns. The input is historical records from the database, and the output is the analysis results.
[1012] Step 8:
[1013] The server generates suggestions to improve driving efficiency and user experience based on the analysis results. It generates appropriate driving advice and feedback for service improvement and displays them on the autonomous vehicle's interface. The input is the analysis results, and the output is the generated suggestions and feedback.
[1014] In this way, each processing step is executed sequentially, creating a system that improves the driving efficiency and user experience of autonomous vehicles.
[1015] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1016] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. In particular, it aims to further improve the quality of interviews by incorporating an emotion engine that recognizes user emotions. The interview recording system based on this invention is constructed as follows.
[1017] System Overview
[1018] 1. User (device)
[1019] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[1020] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application that provides the recording function.
[1021] 2. Server
[1022] The server receives the audio file sent from the terminal and saves it to storage.
[1023] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[1024] The system uses generative AI to extract key points from the converted text data and generate a summary.
[1025] In addition to the generated summary, an emotion engine is used to identify the user's emotions from the audio data, and this information is added to the text data and the generated summary.
[1026] The generated summary and sentiment information are stored in a database.
[1027] We retrieve emotional data from a database along with past records, analyze it using natural language processing techniques, and identify effective interview methods.
[1028] Program processing
[1029] The user records audio and sends it to the server from their device.
[1030] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[1031] The server receives the audio file and saves it locally.
[1032] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[1033] The server converts the audio to text.
[1034] For text conversion, the server sends the audio file to the speech recognition engine, which then generates the acquired text data. This process transcribes the contents of the audio file into text.
[1035] The server generates a summary.
[1036] The generated text data is passed to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. This allows the main discussion points of the interview to be concisely summarized.
[1037] The server recognizes emotions.
[1038] In addition to the generated summary, the user's emotions are analyzed from the audio data using an emotion engine (e.g., emotion recognition AI). The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds the results to the text data and summary.
[1039] The server saves the interview records to the database.
[1040] The generated summary and sentiment information are stored in a database. For example, the format in which it is stored in the "records" table of an SQLite database is as follows:
[1041] sql
[1042] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[1043] The server analyzes past records and emotions.
[1044] The server retrieves past interview records and sentiment data from the database and performs analysis using natural language processing techniques. A comprehensive analysis, including sentiment data, is conducted to identify common keywords and patterns.
[1045] The server suggests effective interview methods.
[1046] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data.
[1047] Specific example
[1048] For example, suppose a user conducts a meeting with a customer. This meeting is recorded and sent to a server via the device. The server converts the audio file into text, extracts the key points of the conversation to generate a summary, and uses an emotion engine to identify emotional information. This information may include, for example, "Many positive emotions of customer satisfaction were recognized." This information and the summary are stored in a database and used for future analysis. In the next meeting, the user can use the emotional information from the previous meeting to communicate more effectively.
[1049] Thus, the system of the present invention automates the recording and analysis of interviews, and by further incorporating emotional information, it is possible to improve work efficiency and reduce overtime.
[1050] The following describes the processing flow.
[1051] Step 1:
[1052] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[1053] Step 2:
[1054] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[1055] Step 3:
[1056] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[1057] Step 4:
[1058] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary storage directory. The save location is, for example, a path like " / path / to / save / meeting_audio.wav".
[1059] Step 5:
[1060] The server sends the stored audio files to a speech recognition engine, which converts them into text. For example, the Google Speech-to-Text API is used as the speech recognition engine. The converted text data is a written record of the interview content.
[1061] Step 6:
[1062] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. Hugging Face's Transformers is used as the generative AI model. As a result, a summary is obtained that concisely summarizes the main discussion points of the interview.
[1063] Step 7:
[1064] Next, the server passes the audio data to the emotion engine to identify the user's emotions. An emotion recognition AI (such as IBM Watson's Empathy API) is used as the emotion engine. The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds this information to the text data and summary.
[1065] Step 8:
[1066] The server stores the generated summary and sentiment information in a database. For example, SQLite is used for the database, and the summary, corresponding sentiment information, and audio file paths are inserted into a table called "records".
[1067] Step 9:
[1068] The server retrieves past interview records and emotional data from the database and analyzes them using natural language processing technology. A comprehensive analysis, including emotional data, is performed to identify common keywords and patterns.
[1069] Step 10:
[1070] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data. This suggestion is provided to the user as a strategic guideline for future interviews.
[1071] (Example 2)
[1072] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1073] Conventional interview recording systems required manual recording of interview content, making it difficult to efficiently organize and store information. Furthermore, they lacked the functionality to analyze emotions from audio data to improve interview quality. This resulted in a lack of information needed to identify effective interview methods, leading to decreased work efficiency and limited communication effectiveness. A new system is needed to address these problems.
[1074] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving and temporarily storing voice data, means for converting voice data into text data, and means for extracting important points from the converted text data using a generated AI model and generating a summary. This enables automatic recording and analysis of interview content, and further enables advanced analysis including user sentiment information. This makes it possible to identify effective interview methods and improve work efficiency.
[1075] "Interview audio data" refers to digital audio files that record the content of conversations conducted in person or remotely.
[1076] A "memory device" refers to hardware or databases used to store digital information such as audio data, text data, and sentiment analysis results.
[1077] "Communication function" refers to the function of transmitting digital data to other devices or servers via a network.
[1078] A "server" is a computer system that processes, stores, and analyzes audio data received from terminals via a network.
[1079] "Text data" refers to text information obtained by transcribing audio data.
[1080] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates summaries and responses based on the input data.
[1081] An "emotion engine" is a system that analyzes and recognizes emotions from voice data.
[1082] "Recording" refers to saving information such as generated summaries and sentiment analysis results so that they can be referenced later.
[1083] A "database" is a software system designed to systematically store and manage information, and to allow for searching and updating as needed.
[1084] "Natural language processing technology" refers to artificial intelligence technology used to analyze, understand, and manipulate human language.
[1085] This invention relates to a system that automatically records and analyzes audio data from interviews and proposes effective interview methods. The following describes a specific implementation of this system.
[1086] System Overview
[1087] This system consists of users (terminals), a server, and a database. Users record interviews at shops or call centers using their terminals and send the recorded data to the server.
[1088] Hardware and software to be used
[1089] 1. Hardware
[1090] Device: The recording device used by the user. This includes smartphones, tablets, and personal computers.
[1091] Server: A computer system for receiving and processing audio data.
[1092] 2. Software
[1093] Voice recording application: An application for recording audio on a device and saving it to local storage.
[1094] Communication protocol: HTTP POST request for sending voice data to a server.
[1095] Speech recognition engine: Converts speech data into text data (e.g., Google Speech-to-Text API).
[1096] Generative AI model: Generates summaries from generated text data (e.g., Transformers for Hugging Face).
[1097] Emotion engine: Software for analyzing emotions from voice data (e.g., emotion recognition AI).
[1098] Database: Stores the generated summary and sentiment information (e.g., an SQLite database).
[1099] Natural language processing technology: A technology for analyzing historical records and sentiment data.
[1100] System operation
[1101] The user uses a voice recording application on their device during the interview. Once recording is complete, the audio data is saved locally on the device and uploaded to the server using the application's transmission function. This transmission is done via an HTTP POST request. The server temporarily stores the received audio data and converts it into text data using a speech recognition engine. The converted text data is then passed to a generation AI model, which extracts key points and generates a summary.
[1102] Simultaneously, the server uses an emotion engine to analyze the user's emotions from the audio data. This emotion information is added to the summary, and the final generated summary and emotion information are stored in a database. The server then uses natural language processing techniques to analyze past interview records and emotion information stored in the database, identifying common keywords and patterns to suggest effective interview methods.
[1103] Specific example
[1104] For example, consider a scenario where a user records a meeting with a customer and sends the audio data from their device to a server. The server converts this audio data into text data, extracts key points, and creates a summary. It also analyzes the customer's emotions from the audio data and adds emotional information to the summary, such as "the customer showed many positive emotions of satisfaction." Finally, this data is stored in a database and used as a reference for future meetings.
[1105] Users can use information from previous meetings to select appropriate questions and topics for subsequent meetings. This improves the quality of meetings, leading to increased work efficiency and reduced overtime.
[1106] Example of a prompt
[1107] For example, you can generate a summary by giving the following prompt to the AI model:
[1108] "Please summarize what we discussed in today's meeting."
[1109] Thus, the present invention is a system that automatically records and analyzes voice data, enabling the proposal of interview methods that also take emotional information into account.
[1110] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1111] System program processing flow
[1112] Step 1:
[1113] Audio recording
[1114] The user records audio data during the interview using the device's audio recording application. Recording starts when the record button is pressed and ends when it is pressed again. The completed audio data (input) is saved to the device's local storage (output). For example, it is saved to the path " / local / path / to / meeting_audio.wav".
[1115] Step 2:
[1116] Sending audio files
[1117] The user clicks the upload button within the voice recording application to send the recorded audio file to the server. The sending process is performed using an HTTP POST request. The input is the audio file in local storage, and the output is the audio file uploaded to the server.
[1118] Step 3:
[1119] Receiving and saving audio files
[1120] The server saves audio files received from the user's terminal to a temporary directory. The input is the audio file sent via an HTTP POST request, and the output is the audio file stored in the server's temporary storage area. For example, the audio data is saved to " / path / to / save / meeting_audio.wav".
[1121] Step 4:
[1122] Speech-to-text conversion
[1123] The server uses the stored audio file to call a speech recognition engine (e.g., Google Speech-to-Text API) and converts the audio data into text data. The input is an audio file, and the output is the converted text data. As a result of the transcription, all statements made during the interview are obtained in text format.
[1124] Step 5:
[1125] Summary generation
[1126] The server passes the acquired text data to an AI model (e.g., Hugging Face's Transformers) to extract key points and generate a summary. The input is the acquired text data, and the output is the generated summary text. This provides a concise overview of the interview content.
[1127] Step 6:
[1128] Recognition of emotions
[1129] The server passes the audio data to an emotion engine (e.g., emotion recognition AI) to analyze the user's emotions. The input is audio data, and the output is the result of the emotion analysis. The analyzed emotion information is added to the summary text of the interview.
[1130] Step 7:
[1131] Preservation of products
[1132] The server stores the generated summary and sentiment information in a database. The input is the summary text and sentiment analysis results, and the stored data is the output. For example, it is stored in the "records" table in an SQLite database in the following format:
[1133] sql
[1134] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[1135] Step 8:
[1136] Analysis of past records and emotions
[1137] The server retrieves past interview records and sentiment information from the database and analyzes them using natural language processing techniques. The input is past interview records and sentiment data from the database, and the output is effective interview methods derived from the analysis. This process identifies common keywords and patterns.
[1138] Step 9:
[1139] Suggestions for effective interview methods
[1140] The server suggests effective interview methods to the user based on the analysis results. The input is the analysis results, and the output is the suggested content. Specifically, it suggests questions and topics that include top keywords, as well as interview procedures that take into account past sentiment data.
[1141] (Application Example 2)
[1142] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1143] While conventional interview recording systems could transcribe audio data and analyze past records, they lacked the functionality to improve the quality of customer service in real time. Furthermore, there was no mechanism to accurately capture customer emotions and provide appropriate advice in real time based on those emotions. Therefore, improving customer satisfaction and response efficiency were identified as challenges.
[1144] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data, means for converting the input voice data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating an effective interview method, means for analyzing the user's emotions and adding the emotional information to the summary data, and means for displaying emotional information in real time during the conversation with the customer and suggesting an appropriate conversation method. This makes it possible to improve responses in real time based on the customer's emotions.
[1145] "Audio data" refers to information that can be recorded, stored, or transmitted in digital format.
[1146] "Means of input" refers to devices or functions that allow users to supply voice data to the system.
[1147] "Text data" refers to information obtained by converting audio data into written text.
[1148] "Means of conversion" refers to the functions of software or hardware used to convert audio data into text data.
[1149] "Key points" refer to the most important information or topics extracted from the interview content.
[1150] A "means for generating summaries" refers to a function that extracts important points from text data and summarizes them concisely.
[1151] "Means of recording and preserving" refers to the functions of storage devices and databases for storing generated summaries and other data.
[1152] "Past records" refer to interview data and summary data that have been saved in the past.
[1153] "Means of analysis" refer to software and algorithms used to analyze past records and identify patterns and trends.
[1154] "Effective interview methods" refer to methods and techniques for optimizing communication, which are proposed based on the analysis results.
[1155] "User emotions" refers to the emotional state a user exhibits during an interview (for example, joy, anger, sadness, etc.).
[1156] "Means of analyzing emotions" refers to software or algorithms that identify a user's emotions from voice data and analyze that information.
[1157] "Emotional information" refers to information about a user's emotional state obtained through means of analyzing emotions.
[1158] "Means of displaying in real time" refers to the functions of displays and software that immediately show acquired emotional information to the user.
[1159] "Means of suggesting dialogue methods" refer to software or algorithms that suggest appropriate communication methods to users based on emotional information and other factors.
[1160] This invention relates to a system that automatically records interviews in shops and call centers, recognizes user emotions, and suggests effective interview methods. The system of this invention includes the following components.
[1161] System Overview
[1162] 1. Input of audio data
[1163] The user records conversations with customers using smart glasses. The recorded audio data is temporarily stored in the smart glasses. After recording is complete, the audio data is automatically sent to a server. This transmission requires an internet connection.
[1164] 2. Converting audio data
[1165] The server converts the received audio data into text data using a speech recognition engine such as the Google Speech-to-Text API. The converted text data is stored in a temporary directory.
[1166] 3. Summarization from text data
[1167] Key points are extracted from text data, and a summary is generated using generative AI models such as Hugging Face's Transformers. The generated summary data is temporarily stored for further analysis.
[1168] 4. Recognition of emotions
[1169] The server identifies the user's emotions from the voice data using an emotion recognition engine such as IBM Watson Tone Analyzer. The identified emotion information is added to the summary data.
[1170] 5. Proposal for a real-time dialogue method
[1171] Based on the generated summary and sentiment information, the server formulates a response plan in real time and displays it on the smart glasses' display. This allows staff to get the appropriate response plan on the spot.
[1172] 6. Recording and Preservation
[1173] The generated summaries and sentiment information are stored in a storage device such as an SQLite database. This makes the data reusable at any time.
[1174] 7. Analysis of past records
[1175] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. By analyzing past conversation content and emotional data, and extracting common patterns and trends, it generates and proposes effective interview methods.
[1176] Specific example
[1177] For example, if a customer asks for a product description, the smart glasses record the conversation and send it to a server. The server converts the recorded audio data into text, generates a summary using a generative AI model, and adds customer sentiment information to the summary. Next, real-time advice such as "The customer is satisfied, so we will suggest further related products" is displayed on the smart glasses' screen, allowing staff to take appropriate action immediately.
[1178] Example of a prompt
[1179] The following prompt messages can be used to efficiently analyze customer interactions and suggest appropriate responses.
[1180] "Could you please explain the latest model in detail?"
[1181] "Are there any other products with similar functions?"
[1182] "Do you have a demo that clearly explains how to use it?"
[1183] As a result, the system of the present invention can improve the quality of communication with customers and enable effective responses.
[1184] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1185] Step 1:
[1186] The user records conversations with customers using smart glasses. The input is the audio of the conversation with the customer, which the smart glasses record as audio data. The output is a temporarily stored audio file. Recording is possible through the recording function within the smart glasses.
[1187] Step 2:
[1188] After recording is complete, the audio data is automatically sent to the server. The input here is the audio file stored in the smart glasses, which is transferred to the server via an internet connection. The output is the audio file stored on the server. This transmission is performed via an HTTP POST request.
[1189] Step 3:
[1190] The server converts the received audio data into text data using the Google Speech-to-Text API. An audio file is used as input, and the output is text data. The server sends the audio data to the API and saves the returned text data to a temporary directory.
[1191] Step 4:
[1192] The server extracts key points from the transformed text data and generates a summary using a generative AI model (e.g., Hugging Face's Transformers). The input is text data, and the output is summary data. The server invokes the generative AI model, generates the summary data, and temporarily stores it.
[1193] Step 5:
[1194] The server identifies the user's emotions from audio data using IBM Watson Tone Analyzer. The input is an audio file, and the output is emotion information. The server sends the audio data to the emotion recognition engine and adds the received emotion information to the summary data.
[1195] Step 6:
[1196] The server proposes a dialogue method to the user in real time based on the generated summary and sentiment information. The input is the summary and sentiment information, and the output is the proposed dialogue method. Based on this data, the server formulates an appropriate response and displays it on the smart glasses' display.
[1197] Step 7:
[1198] The server saves the generated summary and sentiment information to an SQLite database. The input is the summary and sentiment information, and the output is the records stored in the database. The server performs the save process and persists the data.
[1199] Step 8:
[1200] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. The input is historical record data, and the output is the analysis results and effective interview methods. The server reads data from the database and applies algorithms to analyze patterns and trends.
[1201] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1202] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1203] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1204] [Fourth Embodiment]
[1205] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1206] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1207] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1208] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1209] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1210] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1211] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1212] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1213] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1214] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1215] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1216] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1217] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1218] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. The interview recording system based on this invention is constructed as follows.
[1219] System Overview
[1220] 1. User (device)
[1221] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[1222] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application on the device.
[1223] 2. Server
[1224] The server receives the audio file sent from the terminal and saves it to storage.
[1225] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[1226] The system uses generative AI to extract key points from the converted text data and generate a summary.
[1227] The generated summary is saved to the database.
[1228] By retrieving past records from a database and analyzing them using natural language processing techniques, we can identify effective interview methods.
[1229] Program processing
[1230] The user records audio and sends it to the server from their device.
[1231] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[1232] The server receives the audio file and saves it locally.
[1233] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[1234] The server converts the audio to text and generates a summary.
[1235] For text conversion, the server sends the audio file to a speech recognition engine, which then passes the acquired text data to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[1236] The server saves the interview records to the database.
[1237] The generated summary and audio file paths are stored in a database on the server. For example, they might be stored in the "records" table of an SQLite database.
[1238] The server analyzes past records and suggests effective interview methods.
[1239] The server retrieves past interview records from the database and uses natural language processing technology to identify common keywords and patterns. This allows it to generate effective interview methods to improve interview efficiency and provide feedback to the user.
[1240] Specific example
[1241] For example, suppose a user has a meeting with a customer one day. They record this meeting and send it to a server via their device. The server converts the audio file into text, automatically extracts the key points of the conversation, generates a summary, and saves it to a database. A few weeks later, the user analyzes multiple meeting records and identifies frequently occurring keywords, revealing that "customer satisfaction," "product quality," and "service improvement" are important points. This allows them to focus on these keywords in future meetings, leading to more effective communication.
[1242] Thus, the system of the present invention automates the recording and analysis of interviews, thereby improving work efficiency and reducing overtime.
[1243] The following describes the processing flow.
[1244] Step 1:
[1245] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[1246] Step 2:
[1247] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[1248] Step 3:
[1249] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[1250] Step 4:
[1251] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary directory. For example, the file is saved to " / path / to / save / meeting_audio.wav".
[1252] Step 5:
[1253] The server sends the stored audio file to a speech recognition engine, which converts it to text. The server reads the audio file and uses a speech recognition API (such as the Google Speech-to-Text API) to transcribe the audio. As a result, the content of the interview is obtained as text data.
[1254] Step 6:
[1255] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. The generative AI model uses tools such as Hugging Face's Transformers to perform the summarization. As a result, a concise summary is obtained that summarizes the main points of the interview discussion.
[1256] Step 7:
[1257] The server saves the generated summary and the path to the original audio file to a database. For example, SQLite is used for the database, and the summary and the path to the corresponding audio file are inserted into a table called "records".
[1258] Step 8:
[1259] The server retrieves past interview records from the database and analyzes them using natural language processing technology. This identifies common keywords and patterns in the interview content. The analysis results serve as reference material for developing more effective interview methods.
[1260] Step 9:
[1261] Based on the analysis results, the server suggests effective interview methods to the user. Specifically, it suggests incorporating questions and topics that include the top keywords. This suggestion is provided to the user as a strategic guideline for future interviews.
[1262] (Example 1)
[1263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1264] Traditional interview recording systems have the problem of requiring significant time and effort to manually record interview content and summarize key points. Furthermore, there was a lack of concrete methods for effectively analyzing past interview records and using that information to improve future interviews. This led to concerns that interview efficiency would decrease, resulting in a decline in the quality of customer service and overall work.
[1265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1266] In this invention, the server includes means for converting audio data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating effective interview methods, means for the user to record audio data and send it to the server from a terminal, means for the server to temporarily store the received audio data, means for the server to generate a summary using natural language processing technology, and means for the server to save the summary and analyze past records. As a result, the recording and summary generation of interviews and the analysis of past records are automated, making it possible to quickly propose effective interview methods.
[1267] "Audio data" refers to digital audio recordings of conversations and statements made by the user during an interview.
[1268] "Text data" refers to audio data converted into a string of characters, representing the content of the interview in written form.
[1269] A "summary" is a concise compilation of key points extracted from text data, including the main discussion points and conclusions of the interview.
[1270] "User" refers to the person who conducts the interview, starts and stops recording, and sends the data to the server.
[1271] A "terminal" refers to a recording device or computer used by a user, which has the function of recording audio data and sending it to a server.
[1272] A "server" is a central processing unit that receives audio data, stores it, converts it to text data, generates summaries, and analyzes stored historical records.
[1273] "Natural language processing technology" is a technology that enables computers to understand and process human language, and is used for generating summaries of text data and analyzing records.
[1274] A "database" is a system for managing and storing data such as generated summaries and historical records.
[1275] "Temporary storage" refers to the process of temporarily saving audio data between the time it is received and the start of processing.
[1276] A "speech recognition engine" refers to software or services that convert speech data into text data.
[1277] This invention relates to a system for effectively recording and analyzing interviews in shops and call centers. Specific embodiments are described in detail below.
[1278] Hardware and software to be used
[1279] In this invention, a system is built through the collaboration of three parties: the user, the terminal, and the server. The hardware and software used are shown below.
[1280] User device: A device with recording capabilities, such as a smartphone, tablet, or personal computer.
[1281] Server: A device with high-performance computing power and storage, such as a cloud server or on-premises server.
[1282] Speech recognition engine: Software used to convert speech data into text data (e.g., Google Speech-to-Text API).
[1283] Generative AI models: Artificial intelligence models for generating summaries of text data (e.g., Transformers in Hugging Face).
[1284] Database: A database system (e.g., SQLite) for storing generated summaries and interview records.
[1285] Overview of program processing
[1286] The user records the interview and sends it to the server from their device.
[1287] The user records audio during the interview using an application on their device. Once recording is complete, they press the "Stop Recording" button to end the recording and save the audio file to the device's local storage. Then, the user clicks the "Send" button to send the audio file to the server via an HTTP POST request. For example, the audio file might be saved at " / local / path / to / record.wav" and sent to the server.
[1288] The server receives the audio file and saves it locally.
[1289] The server saves the received audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav"). In this step, the server verifies the format and size of the received file and saves only valid audio data.
[1290] Convert audio data to text data
[1291] The server sends the stored audio files to the speech recognition engine. For example, it uses the Google Speech-to-Text API to convert the audio files into text data. Then, it retrieves the converted text data and passes it to a generative AI model.
[1292] Generate a summary from text data
[1293] The server uses a generative AI model (e.g., Hugging Face's Transformers) to extract key points from text data and generate a summary. This summary generation is instructed using prompts. An example prompt is, "Summarize the main points of this text data." The generated summary is a concise text that summarizes the main discussion points and conclusions of the meeting.
[1294] Save the summary and audio file path to the database.
[1295] The server records the generated summary and the path to the audio file in a database. For example, it might use a "records" table in an SQLite database to store the summary and path in association. This allows for easy retrieval from the database later when needed.
[1296] We analyze past records and propose effective interview methods.
[1297] The server retrieves past interview records from the database and analyzes them using natural language processing technology. By analyzing multiple summary data and identifying frequently occurring keywords and patterns, it generates specific suggestions for the next interview. These suggestions allow the user to conduct the interview more effectively.
[1298] Specific example
[1299] For example, suppose a user conducts a meeting with a customer one day and records the conversation. After recording, they click the "Send" button to send the audio file to the server. The server receives the audio file and saves it to a path like " / path / to / save / meeting_audio.wav". Then, a speech recognition engine is used to convert it into text data, and a generative AI model is used to generate a summary. The generated summary might be text like, "This meeting mainly focused on customer satisfaction." This summary and file path are stored in a database and later analyzed along with other meeting records to provide material for suggesting more effective meeting methods.
[1300] Examples of prompt messages include, "Please summarize the main points of this text data."
[1301] The above describes a specific embodiment for carrying out the present invention. This system automates the recording, summarizing, and analysis of interviews, thereby significantly improving work efficiency.
[1302] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1303] Step 1:
[1304] The user records the interview audio.
[1305] Input: Interview audio
[1306] Output: Audio files saved to local storage
[1307] Specific operation: Recording begins when the user launches the application on their device and presses the "Start Recording" button. After the interview ends, the user presses the "Stop Recording" button to end the recording and saves the audio file to the device's local storage. For example, the audio data will be saved to the path " / local / path / to / record.wav".
[1308] Step 2:
[1309] The device sends the audio file to the server.
[1310] Input: Audio files stored in local storage
[1311] Output: Audio file sent to the server
[1312] Specific action: The user clicks the "Send" button on the application. The device reads the audio file from local storage and sends it to the server via an HTTP POST request. For example, an audio file saved at " / local / path / to / record.wav" is sent to the server.
[1313] Step 3:
[1314] The server receives the audio file and saves it to a temporary directory.
[1315] Input: Audio file sent via HTTP POST request
[1316] Output: Audio files saved in a temporary directory
[1317] Specific operation: The server receives an HTTP POST request and verifies the file format and size. After confirming that it is valid audio data, it saves the audio file to a temporary directory (e.g., " / path / to / save / meeting_audio.wav").
[1318] Step 4:
[1319] The server converts the audio data into text data.
[1320] Input: Audio file saved in a temporary directory
[1321] Output: Text data converted from audio data
[1322] Specific operation: The server reads the audio file stored in a temporary directory and sends it to a speech recognition engine (e.g., Google Speech-to-Text API). It receives the text data returned by the speech recognition engine and passes it on to the next process. An example of the converted text would be in the format of "To obtain customer feedback, the following questions were asked."
[1323] Step 5:
[1324] The server generates a summary from the text data it has acquired.
[1325] Input: Text data
[1326] Output: Summary generated from text data
[1327] Specific operation: The server inputs text data into an AI model that generates data (e.g., Hugging Face's Transformers). The prompt is "Summarize the main points of this text data," and the model extracts the key points and generates a summary. An example of a generated summary would be, "This meeting focused on customer satisfaction."
[1328] Step 6:
[1329] The server saves the generated summary and the path to the audio file to the database.
[1330] Input: Generated summary, path to audio file
[1331] Output: Summary and audio file path stored in the database
[1332] Specific operation: The server adds the generated summary and the audio file path to a database (e.g., the "records" table in an SQLite database). Specifically, it uses SQL INSERT statements to record the summary and the corresponding audio file path.
[1333] Step 7:
[1334] The server analyzes past records and generates effective interview methods.
[1335] Input: Past interview records retrieved from the database
[1336] Output: Proposals for effective interview methods
[1337] Specific operation: The server retrieves past interview records from the database and analyzes them using natural language processing technology. It analyzes multiple summary data to identify frequently occurring keywords and patterns. Based on this, it proposes effective interview methods and provides feedback to the user. If the extracted keywords are "customer satisfaction" or "service improvement," it suggests focusing on these points in the next interview.
[1338] (Application Example 1)
[1339] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1340] In autonomous vehicles, there is a lack of effective means to collect and analyze important information regarding communication with passengers and driving conditions. As a result, improvements in driving efficiency and passenger experience are not being fully realized. Furthermore, there is no system that automatically analyzes passenger opinions and feedback to provide appropriate driving advice and feedback for service improvement.
[1341] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1342] In this invention, the server includes means for transmitting voice data to a cloud server and generating suggestions to improve driving efficiency and user experience based on the analysis results, means for providing feedback to the interface of an autonomous vehicle, and means for the cloud server to store the generated summaries in a database. This enables the effective collection, analysis, and summarization of important information regarding communication with passengers and driving conditions, thereby improving driving efficiency and passenger experience.
[1343] "Audio data" refers to data used to record and play back audio information in digital format.
[1344] "Text data" refers to data obtained by analyzing audio data and converting it into textual information.
[1345] A "summary" is a concise compilation of important information extracted from text data.
[1346] A "cloud server" is a remote server accessible via the internet that processes and stores large amounts of data.
[1347] "Driving efficiency" refers to the efficiency of driving a car in terms of fuel consumption, time, and energy consumption.
[1348] "User experience" refers to the overall experience of passengers using autonomous vehicles, including their satisfaction and emotional reactions.
[1349] An "interface" is a device or program that allows humans and machines to exchange information with each other.
[1350] A "database" is a system that systematically stores large amounts of data and allows for efficient searching and updating.
[1351] This invention relates to a system for automatically analyzing conversation recordings collected by autonomous vehicles during operation, extracting important information, and improving driving efficiency and user experience. Details are provided below.
[1352] System configuration and the hardware and software used.
[1353] Hardware to use
[1354] 1. Onboard computer of autonomous vehicle:
[1355] The system collects and partially processes audio data.
[1356] 2. Cloud Server:
[1357] It performs speech analysis, text conversion, database storage, and analysis.
[1358] Software to use
[1359] 1. Speech recognition engine:
[1360] The Google Speech-to-Text API is used to convert speech to text.
[1361] 2. Generative AI models:
[1362] Hugging Face's Transformers are used to extract important information and generate a summary.
[1363] 3. Database system:
[1364] Use MySQL or SQLite to store, search, and analyze large amounts of data.
[1365] System operation
[1366] 1. Collection of audio data:
[1367] Microphones inside the autonomous vehicle record audio, and the data is saved to the onboard computer.
[1368] Send the audio file to the cloud server using an HTTP POST request.
[1369] 2. Text conversion and summary generation of audio data:
[1370] The server saves the received audio files to a temporary directory.
[1371] The speech recognition engine (Google Speech-to-Text API) converts speech to text.
[1372] A generative AI model (Hugging Face's Transformers) is used to generate a summary, which is then saved to a database.
[1373] 3. Analysis of past records and generation of feedback:
[1374] Historical records are retrieved from the database, and common keywords and patterns are identified using NLP (Neuro-Linguistic Programming) techniques.
[1375] It generates effective communication methods and driving advice, and displays them on passenger interfaces such as tablets.
[1376] Specific example of processing
[1377] For example, suppose a passenger is in an autonomous vehicle and is talking about their destination while the vehicle is in motion. Below are some examples of prompts to input into the generating AI model.
[1378] Example of a prompt
[1379] Passenger: How long will it take for this car to reach its destination?
[1380] Vehicle AI: Considering the current traffic conditions, it will take approximately 30 minutes to arrive. Do you have any other questions?
[1381] Passenger: Is it okay if there's traffic?
[1382] Extract and summarize the key points and conclusions of the conversation.
[1383] Based on this prompt, the generative AI model generates the following summary.
[1384] Generated summary
[1385] Key discussion points: Arrival time at destination, traffic conditions, and how to deal with congestion.
[1386] Conclusion: Estimated travel time is approximately 30 minutes, taking traffic conditions into consideration.
[1387] conclusion
[1388] In this way, by applying the invention to autonomous vehicles, it is possible to improve the passenger experience while further increasing driving efficiency. The generated feedback is provided to the driver and passengers through the autonomous vehicle's interface, enabling appropriate advice and service improvements tailored to the driving situation.
[1389] This embodiment of the invention enables effective information collection and analysis in autonomous vehicles, thereby improving both driving efficiency and user experience.
[1390] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1391] Step 1:
[1392] The onboard computer of the autonomous vehicle collects audio data. The onboard computer records audio data using the microphone inside the vehicle, and this data is stored in local storage. The input is conversations between passengers and the driver, and the output is a recorded audio file.
[1393] Step 2:
[1394] The device sends the recorded audio file to the cloud server. The audio file is uploaded to the cloud server using an HTTP POST request. The input is the audio file on the onboard computer, and the output is the audio file stored on the cloud server.
[1395] Step 3:
[1396] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" is used. The input is the uploaded audio file, and the output is the temporary file stored on the server.
[1397] Step 4:
[1398] The server converts the audio file into text data. It uses a speech recognition engine (Google Speech-to-Text API) to convert the audio data into text data. The input is an audio file, and the output is the converted text data.
[1399] Step 5:
[1400] The server extracts key points from text data and generates a summary. It uses a generative AI model (Hugging Face's Transformers) to analyze text data and extract key points. The input is text data, and the output is the generated summary.
[1401] Step 6:
[1402] The server saves the generated summary to a database. The summary is stored in a database (e.g., MySQL or SQLite). The input is the generated summary, and the output is the summary stored in the database.
[1403] Step 7:
[1404] The server retrieves historical records from the database and performs analysis. Natural language processing techniques are used to identify common keywords and patterns. The input is historical records from the database, and the output is the analysis results.
[1405] Step 8:
[1406] The server generates suggestions to improve driving efficiency and user experience based on the analysis results. It generates appropriate driving advice and feedback for service improvement and displays them on the autonomous vehicle's interface. The input is the analysis results, and the output is the generated suggestions and feedback.
[1407] In this way, each processing step is executed sequentially, creating a system that improves the driving efficiency and user experience of autonomous vehicles.
[1408] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1409] This invention relates to a system that automatically records interviews in shops and call centers and suggests effective interview methods. In particular, it aims to further improve the quality of interviews by incorporating an emotion engine that recognizes user emotions. The interview recording system based on this invention is constructed as follows.
[1410] System Overview
[1411] 1. User (device)
[1412] The user records the interview audio on their device. Once recording is complete, the audio data is saved to the device's local storage.
[1413] After recording is complete, the user sends the saved audio file to the server. This transmission is done through the application that provides the recording function.
[1414] 2. Server
[1415] The server receives the audio file sent from the terminal and saves it to storage.
[1416] The saved audio files are converted into text data using a speech recognition engine (for example, Google Speech-to-Text API).
[1417] The system uses generative AI to extract key points from the converted text data and generate a summary.
[1418] In addition to the generated summary, an emotion engine is used to identify the user's emotions from the audio data, and this information is added to the text data and the generated summary.
[1419] The generated summary and sentiment information are stored in a database.
[1420] We retrieve emotional data from a database along with past records, analyze it using natural language processing techniques, and identify effective interview methods.
[1421] Program processing
[1422] The user records audio and sends it to the server from their device.
[1423] The user uses the device's recording function during the interview. Once recording is complete, the user clicks a button to upload the recording file to the server. This operation is performed via an HTTP POST request.
[1424] The server receives the audio file and saves it locally.
[1425] The server saves the received audio file to a temporary directory. For example, a path like " / path / to / save / meeting_audio.wav" might be used.
[1426] The server converts the audio to text.
[1427] For text conversion, the server sends the audio file to the speech recognition engine, which then generates the acquired text data. This process transcribes the contents of the audio file into text.
[1428] The server generates a summary.
[1429] The generated text data is passed to a generative AI model (for example, Hugging Face's Transformers) to generate a summary. This allows the main discussion points of the interview to be concisely summarized.
[1430] The server recognizes emotions.
[1431] In addition to the generated summary, the user's emotions are analyzed from the audio data using an emotion engine (e.g., emotion recognition AI). The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds the results to the text data and summary.
[1432] The server saves the interview records to the database.
[1433] The generated summary and sentiment information are stored in a database. For example, the format in which it is stored in the "records" table of an SQLite database is as follows:
[1434] sql
[1435] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[1436] The server analyzes past records and emotions.
[1437] The server retrieves past interview records and sentiment data from the database and performs analysis using natural language processing techniques. A comprehensive analysis, including sentiment data, is conducted to identify common keywords and patterns.
[1438] The server suggests effective interview methods.
[1439] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data.
[1440] Specific example
[1441] For example, suppose a user conducts a meeting with a customer. This meeting is recorded and sent to a server via the device. The server converts the audio file into text, extracts the key points of the conversation to generate a summary, and uses an emotion engine to identify emotional information. This information may include, for example, "Many positive emotions of customer satisfaction were recognized." This information and the summary are stored in a database and used for future analysis. In the next meeting, the user can use the emotional information from the previous meeting to communicate more effectively.
[1442] Thus, the system of the present invention automates the recording and analysis of interviews, and by further incorporating emotional information, it is possible to improve work efficiency and reduce overtime.
[1443] The following describes the processing flow.
[1444] Step 1:
[1445] The user launches the application on their device and begins recording the interview. By pressing the record button, the device's microphone captures audio data and recording begins. The recorded audio is temporarily saved to the device's local storage.
[1446] Step 2:
[1447] Once the meeting is over, the user stops recording and saves the recording file. The saved recording file is stored in local storage, for example, as "meeting_audio.wav". The user then presses the send button in the application to prepare to send the recording file to the server.
[1448] Step 3:
[1449] When the user presses the send button, the device generates an HTTP POST request and uploads the audio file to the server. This request includes the audio file data and metadata (e.g., file name and interview date and time).
[1450] Step 4:
[1451] The server receives an HTTP POST request from the terminal and saves the audio file to a temporary storage directory. The save location is, for example, a path like " / path / to / save / meeting_audio.wav".
[1452] Step 5:
[1453] The server sends the stored audio files to a speech recognition engine, which converts them into text. For example, the Google Speech-to-Text API is used as the speech recognition engine. The converted text data is a written record of the interview content.
[1454] Step 6:
[1455] Once the text data is generated, the server passes this text to a generative AI model, which extracts the key points and generates a summary. Hugging Face's Transformers is used as the generative AI model. As a result, a summary is obtained that concisely summarizes the main discussion points of the interview.
[1456] Step 7:
[1457] Next, the server passes the audio data to the emotion engine to identify the user's emotions. An emotion recognition AI (such as IBM Watson's Empathy API) is used as the emotion engine. The emotion engine identifies emotions such as anger, joy, and sadness from the audio and adds this information to the text data and summary.
[1458] Step 8:
[1459] The server stores the generated summary and sentiment information in a database. For example, SQLite is used for the database, and the summary, corresponding sentiment information, and audio file paths are inserted into a table called "records".
[1460] Step 9:
[1461] The server retrieves past interview records and emotional data from the database and analyzes them using natural language processing technology. A comprehensive analysis, including emotional data, is performed to identify common keywords and patterns.
[1462] Step 10:
[1463] Based on the analysis results, the server proposes effective interview methods to the user. Specifically, it suggests questions and topics that include top keywords, and further optimizes the interview process by referring to past sentiment data. This suggestion is provided to the user as a strategic guideline for future interviews.
[1464] (Example 2)
[1465] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1466] Conventional interview recording systems required manual recording of interview content, making it difficult to efficiently organize and store information. Furthermore, they lacked the functionality to analyze emotions from audio data to improve interview quality. This resulted in a lack of information needed to identify effective interview methods, leading to decreased work efficiency and limited communication effectiveness. A new system is needed to address these problems.
[1467] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving and temporarily storing voice data, means for converting voice data into text data, and means for extracting important points from the converted text data using a generated AI model and generating a summary. This enables automatic recording and analysis of interview content, and further enables advanced analysis including user sentiment information. This makes it possible to identify effective interview methods and improve work efficiency.
[1468] "Interview audio data" refers to digital audio files that record the content of conversations conducted in person or remotely.
[1469] A "memory device" refers to hardware or databases used to store digital information such as audio data, text data, and sentiment analysis results.
[1470] "Communication function" refers to the function of transmitting digital data to other devices or servers via a network.
[1471] A "server" is a computer system that processes, stores, and analyzes audio data received from terminals via a network.
[1472] "Text data" refers to text information obtained by transcribing audio data.
[1473] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates summaries and responses based on the input data.
[1474] An "emotion engine" is a system that analyzes and recognizes emotions from voice data.
[1475] "Recording" refers to saving information such as generated summaries and sentiment analysis results so that they can be referenced later.
[1476] A "database" is a software system designed to systematically store and manage information, and to allow for searching and updating as needed.
[1477] "Natural language processing technology" refers to artificial intelligence technology used to analyze, understand, and manipulate human language.
[1478] This invention relates to a system that automatically records and analyzes audio data from interviews and proposes effective interview methods. The following describes a specific implementation of this system.
[1479] System Overview
[1480] This system consists of users (terminals), a server, and a database. Users record interviews at shops or call centers using their terminals and send the recorded data to the server.
[1481] Hardware and software to be used
[1482] 1. Hardware
[1483] Device: The recording device used by the user. This includes smartphones, tablets, and personal computers.
[1484] Server: A computer system for receiving and processing audio data.
[1485] 2. Software
[1486] Voice recording application: An application for recording audio on a device and saving it to local storage.
[1487] Communication protocol: HTTP POST request for sending voice data to a server.
[1488] Speech recognition engine: Converts speech data into text data (e.g., Google Speech-to-Text API).
[1489] Generative AI model: Generates summaries from generated text data (e.g., Transformers for Hugging Face).
[1490] Emotion engine: Software for analyzing emotions from voice data (e.g., emotion recognition AI).
[1491] Database: Stores the generated summary and sentiment information (e.g., an SQLite database).
[1492] Natural language processing technology: A technology for analyzing historical records and sentiment data.
[1493] System operation
[1494] The user uses a voice recording application on their device during the interview. Once recording is complete, the audio data is saved locally on the device and uploaded to the server using the application's transmission function. This transmission is done via an HTTP POST request. The server temporarily stores the received audio data and converts it into text data using a speech recognition engine. The converted text data is then passed to a generation AI model, which extracts key points and generates a summary.
[1495] Simultaneously, the server uses an emotion engine to analyze the user's emotions from the audio data. This emotion information is added to the summary, and the final generated summary and emotion information are stored in a database. The server then uses natural language processing techniques to analyze past interview records and emotion information stored in the database, identifying common keywords and patterns to suggest effective interview methods.
[1496] Specific example
[1497] For example, consider a scenario where a user records a meeting with a customer and sends the audio data from their device to a server. The server converts this audio data into text data, extracts key points, and creates a summary. It also analyzes the customer's emotions from the audio data and adds emotional information to the summary, such as "the customer showed many positive emotions of satisfaction." Finally, this data is stored in a database and used as a reference for future meetings.
[1498] Users can use information from previous meetings to select appropriate questions and topics for subsequent meetings. This improves the quality of meetings, leading to increased work efficiency and reduced overtime.
[1499] Example of a prompt
[1500] For example, you can generate a summary by giving the following prompt to the AI model:
[1501] "Please summarize what we discussed in today's meeting."
[1502] Thus, the present invention is a system that automatically records and analyzes voice data, enabling the proposal of interview methods that also take emotional information into account.
[1503] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1504] System program processing flow
[1505] Step 1:
[1506] Audio recording
[1507] The user records audio data during the interview using the device's audio recording application. Recording starts when the record button is pressed and ends when it is pressed again. The completed audio data (input) is saved to the device's local storage (output). For example, it is saved to the path " / local / path / to / meeting_audio.wav".
[1508] Step 2:
[1509] Sending audio files
[1510] The user clicks the upload button within the voice recording application to send the recorded audio file to the server. The sending process is performed using an HTTP POST request. The input is the audio file in local storage, and the output is the audio file uploaded to the server.
[1511] Step 3:
[1512] Receiving and saving audio files
[1513] The server saves audio files received from the user's terminal to a temporary directory. The input is the audio file sent via an HTTP POST request, and the output is the audio file stored in the server's temporary storage area. For example, the audio data is saved to " / path / to / save / meeting_audio.wav".
[1514] Step 4:
[1515] Speech-to-text conversion
[1516] The server uses the stored audio file to call a speech recognition engine (e.g., Google Speech-to-Text API) and converts the audio data into text data. The input is an audio file, and the output is the converted text data. As a result of the transcription, all statements made during the interview are obtained in text format.
[1517] Step 5:
[1518] Summary generation
[1519] The server passes the acquired text data to an AI model (e.g., Hugging Face's Transformers) to extract key points and generate a summary. The input is the acquired text data, and the output is the generated summary text. This provides a concise overview of the interview content.
[1520] Step 6:
[1521] Recognition of emotions
[1522] The server passes the audio data to an emotion engine (e.g., emotion recognition AI) to analyze the user's emotions. The input is audio data, and the output is the result of the emotion analysis. The analyzed emotion information is added to the summary text of the interview.
[1523] Step 7:
[1524] Preservation of products
[1525] The server stores the generated summary and sentiment information in a database. The input is the summary text and sentiment analysis results, and the stored data is the output. For example, it is stored in the "records" table in an SQLite database in the following format:
[1526] sql
[1527] INSERT INTO records (summary, emotion, file_path) VALUES ('Summarized interview content', 'Emotional information', ' / path / to / save / meeting_audio.wav');
[1528] Step 8:
[1529] Analysis of past records and emotions
[1530] The server retrieves past interview records and sentiment information from the database and analyzes them using natural language processing techniques. The input is past interview records and sentiment data from the database, and the output is effective interview methods derived from the analysis. This process identifies common keywords and patterns.
[1531] Step 9:
[1532] Suggestions for effective interview methods
[1533] The server suggests effective interview methods to the user based on the analysis results. The input is the analysis results, and the output is the suggested content. Specifically, it suggests questions and topics that include top keywords, as well as interview procedures that take into account past sentiment data.
[1534] (Application Example 2)
[1535] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1536] While conventional interview recording systems could transcribe audio data and analyze past records, they lacked the functionality to improve the quality of customer service in real time. Furthermore, there was no mechanism to accurately capture customer emotions and provide appropriate advice in real time based on those emotions. Therefore, improving customer satisfaction and response efficiency were identified as challenges.
[1537] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting voice data, means for converting the input voice data into text data, means for extracting important points from the converted text data and generating a summary, means for recording and saving the generated summary, means for analyzing past records and generating an effective interview method, means for analyzing the user's emotions and adding the emotional information to the summary data, and means for displaying emotional information in real time during the conversation with the customer and suggesting an appropriate conversation method. This makes it possible to improve responses in real time based on the customer's emotions.
[1538] "Audio data" refers to information that can be recorded, stored, or transmitted in digital format.
[1539] "Means of input" refers to devices or functions that allow users to supply voice data to the system.
[1540] "Text data" refers to information obtained by converting audio data into written text.
[1541] "Means of conversion" refers to the functions of software or hardware used to convert audio data into text data.
[1542] "Key points" refer to the most important information or topics extracted from the interview content.
[1543] A "means for generating summaries" refers to a function that extracts important points from text data and summarizes them concisely.
[1544] "Means of recording and preserving" refers to the functions of storage devices and databases for storing generated summaries and other data.
[1545] "Past records" refer to interview data and summary data that have been saved in the past.
[1546] "Means of analysis" refer to software and algorithms used to analyze past records and identify patterns and trends.
[1547] "Effective interview methods" refer to methods and techniques for optimizing communication, which are proposed based on the analysis results.
[1548] "User emotions" refers to the emotional state a user exhibits during an interview (for example, joy, anger, sadness, etc.).
[1549] "Means of analyzing emotions" refers to software or algorithms that identify a user's emotions from voice data and analyze that information.
[1550] "Emotional information" refers to information about a user's emotional state obtained through means of analyzing emotions.
[1551] "Means of displaying in real time" refers to the functions of displays and software that immediately show acquired emotional information to the user.
[1552] "Means of suggesting dialogue methods" refer to software or algorithms that suggest appropriate communication methods to users based on emotional information and other factors.
[1553] This invention relates to a system that automatically records interviews in shops and call centers, recognizes user emotions, and suggests effective interview methods. The system of this invention includes the following components.
[1554] System Overview
[1555] 1. Input of audio data
[1556] The user records conversations with customers using smart glasses. The recorded audio data is temporarily stored in the smart glasses. After recording is complete, the audio data is automatically sent to a server. This transmission requires an internet connection.
[1557] 2. Converting audio data
[1558] The server converts the received audio data into text data using a speech recognition engine such as the Google Speech-to-Text API. The converted text data is stored in a temporary directory.
[1559] 3. Summarization from text data
[1560] Key points are extracted from text data, and a summary is generated using generative AI models such as Hugging Face's Transformers. The generated summary data is temporarily stored for further analysis.
[1561] 4. Recognition of emotions
[1562] The server identifies the user's emotions from the voice data using an emotion recognition engine such as IBM Watson Tone Analyzer. The identified emotion information is added to the summary data.
[1563] 5. Proposal for a real-time dialogue method
[1564] Based on the generated summary and sentiment information, the server formulates a response plan in real time and displays it on the smart glasses' display. This allows staff to get the appropriate response plan on the spot.
[1565] 6. Recording and Preservation
[1566] The generated summaries and sentiment information are stored in a storage device such as an SQLite database. This makes the data reusable at any time.
[1567] 7. Analysis of past records
[1568] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. By analyzing past conversation content and emotional data, and extracting common patterns and trends, it generates and proposes effective interview methods.
[1569] Specific example
[1570] For example, if a customer asks for a product description, the smart glasses record the conversation and send it to a server. The server converts the recorded audio data into text, generates a summary using a generative AI model, and adds customer sentiment information to the summary. Next, real-time advice such as "The customer is satisfied, so we will suggest further related products" is displayed on the smart glasses' screen, allowing staff to take appropriate action immediately.
[1571] Example of a prompt
[1572] The following prompt messages can be used to efficiently analyze customer interactions and suggest appropriate responses.
[1573] "Could you please explain the latest model in detail?"
[1574] "Are there any other products with similar functions?"
[1575] "Do you have a demo that clearly explains how to use it?"
[1576] As a result, the system of the present invention can improve the quality of communication with customers and enable effective responses.
[1577] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1578] Step 1:
[1579] The user records conversations with customers using smart glasses. The input is the audio of the conversation with the customer, which the smart glasses record as audio data. The output is a temporarily stored audio file. Recording is possible through the recording function within the smart glasses.
[1580] Step 2:
[1581] After recording is complete, the audio data is automatically sent to the server. The input here is the audio file stored in the smart glasses, which is transferred to the server via an internet connection. The output is the audio file stored on the server. This transmission is performed via an HTTP POST request.
[1582] Step 3:
[1583] The server converts the received audio data into text data using the Google Speech-to-Text API. An audio file is used as input, and the output is text data. The server sends the audio data to the API and saves the returned text data to a temporary directory.
[1584] Step 4:
[1585] The server extracts key points from the transformed text data and generates a summary using a generative AI model (e.g., Hugging Face's Transformers). The input is text data, and the output is summary data. The server invokes the generative AI model, generates the summary data, and temporarily stores it.
[1586] Step 5:
[1587] The server identifies the user's emotions from audio data using IBM Watson Tone Analyzer. The input is an audio file, and the output is emotion information. The server sends the audio data to the emotion recognition engine and adds the received emotion information to the summary data.
[1588] Step 6:
[1589] The server proposes a dialogue method to the user in real time based on the generated summary and sentiment information. The input is the summary and sentiment information, and the output is the proposed dialogue method. Based on this data, the server formulates an appropriate response and displays it on the smart glasses' display.
[1590] Step 7:
[1591] The server saves the generated summary and sentiment information to an SQLite database. The input is the summary and sentiment information, and the output is the records stored in the database. The server performs the save process and persists the data.
[1592] Step 8:
[1593] The server retrieves historical records from stored data and performs analysis using natural language processing techniques. The input is historical record data, and the output is the analysis results and effective interview methods. The server reads data from the database and applies algorithms to analyze patterns and trends.
[1594] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1595] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1596] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1597] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1598] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1599] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1600] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1601] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1602] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1603] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1604] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1605] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1606] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1607] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1608] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1609] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1610] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1611] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1612] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1613] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1614] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1615] The following is further disclosed regarding the embodiments described above.
[1616] (Claim 1)
[1617] A means of inputting audio data,
[1618] A means for converting input audio data into text data,
[1619] A means for extracting key points from converted text data and generating a summary,
[1620] Means for recording and saving the generated summary,
[1621] A means of analyzing past records and generating effective interview methods,
[1622] A system that includes this.
[1623] (Claim 2)
[1624] A means for sending input audio data to a server,
[1625] The system according to claim 1, which includes means for a server to convert audio data into text data.
[1626] (Claim 3)
[1627] A means for the server to save the generated summary to a database,
[1628] The system according to claim 1, comprising means for retrieving and analyzing past records from a database.
[1629] "Example 1"
[1630] (Claim 1)
[1631] A means of inputting audio data,
[1632] A means for converting input audio data into text data,
[1633] A means for extracting key points from converted text data and generating a summary,
[1634] Means for recording and saving the generated summary,
[1635] A means of analyzing past records and generating effective interview methods,
[1636] A means for the user to record audio data and send it to the server from their device,
[1637] A means for the server to temporarily store the audio data it receives,
[1638] A server provides a means for generating a summary using natural language processing technology,
[1639] A server has a means of saving summaries and analyzing past records,
[1640] A system that includes this.
[1641] (Claim 2)
[1642] A means for sending input audio data to a server,
[1643] A means by which the server converts audio data into text data,
[1644] The system according to claim 1.
[1645] (Claim 3)
[1646] A means for the server to save the generated summary to a database,
[1647] A means of retrieving and analyzing historical records from a database,
[1648] The system according to claim 1.
[1649] "Application Example 1"
[1650] (Claim 1)
[1651] A means of inputting audio data,
[1652] A means for converting input audio data into text data,
[1653] A means for extracting key points from converted text data and generating a summary,
[1654] Means for recording and saving the generated summary,
[1655] A means of analyzing past records and generating effective interview methods,
[1656] A means for sending voice data to a cloud server and generating suggestions to improve driving efficiency and user experience based on the analysis results,
[1657] A means of providing feedback to the interface of an autonomous vehicle,
[1658] A system that includes this.
[1659] (Claim 2)
[1660] A means of sending input audio data to a cloud server,
[1661] The system according to claim 1, comprising means for a cloud server to convert audio data into text data and generate a summary.
[1662] (Claim 3)
[1663] A means for the cloud server to save the generated summary to a database,
[1664] The system according to claim 1, comprising means for retrieving and analyzing past records from a database.
[1665] "Example 2 of combining an emotion engine"
[1666] (Claim 1)
[1667] A means of inputting audio data from interviews,
[1668] A means for recording input audio data and saving it to a storage device,
[1669] A means of sending audio data to a server using a communication function,
[1670] A means for the server to receive and temporarily store audio data,
[1671] A means of converting audio data into text data,
[1672] A means for extracting important points from converted text data using a generative AI model and generating a summary,
[1673] A means for analyzing user emotions from audio data and adding the results to text data and a summary,
[1674] A means of recording and storing the generated summary and sentiment information,
[1675] A means of analyzing past records and emotional data to generate effective interview methods,
[1676] A system that includes this.
[1677] (Claim 2)
[1678] A means for sending input audio data to a server,
[1679] The system according to claim 1, which includes means for a server to convert audio data into text data.
[1680] (Claim 3)
[1681] A server provides means for storing generated summaries and sentiment information in a database,
[1682] The system according to claim 1, comprising means for obtaining and analyzing historical records and sentiment data from a database.
[1683] "Application example 2 when combining with an emotional engine"
[1684] (Claim 1)
[1685] A means of inputting audio data,
[1686] A means for converting input audio data into text data,
[1687] A means for extracting key points from converted text data and generating a summary,
[1688] Means for recording and saving the generated summary,
[1689] A means of analyzing past records and generating effective interview methods,
[1690] A means of analyzing user emotions and adding that emotional information to summary data,
[1691] A means of displaying emotional information in real time during customer interactions and suggesting appropriate dialogue methods,
[1692] A system that includes this.
[1693] (Claim 2)
[1694] A means for sending input audio data to a server,
[1695] The system according to claim 1, which includes means for a server to convert audio data into text data.
[1696] (Claim 3)
[1697] A means for the server to save the generated summary to a database,
[1698] The system according to claim 1, comprising means for retrieving and analyzing past records from a database. [Explanation of symbols]
[1699] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of inputting audio data, A means of converting input audio data into text data, A means for extracting key points from converted text data and generating a summary, Means for recording and saving the generated summary, A means of analyzing past records and generating effective interview methods, A system that includes this.
2. A means for sending input audio data to a server, The system according to claim 1, which includes means for a server to convert audio data into text data.
3. A means for the server to save the generated summary to a database, The system according to claim 1, comprising means for obtaining and analyzing past records from a database.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A