system
The system efficiently converts and analyzes interview audio data for frequent keywords and emotion scores, addressing the inefficiencies of existing methods to enhance user and customer experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-10
AI Technical Summary
Existing methods for analyzing interview audio data are time-consuming, labor-intensive, and lack efficiency and accuracy in extracting important information and emotional aspects, making it difficult to improve user and customer experiences.
A system that includes uploading voice data to a server, converting it into text using a voice recognition engine, analyzing the text for frequent keywords and emotion scores, and visualizing the results in a report format, allowing users to quickly grasp important information.
Enables efficient and accurate analysis of large amounts of interview data, extracting useful insights and emotional aspects, thereby improving user and customer experiences.
Smart Images

Figure 2026041536000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Improving user experience (UX) and customer experience (CX) requires conducting numerous interviews and analyzing the results in detail. However, manually analyzing each interview's audio data is extremely time-consuming and labor-intensive. Furthermore, efficiently processing massive amounts of data and extracting important information is a difficult challenge. Furthermore, while there is a need to quickly identify emotional statements and references to specific themes, conventional methods have limitations in terms of accuracy and efficiency. Solutions to these challenges are highly desirable. [Means for solving the problem]
[0005] The present invention provides a system including a means for users to upload voice data, a means for a server to convert the received voice data into text data using a voice recognition engine, a means for the server to analyze the text data and extract frequently occurring keywords, the number of mentions of specific topics, and an emotion score, a means for the server to compile and visualize the analysis results in a report format, and a means for users to view the generated report. This system efficiently analyzes interview voice data and allows users to quickly grasp important information. By using an external voice recognition service as the voice recognition engine, highly accurate transcription is achieved. Furthermore, by generating positive, negative, and neutral emotion scores in the emotion analysis, users can grasp the emotional aspects of speech in detail.
[0006] A "user" is an entity that uploads interview audio data to the system from a terminal and views the analysis results.
[0007] "Server" is a computer system that receives, processes, and analyzes data sent by users.
[0008] "Audio data" refers to recordings of audio collected from interviews, etc.
[0009] A "voice recognition engine" is a technology that analyzes voice data and converts it into text data.
[0010] "Character data" is text information converted from voice data by a voice recognition engine.
[0011] "Analysis" is the process of examining text data in detail and extracting useful information.
[0012] "Frequent keywords" are words or phrases that appear particularly frequently in the character data converted from the voice data.
[0013] "Mentions of a particular topic" refers to counting how many comments are made related to a particular topic or issue.
[0014] The "emotion score" is a numerical representation of the emotional tone of each part of a statement.
[0015] A "report" is a document or graphical display that organizes the results of an analysis and presents them in a format that is easily understandable to the user.
[0016] "Visualization" is the process of visually representing analytical results to enable users to intuitively understand the information.
[0017] A "dashboard" is an interface through which a user can view generated reports and other analytical results. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0040] Audio data upload
[0041] The user collects the interview audio data and uploads it to the server using the file selection interface on the terminal. The audio data format can be a common audio file format (e.g., WAV, MP3, etc.).
[0042] Speech Recognition Processing
[0043] The server receives the voice data uploaded by the user and converts it into text data using a speech recognition engine. The speech recognition engine can use an external speech recognition service (for example, a cloud-based speech recognition API). This allows for highly accurate transcription.
[0044] Text analytics
[0045] The server parses the character data to extract the following information:
[0046] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0047] Counting the number of mentions of a specific topic: By matching a list of keywords related to a pre-defined topic, the number of mentions of that topic is counted.
[0048] Sentiment Analysis: Performs sentiment analysis and calculates positive, negative, and neutral sentiment scores.
[0049] Reporting and Visualization
[0050] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, statistics on the number of mentions of specific themes, and sentiment scores. This information is visualized in a graphical format (e.g., graphs, charts, etc.) to allow users to intuitively understand it.
[0051] View results
[0052] Users access the server using their terminals and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. This system allows users to quickly and efficiently grasp important information without having to relisten to all the interview audio data.
[0053] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights from large amounts of interview data.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data is in a common audio file format such as WAV or MP3.
[0057] Step 2:
[0058] The server receives the audio data sent by the user via an HTTP POST request and saves the files in the server's storage. At this time, each file is assigned a unique identifier for later processing.
[0059] Step 3:
[0060] The server sends the received voice data to a speech recognition engine and begins transcription. By using an external speech recognition service (for example, a cloud-based speech recognition API), highly accurate text data can be obtained.
[0061] Step 4:
[0062] The server receives the text data obtained from the speech recognition engine and stores it in a database. Each text data is assigned an identifier for the corresponding audio file.
[0063] Step 5:
[0064] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, thereby identifying frequently occurring keywords.
[0065] Step 6:
[0066] The server compares the text data with a predefined list of keywords for each topic and tally the number of mentions related to each topic, clarifying how many times a particular topic of interest to the user has been mentioned.
[0067] Step 7:
[0068] The server uses a sentiment analysis library to calculate a sentiment score for each piece of text: positive, negative, and neutral scores, identifying which statements are particularly emotional.
[0069] Step 8:
[0070] The server generates a report based on the analysis, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and a sentiment score, all of which are displayed in a graphical format.
[0071] Step 9:
[0072] The server compiles the generated reports into a dashboard format on the web interface for easy user access. Each analysis result is visualized in graphs and charts, allowing users to intuitively understand the results.
[0073] Step 10:
[0074] Users access the server using their devices and view the generated reports on a dashboard. Users can quickly grasp important information without having to listen to all of the interview audio data again. This series of processes enables users to gain useful insights from large amounts of data in a short amount of time.
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] Previously, there was a lack of efficient means to analyze audio data from interviews conducted with the aim of improving user experience (UX) and customer experience (CX) and extract important information. As a result, it was difficult to quickly extract useful information from massive amounts of interview audio data, and users had to listen to all of the audio data again. This reduced the efficiency of analysis and caused delays in decision-making.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server includes: [means for a user to upload voice data on a terminal; [means for the server to convert the received voice data into text data using a cloud-based voice recognition API; and [means for the server to analyze the text data using a natural language processing library and extract frequently occurring keywords, the number of mentions of specific themes, and sentiment scores.] This makes it possible [to quickly and efficiently extract useful information from massive amounts of interview voice data.]
[0080] "User" refers to a person who uses the system to upload voice data and view the analysis results.
[0081] "Terminal" refers to an electronic device (e.g., a PC or smartphone) that a user uses to upload audio data.
[0082] "Server" refers to a central processing unit that processes voice data received from a user, converts it into text data, and analyzes it.
[0083] "Cloud-based speech recognition API" refers to an external service that converts voice data available via the internet into highly accurate text data.
[0084] "Text data" refers to text information converted from voice data by a cloud-based voice recognition API.
[0085] "Natural language processing library" refers to software tools and algorithms used to analyze text data.
[0086] "Frequent keywords" refer to important words that appear frequently in the character data being analyzed.
[0087] "Number of mentions of a specific theme" refers to the calculated frequency of appearance of keywords related to a predetermined theme.
[0088] "Emotion score" refers to a numerical value that quantitatively evaluates the emotion (positive, negative, neutral) of a sentence in the text data.
[0089] "Report format" refers to a document or graphical display format that summarizes the analysis results in a way that makes them easy for users to understand.
[0090] "Graphical formats" refers to formats such as graphs and charts that visually represent data.
[0091] A "dashboard" refers to a screen that organizes and displays analysis results so that users can check them at a glance.
[0092] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0093] Audio data upload
[0094] Users collect interview audio data on their devices, preferably in common audio file formats (e.g., WAV, MP3, etc.). They then upload these audio files to the server using the device's file selection interface. Specifically, they access the system's upload page using a web browser and click the file selection button to select and upload the audio data.
[0095] Speech Recognition Processing
[0096] The server receives the uploaded voice data and converts it into text using a cloud-based speech recognition API (e.g., Google® Cloud Speech-to-Text API). The server sends the audio file to the API and receives the text it returns. This process results in highly accurate transcription. For example, the audio data "Hello, what would you like to talk about today?" is converted into text.
[0097] Text analytics
[0098] The server analyzes the converted text data using a natural language processing library (e.g., NLTK or spaCy). Specifically, it performs the following process:
[0099] 1. Extraction of frequent keywords: The server extracts all tokens from the text data and calculates their frequency distribution. For example, frequently occurring keywords such as "product" and "quality" are extracted.
[0100] 2. Counting the number of mentions of a specific topic: The server checks against a list of keywords related to a predefined topic and counts the number of mentions of each topic. For example, "price" might be counted 10 times, "service" 15 times, and so on.
[0101] 3. Sentiment Analysis: The server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. For example, the sentence "I am very satisfied" is evaluated as a positive sentiment score.
[0102] Reporting and Visualization
[0103] The server generates reports based on the analysis results and uses tools (such as Matplotlib and D3.js) to visualize them in graphical formats, such as a word cloud of frequently occurring keywords or a bar chart of sentiment scores, allowing users to intuitively understand the analysis results.
[0104] View results
[0105] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to click on frequently occurring keyword clouds and bar charts of the number of mentions by theme to view more detailed information. This system allows users to quickly and efficiently grasp important information without having to relisten to the vast amount of interview audio data.
[0106] Specific examples
[0107] For example, consider a scenario where a user uploads five interview audio files (e.g., file1.wav, file2.wav, file3.mp3, file4.wav, file5.mp3). The server receives these audio files and converts them into text using the Google Cloud Speech-to-Text API. The server then analyzes the text for frequent keywords, mentions of specific themes, and sentiment scores, and compiles the results into a report. Finally, users can view this information in a visual format through a dashboard.
[0108] Examples of prompt statements
[0109] Examples of specific prompts for a generative AI model might include:
[0110] "How do I upload an audio file for transcription and text analysis?"
[0111] "Please explain how to analyze interview audio data to extract frequent keywords and sentiment scores."
[0112] By feeding such prompts into the generative AI model, more specific instructions can be obtained about the system's detailed processing and analysis methods.
[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0114] Step 1:
[0115] The user uploads the audio data from the device to the server.
[0116] Specifically, the user uses the device's file selection interface (e.g., a web browser) to select an interview audio file stored on a PC or smartphone and clicks the upload button. The input is audio data (e.g., WAV or MP3 files), and the output is a notification to the server that the file has been uploaded.
[0117] Step 2:
[0118] The server converts the received voice data into text using a cloud-based speech recognition API.
[0119] Specifically, the server sends the received audio file to a specified cloud-based speech recognition API (for example, Google Cloud Speech-to-Text API). The speech recognition API receives the audio data as input, analyzes it, and converts it into document text data. The output is character data (text format). Through this process, the audio is converted into character data such as "Hello, what would you like to talk about today?"
[0120] Step 3:
[0121] The server analyzes the text data using a natural language processing library and extracts frequently occurring keywords.
[0122] Specifically, the server uses an NLP library (such as NLTK or spaCy) to extract all tokens (words) from the text data and calculate their frequency. The input is text data obtained from a speech recognition API, and the output is a list of keywords with high frequency distribution. For example, important words such as "product" and "quality" are extracted from the text data as frequent keywords.
[0123] Step 4:
[0124] The server analyzes the text data and counts the number of times a particular topic is mentioned.
[0125] Specifically, the server compares the text data with a list of keywords related to pre-set themes and tallies the number of times each theme is mentioned. The input is the text data and a list of theme keywords, and the output is data on the number of times each theme is mentioned. For example, it may be tallied that words related to "price" appear 10 times in the text data, and words related to "service" appear 15 times.
[0126] Step 5:
[0127] The server performs sentiment analysis on the text data and calculates the sentiment score.
[0128] Specifically, the server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. The input is the text data, and the output is a positive, negative, or neutral sentiment score. For example, the sentence "I am very satisfied" receives a positive evaluation, while the sentence "I am not satisfied" receives a negative evaluation.
[0129] Step 6:
[0130] The server generates a report based on the analysis results and visualizes them in a graphical format.
[0131] Specifically, the server generates a report based on the analysis results (e.g., frequently occurring keywords, number of mentions of specific themes, sentiment scores) and visualizes them using a data visualization tool (e.g., Matplotlib or D3.js). The input is the analysis result data, and the output is a graphically represented report. For example, a word cloud of frequently occurring keywords or a bar chart of sentiment scores may be generated.
[0132] Step 7:
[0133] The user views the generated report on their device.
[0134] Specifically, users access the system's dashboard page via a web browser on their device. The input is report data generated and visualized on the server, and the output is a visually displayed report. On the dashboard, users can click on the frequently occurring keyword cloud or the bar chart of the number of mentions by topic to view more detailed information. This system allows users to efficiently analyze massive amounts of voice data and quickly obtain important information.
[0135] (Application example 1)
[0136] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0137] Conventional voice analysis systems aimed at improving user and customer experience rely on uploading and analyzing individually collected voice data as it arrives. This makes it difficult to collect feedback in real time and quickly improve services. Furthermore, uploading and analyzing voice data can be difficult, resulting in inefficiencies due to an inability to handle large volumes of data. For autonomous vehicles in particular, it is important to collect passenger feedback in real time and quickly reflect it in improvements to operation and service quality, but current technology makes this difficult.
[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0139] In this invention, the server includes: [means for a user to upload voice data;] [means for the server to convert the received voice data into text data using a voice recognition engine;] [means for the server to analyze the text data and extract frequently occurring keywords, the number of mentions of specific themes, and an emotion score;] [means for collecting voice data in real time;] [means for uploading the voice data to a cloud server; and [means for collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service. This enables feedback to be collected and analyzed in real time, enabling rapid service improvements.
[0140] "User" refers to a user who collects voice data and provides the data to the system.
[0141] "Audio data" refers to audio files that record what passengers say.
[0142] A "server" is a central processing unit that receives and analyzes voice data.
[0143] A "voice recognition engine" is software for converting voice data into text data.
[0144] "Character data" refers to data obtained by converting voice data into text format.
[0145] "Keywords" are important words or phrases that appear frequently within the audio data.
[0146] A "specific theme" is a predefined topic or issue that is the subject of analysis.
[0147] The "number of mentions" is a numerical value indicating how many times a word related to a particular theme appears in the speech data.
[0148] An "emotion score" is a numerical representation of the emotions in audio data, and is divided into three categories: positive, negative, and neutral.
[0149] A "report" is a report summarizing the results of the analysis of character data.
[0150] "Visualization" means displaying the analysis results in a visually easy-to-understand format such as graphs or charts.
[0151] A "dashboard" is an interface that allows users to view analysis results.
[0152] "Real-time" means that processing occurs with virtually no time delay.
[0153] "Feedback" refers to opinions and ratings provided by passengers.
[0154] "Service Improvements" are specific changes or modifications made to improve the quality of the Services provided.
[0155] A "cloud server" refers to a group of servers on the Internet that can be accessed remotely.
[0156] Basic configuration
[0157] The system includes means for "users to upload voice data," means for "the server converting the received voice data into text data using a voice recognition engine," means for "the server analyzing the text data and extracting frequently occurring keywords, the number of mentions of specific themes, and sentiment scores," means for "the server compiling and visualizing the analysis results in report format," means for "users viewing the generated report," means for "collecting voice data in real time," means for "uploading voice data to a cloud server," and means for "collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service."
[0158] Program processing
[0159] The server first collects audio data in real time using a microphone installed inside the autonomous vehicle. The audio data is recorded as audio data, and the data is uploaded to a cloud server via the on-board computer. The audio file format is a common format such as WAV or MP3.
[0160] When the cloud server receives the voice data, it uses a voice recognition engine to convert the voice data into text data. An external voice recognition service (e.g., Google Cloud Speech-to-Text API) is used as the voice recognition engine. To ensure high accuracy of the conversion, the voice recognition engine uses a cloud-based service.
[0161] The converted text data is then analyzed by the server. During the analysis process, natural language processing (NLP) techniques are used to extract frequently occurring keywords, tally the number of times a particular topic is mentioned, and perform sentiment analysis, which generates a positive, negative, or neutral sentiment score.
[0162] The analysis results are compiled into a report, which is displayed in a visual format such as graphs and charts for easy understanding by the user. The generated report can be viewed by the operation manager via a dedicated mobile app or web dashboard.
[0163] Hardware and software used
[0164] Hardware:
[0165] Microphone for collecting audio data
[0166] On-board computer
[0167] Cloud Server
[0168] software:
[0169] Google Cloud Speech-to-Text API (speech recognition engine)
[0170] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)
[0171] Visualization tools (e.g., Matplotlib, D3.js)
[0172] Web interface (for dashboard display)
[0173] Examples of specific examples and prompts
[0174] For example, if a passenger in an autonomous vehicle says, "I like this route, the scenery is amazing!", the voice is collected by a microphone. The voice data is uploaded to a cloud server and converted into text data using the Google Cloud Speech-to-Text API. It is then analyzed using natural language processing technology to extract sentiment scores and frequently occurring keywords. Finally, the analysis results are visualized as graphs and charts on a dashboard and provided to the operation manager.
[0175] Example prompt sentence:
[0176] Please perform a sentiment analysis on the following Japanese text.
[0177] Text: I love this route, the views are amazing!
[0178] Calculate positive, negative, and neutral sentiment scores and explain why. Also extract frequently occurring keywords and the number of mentions of each theme.
[0179] This allows the system to collect feedback in real time and quickly improve services.
[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0181] Step 1:
[0182] A microphone installed inside the autonomous vehicle collects passenger voice data in real time. The voice data is sent to the on-board computer in WAV or MP3 format. The input is the passenger's speech, and the output is an audio file.
[0183] Step 2:
[0184] The server receives audio data from the onboard computer and uploads it to the cloud server. The input is an audio file, and the output is audio data stored on the cloud. This process sends the audio data to the cloud in real time.
[0185] Step 3:
[0186] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) on a cloud server to convert the voice data into text data. The input is an audio file, and the output is text data. The speech recognition engine analyzes the voice waveform and converts it into text data.
[0187] Step 4:
[0188] The server analyzes the converted text data. It uses natural language processing libraries (e.g., NLTK, spaCy) to extract frequently occurring keywords, tally the number of mentions of specific themes, and perform sentiment analysis (generating a score for positive, negative, or neutral). The input is the text data, and the output is the analysis results. Specific pattern matching and sentiment scoring algorithms are applied.
[0189] Step 5:
[0190] The server generates a report based on the analysis results. It visualizes them as graphs and charts using a graphical visualization tool (e.g., Matplotlib, D3.js). The input is the analysis results, and the output is a visualized report. This process makes it easier for users to intuitively understand the data.
[0191] Step 6:
[0192] Users can view the generated reports using a dedicated mobile app or web dashboard. The input is the visualized report, and the output is new insights including user feedback. Users can check each analysis result on the dashboard and take any necessary actions to improve the service.
[0193] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0194] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information and emotional information with the aim of improving user experience (UX) and customer experience (CX).
[0195] Audio data upload
[0196] Users collect interview audio data and upload it to the server using the file selection interface on their device. The audio data is formatted in common audio file formats such as WAV and MP3. The uploaded audio data is sent to the server and stored in storage.
[0197] Speech Recognition Processing
[0198] The server receives the voice data sent by the user and converts it into text data using a speech recognition engine. The speech recognition engine uses an external speech recognition service (for example, a cloud-based speech recognition API) to achieve highly accurate transcription.
[0199] Text analytics
[0200] The server parses the character data to extract the following information:
[0201] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0202] Counting the number of mentions of a specific topic: Matching a list of keywords for a pre-defined topic and counting the number of mentions of that topic.
[0203] Sentiment Analysis and Sentiment Engine
[0204] The server includes an emotion engine that uses an emotion analysis library to recognize user emotions. The emotion engine calculates positive, negative, and neutral emotion scores for each part of a statement in the text data. The emotion scores are calculated in real time and can be viewed by users on a dashboard. The server also displays the fluctuations in the emotion scores in a time-series graph format, visualizing changes in emotion over time.
[0205] Reporting and Visualization
[0206] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and sentiment scores. This information is displayed in a graphical format (e.g., graphs, charts, etc.) to help users understand it intuitively.
[0207] View results
[0208] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. Fluctuations in emotion scores can also be tracked over time, allowing users to identify when emotional statements were made.
[0209] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights and sentiment information from large amounts of interview data.
[0210] The processing flow will be explained below.
[0211] Step 1:
[0212] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data format is a common audio file format such as WAV or MP3. The user selects the audio file to upload on the file selection screen and presses the send button.
[0213] Step 2:
[0214] The server receives the audio data sent by the user via an HTTP POST request, saves the received file in a temporary directory, and registers meta information (file name, upload date and time, user ID, etc.) in the database.
[0215] Step 3:
[0216] The server sends the received voice data to a voice recognition engine for text conversion. An external voice recognition service (e.g., a cloud-based voice recognition API) is used to convert the voice into text data, and this text data is then retrieved.
[0217] Step 4:
[0218] The server stores the acquired text data in a database, and when storing the text data, associates it with the identifier of the original audio file.
[0219] Step 5:
[0220] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, creating a list of frequently occurring words to identify common keywords.
[0221] Step 6:
[0222] The server compares the text data with a list of keywords for predefined themes and tally the number of mentions related to each theme, allowing the frequency of mentions of specific topics of interest to users to be displayed in graph form.
[0223] Step 7:
[0224] The server uses a sentiment analysis library to calculate a sentiment score for each statement in the text, calculating positive, negative, or neutral sentiment scores to identify which parts are particularly emotional.
[0225] Step 8:
[0226] The server calculates the emotion score in real time and displays it on a dashboard, allowing users to check the emotion associated with each comment in real time.
[0227] Step 9:
[0228] The server tracks the fluctuations in the emotion scores over time and displays them in a graph, allowing users to visually grasp the emotional changes in the comments over time.
[0229] Step 10:
[0230] The server generates a report based on the analysis results and displays it in a graphical format (graphs, charts, etc.). Users can access the server using their devices and view this report from a dashboard, allowing them to quickly grasp important information and emotional information without having to relisten to all the interview audio data.
[0231] The above are the specific processing steps of the system based on the invention that combines the emotion engine.
[0232] Example 2
[0233] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0234] Current systems for improving user experience (UX) and customer experience (CX) face challenges in efficiently analyzing large amounts of interview audio data and quickly extracting important and emotional information. Furthermore, manual data analysis is time-consuming and labor-intensive, resulting in limited insights.
[0235] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to upload voice data using a file selection interface of a terminal; a means for converting the received voice data into text data using a voice recognition engine and storing the text data in a database; a means for analyzing the text data using a natural language processing library and extracting frequently occurring keywords, the number of mentions of a specific theme, and an emotion score; a means for aggregating the analysis results, compiling them in a report format, and visualizing them in a graphical format; and a means for a user to view the generated report on a dashboard via a web interface from their terminal. This makes it possible to quickly and efficiently obtain useful insights and emotion information from a large amount of interview voice data.
[0236] "User" means an end user who uses the System to upload audio data and view analysis results.
[0237] "Terminal" refers to the device that a user uses to upload audio data and view analysis results, such as a PC, tablet, or smartphone.
[0238] A "file selection interface" is a user interface that allows a user to select an audio file on a terminal.
[0239] A "server" is a remote computer system that receives voice data, converts it to text data, analyzes it, and generates reports.
[0240] "Speech recognition engine" means software or a service for analyzing voice data and converting it into text data, including cloud-based speech recognition services.
[0241] "Character data" refers to text-format data converted by a voice recognition engine.
[0242] A "natural language processing library" is a software library for analyzing text data to extract frequently occurring keywords and the number of times a particular topic is mentioned.
[0243] "Keywords" are important words or phrases that are set as the subject of analysis.
[0244] A "theme" is a specific topic or topic that is matched by a pre-defined list of keywords.
[0245] The "emotion score" classifies the emotions of statements contained in text data into three types: positive, negative, and neutral, and quantifies the degree of each.
[0246] A "report" is a document or data that compiles the analysis results and provides them to the user in a graphical format.
[0247] A "dashboard" is a web interface that allows users to view reports and visualizes the analysis results.
[0248] "Visualization" is the process of converting data and analytical results into a format that is easy to understand visually (e.g., graphs, charts, etc.).
[0249] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important and emotional information for the purpose of improving user experience (UX) and customer experience (CX). This system automates the process in which users upload audio data and a server analyzes the data to extract useful insights.
[0250] Hardware and software used
[0251] This system uses the following hardware and software:
[0252] Server: Receives data, converts it to text data, analyzes it, and generates reports.
[0253] Terminal: Used by users to upload audio data and view analysis results.
[0254] Speech recognition engine: Converts voice data into text data using a cloud-based speech recognition service (e.g., Google Speech-to-Text API, IBM Watson® Speech to Text, etc.).
[0255] Natural language processing libraries: Analyze text data and extract frequent keywords or the number of mentions of specific topics (e.g., spaCy, NLTK, etc.).
[0256] Sentiment analysis libraries: Analyze the sentiment of statements in text data and calculate positive, negative, or neutral sentiment scores (e.g., VADER, TextBlob, AWS® Comprehend, etc.).
[0257] Overall system processing flow
[0258] Users collect interview audio data and upload it to the server using a file selection interface on their device. The server converts the received audio data into text using a cloud-based speech recognition engine. The converted text is stored in a database and then analyzed using a natural language processing library.
[0259] The analysis includes the following elements:
[0260] Extraction of frequently occurring keywords
[0261] Counting the number of times a specific topic was mentioned
[0262] Calculating the sentiment score
[0263] This analytical information is aggregated and compiled into reports, which are displayed in graphical formats (e.g., graphs, charts, etc.), and users can view the generated reports in a dashboard via the device's web interface.
[0264] Specific operation example
[0265] For example, a user uploads an MP3 audio file called "User Interview_1.mp3" to the system. This audio file is sent to the server and stored in storage. After that, the speech recognition engine converts the speech into text data, and the generated text data is stored in the database.
[0266] Next, the natural language processing library extracts frequently occurring keywords from the text data, confirming that keywords such as "service improvement" and "customer satisfaction" appear frequently. Furthermore, the sentiment analysis library analyzes each utterance in the text data and finds that there are many positive comments. These analysis results are compiled into a report and displayed in graph format on a dashboard.
[0267] Prompt Sentence Examples
[0268] An example prompt for a generative AI model is:
[0269] "I would like to upload interview audio data, run it through speech recognition to transcribe it, and analyze it for frequent keywords and sentiment scores. Please let me know how this data will be used to improve the user experience."
[0270] As described above, this system provides consistent support from analyzing interview audio data to detecting changes in emotions, and is a mechanism that provides useful information to users.
[0271] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0272] Step 1:
[0273] The user uploads the audio data using the file selection interface on the device. The user collects the interview audio data and uploads it by selecting the audio file on the device. This upload procedure sends the audio file (e.g., WAV or MP3 format) to the server. The input is the audio file on the user's device, and the output is the audio file stored on the server.
[0274] Step 2:
[0275] The server checks the received audio data and saves it in storage. After receiving the audio file sent by the user, the server automatically saves it in a storage system (e.g., cloud storage service). This ensures that the data is stored safely on the server. The input is the audio file uploaded to the server, and the output is the audio file saved in the server's storage.
[0276] Step 3:
[0277] The server starts a speech recognition engine and converts the stored voice data into text data. The server uses a cloud-based speech recognition engine (e.g., Google Speech-to-Text API) and inputs the stored voice data. The speech recognition engine analyzes the voice and generates corresponding text data. The input is the voice data stored in storage, and the output is the text data generated by the speech recognition engine.
[0278] Step 4:
[0279] The server stores the generated text data in a database. The server records the text data output from the speech recognition engine in a database for subsequent analysis processes. The input is the text data generated by the speech recognition engine, and the output is the text data stored in the database.
[0280] Step 5:
[0281] The server invokes a natural language processing library to analyze the text data. The server uses a natural language processing (NLP) library (e.g., spaCy) to input the text data stored in the database. The NLP library performs text analysis to extract frequently occurring keywords and the number of times a particular topic is mentioned. The input is the text data stored in the database, and the output is a list of frequently occurring keywords and the number of times a particular topic is mentioned.
[0282] Step 6:
[0283] The server launches a sentiment analysis library and calculates the sentiment score in the text data. The server uses a sentiment analysis library (e.g., VADER) to calculate the sentiment score for each utterance in the text data. The sentiment score is expressed in three indicators: positive, negative, and neutral. The input is the text data to be analyzed, and the output is a list of sentiment scores.
[0284] Step 7:
[0285] The server aggregates the analysis results and compiles them into a report. The server aggregates the data of the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and generates a report. This report is visualized in a graphical format (e.g., graph, chart, etc.). The input is the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and the output is a graphical report.
[0286] Step 8:
[0287] The user views the generated report on a dashboard via a web interface from their device. The user uses a browser on their device to access the server's web interface and view the generated report. The dashboard displays a list of frequently occurring keywords, the number of times a specific theme is mentioned, and fluctuations in sentiment scores. The input is the generated report, and the output is the analysis results on the dashboard that the user views.
[0288] (Application example 2)
[0289] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0290] Conventional advertising distribution systems have the problem that it is difficult to select optimal advertising information based on user emotions or specific keywords, which makes it difficult to attract user attention and deliver advertisements effectively.
[0291] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading voice data, means for converting the voice data into character data using a voice recognition engine, means for analyzing the character data to extract frequently occurring keywords, the number of mentions of a specific theme, and an emotion score, and means for selecting optimal advertising information based on the emotion score and keywords and displaying it on the user terminal. This makes it possible to deliver optimal advertisements that match the user's emotions in a timely manner.
[0292] "Audio data" refers to audio files that users upload to the server, and are in common audio file formats such as WAV and MP3.
[0293] "Character data" refers to text data converted from voice data using a voice recognition engine.
[0294] A "voice recognition engine" is a program or service that converts voice data into text data, and includes external voice recognition services.
[0295] "Frequent keywords" are words or phrases that appear frequently in the analyzed text data.
[0296] The "number of times a specific theme is mentioned" refers to the number of times that the theme is mentioned, as determined by matching the text data with a keyword list of the theme that is set in advance.
[0297] "Emotion score" refers to the evaluation value of positive, negative, and neutral emotions calculated for statements in text data using an emotion analysis library.
[0298] "Advertising information" refers to advertisements and promotional information selected based on the user's emotional score and keywords.
[0299] A "user terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[0300] "Visualization" refers to displaying analysis results in a graphical format that users can intuitively understand.
[0301] This invention relates to a system that analyzes user voice data, selects optimal advertising information based on emotion scores and keywords, and displays it on the user's terminal. The specific system configuration and processing procedure for implementing this invention are described below.
[0302] System Configuration
[0303] 1. User device: Collect voice data using devices such as smartphones, tablets, and PCs and upload it to the server.
[0304] 2. Server: A server located on the cloud that processes data using a speech recognition engine, natural language processing library, etc.
[0305] 3. Speech recognition engine: A program for converting voice data into text data. It uses an external cloud-based speech recognition service (e.g., Google Speech-to-Text API).
[0306] 4. Natural language processing libraries: Libraries for extracting frequent keywords and the number of mentions of specific topics from text data (e.g., spaCy, NLTK).
[0307] 5. Sentiment analysis library: A library for calculating sentiment scores from text data (e.g., VADER, Google Natural Language API).
[0308] 6. Advertising platform: Selects and delivers optimal advertising information based on sentiment scores and keywords (e.g., Google AdSense).
[0309] Program processing
[0310] The server does the following:
[0311] 1. Audio data collection and upload:
[0312] A user records audio data using a smartphone application.
[0313] The common formats used are WAV and MP3.
[0314] The audio data is uploaded from the application to a cloud server.
[0315] 2. Speech Recognition and Transcription:
[0316] A speech recognition engine (e.g., Google Speech-to-Text API) on a cloud server receives the uploaded voice data and converts it into text data.
[0317] 3. Text Analysis:
[0318] The text data is analyzed on the server using a natural language processing library (e.g., spaCy or NLTK).
[0319] Frequently occurring keywords and the number of mentions of specific topics are extracted.
[0320] 4. Emotion analysis:
[0321] Use a sentiment analysis library (e.g., VADER or Google Natural Language API) to calculate a sentiment score for each piece of text data.
[0322] Sentiment scores are expressed as either positive, negative, or neutral.
[0323] 5. Advertisement Delivery:
[0324] Based on the sentiment score and keywords, the advertising platform (e.g., Google AdSense) selects the most suitable advertising information.
[0325] The selected advertising information is displayed on the user terminal.
[0326] Specific examples
[0327] For example, if a user uploads audio data in which they say, "I want to take a weekend trip," the audio data is converted into text data and keywords such as "trip" and "weekend" are extracted. If the sentiment analysis determines that the comment has a positive sentiment, travel-related advertisements (e.g., hotel discounts or airline ticket promotions) are displayed on the user's device.
[0328] Prompt Sentence Examples
[0329] Analyze the user's spoken voice data and calculate a sentiment score of positive, negative, or neutral. Based on the results, generate prompts that suggest the best ad format and content to display.
[0330] This system allows the most appropriate advertisements to be displayed based on the user's emotional state, maximizing the effectiveness of the advertisements.
[0331] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0332] Step 1:
[0333] A user uses a smartphone application to record audio data. This audio data is saved in WAV or MP3 format. The user then uploads the audio data to a cloud server using an interface within the app. The input is the user's audio data, and the output is an audio file saved on the cloud server.
[0334] Step 2:
[0335] The cloud server sends the received voice data to a voice recognition engine. The server converts the voice data into text data using an external cloud voice recognition service (e.g., Google Speech-to-Text API). The input here is an audio file, and the output is text data converted from the voice data.
[0336] Step 3:
[0337] The server retrieves the text data and begins analyzing it using a natural language processing library (e.g., spaCy or NLTK). Specifically, it extracts frequently occurring keywords from the text data and calculates the number of mentions of a particular topic. The input for this step is the text data, and the output is a list of keywords and a summary of topic mentions.
[0338] Step 4:
[0339] The server uses a sentiment analysis library (e.g., VADER or Google Natural Language API) to analyze each part of the text data and generate a positive, negative, or neutral sentiment score. The sentiment score is quantified according to the intensity of the emotion. The input here is the text data, and the output is a score corresponding to each emotion.
[0340] Step 5:
[0341] The server selects the optimal advertising information using an advertising platform (e.g., Google AdSense) based on the emotion score and keywords. This selection is performed using prompts generated by a generative AI model. An example of such a prompt is, "Analyze the user's spoken voice data and calculate positive, negative, and neutral emotion scores. Based on the calculation results, generate a prompt that provides the ad format and content to display the optimal advertisement." The input is the emotion score and a keyword list, and the output is the optimal advertising information.
[0342] Step 6:
[0343] The selected advertising information is sent from the server to the user terminal. The user terminal displays the received advertising information on a display so that the user can actually see it. The input is the optimal advertising information, and the output is the advertisement displayed on the user terminal.
[0344] The above are the specific processing steps of the system that realizes the application example. This series of processing provides optimal advertising information based on the user's emotions.
[0345] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0346] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0347] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0348] [Second embodiment]
[0349] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0350] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0351] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0352] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0353] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0354] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0355] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0356] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0357] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0358] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0359] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0360] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0361] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0362] Audio data upload
[0363] The user collects the interview audio data and uploads it to the server using the file selection interface on the terminal. The audio data format can be a common audio file format (e.g., WAV, MP3, etc.).
[0364] Speech Recognition Processing
[0365] The server receives the voice data uploaded by the user and converts it into text data using a speech recognition engine. The speech recognition engine can use an external speech recognition service (for example, a cloud-based speech recognition API). This allows for highly accurate transcription.
[0366] Text analytics
[0367] The server parses the character data to extract the following information:
[0368] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0369] Counting the number of mentions of a specific topic: By matching a list of keywords related to a pre-defined topic, the number of mentions of that topic is counted.
[0370] Sentiment Analysis: Performs sentiment analysis and calculates positive, negative, and neutral sentiment scores.
[0371] Reporting and Visualization
[0372] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, statistics on the number of mentions of specific themes, and sentiment scores. This information is visualized in a graphical format (e.g., graphs, charts, etc.) to allow users to intuitively understand it.
[0373] View results
[0374] Users access the server using their terminals and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. This system allows users to quickly and efficiently grasp important information without having to relisten to all the interview audio data.
[0375] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights from large amounts of interview data.
[0376] The processing flow will be explained below.
[0377] Step 1:
[0378] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data is in a common audio file format such as WAV or MP3.
[0379] Step 2:
[0380] The server receives the audio data sent by the user via an HTTP POST request and saves the files in the server's storage. At this time, each file is assigned a unique identifier for later processing.
[0381] Step 3:
[0382] The server sends the received voice data to a speech recognition engine and begins transcription. By using an external speech recognition service (for example, a cloud-based speech recognition API), highly accurate text data can be obtained.
[0383] Step 4:
[0384] The server receives the text data obtained from the speech recognition engine and stores it in a database. Each text data is assigned an identifier for the corresponding audio file.
[0385] Step 5:
[0386] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, thereby identifying frequently occurring keywords.
[0387] Step 6:
[0388] The server compares the text data with a predefined list of keywords for each topic and tally the number of mentions related to each topic, clarifying how many times a particular topic of interest to the user has been mentioned.
[0389] Step 7:
[0390] The server uses a sentiment analysis library to calculate a sentiment score for each piece of text: positive, negative, and neutral scores, identifying which statements are particularly emotional.
[0391] Step 8:
[0392] The server generates a report based on the analysis, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and a sentiment score, all of which are displayed in a graphical format.
[0393] Step 9:
[0394] The server compiles the generated reports into a dashboard format on the web interface for easy user access. Each analysis result is visualized in graphs and charts, allowing users to intuitively understand the results.
[0395] Step 10:
[0396] Users access the server using their devices and view the generated reports on a dashboard. Users can quickly grasp important information without having to listen to all of the interview audio data again. This series of processes enables users to gain useful insights from large amounts of data in a short amount of time.
[0397] Example 1
[0398] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0399] Previously, there was a lack of efficient means to analyze audio data from interviews conducted with the aim of improving user experience (UX) and customer experience (CX) and extract important information. As a result, it was difficult to quickly extract useful information from massive amounts of interview audio data, and users had to listen to all of the audio data again. This reduced the efficiency of analysis and caused delays in decision-making.
[0400] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0401] In this invention, the server includes: [means for a user to upload voice data on a terminal; [means for the server to convert the received voice data into text data using a cloud-based voice recognition API; and [means for the server to analyze the text data using a natural language processing library and extract frequently occurring keywords, the number of mentions of specific themes, and sentiment scores.] This makes it possible [to quickly and efficiently extract useful information from massive amounts of interview voice data.]
[0402] "User" refers to a person who uses the system to upload voice data and view the analysis results.
[0403] "Terminal" refers to an electronic device (e.g., a PC or smartphone) that a user uses to upload audio data.
[0404] "Server" refers to a central processing unit that processes voice data received from a user, converts it into text data, and analyzes it.
[0405] "Cloud-based speech recognition API" refers to an external service that converts voice data available via the internet into highly accurate text data.
[0406] "Text data" refers to text information converted from voice data by a cloud-based voice recognition API.
[0407] "Natural language processing library" refers to software tools and algorithms used to analyze text data.
[0408] "Frequent keywords" refer to important words that appear frequently in the character data being analyzed.
[0409] "Number of mentions of a specific theme" refers to the calculated frequency of appearance of keywords related to a predetermined theme.
[0410] "Emotion score" refers to a numerical value that quantitatively evaluates the emotion (positive, negative, neutral) of a sentence in the text data.
[0411] "Report format" refers to a document or graphical display format that summarizes the analysis results in a way that makes them easy for users to understand.
[0412] "Graphical formats" refers to formats such as graphs and charts that visually represent data.
[0413] A "dashboard" refers to a screen that organizes and displays analysis results so that users can check them at a glance.
[0414] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0415] Audio data upload
[0416] Users collect interview audio data on their devices, preferably in common audio file formats (e.g., WAV, MP3, etc.). They then upload these audio files to the server using the device's file selection interface. Specifically, they access the system's upload page using a web browser and click the file selection button to select and upload the audio data.
[0417] Speech Recognition Processing
[0418] The server receives the uploaded audio data and converts it into text using a cloud-based speech recognition API (e.g., Google Cloud Speech-to-Text API). The server sends the audio file to the API and receives the text it returns. This process results in highly accurate transcription. For example, the audio data "Hello, what would you like to talk about today?" is converted into text.
[0419] Text analytics
[0420] The server analyzes the converted text data using a natural language processing library (e.g., NLTK or spaCy). Specifically, it performs the following process:
[0421] 1. Extraction of frequent keywords: The server extracts all tokens from the text data and calculates their frequency distribution. For example, frequently occurring keywords such as "product" and "quality" are extracted.
[0422] 2. Counting the number of mentions of a specific topic: The server checks against a list of keywords related to a predefined topic and counts the number of mentions of each topic. For example, "price" might be counted 10 times, "service" 15 times, and so on.
[0423] 3. Sentiment Analysis: The server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. For example, the sentence "I am very satisfied" is evaluated as a positive sentiment score.
[0424] Reporting and Visualization
[0425] The server generates reports based on the analysis results and uses tools (such as Matplotlib and D3.js) to visualize them in graphical formats, such as a word cloud of frequently occurring keywords or a bar chart of sentiment scores, allowing users to intuitively understand the analysis results.
[0426] View results
[0427] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to click on frequently occurring keyword clouds and bar charts of the number of mentions by theme to view more detailed information. This system allows users to quickly and efficiently grasp important information without having to relisten to the vast amount of interview audio data.
[0428] Specific examples
[0429] For example, consider a scenario where a user uploads five interview audio files (e.g., file1.wav, file2.wav, file3.mp3, file4.wav, file5.mp3). The server receives these audio files and converts them into text using the Google Cloud Speech-to-Text API. The server then analyzes the text for frequent keywords, mentions of specific themes, and sentiment scores, and compiles the results into a report. Finally, users can view this information in a visual format through a dashboard.
[0430] Examples of prompt statements
[0431] Examples of specific prompts for a generative AI model might include:
[0432] "How do I upload an audio file for transcription and text analysis?"
[0433] "Please explain how to analyze interview audio data to extract frequent keywords and sentiment scores."
[0434] By feeding such prompts into the generative AI model, more specific instructions can be obtained about the system's detailed processing and analysis methods.
[0435] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0436] Step 1:
[0437] The user uploads the audio data from the device to the server.
[0438] Specifically, the user uses the device's file selection interface (e.g., a web browser) to select an interview audio file stored on a PC or smartphone and clicks the upload button. The input is audio data (e.g., WAV or MP3 files), and the output is a notification to the server that the file has been uploaded.
[0439] Step 2:
[0440] The server converts the received voice data into text using a cloud-based speech recognition API.
[0441] Specifically, the server sends the received audio file to a specified cloud-based speech recognition API (for example, Google Cloud Speech-to-Text API). The speech recognition API receives the audio data as input, analyzes it, and converts it into document text data. The output is character data (text format). Through this process, the audio is converted into character data such as "Hello, what would you like to talk about today?"
[0442] Step 3:
[0443] The server analyzes the text data using a natural language processing library and extracts frequently occurring keywords.
[0444] Specifically, the server uses an NLP library (such as NLTK or spaCy) to extract all tokens (words) from the text data and calculate their frequency. The input is text data obtained from a speech recognition API, and the output is a list of keywords with high frequency distribution. For example, important words such as "product" and "quality" are extracted from the text data as frequent keywords.
[0445] Step 4:
[0446] The server analyzes the text data and counts the number of times a particular topic is mentioned.
[0447] Specifically, the server compares the text data with a list of keywords related to pre-set themes and tallies the number of times each theme is mentioned. The input is the text data and a list of theme keywords, and the output is data on the number of times each theme is mentioned. For example, it may be tallied that words related to "price" appear 10 times in the text data, and words related to "service" appear 15 times.
[0448] Step 5:
[0449] The server performs sentiment analysis on the text data and calculates the sentiment score.
[0450] Specifically, the server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. The input is the text data, and the output is a positive, negative, or neutral sentiment score. For example, the sentence "I am very satisfied" receives a positive evaluation, while the sentence "I am not satisfied" receives a negative evaluation.
[0451] Step 6:
[0452] The server generates a report based on the analysis results and visualizes them in a graphical format.
[0453] Specifically, the server generates a report based on the analysis results (e.g., frequently occurring keywords, number of mentions of specific themes, sentiment scores) and visualizes them using a data visualization tool (e.g., Matplotlib or D3.js). The input is the analysis result data, and the output is a graphically represented report. For example, a word cloud of frequently occurring keywords or a bar chart of sentiment scores may be generated.
[0454] Step 7:
[0455] The user views the generated report on their device.
[0456] Specifically, users access the system's dashboard page via a web browser on their device. The input is report data generated and visualized on the server, and the output is a visually displayed report. On the dashboard, users can click on the frequently occurring keyword cloud or the bar chart of the number of mentions by topic to view more detailed information. This system allows users to efficiently analyze massive amounts of voice data and quickly obtain important information.
[0457] (Application example 1)
[0458] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0459] Conventional voice analysis systems aimed at improving user and customer experience rely on uploading and analyzing individually collected voice data as it arrives. This makes it difficult to collect feedback in real time and quickly improve services. Furthermore, uploading and analyzing voice data can be difficult, resulting in inefficiencies due to an inability to handle large volumes of data. For autonomous vehicles in particular, it is important to collect passenger feedback in real time and quickly reflect it in improvements to operation and service quality, but current technology makes this difficult.
[0460] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0461] In this invention, the server includes: [means for a user to upload voice data;] [means for the server to convert the received voice data into text data using a voice recognition engine;] [means for the server to analyze the text data and extract frequently occurring keywords, the number of mentions of specific themes, and an emotion score;] [means for collecting voice data in real time;] [means for uploading the voice data to a cloud server; and [means for collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service. This enables feedback to be collected and analyzed in real time, enabling rapid service improvements.
[0462] "User" refers to a user who collects voice data and provides the data to the system.
[0463] "Audio data" refers to audio files that record what passengers say.
[0464] A "server" is a central processing unit that receives and analyzes voice data.
[0465] A "voice recognition engine" is software for converting voice data into text data.
[0466] "Character data" refers to data obtained by converting voice data into text format.
[0467] "Keywords" are important words or phrases that appear frequently within the audio data.
[0468] A "specific theme" is a predefined topic or issue that is the subject of analysis.
[0469] The "number of mentions" is a numerical value indicating how many times a word related to a particular theme appears in the speech data.
[0470] An "emotion score" is a numerical representation of the emotions in audio data, and is divided into three categories: positive, negative, and neutral.
[0471] A "report" is a report summarizing the results of the analysis of character data.
[0472] "Visualization" means displaying the analysis results in a visually easy-to-understand format such as graphs or charts.
[0473] A "dashboard" is an interface that allows users to view analysis results.
[0474] "Real-time" means that processing occurs with virtually no time delay.
[0475] "Feedback" refers to opinions and ratings provided by passengers.
[0476] "Service Improvements" are specific changes or modifications made to improve the quality of the Services provided.
[0477] A "cloud server" refers to a group of servers on the Internet that can be accessed remotely.
[0478] Basic configuration
[0479] The system includes means for "users to upload voice data," means for "the server converting the received voice data into text data using a voice recognition engine," means for "the server analyzing the text data and extracting frequently occurring keywords, the number of mentions of specific themes, and sentiment scores," means for "the server compiling and visualizing the analysis results in report format," means for "users viewing the generated report," means for "collecting voice data in real time," means for "uploading voice data to a cloud server," and means for "collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service."
[0480] Program processing
[0481] The server first collects audio data in real time using a microphone installed inside the autonomous vehicle. The audio data is recorded as audio data, and the data is uploaded to a cloud server via the on-board computer. The audio file format is a common format such as WAV or MP3.
[0482] When the cloud server receives the voice data, it uses a voice recognition engine to convert the voice data into text data. An external voice recognition service (e.g., Google Cloud Speech-to-Text API) is used as the voice recognition engine. To ensure high accuracy of the conversion, the voice recognition engine uses a cloud-based service.
[0483] The converted text data is then analyzed by the server. During the analysis process, natural language processing (NLP) techniques are used to extract frequently occurring keywords, tally the number of times a particular topic is mentioned, and perform sentiment analysis, which generates a positive, negative, or neutral sentiment score.
[0484] The analysis results are compiled into a report, which is displayed in a visual format such as graphs and charts for easy understanding by the user. The generated report can be viewed by the operation manager via a dedicated mobile app or web dashboard.
[0485] Hardware and software used
[0486] Hardware:
[0487] Microphone for collecting audio data
[0488] On-board computer
[0489] Cloud Server
[0490] software:
[0491] Google Cloud Speech-to-Text API (speech recognition engine)
[0492] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)
[0493] Visualization tools (e.g., Matplotlib, D3.js)
[0494] Web interface (for dashboard display)
[0495] Examples of specific examples and prompts
[0496] For example, if a passenger in an autonomous vehicle says, "I like this route, the scenery is amazing!", the voice is collected by a microphone. The voice data is uploaded to a cloud server and converted into text data using the Google Cloud Speech-to-Text API. It is then analyzed using natural language processing technology to extract sentiment scores and frequently occurring keywords. Finally, the analysis results are visualized as graphs and charts on a dashboard and provided to the operation manager.
[0497] Example prompt sentence:
[0498] Please perform a sentiment analysis on the following Japanese text.
[0499] Text: I love this route, the views are amazing!
[0500] Calculate positive, negative, and neutral sentiment scores and explain why. Also extract frequently occurring keywords and the number of mentions of each theme.
[0501] This allows the system to collect feedback in real time and quickly improve services.
[0502] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0503] Step 1:
[0504] A microphone installed inside the autonomous vehicle collects passenger voice data in real time. The voice data is sent to the on-board computer in WAV or MP3 format. The input is the passenger's speech, and the output is an audio file.
[0505] Step 2:
[0506] The server receives audio data from the onboard computer and uploads it to the cloud server. The input is an audio file, and the output is audio data stored on the cloud. This process sends the audio data to the cloud in real time.
[0507] Step 3:
[0508] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) on a cloud server to convert the voice data into text data. The input is an audio file, and the output is text data. The speech recognition engine analyzes the voice waveform and converts it into text data.
[0509] Step 4:
[0510] The server analyzes the converted text data. It uses natural language processing libraries (e.g., NLTK, spaCy) to extract frequently occurring keywords, tally the number of mentions of specific themes, and perform sentiment analysis (generating a score for positive, negative, or neutral). The input is the text data, and the output is the analysis results. Specific pattern matching and sentiment scoring algorithms are applied.
[0511] Step 5:
[0512] The server generates a report based on the analysis results. It visualizes them as graphs and charts using a graphical visualization tool (e.g., Matplotlib, D3.js). The input is the analysis results, and the output is a visualized report. This process makes it easier for users to intuitively understand the data.
[0513] Step 6:
[0514] Users can view the generated reports using a dedicated mobile app or web dashboard. The input is the visualized report, and the output is new insights including user feedback. Users can check each analysis result on the dashboard and take any necessary actions to improve the service.
[0515] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0516] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information and emotional information with the aim of improving user experience (UX) and customer experience (CX).
[0517] Audio data upload
[0518] Users collect interview audio data and upload it to the server using the file selection interface on their device. The audio data is formatted in common audio file formats such as WAV and MP3. The uploaded audio data is sent to the server and stored in storage.
[0519] Speech Recognition Processing
[0520] The server receives the voice data sent by the user and converts it into text data using a speech recognition engine. The speech recognition engine uses an external speech recognition service (for example, a cloud-based speech recognition API) to achieve highly accurate transcription.
[0521] Text analytics
[0522] The server parses the character data to extract the following information:
[0523] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0524] Counting the number of mentions of a specific topic: Matching a list of keywords for a pre-defined topic and counting the number of mentions of that topic.
[0525] Sentiment Analysis and Sentiment Engine
[0526] The server includes an emotion engine that uses an emotion analysis library to recognize user emotions. The emotion engine calculates positive, negative, and neutral emotion scores for each part of a statement in the text data. The emotion scores are calculated in real time and can be viewed by users on a dashboard. The server also displays the fluctuations in the emotion scores in a time-series graph format, visualizing changes in emotion over time.
[0527] Reporting and Visualization
[0528] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and sentiment scores. This information is displayed in a graphical format (e.g., graphs, charts, etc.) to help users understand it intuitively.
[0529] View results
[0530] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. Fluctuations in emotion scores can also be tracked over time, allowing users to identify when emotional statements were made.
[0531] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights and sentiment information from large amounts of interview data.
[0532] The processing flow will be explained below.
[0533] Step 1:
[0534] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data format is a common audio file format such as WAV or MP3. The user selects the audio file to upload on the file selection screen and presses the send button.
[0535] Step 2:
[0536] The server receives the audio data sent by the user via an HTTP POST request, saves the received file in a temporary directory, and registers meta information (file name, upload date and time, user ID, etc.) in the database.
[0537] Step 3:
[0538] The server sends the received voice data to a voice recognition engine for text conversion. An external voice recognition service (e.g., a cloud-based voice recognition API) is used to convert the voice into text data, and this text data is then retrieved.
[0539] Step 4:
[0540] The server stores the acquired text data in a database, and when storing the text data, associates it with the identifier of the original audio file.
[0541] Step 5:
[0542] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, creating a list of frequently occurring words to identify common keywords.
[0543] Step 6:
[0544] The server compares the text data with a list of keywords for predefined themes and tally the number of mentions related to each theme, allowing the frequency of mentions of specific topics of interest to users to be displayed in graph form.
[0545] Step 7:
[0546] The server uses a sentiment analysis library to calculate a sentiment score for each statement in the text, calculating positive, negative, or neutral sentiment scores to identify which parts are particularly emotional.
[0547] Step 8:
[0548] The server calculates the emotion score in real time and displays it on a dashboard, allowing users to check the emotion associated with each comment in real time.
[0549] Step 9:
[0550] The server tracks the fluctuations in the emotion scores over time and displays them in a graph, allowing users to visually grasp the emotional changes in the comments over time.
[0551] Step 10:
[0552] The server generates a report based on the analysis results and displays it in a graphical format (graphs, charts, etc.). Users can access the server using their devices and view this report from a dashboard, allowing them to quickly grasp important information and emotional information without having to relisten to all the interview audio data.
[0553] The above are the specific processing steps of the system based on the invention that combines the emotion engine.
[0554] Example 2
[0555] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0556] Current systems for improving user experience (UX) and customer experience (CX) face challenges in efficiently analyzing large amounts of interview audio data and quickly extracting important and emotional information. Furthermore, manual data analysis is time-consuming and labor-intensive, resulting in limited insights.
[0557] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to upload voice data using a file selection interface of a terminal; a means for converting the received voice data into text data using a voice recognition engine and storing the text data in a database; a means for analyzing the text data using a natural language processing library and extracting frequently occurring keywords, the number of mentions of a specific theme, and an emotion score; a means for aggregating the analysis results, compiling them in a report format, and visualizing them in a graphical format; and a means for a user to view the generated report on a dashboard via a web interface from their terminal. This makes it possible to quickly and efficiently obtain useful insights and emotion information from a large amount of interview voice data.
[0558] "User" means an end user who uses the System to upload audio data and view analysis results.
[0559] "Terminal" refers to the device that a user uses to upload audio data and view analysis results, such as a PC, tablet, or smartphone.
[0560] A "file selection interface" is a user interface that allows a user to select an audio file on a terminal.
[0561] A "server" is a remote computer system that receives voice data, converts it to text data, analyzes it, and generates reports.
[0562] "Speech recognition engine" means software or a service for analyzing voice data and converting it into text data, including cloud-based speech recognition services.
[0563] "Character data" refers to text-format data converted by a voice recognition engine.
[0564] A "natural language processing library" is a software library for analyzing text data to extract frequently occurring keywords and the number of times a particular topic is mentioned.
[0565] "Keywords" are important words or phrases that are set as the subject of analysis.
[0566] A "theme" is a specific topic or topic that is matched by a pre-defined list of keywords.
[0567] The "emotion score" classifies the emotions of statements contained in text data into three types: positive, negative, and neutral, and quantifies the degree of each.
[0568] A "report" is a document or data that compiles the analysis results and provides them to the user in a graphical format.
[0569] A "dashboard" is a web interface that allows users to view reports and visualizes the analysis results.
[0570] "Visualization" is the process of converting data and analytical results into a format that is easy to understand visually (e.g., graphs, charts, etc.).
[0571] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important and emotional information for the purpose of improving user experience (UX) and customer experience (CX). This system automates the process in which users upload audio data and a server analyzes the data to extract useful insights.
[0572] Hardware and software used
[0573] This system uses the following hardware and software:
[0574] Server: Receives data, converts it to text data, analyzes it, and generates reports.
[0575] Terminal: Used by users to upload audio data and view analysis results.
[0576] Speech recognition engine: Converts voice data into text using a cloud-based speech recognition service (e.g., Google Speech-to-Text API, IBM Watson Speech to Text, etc.).
[0577] Natural language processing libraries: Analyze text data and extract frequent keywords or the number of mentions of specific topics (e.g., spaCy, NLTK, etc.).
[0578] Sentiment analysis libraries: Analyze the sentiment of statements in text data and calculate positive, negative, or neutral sentiment scores (e.g., VADER, TextBlob, AWS Comprehend, etc.).
[0579] Overall system processing flow
[0580] Users collect interview audio data and upload it to the server using a file selection interface on their device. The server converts the received audio data into text using a cloud-based speech recognition engine. The converted text is stored in a database and then analyzed using a natural language processing library.
[0581] The analysis includes the following elements:
[0582] Extraction of frequently occurring keywords
[0583] Counting the number of times a specific topic was mentioned
[0584] Calculating the sentiment score
[0585] This analytical information is aggregated and compiled into reports, which are displayed in graphical formats (e.g., graphs, charts, etc.), and users can view the generated reports in a dashboard via the device's web interface.
[0586] Specific operation example
[0587] For example, a user uploads an MP3 audio file called "User Interview_1.mp3" to the system. This audio file is sent to the server and stored in storage. After that, the speech recognition engine converts the speech into text data, and the generated text data is stored in the database.
[0588] Next, the natural language processing library extracts frequently occurring keywords from the text data, confirming that keywords such as "service improvement" and "customer satisfaction" appear frequently. Furthermore, the sentiment analysis library analyzes each utterance in the text data and finds that there are many positive comments. These analysis results are compiled into a report and displayed in graph format on a dashboard.
[0589] Prompt Sentence Examples
[0590] An example prompt for a generative AI model is:
[0591] "I would like to upload interview audio data, run it through speech recognition to transcribe it, and analyze it for frequent keywords and sentiment scores. Please let me know how this data will be used to improve the user experience."
[0592] As described above, this system provides consistent support from analyzing interview audio data to detecting changes in emotions, and is a mechanism that provides useful information to users.
[0593] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0594] Step 1:
[0595] The user uploads the audio data using the file selection interface on the device. The user collects the interview audio data and uploads it by selecting the audio file on the device. This upload procedure sends the audio file (e.g., WAV or MP3 format) to the server. The input is the audio file on the user's device, and the output is the audio file stored on the server.
[0596] Step 2:
[0597] The server checks the received audio data and saves it in storage. After receiving the audio file sent by the user, the server automatically saves it in a storage system (e.g., cloud storage service). This ensures that the data is stored safely on the server. The input is the audio file uploaded to the server, and the output is the audio file saved in the server's storage.
[0598] Step 3:
[0599] The server starts a speech recognition engine and converts the stored voice data into text data. The server uses a cloud-based speech recognition engine (e.g., Google Speech-to-Text API) and inputs the stored voice data. The speech recognition engine analyzes the voice and generates corresponding text data. The input is the voice data stored in storage, and the output is the text data generated by the speech recognition engine.
[0600] Step 4:
[0601] The server stores the generated text data in a database. The server records the text data output from the speech recognition engine in a database for subsequent analysis processes. The input is the text data generated by the speech recognition engine, and the output is the text data stored in the database.
[0602] Step 5:
[0603] The server invokes a natural language processing library to analyze the text data. The server uses a natural language processing (NLP) library (e.g., spaCy) to input the text data stored in the database. The NLP library performs text analysis to extract frequently occurring keywords and the number of times a particular topic is mentioned. The input is the text data stored in the database, and the output is a list of frequently occurring keywords and the number of times a particular topic is mentioned.
[0604] Step 6:
[0605] The server launches a sentiment analysis library and calculates the sentiment score in the text data. The server uses a sentiment analysis library (e.g., VADER) to calculate the sentiment score for each utterance in the text data. The sentiment score is expressed in three indicators: positive, negative, and neutral. The input is the text data to be analyzed, and the output is a list of sentiment scores.
[0606] Step 7:
[0607] The server aggregates the analysis results and compiles them into a report. The server aggregates the data of the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and generates a report. This report is visualized in a graphical format (e.g., graph, chart, etc.). The input is the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and the output is a graphical report.
[0608] Step 8:
[0609] The user views the generated report on a dashboard via a web interface from their device. The user uses a browser on their device to access the server's web interface and view the generated report. The dashboard displays a list of frequently occurring keywords, the number of times a specific theme is mentioned, and fluctuations in sentiment scores. The input is the generated report, and the output is the analysis results on the dashboard that the user views.
[0610] (Application example 2)
[0611] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0612] Conventional advertising distribution systems have the problem that it is difficult to select optimal advertising information based on user emotions or specific keywords, which makes it difficult to attract user attention and deliver advertisements effectively.
[0613] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading voice data, means for converting the voice data into character data using a voice recognition engine, means for analyzing the character data to extract frequently occurring keywords, the number of mentions of a specific theme, and an emotion score, and means for selecting optimal advertising information based on the emotion score and keywords and displaying it on the user terminal. This makes it possible to deliver optimal advertisements that match the user's emotions in a timely manner.
[0614] "Audio data" refers to audio files that users upload to the server, and are in common audio file formats such as WAV and MP3.
[0615] "Character data" refers to text data converted from voice data using a voice recognition engine.
[0616] A "voice recognition engine" is a program or service that converts voice data into text data, and includes external voice recognition services.
[0617] "Frequent keywords" are words or phrases that appear frequently in the analyzed text data.
[0618] The "number of times a specific theme is mentioned" refers to the number of times that the theme is mentioned, as determined by matching the text data with a keyword list of the theme that is set in advance.
[0619] "Emotion score" refers to the evaluation value of positive, negative, and neutral emotions calculated for statements in text data using an emotion analysis library.
[0620] "Advertising information" refers to advertisements and promotional information selected based on the user's emotional score and keywords.
[0621] A "user terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[0622] "Visualization" refers to displaying analysis results in a graphical format that users can intuitively understand.
[0623] This invention relates to a system that analyzes user voice data, selects optimal advertising information based on emotion scores and keywords, and displays it on the user's terminal. The specific system configuration and processing procedure for implementing this invention are described below.
[0624] System Configuration
[0625] 1. User device: Collect voice data using devices such as smartphones, tablets, and PCs and upload it to the server.
[0626] 2. Server: A server located on the cloud that processes data using a speech recognition engine, natural language processing library, etc.
[0627] 3. Speech recognition engine: A program for converting voice data into text data. It uses an external cloud-based speech recognition service (e.g., Google Speech-to-Text API).
[0628] 4. Natural language processing libraries: Libraries for extracting frequent keywords and the number of mentions of specific topics from text data (e.g., spaCy, NLTK).
[0629] 5. Sentiment analysis library: A library for calculating sentiment scores from text data (e.g., VADER, Google Natural Language API).
[0630] 6. Advertising platform: Selects and delivers optimal advertising information based on sentiment scores and keywords (e.g., Google AdSense).
[0631] Program processing
[0632] The server does the following:
[0633] 1. Audio data collection and upload:
[0634] A user records audio data using a smartphone application.
[0635] The common formats used are WAV and MP3.
[0636] The audio data is uploaded from the application to a cloud server.
[0637] 2. Speech Recognition and Transcription:
[0638] A speech recognition engine (e.g., Google Speech-to-Text API) on a cloud server receives the uploaded voice data and converts it into text data.
[0639] 3. Text Analysis:
[0640] The text data is analyzed on the server using a natural language processing library (e.g., spaCy or NLTK).
[0641] Frequently occurring keywords and the number of mentions of specific topics are extracted.
[0642] 4. Emotion analysis:
[0643] Use a sentiment analysis library (e.g., VADER or Google Natural Language API) to calculate a sentiment score for each piece of text data.
[0644] Sentiment scores are expressed as either positive, negative, or neutral.
[0645] 5. Advertisement Delivery:
[0646] Based on the sentiment score and keywords, the advertising platform (e.g., Google AdSense) selects the most suitable advertising information.
[0647] The selected advertising information is displayed on the user terminal.
[0648] Specific examples
[0649] For example, if a user uploads audio data in which they say, "I want to take a weekend trip," the audio data is converted into text data and keywords such as "trip" and "weekend" are extracted. If the sentiment analysis determines that the comment has a positive sentiment, travel-related advertisements (e.g., hotel discounts or airline ticket promotions) are displayed on the user's device.
[0650] Prompt Sentence Examples
[0651] Analyze the user's spoken voice data and calculate a sentiment score of positive, negative, or neutral. Based on the results, generate prompts that suggest the ad format and content to display the optimal ad.
[0652] This system allows the most appropriate advertisements to be displayed based on the user's emotional state, maximizing the effectiveness of the advertisements.
[0653] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0654] Step 1:
[0655] A user uses a smartphone application to record audio data. This audio data is saved in WAV or MP3 format. The user then uploads the audio data to a cloud server using an interface within the app. The input is the user's audio data, and the output is an audio file saved on the cloud server.
[0656] Step 2:
[0657] The cloud server sends the received voice data to a voice recognition engine. The server converts the voice data into text data using an external cloud voice recognition service (e.g., Google Speech-to-Text API). The input here is an audio file, and the output is text data converted from the voice data.
[0658] Step 3:
[0659] The server retrieves the text data and begins analyzing it using a natural language processing library (e.g., spaCy or NLTK). Specifically, it extracts frequently occurring keywords from the text data and calculates the number of mentions of a particular topic. The input for this step is the text data, and the output is a list of keywords and a summary of topic mentions.
[0660] Step 4:
[0661] The server uses a sentiment analysis library (e.g., VADER or Google Natural Language API) to analyze each part of the text data and generate a positive, negative, or neutral sentiment score. The sentiment score is quantified according to the intensity of the emotion. The input here is the text data, and the output is a score corresponding to each emotion.
[0662] Step 5:
[0663] The server selects the optimal advertising information using an advertising platform (e.g., Google AdSense) based on the emotion score and keywords. This selection is performed using prompts generated by a generative AI model. An example of such a prompt is, "Analyze the user's spoken voice data and calculate positive, negative, and neutral emotion scores. Based on the calculation results, generate a prompt that provides the ad format and content to display the optimal advertisement." The input is the emotion score and a keyword list, and the output is the optimal advertising information.
[0664] Step 6:
[0665] The selected advertising information is sent from the server to the user terminal. The user terminal displays the received advertising information on a display so that the user can actually see it. The input is the optimal advertising information, and the output is the advertisement displayed on the user terminal.
[0666] The above are the specific processing steps of the system that realizes the application example. This series of processing provides optimal advertising information based on the user's emotions.
[0667] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0668] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0669] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0670] [Third embodiment]
[0671] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0672] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0673] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0674] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0675] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0676] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0677] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0678] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0679] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0680] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0681] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0682] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0683] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0684] Audio data upload
[0685] The user collects the interview audio data and uploads it to the server using the file selection interface on the terminal. The audio data format can be a common audio file format (e.g., WAV, MP3, etc.).
[0686] Speech Recognition Processing
[0687] The server receives the voice data uploaded by the user and converts it into text data using a speech recognition engine. The speech recognition engine can use an external speech recognition service (for example, a cloud-based speech recognition API). This allows for highly accurate transcription.
[0688] Text analytics
[0689] The server parses the character data to extract the following information:
[0690] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0691] Counting the number of mentions of a specific topic: By matching a list of keywords related to a pre-defined topic, the number of mentions of that topic is counted.
[0692] Sentiment Analysis: Performs sentiment analysis and calculates positive, negative, and neutral sentiment scores.
[0693] Reporting and Visualization
[0694] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, statistics on the number of mentions of specific themes, and sentiment scores. This information is visualized in a graphical format (e.g., graphs, charts, etc.) to allow users to intuitively understand it.
[0695] View results
[0696] Users access the server using their terminals and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. This system allows users to quickly and efficiently grasp important information without having to relisten to all the interview audio data.
[0697] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights from large amounts of interview data.
[0698] The processing flow will be explained below.
[0699] Step 1:
[0700] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data is in a common audio file format such as WAV or MP3.
[0701] Step 2:
[0702] The server receives the audio data sent by the user via an HTTP POST request and saves the files in the server's storage. At this time, each file is assigned a unique identifier for later processing.
[0703] Step 3:
[0704] The server sends the received voice data to a speech recognition engine and begins transcription. By using an external speech recognition service (for example, a cloud-based speech recognition API), highly accurate text data can be obtained.
[0705] Step 4:
[0706] The server receives the text data obtained from the speech recognition engine and stores it in a database. Each text data is assigned an identifier for the corresponding audio file.
[0707] Step 5:
[0708] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, thereby identifying frequently occurring keywords.
[0709] Step 6:
[0710] The server compares the text data with a predefined list of keywords for each topic and tally the number of mentions related to each topic, clarifying how many times a particular topic of interest to the user has been mentioned.
[0711] Step 7:
[0712] The server uses a sentiment analysis library to calculate a sentiment score for each piece of text: positive, negative, and neutral scores, identifying which statements are particularly emotional.
[0713] Step 8:
[0714] The server generates a report based on the analysis, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and a sentiment score, all of which are displayed in a graphical format.
[0715] Step 9:
[0716] The server compiles the generated reports into a dashboard format on the web interface for easy user access. Each analysis result is visualized in graphs and charts, allowing users to intuitively understand the results.
[0717] Step 10:
[0718] Users access the server using their devices and view the generated reports on a dashboard. Users can quickly grasp important information without having to listen to all of the interview audio data again. This series of processes enables users to gain useful insights from large amounts of data in a short amount of time.
[0719] Example 1
[0720] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0721] Previously, there was a lack of efficient means to analyze audio data from interviews conducted with the aim of improving user experience (UX) and customer experience (CX) and extract important information. As a result, it was difficult to quickly extract useful information from massive amounts of interview audio data, and users had to listen to all of the audio data again. This reduced the efficiency of analysis and caused delays in decision-making.
[0722] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0723] In this invention, the server includes: [means for a user to upload voice data on a terminal; [means for the server to convert the received voice data into text data using a cloud-based voice recognition API; and [means for the server to analyze the text data using a natural language processing library and extract frequently occurring keywords, the number of mentions of specific themes, and sentiment scores.] This makes it possible [to quickly and efficiently extract useful information from massive amounts of interview voice data.]
[0724] "User" refers to a person who uses the system to upload voice data and view the analysis results.
[0725] "Terminal" refers to an electronic device (e.g., a PC or smartphone) that a user uses to upload audio data.
[0726] "Server" refers to a central processing unit that processes voice data received from a user, converts it into text data, and analyzes it.
[0727] "Cloud-based speech recognition API" refers to an external service that converts voice data available via the internet into highly accurate text data.
[0728] "Text data" refers to text information converted from voice data by a cloud-based voice recognition API.
[0729] "Natural language processing library" refers to software tools and algorithms used to analyze text data.
[0730] "Frequent keywords" refer to important words that appear frequently in the character data being analyzed.
[0731] "Number of mentions of a specific theme" refers to the calculated frequency of appearance of keywords related to a predetermined theme.
[0732] "Emotion score" refers to a numerical value that quantitatively evaluates the emotion (positive, negative, neutral) of a sentence in the text data.
[0733] "Report format" refers to a document or graphical display format that summarizes the analysis results in a way that makes them easy for users to understand.
[0734] "Graphical formats" refers to formats such as graphs and charts that visually represent data.
[0735] A "dashboard" refers to a screen that organizes and displays analysis results so that users can check them at a glance.
[0736] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[0737] Audio data upload
[0738] Users collect interview audio data on their devices, preferably in common audio file formats (e.g., WAV, MP3, etc.). They then upload these audio files to the server using the device's file selection interface. Specifically, they access the system's upload page using a web browser and click the file selection button to select and upload the audio data.
[0739] Speech Recognition Processing
[0740] The server receives the uploaded audio data and converts it into text using a cloud-based speech recognition API (e.g., Google Cloud Speech-to-Text API). The server sends the audio file to the API and receives the text it returns. This process results in highly accurate transcription. For example, the audio data "Hello, what would you like to talk about today?" is converted into text.
[0741] Text analytics
[0742] The server analyzes the converted text data using a natural language processing library (e.g., NLTK or spaCy). Specifically, it performs the following process:
[0743] 1. Extraction of frequent keywords: The server extracts all tokens from the text data and calculates their frequency distribution. For example, frequently occurring keywords such as "product" and "quality" are extracted.
[0744] 2. Counting the number of mentions of a specific topic: The server checks against a list of keywords related to a predefined topic and counts the number of mentions of each topic. For example, "price" might be counted 10 times, "service" 15 times, and so on.
[0745] 3. Sentiment Analysis: The server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. For example, the sentence "I am very satisfied" is evaluated as a positive sentiment score.
[0746] Reporting and Visualization
[0747] The server generates reports based on the analysis results and uses tools (such as Matplotlib and D3.js) to visualize them in graphical formats, such as a word cloud of frequently occurring keywords or a bar chart of sentiment scores, allowing users to intuitively understand the analysis results.
[0748] View results
[0749] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to click on frequently occurring keyword clouds and bar charts of the number of mentions by theme to view more detailed information. This system allows users to quickly and efficiently grasp important information without having to relisten to the vast amount of interview audio data.
[0750] Specific examples
[0751] For example, consider a scenario where a user uploads five interview audio files (e.g., file1.wav, file2.wav, file3.mp3, file4.wav, file5.mp3). The server receives these audio files and converts them into text using the Google Cloud Speech-to-Text API. The server then analyzes the text for frequent keywords, mentions of specific themes, and sentiment scores, and compiles the results into a report. Finally, users can view this information in a visual format through a dashboard.
[0752] Examples of prompt statements
[0753] Examples of specific prompts for a generative AI model might include:
[0754] "How do I upload an audio file for transcription and text analysis?"
[0755] "Please explain how to analyze interview audio data to extract frequent keywords and sentiment scores."
[0756] By feeding such prompts into the generative AI model, more specific instructions can be obtained about the system's detailed processing and analysis methods.
[0757] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0758] Step 1:
[0759] The user uploads the audio data from the device to the server.
[0760] Specifically, the user uses the device's file selection interface (e.g., a web browser) to select an interview audio file stored on a PC or smartphone and clicks the upload button. The input is audio data (e.g., WAV or MP3 files), and the output is a notification to the server that the file has been uploaded.
[0761] Step 2:
[0762] The server converts the received voice data into text using a cloud-based speech recognition API.
[0763] Specifically, the server sends the received audio file to a specified cloud-based speech recognition API (for example, Google Cloud Speech-to-Text API). The speech recognition API receives the audio data as input, analyzes it, and converts it into document text data. The output is character data (text format). Through this process, the audio is converted into character data such as "Hello, what would you like to talk about today?"
[0764] Step 3:
[0765] The server analyzes the text data using a natural language processing library and extracts frequently occurring keywords.
[0766] Specifically, the server uses an NLP library (such as NLTK or spaCy) to extract all tokens (words) from the text data and calculate their frequency. The input is text data obtained from a speech recognition API, and the output is a list of keywords with high frequency distribution. For example, important words such as "product" and "quality" are extracted from the text data as frequent keywords.
[0767] Step 4:
[0768] The server analyzes the text data and counts the number of times a particular topic is mentioned.
[0769] Specifically, the server compares the text data with a list of keywords related to pre-set themes and tallies the number of times each theme is mentioned. The input is the text data and a list of theme keywords, and the output is data on the number of times each theme is mentioned. For example, it may be tallied that words related to "price" appear 10 times in the text data, and words related to "service" appear 15 times.
[0770] Step 5:
[0771] The server performs sentiment analysis on the text data and calculates the sentiment score.
[0772] Specifically, the server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. The input is the text data, and the output is a positive, negative, or neutral sentiment score. For example, the sentence "I am very satisfied" receives a positive evaluation, while the sentence "I am not satisfied" receives a negative evaluation.
[0773] Step 6:
[0774] The server generates a report based on the analysis results and visualizes them in a graphical format.
[0775] Specifically, the server generates a report based on the analysis results (e.g., frequently occurring keywords, number of mentions of specific themes, sentiment scores) and visualizes them using a data visualization tool (e.g., Matplotlib or D3.js). The input is the analysis result data, and the output is a graphically represented report. For example, a word cloud of frequently occurring keywords or a bar chart of sentiment scores may be generated.
[0776] Step 7:
[0777] The user views the generated report on their device.
[0778] Specifically, users access the system's dashboard page via a web browser on their device. The input is report data generated and visualized on the server, and the output is a visually displayed report. On the dashboard, users can click on the frequently occurring keyword cloud or the bar chart of the number of mentions by topic to view more detailed information. This system allows users to efficiently analyze massive amounts of voice data and quickly obtain important information.
[0779] (Application example 1)
[0780] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0781] Conventional voice analysis systems aimed at improving user and customer experience rely on uploading and analyzing individually collected voice data as it arrives. This makes it difficult to collect feedback in real time and quickly improve services. Furthermore, uploading and analyzing voice data can be difficult, resulting in inefficiencies due to an inability to handle large volumes of data. For autonomous vehicles in particular, it is important to collect passenger feedback in real time and quickly reflect it in improvements to operation and service quality, but current technology makes this difficult.
[0782] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0783] In this invention, the server includes: [means for a user to upload voice data;] [means for the server to convert the received voice data into text data using a voice recognition engine;] [means for the server to analyze the text data and extract frequently occurring keywords, the number of mentions of specific themes, and an emotion score;] [means for collecting voice data in real time;] [means for uploading the voice data to a cloud server; and [means for collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service. This enables feedback to be collected and analyzed in real time, enabling rapid service improvements.
[0784] "User" refers to a user who collects voice data and provides the data to the system.
[0785] "Audio data" refers to audio files that record what passengers say.
[0786] A "server" is a central processing unit that receives and analyzes voice data.
[0787] A "voice recognition engine" is software for converting voice data into text data.
[0788] "Character data" refers to data obtained by converting voice data into text format.
[0789] "Keywords" are important words or phrases that appear frequently within the audio data.
[0790] A "specific theme" is a predefined topic or issue that is the subject of analysis.
[0791] The "number of mentions" is a numerical value indicating how many times a word related to a particular theme appears in the speech data.
[0792] An "emotion score" is a numerical representation of the emotions in audio data, and is divided into three categories: positive, negative, and neutral.
[0793] A "report" is a report summarizing the results of the analysis of character data.
[0794] "Visualization" means displaying the analysis results in a visually easy-to-understand format such as graphs or charts.
[0795] A "dashboard" is an interface that allows users to view analysis results.
[0796] "Real-time" means that processing occurs with virtually no time delay.
[0797] "Feedback" refers to opinions and ratings provided by passengers.
[0798] "Service Improvements" are specific changes or modifications made to improve the quality of the Services provided.
[0799] A "cloud server" refers to a group of servers on the Internet that can be accessed remotely.
[0800] Basic configuration
[0801] The system includes means for "users to upload voice data," means for "the server converting the received voice data into text data using a voice recognition engine," means for "the server analyzing the text data and extracting frequently occurring keywords, the number of mentions of specific themes, and sentiment scores," means for "the server compiling and visualizing the analysis results in report format," means for "users viewing the generated report," means for "collecting voice data in real time," means for "uploading voice data to a cloud server," and means for "collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service."
[0802] Program processing
[0803] The server first collects audio data in real time using a microphone installed inside the autonomous vehicle. The audio data is recorded as audio data, and the data is uploaded to a cloud server via the on-board computer. The audio file format is a common format such as WAV or MP3.
[0804] When the cloud server receives the voice data, it uses a voice recognition engine to convert the voice data into text data. An external voice recognition service (e.g., Google Cloud Speech-to-Text API) is used as the voice recognition engine. To ensure high accuracy of the conversion, the voice recognition engine uses a cloud-based service.
[0805] The converted text data is then analyzed by the server. During the analysis process, natural language processing (NLP) techniques are used to extract frequently occurring keywords, tally the number of times a particular topic is mentioned, and perform sentiment analysis, which generates a positive, negative, or neutral sentiment score.
[0806] The analysis results are compiled into a report, which is displayed in a visual format such as graphs and charts for easy understanding by the user. The generated report can be viewed by the operation manager via a dedicated mobile app or web dashboard.
[0807] Hardware and software used
[0808] Hardware:
[0809] Microphone for collecting audio data
[0810] On-board computer
[0811] Cloud Server
[0812] software:
[0813] Google Cloud Speech-to-Text API (speech recognition engine)
[0814] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)
[0815] Visualization tools (e.g., Matplotlib, D3.js)
[0816] Web interface (for dashboard display)
[0817] Examples of specific examples and prompts
[0818] For example, if a passenger in an autonomous vehicle says, "I like this route, the scenery is amazing!", the voice is collected by a microphone. The voice data is uploaded to a cloud server and converted into text data using the Google Cloud Speech-to-Text API. It is then analyzed using natural language processing technology to extract sentiment scores and frequently occurring keywords. Finally, the analysis results are visualized as graphs and charts on a dashboard and provided to the operation manager.
[0819] Example prompt sentence:
[0820] Please perform a sentiment analysis on the following Japanese text.
[0821] Text: I love this route, the views are amazing!
[0822] Calculate positive, negative, and neutral sentiment scores and explain why. Also extract frequently occurring keywords and the number of mentions of each theme.
[0823] This allows the system to collect feedback in real time and quickly improve services.
[0824] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0825] Step 1:
[0826] A microphone installed inside the autonomous vehicle collects passenger voice data in real time. The voice data is sent to the on-board computer in WAV or MP3 format. The input is the passenger's speech, and the output is an audio file.
[0827] Step 2:
[0828] The server receives audio data from the onboard computer and uploads it to the cloud server. The input is an audio file, and the output is audio data stored on the cloud. This process sends the audio data to the cloud in real time.
[0829] Step 3:
[0830] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) on a cloud server to convert the voice data into text data. The input is an audio file, and the output is text data. The speech recognition engine analyzes the voice waveform and converts it into text data.
[0831] Step 4:
[0832] The server analyzes the converted text data. It uses natural language processing libraries (e.g., NLTK, spaCy) to extract frequently occurring keywords, tally the number of mentions of specific themes, and perform sentiment analysis (generating a score for positive, negative, or neutral). The input is the text data, and the output is the analysis results. Specific pattern matching and sentiment scoring algorithms are applied.
[0833] Step 5:
[0834] The server generates a report based on the analysis results. It visualizes them as graphs and charts using a graphical visualization tool (e.g., Matplotlib, D3.js). The input is the analysis results, and the output is a visualized report. This process makes it easier for users to intuitively understand the data.
[0835] Step 6:
[0836] Users can view the generated reports using a dedicated mobile app or web dashboard. The input is the visualized report, and the output is new insights including user feedback. Users can check each analysis result on the dashboard and take any necessary actions to improve the service.
[0837] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0838] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information and emotional information with the aim of improving user experience (UX) and customer experience (CX).
[0839] Audio data upload
[0840] Users collect interview audio data and upload it to the server using the file selection interface on their device. The audio data is formatted in common audio file formats such as WAV and MP3. The uploaded audio data is sent to the server and stored in storage.
[0841] Speech Recognition Processing
[0842] The server receives the voice data sent by the user and converts it into text data using a speech recognition engine. The speech recognition engine uses an external speech recognition service (for example, a cloud-based speech recognition API) to achieve highly accurate transcription.
[0843] Text analytics
[0844] The server parses the character data to extract the following information:
[0845] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[0846] Counting the number of mentions of a specific topic: Matching a list of keywords for a pre-defined topic and counting the number of mentions of that topic.
[0847] Sentiment Analysis and Sentiment Engine
[0848] The server includes an emotion engine that uses an emotion analysis library to recognize user emotions. The emotion engine calculates positive, negative, and neutral emotion scores for each part of a statement in the text data. The emotion scores are calculated in real time and can be viewed by users on a dashboard. The server also displays the fluctuations in the emotion scores in a time-series graph format, visualizing changes in emotion over time.
[0849] Reporting and Visualization
[0850] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and sentiment scores. This information is displayed in a graphical format (e.g., graphs, charts, etc.) to help users understand it intuitively.
[0851] View results
[0852] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. Fluctuations in emotion scores can also be tracked over time, allowing users to identify when emotional statements were made.
[0853] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights and sentiment information from large amounts of interview data.
[0854] The processing flow will be explained below.
[0855] Step 1:
[0856] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data format is a common audio file format such as WAV or MP3. The user selects the audio file to upload on the file selection screen and presses the send button.
[0857] Step 2:
[0858] The server receives the audio data sent by the user via an HTTP POST request, saves the received file in a temporary directory, and registers meta information (file name, upload date and time, user ID, etc.) in the database.
[0859] Step 3:
[0860] The server sends the received voice data to a voice recognition engine for text conversion. An external voice recognition service (e.g., a cloud-based voice recognition API) is used to convert the voice into text data, and this text data is then retrieved.
[0861] Step 4:
[0862] The server stores the acquired text data in a database, and when storing the text data, associates it with the identifier of the original audio file.
[0863] Step 5:
[0864] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, creating a list of frequently occurring words to identify common keywords.
[0865] Step 6:
[0866] The server compares the text data with a list of keywords for predefined themes and tally the number of mentions related to each theme, allowing the frequency of mentions of specific topics of interest to users to be displayed in graph form.
[0867] Step 7:
[0868] The server uses a sentiment analysis library to calculate a sentiment score for each statement in the text, calculating positive, negative, or neutral sentiment scores to identify which parts are particularly emotional.
[0869] Step 8:
[0870] The server calculates the emotion score in real time and displays it on a dashboard, allowing users to check the emotion associated with each comment in real time.
[0871] Step 9:
[0872] The server tracks the fluctuations in the emotion scores over time and displays them in a graph, allowing users to visually grasp the emotional changes in the comments over time.
[0873] Step 10:
[0874] The server generates a report based on the analysis results and displays it in a graphical format (graphs, charts, etc.). Users can access the server using their devices and view this report from a dashboard, allowing them to quickly grasp important information and emotional information without having to relisten to all the interview audio data.
[0875] The above are the specific processing steps of the system based on the invention that combines the emotion engine.
[0876] Example 2
[0877] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0878] Current systems for improving user experience (UX) and customer experience (CX) face challenges in efficiently analyzing large amounts of interview audio data and quickly extracting important and emotional information. Furthermore, manual data analysis is time-consuming and labor-intensive, resulting in limited insights.
[0879] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to upload voice data using a file selection interface of a terminal; a means for converting the received voice data into text data using a voice recognition engine and storing the text data in a database; a means for analyzing the text data using a natural language processing library and extracting frequently occurring keywords, the number of mentions of a specific theme, and an emotion score; a means for aggregating the analysis results, compiling them in a report format, and visualizing them in a graphical format; and a means for a user to view the generated report on a dashboard via a web interface from their terminal. This makes it possible to quickly and efficiently obtain useful insights and emotion information from a large amount of interview voice data.
[0880] "User" means an end user who uses the System to upload audio data and view analysis results.
[0881] "Terminal" refers to the device that a user uses to upload audio data and view analysis results, such as a PC, tablet, or smartphone.
[0882] A "file selection interface" is a user interface that allows a user to select an audio file on a terminal.
[0883] A "server" is a remote computer system that receives voice data, converts it to text data, analyzes it, and generates reports.
[0884] "Speech recognition engine" means software or a service for analyzing voice data and converting it into text data, including cloud-based speech recognition services.
[0885] "Character data" refers to text-format data converted by a voice recognition engine.
[0886] A "natural language processing library" is a software library for analyzing text data to extract frequently occurring keywords and the number of times a particular topic is mentioned.
[0887] "Keywords" are important words or phrases that are set as the subject of analysis.
[0888] A "theme" is a specific topic or topic that is matched by a pre-defined list of keywords.
[0889] The "emotion score" classifies the emotions of statements contained in text data into three types: positive, negative, and neutral, and quantifies the degree of each.
[0890] A "report" is a document or data that compiles the analysis results and provides them to the user in a graphical format.
[0891] A "dashboard" is a web interface that allows users to view reports and visualizes the analysis results.
[0892] "Visualization" is the process of converting data and analytical results into a format that is easy to understand visually (e.g., graphs, charts, etc.).
[0893] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important and emotional information for the purpose of improving user experience (UX) and customer experience (CX). This system automates the process in which users upload audio data and a server analyzes the data to extract useful insights.
[0894] Hardware and software used
[0895] This system uses the following hardware and software:
[0896] Server: Receives data, converts it to text data, analyzes it, and generates reports.
[0897] Terminal: Used by users to upload audio data and view analysis results.
[0898] Speech recognition engine: Converts voice data into text using a cloud-based speech recognition service (e.g., Google Speech-to-Text API, IBM Watson Speech to Text, etc.).
[0899] Natural language processing libraries: Analyze text data and extract frequent keywords or the number of mentions of specific topics (e.g., spaCy, NLTK, etc.).
[0900] Sentiment analysis libraries: Analyze the sentiment of statements in text data and calculate positive, negative, or neutral sentiment scores (e.g., VADER, TextBlob, AWS Comprehend, etc.).
[0901] Overall system processing flow
[0902] Users collect interview audio data and upload it to the server using a file selection interface on their device. The server converts the received audio data into text using a cloud-based speech recognition engine. The converted text is stored in a database and then analyzed using a natural language processing library.
[0903] The analysis includes the following elements:
[0904] Extraction of frequently occurring keywords
[0905] Counting the number of times a specific topic was mentioned
[0906] Calculating the sentiment score
[0907] This analytical information is aggregated and compiled into reports, which are displayed in graphical formats (e.g., graphs, charts, etc.), and users can view the generated reports in a dashboard via the device's web interface.
[0908] Specific operation example
[0909] For example, a user uploads an MP3 audio file called "User Interview_1.mp3" to the system. This audio file is sent to the server and stored in storage. After that, the speech recognition engine converts the speech into text data, and the generated text data is stored in the database.
[0910] Next, the natural language processing library extracts frequently occurring keywords from the text data, confirming that keywords such as "service improvement" and "customer satisfaction" appear frequently. Furthermore, the sentiment analysis library analyzes each utterance in the text data and finds that there are many positive comments. These analysis results are compiled into a report and displayed in graph format on a dashboard.
[0911] Prompt Sentence Examples
[0912] An example prompt for a generative AI model is:
[0913] "I would like to upload interview audio data, run it through speech recognition to transcribe it, and analyze it for frequent keywords and sentiment scores. Please let me know how this data will be used to improve the user experience."
[0914] As described above, this system provides consistent support from analyzing interview audio data to detecting changes in emotions, and is a mechanism that provides useful information to users.
[0915] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0916] Step 1:
[0917] The user uploads the audio data using the file selection interface on the device. The user collects the interview audio data and uploads it by selecting the audio file on the device. This upload procedure sends the audio file (e.g., WAV or MP3 format) to the server. The input is the audio file on the user's device, and the output is the audio file stored on the server.
[0918] Step 2:
[0919] The server checks the received audio data and saves it in storage. After receiving the audio file sent by the user, the server automatically saves it in a storage system (e.g., cloud storage service). This ensures that the data is stored safely on the server. The input is the audio file uploaded to the server, and the output is the audio file saved in the server's storage.
[0920] Step 3:
[0921] The server starts a speech recognition engine and converts the stored voice data into text data. The server uses a cloud-based speech recognition engine (e.g., Google Speech-to-Text API) and inputs the stored voice data. The speech recognition engine analyzes the voice and generates corresponding text data. The input is the voice data stored in storage, and the output is the text data generated by the speech recognition engine.
[0922] Step 4:
[0923] The server stores the generated text data in a database. The server records the text data output from the speech recognition engine in a database for subsequent analysis processes. The input is the text data generated by the speech recognition engine, and the output is the text data stored in the database.
[0924] Step 5:
[0925] The server invokes a natural language processing library to analyze the text data. The server uses a natural language processing (NLP) library (e.g., spaCy) to input the text data stored in the database. The NLP library performs text analysis to extract frequently occurring keywords and the number of times a particular topic is mentioned. The input is the text data stored in the database, and the output is a list of frequently occurring keywords and the number of times a particular topic is mentioned.
[0926] Step 6:
[0927] The server launches a sentiment analysis library and calculates the sentiment score in the text data. The server uses a sentiment analysis library (e.g., VADER) to calculate the sentiment score for each utterance in the text data. The sentiment score is expressed in three indicators: positive, negative, and neutral. The input is the text data to be analyzed, and the output is a list of sentiment scores.
[0928] Step 7:
[0929] The server aggregates the analysis results and compiles them into a report. The server aggregates the data of the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and generates a report. This report is visualized in a graphical format (e.g., graph, chart, etc.). The input is the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and the output is a graphical report.
[0930] Step 8:
[0931] The user views the generated report on a dashboard via a web interface from their device. The user uses a browser on their device to access the server's web interface and view the generated report. The dashboard displays a list of frequently occurring keywords, the number of times a specific theme is mentioned, and fluctuations in sentiment scores. The input is the generated report, and the output is the analysis results on the dashboard that the user views.
[0932] (Application example 2)
[0933] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0934] Conventional advertising distribution systems have the problem that it is difficult to select optimal advertising information based on user emotions or specific keywords, which makes it difficult to attract user attention and deliver advertisements effectively.
[0935] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading voice data, means for converting the voice data into character data using a voice recognition engine, means for analyzing the character data to extract frequently occurring keywords, the number of mentions of a specific theme, and an emotion score, and means for selecting optimal advertising information based on the emotion score and keywords and displaying it on the user terminal. This makes it possible to deliver optimal advertisements that match the user's emotions in a timely manner.
[0936] "Audio data" refers to audio files that users upload to the server, and are in common audio file formats such as WAV and MP3.
[0937] "Character data" refers to text data converted from voice data using a voice recognition engine.
[0938] A "voice recognition engine" is a program or service that converts voice data into text data, and includes external voice recognition services.
[0939] "Frequent keywords" are words or phrases that appear frequently in the analyzed text data.
[0940] The "number of times a specific theme is mentioned" refers to the number of times that the theme is mentioned, as determined by matching the text data with a keyword list of the theme that is set in advance.
[0941] "Emotion score" refers to the evaluation value of positive, negative, and neutral emotions calculated for statements in text data using an emotion analysis library.
[0942] "Advertising information" refers to advertisements and promotional information selected based on the user's emotional score and keywords.
[0943] A "user terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[0944] "Visualization" refers to displaying analysis results in a graphical format that users can intuitively understand.
[0945] This invention relates to a system that analyzes user voice data, selects optimal advertising information based on emotion scores and keywords, and displays it on the user's terminal. The specific system configuration and processing procedure for implementing this invention are described below.
[0946] System Configuration
[0947] 1. User device: Collect voice data using devices such as smartphones, tablets, and PCs and upload it to the server.
[0948] 2. Server: A server located on the cloud that processes data using a speech recognition engine, natural language processing library, etc.
[0949] 3. Speech recognition engine: A program for converting voice data into text data. It uses an external cloud-based speech recognition service (e.g., Google Speech-to-Text API).
[0950] 4. Natural language processing libraries: Libraries for extracting frequent keywords and the number of mentions of specific topics from text data (e.g., spaCy, NLTK).
[0951] 5. Sentiment analysis library: A library for calculating sentiment scores from text data (e.g., VADER, Google Natural Language API).
[0952] 6. Advertising platform: Selects and delivers optimal advertising information based on sentiment scores and keywords (e.g., Google AdSense).
[0953] Program processing
[0954] The server does the following:
[0955] 1. Audio data collection and upload:
[0956] A user records audio data using a smartphone application.
[0957] The common formats used are WAV and MP3.
[0958] The audio data is uploaded from the application to a cloud server.
[0959] 2. Speech Recognition and Transcription:
[0960] A speech recognition engine (e.g., Google Speech-to-Text API) on a cloud server receives the uploaded voice data and converts it into text data.
[0961] 3. Text Analysis:
[0962] The text data is analyzed on the server using a natural language processing library (e.g., spaCy or NLTK).
[0963] Frequently occurring keywords and the number of mentions of specific topics are extracted.
[0964] 4. Emotion analysis:
[0965] Use a sentiment analysis library (e.g., VADER or Google Natural Language API) to calculate a sentiment score for each piece of text data.
[0966] Sentiment scores are expressed as either positive, negative, or neutral.
[0967] 5. Advertisement Delivery:
[0968] Based on the sentiment score and keywords, the advertising platform (e.g., Google AdSense) selects the most suitable advertising information.
[0969] The selected advertising information is displayed on the user terminal.
[0970] Specific examples
[0971] For example, if a user uploads audio data in which they say, "I want to take a weekend trip," the audio data is converted into text data and keywords such as "trip" and "weekend" are extracted. If the sentiment analysis determines that the comment has a positive sentiment, travel-related advertisements (e.g., hotel discounts or airline ticket promotions) are displayed on the user's device.
[0972] Prompt Sentence Examples
[0973] Analyze the user's spoken voice data and calculate a sentiment score of positive, negative, or neutral. Based on the results, generate prompts that suggest the ad format and content to display the optimal ad.
[0974] This system allows the most appropriate advertisements to be displayed based on the user's emotional state, maximizing the effectiveness of the advertisements.
[0975] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0976] Step 1:
[0977] A user uses a smartphone application to record audio data. This audio data is saved in WAV or MP3 format. The user then uploads the audio data to a cloud server using an interface within the app. The input is the user's audio data, and the output is an audio file saved on the cloud server.
[0978] Step 2:
[0979] The cloud server sends the received voice data to a voice recognition engine. The server converts the voice data into text data using an external cloud voice recognition service (e.g., Google Speech-to-Text API). The input here is an audio file, and the output is text data converted from the voice data.
[0980] Step 3:
[0981] The server retrieves the text data and begins analyzing it using a natural language processing library (e.g., spaCy or NLTK). Specifically, it extracts frequently occurring keywords from the text data and calculates the number of mentions of a particular topic. The input for this step is the text data, and the output is a list of keywords and a summary of topic mentions.
[0982] Step 4:
[0983] The server uses a sentiment analysis library (e.g., VADER or Google Natural Language API) to analyze each part of the text data and generate a positive, negative, or neutral sentiment score. The sentiment score is quantified according to the intensity of the emotion. The input here is the text data, and the output is a score corresponding to each emotion.
[0984] Step 5:
[0985] The server selects the optimal advertising information using an advertising platform (e.g., Google AdSense) based on the emotion score and keywords. This selection is performed using prompts generated by a generative AI model. An example of such a prompt is, "Analyze the user's spoken voice data and calculate positive, negative, and neutral emotion scores. Based on the calculation results, generate a prompt that provides the ad format and content to display the optimal advertisement." The input is the emotion score and a keyword list, and the output is the optimal advertising information.
[0986] Step 6:
[0987] The selected advertising information is sent from the server to the user terminal. The user terminal displays the received advertising information on a display so that the user can actually see it. The input is the optimal advertising information, and the output is the advertisement displayed on the user terminal.
[0988] The above are the specific processing steps of the system that realizes the application example. This series of processing provides optimal advertising information based on the user's emotions.
[0989] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0990] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0991] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0992] [Fourth embodiment]
[0993] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0994] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0995] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0996] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0997] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0998] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0999] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1000] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1001] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1002] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1003] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1004] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1005] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1006] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[1007] Audio data upload
[1008] The user collects the interview audio data and uploads it to the server using the file selection interface on the terminal. The audio data format can be a common audio file format (e.g., WAV, MP3, etc.).
[1009] Speech Recognition Processing
[1010] The server receives the voice data uploaded by the user and converts it into text data using a speech recognition engine. The speech recognition engine can use an external speech recognition service (for example, a cloud-based speech recognition API). This allows for highly accurate transcription.
[1011] Text analytics
[1012] The server parses the character data to extract the following information:
[1013] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[1014] Counting the number of mentions of a specific topic: By matching a list of keywords related to a pre-defined topic, the number of mentions of that topic is counted.
[1015] Sentiment Analysis: Performs sentiment analysis and calculates positive, negative, and neutral sentiment scores.
[1016] Reporting and Visualization
[1017] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, statistics on the number of mentions of specific themes, and sentiment scores. This information is visualized in a graphical format (e.g., graphs, charts, etc.) to allow users to intuitively understand it.
[1018] View results
[1019] Users access the server using their terminals and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. This system allows users to quickly and efficiently grasp important information without having to relisten to all the interview audio data.
[1020] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights from large amounts of interview data.
[1021] The processing flow will be explained below.
[1022] Step 1:
[1023] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data is in a common audio file format such as WAV or MP3.
[1024] Step 2:
[1025] The server receives the audio data sent by the user via an HTTP POST request and saves the files in the server's storage. At this time, each file is assigned a unique identifier for later processing.
[1026] Step 3:
[1027] The server sends the received voice data to a speech recognition engine and begins transcription. By using an external speech recognition service (for example, a cloud-based speech recognition API), highly accurate text data can be obtained.
[1028] Step 4:
[1029] The server receives the text data obtained from the speech recognition engine and stores it in a database. Each text data is assigned an identifier for the corresponding audio file.
[1030] Step 5:
[1031] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, thereby identifying frequently occurring keywords.
[1032] Step 6:
[1033] The server compares the text data with a predefined list of keywords for each topic and tally the number of mentions related to each topic, clarifying how many times a particular topic of interest to the user has been mentioned.
[1034] Step 7:
[1035] The server uses a sentiment analysis library to calculate a sentiment score for each piece of text: positive, negative, and neutral scores, identifying which statements are particularly emotional.
[1036] Step 8:
[1037] The server generates a report based on the analysis, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and a sentiment score, all of which are displayed in a graphical format.
[1038] Step 9:
[1039] The server compiles the generated reports into a dashboard format on the web interface for easy user access. Each analysis result is visualized in graphs and charts, allowing users to intuitively understand the results.
[1040] Step 10:
[1041] Users access the server using their devices and view the generated reports on a dashboard. Users can quickly grasp important information without having to listen to all of the interview audio data again. This series of processes enables users to gain useful insights from large amounts of data in a short amount of time.
[1042] Example 1
[1043] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1044] Previously, there was a lack of efficient means to analyze audio data from interviews conducted with the aim of improving user experience (UX) and customer experience (CX) and extract important information. As a result, it was difficult to quickly extract useful information from massive amounts of interview audio data, and users had to listen to all of the audio data again. This reduced the efficiency of analysis and caused delays in decision-making.
[1045] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1046] In this invention, the server includes: [means for a user to upload voice data on a terminal; [means for the server to convert the received voice data into text data using a cloud-based voice recognition API; and [means for the server to analyze the text data using a natural language processing library and extract frequently occurring keywords, the number of mentions of specific themes, and sentiment scores.] This makes it possible [to quickly and efficiently extract useful information from massive amounts of interview voice data.]
[1047] "User" refers to a person who uses the system to upload voice data and view the analysis results.
[1048] "Terminal" refers to an electronic device (e.g., a PC or smartphone) that a user uses to upload audio data.
[1049] "Server" refers to a central processing unit that processes voice data received from a user, converts it into text data, and analyzes it.
[1050] "Cloud-based speech recognition API" refers to an external service that converts voice data available via the internet into highly accurate text data.
[1051] "Text data" refers to text information converted from voice data by a cloud-based voice recognition API.
[1052] "Natural language processing library" refers to software tools and algorithms used to analyze text data.
[1053] "Frequent keywords" refer to important words that appear frequently in the character data being analyzed.
[1054] "Number of mentions of a specific theme" refers to the calculated frequency of appearance of keywords related to a predetermined theme.
[1055] "Emotion score" refers to a numerical value that quantitatively evaluates the emotion (positive, negative, neutral) of a sentence in the text data.
[1056] "Report format" refers to a document or graphical display format that summarizes the analysis results in a way that makes them easy for users to understand.
[1057] "Graphical formats" refers to formats such as graphs and charts that visually represent data.
[1058] A "dashboard" refers to a screen that organizes and displays analysis results so that users can check them at a glance.
[1059] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information with the aim of improving user experience (UX) and customer experience (CX).
[1060] Audio data upload
[1061] Users collect interview audio data on their devices, preferably in common audio file formats (e.g., WAV, MP3, etc.). They then upload these audio files to the server using the device's file selection interface. Specifically, they access the system's upload page using a web browser and click the file selection button to select and upload the audio data.
[1062] Speech Recognition Processing
[1063] The server receives the uploaded audio data and converts it into text using a cloud-based speech recognition API (e.g., Google Cloud Speech-to-Text API). The server sends the audio file to the API and receives the text it returns. This process results in highly accurate transcription. For example, the audio data "Hello, what would you like to talk about today?" is converted into text.
[1064] Text analytics
[1065] The server analyzes the converted text data using a natural language processing library (e.g., NLTK or spaCy). Specifically, it performs the following process:
[1066] 1. Extraction of frequent keywords: The server extracts all tokens from the text data and calculates their frequency distribution. For example, frequently occurring keywords such as "product" and "quality" are extracted.
[1067] 2. Counting the number of mentions of a specific topic: The server checks against a list of keywords related to a predefined topic and counts the number of mentions of each topic. For example, "price" might be counted 10 times, "service" 15 times, and so on.
[1068] 3. Sentiment Analysis: The server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. For example, the sentence "I am very satisfied" is evaluated as a positive sentiment score.
[1069] Reporting and Visualization
[1070] The server generates reports based on the analysis results and uses tools (such as Matplotlib and D3.js) to visualize them in graphical formats, such as a word cloud of frequently occurring keywords or a bar chart of sentiment scores, allowing users to intuitively understand the analysis results.
[1071] View results
[1072] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to click on frequently occurring keyword clouds and bar charts of the number of mentions by theme to view more detailed information. This system allows users to quickly and efficiently grasp important information without having to relisten to the vast amount of interview audio data.
[1073] Specific examples
[1074] For example, consider a scenario where a user uploads five interview audio files (e.g., file1.wav, file2.wav, file3.mp3, file4.wav, file5.mp3). The server receives these audio files and converts them into text using the Google Cloud Speech-to-Text API. The server then analyzes the text for frequent keywords, mentions of specific themes, and sentiment scores, and compiles the results into a report. Finally, users can view this information in a visual format through a dashboard.
[1075] Examples of prompt statements
[1076] Examples of specific prompts for a generative AI model might include:
[1077] "How do I upload an audio file for transcription and text analysis?"
[1078] "Please explain how to analyze interview audio data to extract frequent keywords and sentiment scores."
[1079] By feeding such prompts into the generative AI model, more specific instructions can be obtained about the system's detailed processing and analysis methods.
[1080] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1081] Step 1:
[1082] The user uploads the audio data from the device to the server.
[1083] Specifically, the user uses the device's file selection interface (e.g., a web browser) to select an interview audio file stored on a PC or smartphone and clicks the upload button. The input is audio data (e.g., WAV or MP3 files), and the output is a notification to the server that the file has been uploaded.
[1084] Step 2:
[1085] The server converts the received voice data into text using a cloud-based speech recognition API.
[1086] Specifically, the server sends the received audio file to a specified cloud-based speech recognition API (for example, Google Cloud Speech-to-Text API). The speech recognition API receives the audio data as input, analyzes it, and converts it into document text data. The output is character data (text format). Through this process, the audio is converted into character data such as "Hello, what would you like to talk about today?"
[1087] Step 3:
[1088] The server analyzes the text data using a natural language processing library and extracts frequently occurring keywords.
[1089] Specifically, the server uses an NLP library (such as NLTK or spaCy) to extract all tokens (words) from the text data and calculate their frequency. The input is text data obtained from a speech recognition API, and the output is a list of keywords with high frequency distribution. For example, important words such as "product" and "quality" are extracted from the text data as frequent keywords.
[1090] Step 4:
[1091] The server analyzes the text data and counts the number of times a particular topic is mentioned.
[1092] Specifically, the server compares the text data with a list of keywords related to pre-set themes and tallies the number of times each theme is mentioned. The input is the text data and a list of theme keywords, and the output is data on the number of times each theme is mentioned. For example, it may be tallied that words related to "price" appear 10 times in the text data, and words related to "service" appear 15 times.
[1093] Step 5:
[1094] The server performs sentiment analysis on the text data and calculates the sentiment score.
[1095] Specifically, the server uses a sentiment analysis library (e.g., TextBlob or VADER) to calculate the sentiment score of the text data. The input is the text data, and the output is a positive, negative, or neutral sentiment score. For example, the sentence "I am very satisfied" receives a positive evaluation, while the sentence "I am not satisfied" receives a negative evaluation.
[1096] Step 6:
[1097] The server generates a report based on the analysis results and visualizes them in a graphical format.
[1098] Specifically, the server generates a report based on the analysis results (e.g., frequently occurring keywords, number of mentions of specific themes, sentiment scores) and visualizes them using a data visualization tool (e.g., Matplotlib or D3.js). The input is the analysis result data, and the output is a graphically represented report. For example, a word cloud of frequently occurring keywords or a bar chart of sentiment scores may be generated.
[1099] Step 7:
[1100] The user views the generated report on their device.
[1101] Specifically, users access the system's dashboard page via a web browser on their device. The input is report data generated and visualized on the server, and the output is a visually displayed report. On the dashboard, users can click on the frequently occurring keyword cloud or the bar chart of the number of mentions by topic to view more detailed information. This system allows users to efficiently analyze massive amounts of voice data and quickly obtain important information.
[1102] (Application example 1)
[1103] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1104] Conventional voice analysis systems aimed at improving user and customer experience rely on uploading and analyzing individually collected voice data as it arrives. This makes it difficult to collect feedback in real time and quickly improve services. Furthermore, uploading and analyzing voice data can be difficult, resulting in inefficiencies due to an inability to handle large volumes of data. For autonomous vehicles in particular, it is important to collect passenger feedback in real time and quickly reflect it in improvements to operation and service quality, but current technology makes this difficult.
[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1106] In this invention, the server includes: [means for a user to upload voice data;] [means for the server to convert the received voice data into text data using a voice recognition engine;] [means for the server to analyze the text data and extract frequently occurring keywords, the number of mentions of specific themes, and an emotion score;] [means for collecting voice data in real time;] [means for uploading the voice data to a cloud server; and [means for collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service. This enables feedback to be collected and analyzed in real time, enabling rapid service improvements.
[1107] "User" refers to a user who collects voice data and provides the data to the system.
[1108] "Audio data" refers to audio files that record what passengers say.
[1109] A "server" is a central processing unit that receives and analyzes voice data.
[1110] A "voice recognition engine" is software for converting voice data into text data.
[1111] "Character data" refers to data obtained by converting voice data into text format.
[1112] "Keywords" are important words or phrases that appear frequently within the audio data.
[1113] A "specific theme" is a predefined topic or issue that is the subject of analysis.
[1114] The "number of mentions" is a numerical value indicating how many times a word related to a particular theme appears in the speech data.
[1115] An "emotion score" is a numerical representation of the emotions in audio data, and is divided into three categories: positive, negative, and neutral.
[1116] A "report" is a report summarizing the results of the analysis of character data.
[1117] "Visualization" means displaying the analysis results in a visually easy-to-understand format such as graphs or charts.
[1118] A "dashboard" is an interface that allows users to view analysis results.
[1119] "Real-time" means that processing occurs with virtually no time delay.
[1120] "Feedback" refers to opinions and ratings provided by passengers.
[1121] "Service Improvements" are specific changes or modifications made to improve the quality of the Services provided.
[1122] A "cloud server" refers to a group of servers on the Internet that can be accessed remotely.
[1123] Basic configuration
[1124] The system includes means for "users to upload voice data," means for "the server converting the received voice data into text data using a voice recognition engine," means for "the server analyzing the text data and extracting frequently occurring keywords, the number of mentions of specific themes, and sentiment scores," means for "the server compiling and visualizing the analysis results in report format," means for "users viewing the generated report," means for "collecting voice data in real time," means for "uploading voice data to a cloud server," and means for "collecting passenger feedback in real time and quickly identifying areas for improvement in operation and service."
[1125] Program processing
[1126] The server first collects audio data in real time using a microphone installed inside the autonomous vehicle. The audio data is recorded as audio data, and the data is uploaded to a cloud server via the on-board computer. The audio file format is a common format such as WAV or MP3.
[1127] When the cloud server receives the voice data, it uses a voice recognition engine to convert the voice data into text data. An external voice recognition service (e.g., Google Cloud Speech-to-Text API) is used as the voice recognition engine. To ensure high accuracy of the conversion, the voice recognition engine uses a cloud-based service.
[1128] The converted text data is then analyzed by the server. During the analysis process, natural language processing (NLP) techniques are used to extract frequently occurring keywords, tally the number of times a particular topic is mentioned, and perform sentiment analysis, which generates a positive, negative, or neutral sentiment score.
[1129] The analysis results are compiled into a report, which is displayed in a visual format such as graphs and charts for easy understanding by the user. The generated report can be viewed by the operation manager via a dedicated mobile app or web dashboard.
[1130] Hardware and software used
[1131] Hardware:
[1132] Microphone for collecting audio data
[1133] On-board computer
[1134] Cloud Server
[1135] software:
[1136] Google Cloud Speech-to-Text API (speech recognition engine)
[1137] Natural Language Processing (NLP) libraries (e.g., NLTK, spaCy)
[1138] Visualization tools (e.g., Matplotlib, D3.js)
[1139] Web interface (for dashboard display)
[1140] Examples of specific examples and prompts
[1141] For example, if a passenger in an autonomous vehicle says, "I like this route, the scenery is amazing!", the voice is collected by a microphone. The voice data is uploaded to a cloud server and converted into text data using the Google Cloud Speech-to-Text API. It is then analyzed using natural language processing technology to extract sentiment scores and frequently occurring keywords. Finally, the analysis results are visualized as graphs and charts on a dashboard and provided to the operation manager.
[1142] Example prompt sentence:
[1143] Please perform a sentiment analysis on the following Japanese text.
[1144] Text: I love this route, the views are amazing!
[1145] Calculate positive, negative, and neutral sentiment scores and explain why. Also extract frequently occurring keywords and the number of mentions of each theme.
[1146] This allows the system to collect feedback in real time and quickly improve services.
[1147] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1148] Step 1:
[1149] A microphone installed inside the autonomous vehicle collects passenger voice data in real time. The voice data is sent to the on-board computer in WAV or MP3 format. The input is the passenger's speech, and the output is an audio file.
[1150] Step 2:
[1151] The server receives audio data from the onboard computer and uploads it to the cloud server. The input is an audio file, and the output is audio data stored on the cloud. This process sends the audio data to the cloud in real time.
[1152] Step 3:
[1153] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text API) on a cloud server to convert the voice data into text data. The input is an audio file, and the output is text data. The speech recognition engine analyzes the voice waveform and converts it into text data.
[1154] Step 4:
[1155] The server analyzes the converted text data. It uses natural language processing libraries (e.g., NLTK, spaCy) to extract frequently occurring keywords, tally the number of mentions of specific themes, and perform sentiment analysis (generating a score for positive, negative, or neutral). The input is the text data, and the output is the analysis results. Specific pattern matching and sentiment scoring algorithms are applied.
[1156] Step 5:
[1157] The server generates a report based on the analysis results. It visualizes them as graphs and charts using a graphical visualization tool (e.g., Matplotlib, D3.js). The input is the analysis results, and the output is a visualized report. This process makes it easier for users to intuitively understand the data.
[1158] Step 6:
[1159] Users can view the generated reports using a dedicated mobile app or web dashboard. The input is the visualized report, and the output is new insights including user feedback. Users can check each analysis result on the dashboard and take any necessary actions to improve the service.
[1160] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1161] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important information and emotional information with the aim of improving user experience (UX) and customer experience (CX).
[1162] Audio data upload
[1163] Users collect interview audio data and upload it to the server using the file selection interface on their device. The audio data is formatted in common audio file formats such as WAV and MP3. The uploaded audio data is sent to the server and stored in storage.
[1164] Speech Recognition Processing
[1165] The server receives the voice data sent by the user and converts it into text data using a speech recognition engine. The speech recognition engine uses an external speech recognition service (for example, a cloud-based speech recognition API) to achieve highly accurate transcription.
[1166] Text analytics
[1167] The server parses the character data to extract the following information:
[1168] Extracting frequent keywords: Using a natural language processing (NLP) library, extract all tokens and calculate their frequency distribution.
[1169] Counting the number of mentions of a specific topic: Matching a list of keywords for a pre-defined topic and counting the number of mentions of that topic.
[1170] Sentiment Analysis and Sentiment Engine
[1171] The server includes an emotion engine that uses an emotion analysis library to recognize user emotions. The emotion engine calculates positive, negative, and neutral emotion scores for each part of a statement in the text data. The emotion scores are calculated in real time and can be viewed by users on a dashboard. The server also displays the fluctuations in the emotion scores in a time-series graph format, visualizing changes in emotion over time.
[1172] Reporting and Visualization
[1173] The server generates a report based on the analysis results, which includes a list of frequently occurring keywords, the number of mentions of specific themes, and sentiment scores. This information is displayed in a graphical format (e.g., graphs, charts, etc.) to help users understand it intuitively.
[1174] View results
[1175] Users access the server using their devices and view the generated reports in dashboard format. The dashboard is provided via a web interface, allowing users to easily check each analysis result. Fluctuations in emotion scores can also be tracked over time, allowing users to identify when emotional statements were made.
[1176] For example, if a user uploads five interview audio files, the server receives the audio data, transcribes each file, and analyzes it for frequently occurring keywords, the number of times specific themes are mentioned, and sentiment scores. The analysis results are then compiled into a report, and users can view this information in a visualized format through a dashboard. This process allows users to quickly extract useful insights and sentiment information from large amounts of interview data.
[1177] The processing flow will be explained below.
[1178] Step 1:
[1179] The user collects the interview audio data and uploads it to the server using the file selection interface on the device. The audio data format is a common audio file format such as WAV or MP3. The user selects the audio file to upload on the file selection screen and presses the send button.
[1180] Step 2:
[1181] The server receives the audio data sent by the user via an HTTP POST request, saves the received file in a temporary directory, and registers meta information (file name, upload date and time, user ID, etc.) in the database.
[1182] Step 3:
[1183] The server sends the received voice data to a voice recognition engine for text conversion. An external voice recognition service (e.g., a cloud-based voice recognition API) is used to convert the voice into text data, and this text data is then retrieved.
[1184] Step 4:
[1185] The server stores the acquired text data in a database, and when storing the text data, associates it with the identifier of the original audio file.
[1186] Step 5:
[1187] To analyze the text data, the server uses a natural language processing (NLP) library to extract all tokens and calculate their frequency distribution, creating a list of frequently occurring words to identify common keywords.
[1188] Step 6:
[1189] The server compares the text data with a list of keywords for predefined themes and tally the number of mentions related to each theme, allowing the frequency of mentions of specific topics of interest to users to be displayed in graph form.
[1190] Step 7:
[1191] The server uses a sentiment analysis library to calculate a sentiment score for each statement in the text, calculating positive, negative, or neutral sentiment scores to identify which parts are particularly emotional.
[1192] Step 8:
[1193] The server calculates the emotion score in real time and displays it on a dashboard, allowing users to check the emotion associated with each comment in real time.
[1194] Step 9:
[1195] The server tracks the fluctuations in the emotion scores over time and displays them in a graph, allowing users to visually grasp the emotional changes in the comments over time.
[1196] Step 10:
[1197] The server generates a report based on the analysis results and displays it in a graphical format (graphs, charts, etc.). Users can access the server using their devices and view this report from a dashboard, allowing them to quickly grasp important information and emotional information without having to relisten to all the interview audio data.
[1198] The above are the specific processing steps of the system based on the invention that combines the emotion engine.
[1199] Example 2
[1200] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1201] Current systems for improving user experience (UX) and customer experience (CX) face challenges in efficiently analyzing large amounts of interview audio data and quickly extracting important and emotional information. Furthermore, manual data analysis is time-consuming and labor-intensive, resulting in limited insights.
[1202] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: a means for a user to upload voice data using a file selection interface of a terminal; a means for converting the received voice data into text data using a voice recognition engine and storing the text data in a database; a means for analyzing the text data using a natural language processing library and extracting frequently occurring keywords, the number of mentions of a specific theme, and an emotion score; a means for aggregating the analysis results, compiling them in a report format, and visualizing them in a graphical format; and a means for a user to view the generated report on a dashboard via a web interface from their terminal. This makes it possible to quickly and efficiently obtain useful insights and emotion information from a large amount of interview voice data.
[1203] "User" means an end user who uses the System to upload audio data and view analysis results.
[1204] "Terminal" refers to the device that a user uses to upload audio data and view analysis results, such as a PC, tablet, or smartphone.
[1205] A "file selection interface" is a user interface that allows a user to select an audio file on a terminal.
[1206] A "server" is a remote computer system that receives voice data, converts it to text data, analyzes it, and generates reports.
[1207] "Speech recognition engine" means software or a service for analyzing voice data and converting it into text data, including cloud-based speech recognition services.
[1208] "Character data" refers to text-format data converted by a voice recognition engine.
[1209] A "natural language processing library" is a software library for analyzing text data to extract frequently occurring keywords and the number of times a particular topic is mentioned.
[1210] "Keywords" are important words or phrases that are set as the subject of analysis.
[1211] A "theme" is a specific topic or topic that is matched by a pre-defined list of keywords.
[1212] The "emotion score" classifies the emotions of statements contained in text data into three types: positive, negative, and neutral, and quantifies the degree of each.
[1213] A "report" is a document or data that compiles the analysis results and provides them to the user in a graphical format.
[1214] A "dashboard" is a web interface that allows users to view reports and visualizes the analysis results.
[1215] "Visualization" is the process of converting data and analytical results into a format that is easy to understand visually (e.g., graphs, charts, etc.).
[1216] The present invention relates to a system that efficiently analyzes a large amount of interview audio data and extracts important and emotional information for the purpose of improving user experience (UX) and customer experience (CX). This system automates the process in which users upload audio data and a server analyzes the data to extract useful insights.
[1217] Hardware and software used
[1218] This system uses the following hardware and software:
[1219] Server: Receives data, converts it to text data, analyzes it, and generates reports.
[1220] Terminal: Used by users to upload audio data and view analysis results.
[1221] Speech recognition engine: Converts voice data into text using a cloud-based speech recognition service (e.g., Google Speech-to-Text API, IBM Watson Speech to Text, etc.).
[1222] Natural language processing libraries: Analyze text data and extract frequent keywords or the number of mentions of specific topics (e.g., spaCy, NLTK, etc.).
[1223] Sentiment analysis libraries: Analyze the sentiment of statements in text data and calculate positive, negative, or neutral sentiment scores (e.g., VADER, TextBlob, AWS Comprehend, etc.).
[1224] Overall system processing flow
[1225] Users collect interview audio data and upload it to the server using a file selection interface on their device. The server converts the received audio data into text using a cloud-based speech recognition engine. The converted text is stored in a database and then analyzed using a natural language processing library.
[1226] The analysis includes the following elements:
[1227] Extraction of frequently occurring keywords
[1228] Counting the number of times a specific topic was mentioned
[1229] Calculating the sentiment score
[1230] This analytical information is aggregated and compiled into reports, which are displayed in graphical formats (e.g., graphs, charts, etc.), and users can view the generated reports in a dashboard via the device's web interface.
[1231] Specific operation example
[1232] For example, a user uploads an MP3 audio file called "User Interview_1.mp3" to the system. This audio file is sent to the server and stored in storage. After that, the speech recognition engine converts the speech into text data, and the generated text data is stored in the database.
[1233] Next, the natural language processing library extracts frequently occurring keywords from the text data, confirming that keywords such as "service improvement" and "customer satisfaction" appear frequently. Furthermore, the sentiment analysis library analyzes each utterance in the text data and finds that there are many positive comments. These analysis results are compiled into a report and displayed in graph format on a dashboard.
[1234] Prompt Sentence Examples
[1235] An example prompt for a generative AI model is:
[1236] "I would like to upload interview audio data, run it through speech recognition to transcribe it, and analyze it for frequent keywords and sentiment scores. Please let me know how this data will be used to improve the user experience."
[1237] As described above, this system provides consistent support from analyzing interview audio data to detecting changes in emotions, and is a mechanism that provides useful information to users.
[1238] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1239] Step 1:
[1240] The user uploads the audio data using the file selection interface on the device. The user collects the interview audio data and uploads it by selecting the audio file on the device. This upload procedure sends the audio file (e.g., WAV or MP3 format) to the server. The input is the audio file on the user's device, and the output is the audio file stored on the server.
[1241] Step 2:
[1242] The server checks the received audio data and saves it in storage. After receiving the audio file sent by the user, the server automatically saves it in a storage system (e.g., cloud storage service). This ensures that the data is stored safely on the server. The input is the audio file uploaded to the server, and the output is the audio file saved in the server's storage.
[1243] Step 3:
[1244] The server starts a speech recognition engine and converts the stored voice data into text data. The server uses a cloud-based speech recognition engine (e.g., Google Speech-to-Text API) and inputs the stored voice data. The speech recognition engine analyzes the voice and generates corresponding text data. The input is the voice data stored in storage, and the output is the text data generated by the speech recognition engine.
[1245] Step 4:
[1246] The server stores the generated text data in a database. The server records the text data output from the speech recognition engine in a database for subsequent analysis processes. The input is the text data generated by the speech recognition engine, and the output is the text data stored in the database.
[1247] Step 5:
[1248] The server invokes a natural language processing library to analyze the text data. The server uses a natural language processing (NLP) library (e.g., spaCy) to input the text data stored in the database. The NLP library performs text analysis to extract frequently occurring keywords and the number of times a particular topic is mentioned. The input is the text data stored in the database, and the output is a list of frequently occurring keywords and the number of times a particular topic is mentioned.
[1249] Step 6:
[1250] The server launches a sentiment analysis library and calculates the sentiment score in the text data. The server uses a sentiment analysis library (e.g., VADER) to calculate the sentiment score for each utterance in the text data. The sentiment score is expressed in three indicators: positive, negative, and neutral. The input is the text data to be analyzed, and the output is a list of sentiment scores.
[1251] Step 7:
[1252] The server aggregates the analysis results and compiles them into a report. The server aggregates the data of the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and generates a report. This report is visualized in a graphical format (e.g., graph, chart, etc.). The input is the list of frequently occurring keywords, the number of times a specific theme is mentioned, and the sentiment score, and the output is a graphical report.
[1253] Step 8:
[1254] The user views the generated report on a dashboard via a web interface from their device. The user uses a browser on their device to access the server's web interface and view the generated report. The dashboard displays a list of frequently occurring keywords, the number of times a specific theme is mentioned, and fluctuations in sentiment scores. The input is the generated report, and the output is the analysis results on the dashboard that the user views.
[1255] (Application example 2)
[1256] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1257] Conventional advertising distribution systems have the problem that it is difficult to select optimal advertising information based on user emotions or specific keywords, which makes it difficult to attract user attention and deliver advertisements effectively.
[1258] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading voice data, means for converting the voice data into character data using a voice recognition engine, means for analyzing the character data to extract frequently occurring keywords, the number of mentions of a specific theme, and an emotion score, and means for selecting optimal advertising information based on the emotion score and keywords and displaying it on the user terminal. This makes it possible to deliver optimal advertisements that match the user's emotions in a timely manner.
[1259] "Audio data" refers to audio files that users upload to the server, and are in common audio file formats such as WAV and MP3.
[1260] "Character data" refers to text data converted from voice data using a voice recognition engine.
[1261] A "voice recognition engine" is a program or service that converts voice data into text data, and includes external voice recognition services.
[1262] "Frequent keywords" are words or phrases that appear frequently in the analyzed text data.
[1263] The "number of times a specific theme is mentioned" refers to the number of times that the theme is mentioned, as determined by matching the text data with a keyword list of the theme that is set in advance.
[1264] "Emotion score" refers to the evaluation value of positive, negative, and neutral emotions calculated for statements in text data using an emotion analysis library.
[1265] "Advertising information" refers to advertisements and promotional information selected based on the user's emotional score and keywords.
[1266] A "user terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[1267] "Visualization" refers to displaying analysis results in a graphical format that users can intuitively understand.
[1268] This invention relates to a system that analyzes user voice data, selects optimal advertising information based on emotion scores and keywords, and displays it on the user's terminal. The specific system configuration and processing procedure for implementing this invention are described below.
[1269] System Configuration
[1270] 1. User device: Collect voice data using devices such as smartphones, tablets, and PCs and upload it to the server.
[1271] 2. Server: A server located on the cloud that processes data using a speech recognition engine, natural language processing library, etc.
[1272] 3. Speech recognition engine: A program for converting voice data into text data. It uses an external cloud-based speech recognition service (e.g., Google Speech-to-Text API).
[1273] 4. Natural language processing libraries: Libraries for extracting frequent keywords and the number of mentions of specific topics from text data (e.g., spaCy, NLTK).
[1274] 5. Sentiment analysis library: A library for calculating sentiment scores from text data (e.g., VADER, Google Natural Language API).
[1275] 6. Advertising platform: Selects and delivers optimal advertising information based on sentiment scores and keywords (e.g., Google AdSense).
[1276] Program processing
[1277] The server does the following:
[1278] 1. Audio data collection and upload:
[1279] A user records audio data using a smartphone application.
[1280] The common formats used are WAV and MP3.
[1281] The audio data is uploaded from the application to a cloud server.
[1282] 2. Speech Recognition and Transcription:
[1283] A speech recognition engine (e.g., Google Speech-to-Text API) on a cloud server receives the uploaded voice data and converts it into text data.
[1284] 3. Text Analysis:
[1285] The text data is analyzed on the server using a natural language processing library (e.g., spaCy or NLTK).
[1286] Frequently occurring keywords and the number of mentions of specific topics are extracted.
[1287] 4. Emotion analysis:
[1288] Use a sentiment analysis library (e.g., VADER or Google Natural Language API) to calculate a sentiment score for each piece of text data.
[1289] Sentiment scores are expressed as either positive, negative, or neutral.
[1290] 5. Advertisement Delivery:
[1291] Based on the sentiment score and keywords, the advertising platform (e.g., Google AdSense) selects the most suitable advertising information.
[1292] The selected advertising information is displayed on the user terminal.
[1293] Specific examples
[1294] For example, if a user uploads audio data in which they say, "I want to take a weekend trip," the audio data is converted into text data and keywords such as "trip" and "weekend" are extracted. If the sentiment analysis determines that the comment has a positive sentiment, travel-related advertisements (e.g., hotel discounts or airline ticket promotions) are displayed on the user's device.
[1295] Prompt Sentence Examples
[1296] Analyze the user's spoken voice data and calculate a sentiment score of positive, negative, or neutral. Based on the results, generate prompts that suggest the ad format and content to display the optimal ad.
[1297] This system allows the most appropriate advertisements to be displayed based on the user's emotional state, maximizing the effectiveness of the advertisements.
[1298] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1299] Step 1:
[1300] A user uses a smartphone application to record audio data. This audio data is saved in WAV or MP3 format. The user then uploads the audio data to a cloud server using an interface within the app. The input is the user's audio data, and the output is an audio file saved on the cloud server.
[1301] Step 2:
[1302] The cloud server sends the received voice data to a voice recognition engine. The server converts the voice data into text data using an external cloud voice recognition service (e.g., Google Speech-to-Text API). The input here is an audio file, and the output is text data converted from the voice data.
[1303] Step 3:
[1304] The server retrieves the text data and begins analyzing it using a natural language processing library (e.g., spaCy or NLTK). Specifically, it extracts frequently occurring keywords from the text data and calculates the number of mentions of a particular topic. The input for this step is the text data, and the output is a list of keywords and a summary of topic mentions.
[1305] Step 4:
[1306] The server uses a sentiment analysis library (e.g., VADER or Google Natural Language API) to analyze each part of the text data and generate a positive, negative, or neutral sentiment score. The sentiment score is quantified according to the intensity of the emotion. The input here is the text data, and the output is a score corresponding to each emotion.
[1307] Step 5:
[1308] The server selects the optimal advertising information using an advertising platform (e.g., Google AdSense) based on the emotion score and keywords. This selection is performed using prompts generated by a generative AI model. An example of such a prompt is, "Analyze the user's spoken voice data and calculate positive, negative, and neutral emotion scores. Based on the calculation results, generate a prompt that provides the ad format and content to display the optimal advertisement." The input is the emotion score and a keyword list, and the output is the optimal advertising information.
[1309] Step 6:
[1310] The selected advertising information is sent from the server to the user terminal. The user terminal displays the received advertising information on a display so that the user can actually see it. The input is the optimal advertising information, and the output is the advertisement displayed on the user terminal.
[1311] The above are the specific processing steps of the system that realizes the application example. This series of processing provides optimal advertising information based on the user's emotions.
[1312] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1313] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1314] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1315] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1316] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1317] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1318] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1319] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1320] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1321] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1322] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1323] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1324] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1325] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1326] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1327] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1328] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1329] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1330] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1331] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1332] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1333] The following is further disclosed regarding the above embodiment.
[1334] (Claim 1)
[1335] [A means for users to upload audio data;
[1336] [Means for converting the received voice data into text data using a voice recognition engine;
[1337] [Means for the server to analyze the text data and extract frequently occurring keywords, the number of times a particular topic is mentioned, and a sentiment score;
[1338] [The server compiles the analysis results into a report and visualizes them.
[1339] [A means for users to view the generated reports; and
[1340] A system including:
[1341] (Claim 2)
[1342] [The system according to claim 1, wherein an external speech recognition service is used as the speech recognition engine.
[1343] (Claim 3)
[1344] The system of claim 1, wherein the sentiment analysis generates positive, negative, and neutral sentiment scores.
[1345] "Example 1"
[1346] (Claim 1)
[1347] [A means for users to upload audio data on their devices;
[1348] [Means for converting the received voice data into text data using a cloud-based voice recognition API;
[1349] [The server analyzes the text data using a natural language processing library to extract frequently occurring keywords, the number of mentions of specific topics, and sentiment scores;
[1350] [The server compiles the analysis results into a report and visualizes them in a graphical format.
[1351] [means for a user to view the generated report using a terminal; and
[1352] A system including:
[1353] (Claim 2)
[1354] [The system of claim 1, wherein the speech recognition engine uses an external speech recognition service.
[1355] (Claim 3)
[1356] The sentiment analysis system of claim 1, which generates positive, negative, and neutral sentiment scores.
[1357] "Application Example 1"
[1358] (Claim 1)
[1359] [A means for users to upload audio data;
[1360] [Means for converting the received voice data into text data using a voice recognition engine;
[1361] [Means for the server to analyze the text data and extract frequently occurring keywords, the number of times a particular topic is mentioned, and a sentiment score;
[1362] [The server compiles the analysis results into a report and visualizes them.
[1363] [A means for users to view the generated reports; and
[1364] [Means for collecting voice data in real time;
[1365] [Means for uploading audio data to a cloud server;
[1366] [Means of collecting passenger feedback in real time and quickly identifying operational and service improvements; and
[1367] A system including:
[1368] (Claim 2)
[1369] [The system according to claim 1, wherein an external speech recognition service is used as the speech recognition engine.
[1370] (Claim 3)
[1371] The system of claim 1, wherein the sentiment analysis generates positive, negative, and neutral sentiment scores.
[1372] "Example 2: Combining Emotion Engines"
[1373] (Claim 1)
[1374] [A means for a user to upload audio data using a file selection interface on the device; and
[1375] [Means for converting the voice data received by the server into text data using a voice recognition engine and storing the data in a database;
[1376] [Means for the server to analyze the text data using a natural language processing library and extract frequently occurring keywords, the number of times a specific theme is mentioned, and a sentiment score;
[1377] [The server aggregates the analysis results, compiles them into a report, and visualizes them in a graphical format.
[1378] [A means for users to view the generated reports on a dashboard via a web interface from their device; and
[1379] A system including:
[1380] (Claim 2)
[1381] [The system of claim 1, which uses a cloud-based speech recognition service as the speech recognition engine.
[1382] (Claim 3)
[1383] The system of claim 1, wherein a sentiment analysis library is used to generate positive, negative, and neutral sentiment scores and visualize them in a time-series graph format.
[1384] "Application example 2 when combining emotion engines"
[1385] (Claim 1)
[1386] [A means for users to upload audio data;
[1387] [Means for converting the received voice data into text data using a voice recognition engine;
[1388] [Means for the server to analyze the text data and extract frequently occurring keywords, the number of times a particular topic is mentioned, and a sentiment score;
[1389] [The server compiles the analysis results into a report and visualizes them.
[1390] [A means for users to view the generated reports; and
[1391] [Means for selecting optimal advertising information based on emotion scores and keywords and displaying it on the user's device;
[1392] A system including:
[1393] (Claim 2)
[1394] [The system according to claim 1, wherein an external speech recognition service is used as the speech recognition engine.
[1395] (Claim 3)
[1396] The system of claim 1, wherein the sentiment analysis generates positive, negative, and neutral sentiment scores. [Explanation of symbols]
[1397] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for a user to upload audio data; A means for converting the voice data received by the server into character data using a voice recognition engine; A server analyzes the text data and extracts frequently occurring keywords, the number of times a particular topic is mentioned, and a sentiment score; The server compiles the analysis results into a report and visualizes them. a means for a user to view the generated report; A system including:
2. 2. The system according to claim 1, wherein an external voice recognition service is used as the voice recognition engine.
3. The system of claim 1 , wherein the sentiment analysis generates positive, negative, and neutral sentiment scores.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A