System
The system automates the conversion of meeting audio into concise, visually understandable summaries, addressing the inefficiencies of existing methods by providing intuitive graphical displays.
Patent Information
- Application Number
- JP2024123832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Existing systems struggle to efficiently convert speech from meetings into concise, visually understandable summaries, requiring skilled experts and are not universally applicable.
A system that includes a user terminal, server, and storage device for uploading audio, performing speech recognition, summarization, and generating graphical displays to summarize and visualize meeting content.
Enables users to quickly grasp meeting content, facilitating decision-making and next actions with efficient, intuitive visualizations.
Smart Images

Figure 2026022315000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When the content of a meeting is diverse, it can be difficult for participants to understand the content, especially if there are a lack of visual elements. To solve this problem, there is a technique called "graphic recording" that visualizes the content of meetings, lectures, etc. in real time. However, this method requires skilled experts, and it is difficult to use in all meetings. Therefore, there is a need for automated technology to efficiently understand the content of meetings. [Means for solving the problem]
[0005] This invention provides a system including a means for a user to upload an audio file, a server to receive and store the audio file, a server to generate text from the audio file using speech recognition, a system to summarize the generated text, a system to graphically display the summary, and a system to provide the user with the graphically displayed summary. This system automatically summarizes and visualizes the contents of a meeting, allowing participants to easily understand the contents and quickly make decisions or take the next action.
[0006] Below are definitions of important terms included in the claims.
[0007] An "audio file" is a digital or analog file that contains recorded audio data.
[0008] "Upload" is an operation in which a user sends data from their own terminal to a server.
[0009] "Receiving" is the operation of the server taking in data sent over the network.
[0010] "Storage" means recording the received data in a designated storage device.
[0011] "Speech recognition" is a technology that extracts linguistic information from voice data and converts it into text.
[0012] "Text generation" refers to the process of expressing linguistic information extracted by speech recognition as a string of characters.
[0013] "Summarization" is the process of compacting the original text information and creating a document that extracts only the important parts.
[0014] "Graphical display" refers to the use of charts, illustrations, diagrams, and other visual elements to visually represent information.
[0015] "Providing" refers to making the generated graphical display available to a user.
[0016] A "system" is a whole in which multiple components work together to achieve a set of functions. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] To implement this invention, a system consisting of three main components is utilized: a user terminal, a server, and a storage device. The present invention involves a process in which a user uploads an audio file, the server recognizes and summarizes the content, and then displays the summary graphically.
[0039] System Overview
[0040] 1. Upload your audio file
[0041] Users select audio files from their devices and upload them to the server, typically via a web interface.
[0042] The server receives the audio file and stores it on a storage device.
[0043] 2. Speech Recognition and Summarization
[0044] The server loads the audio file stored on the storage device and performs speech recognition to convert it into text, preferably using a high-precision speech recognition library.
[0045] The converted text is input into a summary generation model to generate a summary, which uses a natural language processing algorithm.
[0046] 3. Generating the Graphical Display
[0047] The server generates a graphical display from the summarized text, which can be in the form of diagrams and text boxes in a visually friendly format.
[0048] The graphical representation is saved as an image file and presented to the user.
[0049] Program processing overview
[0050] Uploading and saving audio files
[0051] The user uses the device's web interface to select and upload audio files to the server, which receives the uploaded audio files and stores them on a designated storage device.
[0052] Speech Recognition and Text Generation
[0053] The server loads the audio file from the storage device and uses a speech recognition library to convert the speech into text, which reflects the content of the meeting.
[0054] Generate a summary
[0055] The generated text is then summarized using a summary generation algorithm, which reconstructs long texts into short, meaningful summaries.
[0056] Generating a Graphical Display
[0057] The server creates a graphical display from the summarized text using visually friendly formats (e.g., text boxes and diagrams) and stores the generated image files.
[0058] Providing results
[0059] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files using a terminal.
[0060] Specific examples
[0061] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and stores this file, it performs speech recognition and converts it into text. Suppose the converted text is "Today's meeting reviewed project progress and assigned new tasks." This text is summarized to generate a summary: "Review progress and assigned new tasks." A graphical representation is created based on this summary and provided to the user as an image file.
[0062] This system allows users to easily and quickly grasp the important points of a meeting, enabling them to smoothly proceed with decision-making and next actions.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0066] Step 2:
[0067] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0068] Step 3:
[0069] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0070] Step 4:
[0071] The converted text is input to a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is retrieved.
[0072] Step 5:
[0073] Configure the server to create a graphical display of the generated summary text, visualizing the summary content in a visually understandable format (e.g., text boxes).
[0074] Step 6:
[0075] The server saves the graphical representation as an image file (e.g. "summary_visual.png") and retrieves the path of the saved image file.
[0076] Step 7:
[0077] The server generates a response including a path to the image file for providing the generated image file of the graphical representation to the user, and sends the response to the user terminal.
[0078] Step 8:
[0079] The user receives the response from the server on the device, and uses the image file path included in the response to download and view the generated graphical display.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] In today's business environment, it is important to efficiently process, summarize, and share audio recordings of meetings and conferences. However, conventional methods require converting speech to text, summarizing it, and providing it in a visually understandable format, which is time-consuming and inefficient. Furthermore, there is a lack of systems that can consistently perform highly accurate speech recognition, generate summaries, and provide intuitively understandable visualizations, making it difficult for users to easily grasp the content of meetings.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for a user to upload voice data, means for the server to receive and store the voice data, means for the server to generate character data from the voice data using voice recognition technology, means for summarizing the generated character data using a summary generation algorithm, means for generating a graphical representation of the summarized character data, and means for providing the summarized graphical representation to the user. This makes it possible to efficiently generate a summary from voice data and further provide the summary in an intuitively easy-to-understand form.
[0085] "User" means a person or entity that uses the system to upload audio data and is responsible for receiving the results.
[0086] "Audio data" refers to digital files in audio format that users upload to the system, including records of meetings, conferences, etc.
[0087] A "server" is a computer system that receives, stores, analyzes, and provides results from audio data.
[0088] "Speech recognition technology" is a technology for converting voice data into text data, analyzing the content of the voice, and outputting it in text format.
[0089] "Text data" is text-format data generated by speech recognition technology, and is a recording of the contents of speech.
[0090] A "summary generation algorithm" is an algorithm for shortening long text data and extracting only the main points to create a concise summary.
[0091] "Graphical display" refers to converting summarized text data into a visually understandable format and displaying it in the form of charts, tables, text boxes, etc.
[0092] A "generative AI model" is an artificial intelligence model used to perform complex tasks such as natural language processing and summary generation.
[0093] A "prompt" is an instruction to be input to a generative AI model, which serves as a guideline for the model to generate appropriate output.
[0094] "Storage Device" means a storage device used by the server to store audio data and other data.
[0095] This invention is implemented using a system that mainly consists of a user terminal, a server, and a storage device. The system shows a series of processes in which a user uploads voice data, the server performs speech recognition and summarization of the data, and provides a graphical display of the results.
[0096] Uploading audio data
[0097] Users select audio data from their devices and upload it to the server via a web interface. The server receives the data and stores it on a designated storage device. The storage device can be a standard cloud storage device (e.g., an Amazon S3 bucket).
[0098] Speech Recognition and Text Generation
[0099] The server loads the saved audio data and converts it into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API). This text data is a recorded text of the meeting or discussion.
[0100] Text summary generation
[0101] The server inputs the generated text data into a natural language processing model (e.g., Hugging Face's BART or T5 model) and runs a summary generation algorithm. The generated text data is shortened to generate a summary that includes only the key points.
[0102] Generating a Graphical Display
[0103] The server creates a visually easy-to-understand graphical representation based on the summary content. This representation uses a chart generation library (e.g., D3.js or Google Chart API) to display the summary in an easy-to-understand format, such as a text box or diagram. The generated graphical representation is saved to the storage device as an image file (e.g., PNG format).
[0104] Results provided and downloaded
[0105] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files through a web interface using their terminal.
[0106] Specific examples
[0107] For example, consider a case where a user uploads a meeting recording file called "meeting_audio.wav." This file is received by the server and saved to a storage device. The server then converts this audio file into text using a speech recognition library, resulting in the text "Today's meeting reviewed project progress and assigned new tasks." This text is then summarized using a natural language processing model to generate a summary: "Reviewed progress and assigned new tasks." A visually friendly graphical representation is then generated from this summary and served to the user as an image file.
[0108] Prompt Sentence Examples
[0109] A user has uploaded a meeting recording file, "meeting_audio.wav." Convert the contents of this audio file into text and generate a summary. Then, display the summary graphically.
[0110] Audio: Today's meeting reviewed project progress and assigned new tasks.
[0111] Summary: Check progress and assign new tasks
[0112] The generated image file of the graphical display will be provided to the user, so please name the file "summary_graphic.png".
[0113] This invention allows users to quickly and efficiently grasp important information about meetings and conferences, which can lead to decision-making and next actions.
[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0115] Step 1: Select and upload audio data
[0116] The user accesses the web interface from their device and selects the audio data (e.g., "meeting_audio.wav").
[0117] Input: Audio data file
[0118] When the user clicks the "Upload" button, the terminal sends the audio data to the server.
[0119] Output: Audio data file sent to the server
[0120] Step 2: Receiving and storing audio data
[0121] The server receives the voice data sent from the user's terminal.
[0122] Input: Audio data file sent from the user's device
[0123] The server stores the received audio data in a designated storage device.
[0124] Output: An audio data file saved on your storage device (e.g. "meeting_audio.wav")
[0125] Step 3: Loading audio data and recognizing it
[0126] The server loads the audio data from the storage device where it is stored.
[0127] Input: Audio data file saved on a storage device
[0128] The server converts the audio data into text data using a speech recognition library such as the Google Speech-to-Text API.
[0129] Output: Converted text data (e.g. "In today's meeting, project progress was reviewed and new tasks were assigned.")
[0130] Step 4: Text summary generation
[0131] The server inputs the generated text data into a natural language processing library (e.g., Hugging Face's BART model).
[0132] Input: Generated character data
[0133] The server executes a summary generation algorithm to summarize the text data.
[0134] Output: Summary text data (e.g. "Check progress and assign new tasks")
[0135] Step 5: Generate the graphical display
[0136] The server creates a visually easy-to-understand graphical display based on the summarized character data.
[0137] Input: summarized character data
[0138] The server uses a chart generation library (e.g. D3.js or Google Chart API) to display the data as text boxes or diagrams.
[0139] Output: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0140] Step 6: Submit and download results
[0141] The server provides the generated image file of the graphical representation to the user.
[0142] Input: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0143] A user uses a terminal to download and view image files from a web interface.
[0144] Output: Image file downloaded to the user
[0145] (Application example 1)
[0146] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0147] To improve the efficiency and accuracy of communication between workers and robots in factories, a means is needed to quickly understand voice instructions and transmit them to the robot as accurate work instructions. It is also necessary to improve work efficiency by visually monitoring progress. Currently, there is a lack of systems that meet these requirements, making it difficult to provide an efficient work environment.
[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0149] In this invention, the server includes means for a user to upload an audio file, means for the server to receive and store the audio file, means for the server to generate text from the audio file using speech recognition, means for summarizing the generated text, means for graphically displaying the summary, means for providing the summarized graphical display to the user, means for extracting instructions from the audio file and transmitting them to the robot as work instructions, and means for providing a dashboard for visually monitoring the instructions. This enables the voice instructions of workers in a factory to be efficiently and accurately transmitted to robots, and progress to be monitored in real time.
[0150] An "audio file" is digital data in which a user has recorded their voice.
[0151] "Upload" is the act of a user sending data from their own terminal to a server.
[0152] A "server" is a computer system that provides various services over a network.
[0153] "Speech recognition" is a technology that analyzes voice data and converts the content into text.
[0154] "Text" is a digital string of characters generated by speech recognition.
[0155] A "summary" refers to text data that has been shortened and only important information has been extracted.
[0156] A "graphical display" is a digital representation that includes pictures and text boxes to visually display the summarized text.
[0157] "User" refers to the entity that uses the system.
[0158] A "work instruction" is a command that indicates the specific actions or tasks that a robot should perform.
[0159] A "robot" is a mechanical device that operates autonomously or semi-autonomously.
[0160] A "dashboard" is an interface that visually displays system information and monitors progress and key indicators.
[0161] To implement the present invention, the following system configuration and algorithm are used: The system mainly consists of a user terminal, a server, and a storage device.
[0162] Detailed system configuration
[0163] 1. Upload and save audio files
[0164] Users select and upload audio files to the server using a web interface or a dedicated smartphone application. The audio files are stored as digital data and received and saved by the server. This allows the user's work instructions to be stored on the server as audio files.
[0165] 2. Speech Recognition and Text Generation
[0166] The server loads the received audio file from the storage device and converts the audio data into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API or IBM Watson Speech to Text). In this step, the instructions in the audio file are extracted as text data.
[0167] 3. Summary Generation
[0168] The generated text is then summarized using a natural language processing model (e.g., BART or T5) using Hugging Face's transformers library. The summarization model shortens the text data and outputs a shortened version that extracts only the important information. This process eliminates redundant information and summarizes the key instructions.
[0169] 4. Generating the Graphical Display
[0170] The server uses a diagram generation library such as matplotlib to create a graphical representation of the summarized text. The summary is displayed in a visually understandable format and generated as an image file. This image file allows the user to intuitively understand the work instructions.
[0171] 5. Transferring work instructions to robots and providing dashboards
[0172] The instructions in the audio file are extracted as text, a summary is generated, and the text is sent to the robot as work instructions. A dashboard is also provided for visually monitoring the instructions and progress. This dashboard is used to check the progress of work in real time, enabling efficient work management within the factory.
[0173] Specific examples
[0174] For example, a user might use their smartphone to voice instructions such as, "After completing the bearing replacement on Line 1, please adjust the machine on Line 2." This voice file is then uploaded to the server, which converts the voice data into text and further summarizes this text to "Line 1: Bearing replacement completed, Line 2: Machine adjustment." The summarized text is then displayed graphically and provided as an image file. The robot receives the summarized work instructions, and real-time progress is displayed on a dashboard.
[0175] Example prompt sentence:
[0176] "Please input operator voice instructions: After completing the bearing replacement on Line 1, please adjust the machine on Line 2."
[0177] In this way, the present invention can significantly improve work efficiency within a factory.
[0178] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0179] Step 1:
[0180] The user uploads the audio file via their smartphone or web interface. Specifically, the user launches the application, selects the recorded audio file (e.g., "Work Instructions_audio.wav"), and presses the "Upload" button. At this time, the audio file is sent to the server.
[0181] Input: Audio file
[0182] Output: Audio file uploaded to the server
[0183] Step 2:
[0184] The server receives the audio file and saves it to a storage device. The server receives the audio file through an HTTP request and saves it in a specified folder in the file system for further processing.
[0185] Input: Uploaded audio file
[0186] Output: Audio file saved on your storage device
[0187] Step 3:
[0188] The server loads the audio file from the storage device and converts the audio into text using a highly accurate speech recognition library. The server then reads the audio file and converts the audio data into text using the Google Speech-to-Text API or IBM Watson Speech to Text. During this process, the instructions in the audio are extracted as a string.
[0189] Input: Audio file saved on your storage device
[0190] Output: Converted text data
[0191] Step 4:
[0192] The server summarizes the generated text using a natural language processing model. The server uses Hugging Face's transformers library to preprocess the text data before inputting it into the summarization model. Summarization is performed using BART or T5 models, and important instructions are extracted without redundant information.
[0193] Input: Converted text data
[0194] Output: Summarized text data
[0195] Step 5:
[0196] The server creates a graphical display to visualize the summarized text. The server uses a chart generation library such as matplotlib to create an image to graphically represent the summarized text. This image is then provided to a dashboard or other display medium.
[0197] Input: Summarized text data
[0198] Output: Image file of the graphical display
[0199] Step 6:
[0200] The server sends the summarized instructions to the robot as work instructions, and the server sends messages to the robot via the network, and the robot performs the work based on these instructions.
[0201] Input: Summarized text data
[0202] Output: Instruction message to the robot
[0203] Step 7:
[0204] The server provides a dashboard for visually monitoring the instructions and progress of work. The server updates data in real time through the dashboard application, allowing users to monitor the progress. This allows users to check the progress of work at a glance.
[0205] Input: Feedback data from the robot
[0206] Output: Real-time updated dashboard
[0207] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0208] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0209] System Overview
[0210] 1. Upload and save audio files:
[0211] The user selects an audio file from their device and uploads it to the server. The server receives the audio file sent by the user and saves it in the specified storage.
[0212] 2. Speech Recognition and Text Generation:
[0213] The server loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library.
[0214] 3. Emotion recognition:
[0215] Apart from the text data recognized from the audio file, the server also analyzes the user's emotions using an emotion engine, which uses multiple audio features such as intonation, stress, and speed as input.
[0216] 4. Summary generation:
[0217] The converted text is then fed into a summary generation algorithm to generate a summary, which uses a natural language processing model to extract only the important parts.
[0218] 5. Emotion-based regulation:
[0219] After the summary is generated, the server adjusts the summary content and graphical display format based on the perceived user emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[0220] 6. Generate the graphical display:
[0221] The server creates a graphical display based on the adjusted summary, using text boxes and diagrams for visual comprehension, and saves this display as an image file.
[0222] 7. Providing Results:
[0223] The server provides the user with a graphical representation stored as an image file, which the user can download and view through their terminal.
[0224] Program processing overview
[0225] Uploading and saving audio files
[0226] The user selects an audio file from the device's web interface and uploads it to the server, which then stores the received audio file in its storage.
[0227] Speech Recognition and Text Generation
[0228] The server loads the saved audio files and performs speech recognition to generate text that contains the full content of the meeting.
[0229] emotion recognition
[0230] The server uses an emotion engine to recognize the user's emotions from the audio file. The emotion engine analyzes the voice characteristics and infers the user's emotional state, for example, stress or satisfaction based on changes in speech rate and volume.
[0231] Summary Generation
[0232] The generated text is then summarized using a natural language processing model, which turns long texts into short, important summaries.
[0233] Emotion-Based Adjustment
[0234] The server adjusts the summary content based on the user's recognized emotions, for example, by making the summary more concise if the user is in a hurry.
[0235] Generating a Graphical Display
[0236] The server creates a graphical display from the tailored summary, which may include text boxes and diagrams to present the information in a visually easy-to-understand format.
[0237] Providing results
[0238] Finally, the server saves the generated graphical display as an image file and provides the path to the file to the user, who can then download the image file to their terminal and view its contents.
[0239] Specific examples
[0240] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0241] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0242] The processing flow will be explained below.
[0243] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The purpose of this invention is to provide an efficient understanding of meetings and information that adapts to the user's emotions.
[0244] Processing step details
[0245] Step 1:
[0246] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0247] Step 2:
[0248] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0249] Step 3:
[0250] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0251] Step 4:
[0252] The server analyzes the tone, speed, volume, etc. of the voice in the audio file and uses an emotion engine to recognize the user's emotion. This recognition uses an emotion analysis algorithm to evaluate the user's emotional state (e.g., stress, joy, fatigue, etc.).
[0253] Step 5:
[0254] The server inputs the converted text into a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is then retrieved.
[0255] Step 6:
[0256] The server adjusts the summary content and the format of the graphical display based on the user's perceived emotions, for example, if the user is tired, the summary text will be more concise and displayed in a format that is easy to understand visually.
[0257] Step 7:
[0258] The server then creates a graphical display based on the emotion-adjusted summary, including text boxes and diagrams that present the information in a visually accessible format.
[0259] Step 8:
[0260] The server saves the generated graphical representation image file (e.g. "summary_visual.png") and obtains its path.
[0261] Step 9:
[0262] The server sends a response containing the path of the image file of the generated graphical representation to the user terminal, and the user receives the response and uses the path of the image file to download and view the graphical representation.
[0263] Specific examples
[0264] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0265] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0266] Example 2
[0267] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0268] In conventional meeting recording systems, simply converting audio data into text and summarizing it was difficult to fully support users' understanding of the meeting content. Furthermore, since information was not provided according to the user's emotional state, there were issues with differences in how information was received and the level of understanding.
[0269] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving and saving an audio file, means for generating text from the audio file using speech recognition technology, means for summarizing the generated text using a natural language processing model, means for using an emotion recognition engine to analyze the user's emotional state, means for adjusting the summary content based on the recognized user emotion, means for graphically displaying the summary content in a visually easy-to-understand format, and means for providing the generated graphical display content to the user. This allows the user to grasp the meeting content efficiently and in an easy-to-understand format, and makes it possible to provide information adapted to the user's emotions.
[0270] An "audio file" is a digital file containing audio data that can be uploaded by a user.
[0271] A "server" is a computer system that stores received audio files and performs speech recognition, natural language processing, emotion recognition, and graphical display.
[0272] "Voice recognition technology" is a technology for analyzing voice data from an audio file and converting it into text data.
[0273] A "natural language processing model" is an algorithm or machine learning model for analyzing generated text and performing summarization or other language processing.
[0274] An "emotion recognition engine" is a technology for analyzing and estimating a user's emotional state from voice data.
[0275] A "summary" is a text that extracts important information from the original text data and summarizes it in a short, concise format.
[0276] A "graphical display" is a display format that includes graphic elements to display text or data in a visually understandable format.
[0277] A "visualization library" is a program library for displaying text and data graphically.
[0278] This invention relates to a system that allows users to upload audio files to a server and then performs speech recognition, emotion recognition, summary generation, and graphical display of the content. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0279] First, the user selects an audio file through the device's web interface and uploads it to the server. The uploaded audio file is received by the server and saved in the specified storage. For example, let's assume that the user uploads an audio file called "meeting_audio.wav."
[0280] The server then loads the saved audio file and converts it into text using speech recognition technology, such as a reliable speech recognition library like the Google Speech-to-Text API. The audio file, "In today's meeting, project progress was reviewed and new tasks were assigned," is converted into text using this technology.
[0281] Furthermore, the server uses an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the audio data extracted from the audio file and recognize the user's emotions. The emotion recognition engine analyzes speech characteristics such as speed, stress, and intonation to determine the user's emotional state. Through this process, it can be recognized that the user is tired, for example.
[0282] The generated text is then fed into a natural language processing model (e.g., the OpenAI GPT-3 model) to generate a summary. The generated text, "In today's meeting, project progress was reviewed and new tasks were assigned," is summarized by the summary generation algorithm as "Progress review and new tasks assigned."
[0283] Based on the user's emotions, the server adjusts the summary. For example, if the server detects that the user is tired, it will make the summary more concise and easier to understand. The adjusted summary is then presented as "progress and new tasks."
[0284] Based on the tailored summary, the server creates a graphical display, using text boxes and diagrams to arrange the information in a visually understandable way, and saves the display as an image file.
[0285] Finally, the server provides the generated graphical display to the user, who can download the image file through his / her terminal and visually confirm the important points of the meeting.
[0286] Specific examples
[0287] For example, if a user uploads a meeting recording file called "meeting_audio.wav," the system works as follows: First, the server receives the audio file and saves it in storage. Then, the Google Speech-to-Text API is used to convert the audio into text, generating the following: "In today's meeting, project progress was reviewed and new tasks were assigned." At the same time, an emotion recognition engine recognizes the user's fatigue, and a summary generation algorithm summarizes the text as "Progress review and new tasks assigned." Finally, the summary is adjusted and presented to the user as a graphical representation of "progress and new tasks."
[0288] Prompt Sentence Examples
[0289] "Perform speech recognition on the meeting recording file 'meeting_audio.wav' and generate a graphical representation of the results of the summary and sentiment analysis."
[0290] This system allows users to grasp the contents of a meeting efficiently and in an easy-to-understand format, and provides information that adapts to the user's emotions.
[0291] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0292] Step 1: Upload and save your audio file
[0293] Specific operation: The user selects an audio file using their own device and uploads it to the server via a web interface. Specifically, the user opens a browser and accesses the system's upload page. They click the file selection button, select "meeting_audio.wav" from the device's file system, and click the upload button to send the audio file to the server. The server receives the file and saves it in a specific directory.
[0294] Input: An audio file selected and uploaded by the user (e.g. "meeting_audio.wav").
[0295] Output: Audio file saved on the server.
[0296] Step 2: Speech recognition and text generation
[0297] Specific operation: The server loads the audio file from storage. Then, it calls a speech recognition library (e.g., Google Speech-to-Text API) to convert the audio data into text. This text contains the entire content of the meeting. The converted text is stored in a temporary variable. For example, the audio "In today's meeting, the project progress was reviewed and new tasks were assigned." is converted to text.
[0298] Input: The audio file saved on the server in step 1.
[0299] Output: Text data generated by the server using speech recognition technology.
[0300] Step 3: Emotion Recognition
[0301] Specific operation: The server inputs the voice data extracted from the audio file into an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the user's emotions. The emotion recognition engine analyzes the speech rate, stress, intonation, etc. to determine the user's emotional state. In this process, it may recognize, for example, that the user is tired.
[0302] Input: The audio file saved on the server in step 1.
[0303] Output: Data about the user's emotional state (e.g., fatigue, satisfaction).
[0304] Step 4: Summary generation
[0305] How it works: The server inputs the generated text into a natural language processing model (e.g., OpenAI GPT-3 model) to create a summary. This converts long text into a short summary that includes only the key points. For example, "At today's meeting, project progress was reviewed and new tasks were assigned." is converted into the summary "Progress review and new tasks assigned."
[0306] Input: The text data generated in step 2.
[0307] Output: Summarized text data.
[0308] Step 5: Emotional Adjustment
[0309] Specific behavior: The server adjusts the summary content based on the user's recognized emotions. For example, if the user is tired, the summary text will be more concise and easy to understand. After this adjustment, the summary will be provided as "progress and new tasks."
[0310] Input: The user's emotional state recognized in step 3, and the summary text generated in step 4.
[0311] Output: A summary text adjusted according to the user's sentiment.
[0312] Step 6: Generate the graphical display
[0313] Specific operation: The server creates a graphical display based on the adjusted summary. For example, it arranges the information in a visually easy-to-understand manner using text boxes and diagrams. The generated graphical display is saved as an image file.
[0314] Input: Summary text adjusted in step 5.
[0315] Output: An image file that graphically displays the summary in a visually easy-to-understand format.
[0316] Step 7: Delivering results
[0317] Specific operation: The server provides the generated graphical display to the user. Specifically, the server generates a path to the image file of the graphical display and notifies the user of this path or provides a direct download link. The user clicks the notified link to download the image file to their device and check its contents.
[0318] Input: The image file of the graphical display generated in step 6.
[0319] Output: An image file of the graphical representation that is presented to the user.
[0320] (Application example 2)
[0321] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0322] In order to improve employee customer service skills and customer satisfaction, it is essential to accurately understand the content of conversations between employees and customers and provide feedback. However, current methods require manual analysis of conversation content and emotion recognition, which is time-consuming and labor-intensive. In addition, there are few tools available for providing appropriate feedback based on emotions. This makes it difficult to efficiently support the improvement of employee customer service skills.
[0323] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to upload an audio file; means for the server to receive and store the audio file; means for the server to generate text from the audio file using speech recognition; means for summarizing the generated text; means for graphically displaying the summary; means for providing the summarized graphical display content to the user; means for the server to recognize emotions from the audio file; means for adjusting the summary content based on the emotion recognition; and means for graphically displaying the adjusted summary content. This makes it possible to efficiently analyze the content of a conversation between an employee and a customer and provide appropriate feedback based on emotions.
[0324] "User" refers to a person who uses this system.
[0325] "Audio file" refers to a data file that digitally records sound.
[0326] "Server" refers to a computing device for receiving and processing audio files.
[0327] "Speech recognition" refers to the technology of converting the speech contained in an audio file into text.
[0328] "Text" refers to character string data generated by speech recognition.
[0329] A "summary" refers to a concise summary that extracts only the important parts of the generated text.
[0330] "Graphical display" means a visual representation of the summarized content, provided in the form of an image or chart.
[0331] "Emotion recognition" refers to the technology of analyzing a person's emotional state from the audio features contained in an audio file.
[0332] "Adjustment" refers to changing or adapting the summary content based on the results of emotion recognition.
[0333] The present invention relates to a system for recording and analyzing conversations between store employees while they are serving customers, with the aim of improving the customer service skills of store employees. Specific embodiments of the system are described below.
[0334] This system allows users (employees) to upload audio files while serving customers, and provides the content with speech recognition, emotion recognition, summary generation, and graphical display.
[0335] First, the user uploads the audio file of the customer service session from their own device (such as a smartphone) to the server. The server then stores the received audio file in the specified storage.
[0336] The server then loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library (e.g., the SpeechRecognition library).
[0337] Apart from the converted text data, the server also performs emotion recognition on the audio file using an emotion engine, which uses multiple audio features as input, such as intonation, stress, and speed, using a generative AI model (e.g., Transformers emotion recognition model).
[0338] The converted text is then fed into a summary generation algorithm to generate a summary that extracts only the important parts. A natural language processing model (e.g., Transformers' summary generation model) is used for summary generation.
[0339] After the summary is generated, the server adjusts the summary content and graphical display format based on the recognized emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[0340] Finally, the server creates a graphical representation of the adjusted summary, using text boxes and diagrams for easy visual interpretation, and saves the representation as an image file that users can download to their devices to view the content.
[0341] As a concrete example, suppose a user uploads a recording file of a customer service session called "meeting_audio.wav." After receiving and saving this file, the server performs speech recognition to convert it into text. The converted text is "Today's conversation confirmed the customer's request and explained additional options." If this text is input into a summary generation algorithm, a summary of "Confirmation of customer request and explanation of additional options" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server further simplifies the summary and presents it as a graphical representation such as "Customer request and options." The user can view this graphical representation and easily understand the key points of the conversation.
[0342] An example prompt is, "Please upload an audio file of your conversation. We will convert the audio into text and perform sentiment analysis. We will also summarize the conversation and provide a graphical representation."
[0343] This system allows employees to efficiently analyze conversations during customer service and provide appropriate feedback based on customer sentiment.
[0344] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0345] Step 1:
[0346] The user uploads the audio file recorded during customer service from a device (such as a smartphone or PC) to the server. The audio file (e.g., "meeting_audio.wav") is used as input. The audio file is saved on the server as output. Specifically, the user selects the file from the web interface and clicks the upload button.
[0347] Step 2:
[0348] The server saves the received audio file in the specified storage. At this time, the audio file received from the user is used as input. The path to the audio file saved in the storage is obtained as output. Specifically, the server analyzes the HTTP request and writes the audio file to the file system.
[0349] Step 3:
[0350] The server loads the saved audio file and performs speech recognition, using the saved audio file as input and generating text data as output. Specifically, the server uses a speech recognition library to analyze the audio file and convert it into text.
[0351] Step 4:
[0352] The server inputs the generated text data into an emotion recognition engine to recognize emotions. The generated text data is used as input, and the emotion analysis results are obtained as output. Specifically, the server applies an emotion recognition AI model to the text data and analyzes the main emotional state.
[0353] Step 5:
[0354] The server inputs the generated text data into a summary generation algorithm to create a summary. The generated text data is used as input, and a summary text is generated as output. Specifically, the server uses a natural language processing model to extract important parts from the long text.
[0355] Step 6:
[0356] The server adjusts the summary content based on the emotion recognition results. The inputs are the emotion analysis results and the summary text. The output is the adjusted summary text. Specifically, the server simplifies and emphasizes the text according to the emotional state based on rules.
[0357] Step 7:
[0358] The server generates a graphical display based on the adjusted summary text. The adjusted summary text is used as input. A visual graphical display (e.g., an image file) is generated as output. Specifically, the server uses WordCloud or Matplotlib to visually represent the words in the summary text.
[0359] Step 8:
[0360] The server saves the generated graphical representation as an image file and provides the path to the file to the user. At this time, the generated graphical representation is used as input. As output, a URL of the image file that the user can access is obtained. Specifically, the server saves the image in the file system and notifies the user of the path.
[0361] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0363] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0364] [Second embodiment]
[0365] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0366] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0367] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0368] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0369] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0370] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0371] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0372] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0373] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0374] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0375] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0376] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0377] To implement this invention, a system consisting of three main components is utilized: a user terminal, a server, and a storage device. The present invention involves a process in which a user uploads an audio file, the server recognizes and summarizes the content, and then displays the summary graphically.
[0378] System Overview
[0379] 1. Upload your audio file
[0380] Users select audio files from their devices and upload them to the server, typically via a web interface.
[0381] The server receives the audio file and stores it on a storage device.
[0382] 2. Speech Recognition and Summarization
[0383] The server loads the audio file stored on the storage device and performs speech recognition to convert it into text, preferably using a high-precision speech recognition library.
[0384] The converted text is input into a summary generation model to generate a summary, which uses a natural language processing algorithm.
[0385] 3. Generating the Graphical Display
[0386] The server generates a graphical display from the summarized text, which can be in the form of diagrams and text boxes in a visually friendly format.
[0387] The graphical representation is saved as an image file and presented to the user.
[0388] Program processing overview
[0389] Uploading and saving audio files
[0390] The user uses the device's web interface to select and upload audio files to the server, which receives the uploaded audio files and stores them on a designated storage device.
[0391] Speech Recognition and Text Generation
[0392] The server loads the audio file from the storage device and uses a speech recognition library to convert the speech into text, which reflects the content of the meeting.
[0393] Generate a summary
[0394] The generated text is then summarized using a summary generation algorithm, which reconstructs long texts into short, meaningful summaries.
[0395] Generating a Graphical Display
[0396] The server creates a graphical display from the summarized text using visually friendly formats (e.g., text boxes and diagrams) and stores the generated image files.
[0397] Providing results
[0398] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files using a terminal.
[0399] Specific examples
[0400] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and stores this file, it performs speech recognition and converts it into text. Suppose the converted text is "Today's meeting reviewed project progress and assigned new tasks." This text is summarized to generate a summary: "Review progress and assigned new tasks." A graphical representation is created based on this summary and provided to the user as an image file.
[0401] This system allows users to easily and quickly grasp the important points of a meeting, enabling them to smoothly proceed with decision-making and next actions.
[0402] The processing flow will be explained below.
[0403] Step 1:
[0404] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0405] Step 2:
[0406] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0407] Step 3:
[0408] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0409] Step 4:
[0410] The converted text is input to a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is retrieved.
[0411] Step 5:
[0412] Configure the server to create a graphical display of the generated summary text, visualizing the summary content in a visually understandable format (e.g., text boxes).
[0413] Step 6:
[0414] The server saves the graphical representation as an image file (e.g. "summary_visual.png") and retrieves the path of the saved image file.
[0415] Step 7:
[0416] The server generates a response including a path to the image file for providing the generated image file of the graphical representation to the user, and sends the response to the user terminal.
[0417] Step 8:
[0418] The user receives the response from the server on the device, and uses the image file path included in the response to download and view the generated graphical display.
[0419] Example 1
[0420] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0421] In today's business environment, it is important to efficiently process, summarize, and share audio recordings of meetings and conferences. However, conventional methods require converting speech to text, summarizing it, and providing it in a visually understandable format, which is time-consuming and inefficient. Furthermore, there is a lack of systems that can consistently perform highly accurate speech recognition, generate summaries, and provide intuitively understandable visualizations, making it difficult for users to easily grasp the content of meetings.
[0422] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0423] In this invention, the server includes means for a user to upload voice data, means for the server to receive and store the voice data, means for the server to generate character data from the voice data using voice recognition technology, means for summarizing the generated character data using a summary generation algorithm, means for generating a graphical representation of the summarized character data, and means for providing the summarized graphical representation to the user. This makes it possible to efficiently generate a summary from voice data and further provide the summary in an intuitively easy-to-understand form.
[0424] "User" means a person or entity that uses the system to upload audio data and is responsible for receiving the results.
[0425] "Audio data" refers to digital files in audio format that users upload to the system, including records of meetings, conferences, etc.
[0426] A "server" is a computer system that receives, stores, analyzes, and provides results from audio data.
[0427] "Speech recognition technology" is a technology for converting voice data into text data, analyzing the content of the voice, and outputting it in text format.
[0428] "Text data" is text-format data generated by speech recognition technology, and is a recording of the contents of speech.
[0429] A "summary generation algorithm" is an algorithm for shortening long text data and extracting only the main points to create a concise summary.
[0430] "Graphical display" refers to converting summarized text data into a visually understandable format and displaying it in the form of charts, tables, text boxes, etc.
[0431] A "generative AI model" is an artificial intelligence model used to perform complex tasks such as natural language processing and summary generation.
[0432] A "prompt" is an instruction to be input to a generative AI model, which serves as a guideline for the model to generate appropriate output.
[0433] "Storage Device" means a storage device used by the server to store audio data and other data.
[0434] This invention is implemented using a system that mainly consists of a user terminal, a server, and a storage device. The system shows a series of processes in which a user uploads voice data, the server performs speech recognition and summarization of the data, and provides a graphical display of the results.
[0435] Uploading audio data
[0436] Users select audio data from their devices and upload it to the server via a web interface. The server receives the data and stores it on a designated storage device. The storage device can be a standard cloud storage device (e.g., an Amazon S3 bucket).
[0437] Speech Recognition and Text Generation
[0438] The server loads the saved audio data and converts it into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API). This text data is a recorded text of the meeting or discussion.
[0439] Text summary generation
[0440] The server inputs the generated text data into a natural language processing model (e.g., Hugging Face's BART or T5 model) and runs a summary generation algorithm. The generated text data is shortened to generate a summary that includes only the key points.
[0441] Generating a Graphical Display
[0442] The server creates a visually easy-to-understand graphical representation based on the summary content. This representation uses a chart generation library (e.g., D3.js or Google Chart API) to display the summary in an easy-to-understand format, such as a text box or diagram. The generated graphical representation is saved to the storage device as an image file (e.g., PNG format).
[0443] Results provided and downloaded
[0444] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files through a web interface using their terminal.
[0445] Specific examples
[0446] For example, consider a case where a user uploads a meeting recording file called "meeting_audio.wav." This file is received by the server and saved to a storage device. The server then converts this audio file into text using a speech recognition library, resulting in the text "Today's meeting reviewed project progress and assigned new tasks." This text is then summarized using a natural language processing model to generate a summary: "Reviewed progress and assigned new tasks." A visually friendly graphical representation is then generated from this summary and served to the user as an image file.
[0447] Prompt Sentence Examples
[0448] A user has uploaded a meeting recording file, "meeting_audio.wav." Convert the contents of this audio file into text and generate a summary. Then, display the summary graphically.
[0449] Audio: Today's meeting reviewed project progress and assigned new tasks.
[0450] Summary: Check progress and assign new tasks
[0451] The generated image file of the graphical display will be provided to the user, so please name the file "summary_graphic.png".
[0452] This invention allows users to quickly and efficiently grasp important information about meetings and conferences, which can lead to decision-making and next actions.
[0453] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0454] Step 1: Select and upload audio data
[0455] The user accesses the web interface from their device and selects the audio data (e.g., "meeting_audio.wav").
[0456] Input: Audio data file
[0457] When the user clicks the "Upload" button, the terminal sends the audio data to the server.
[0458] Output: Audio data file sent to the server
[0459] Step 2: Receiving and storing audio data
[0460] The server receives the voice data sent from the user's terminal.
[0461] Input: Audio data file sent from the user's device
[0462] The server stores the received audio data in a designated storage device.
[0463] Output: An audio data file saved on your storage device (e.g. "meeting_audio.wav")
[0464] Step 3: Loading audio data and recognizing it
[0465] The server loads the audio data from the storage device where it is stored.
[0466] Input: Audio data file saved on a storage device
[0467] The server converts the audio data into text data using a speech recognition library such as the Google Speech-to-Text API.
[0468] Output: Converted text data (e.g. "In today's meeting, project progress was reviewed and new tasks were assigned.")
[0469] Step 4: Text summary generation
[0470] The server inputs the generated text data into a natural language processing library (e.g., Hugging Face's BART model).
[0471] Input: Generated character data
[0472] The server executes a summary generation algorithm to summarize the text data.
[0473] Output: Summary text data (e.g. "Check progress and assign new tasks")
[0474] Step 5: Generate the graphical display
[0475] The server creates a visually easy-to-understand graphical display based on the summarized character data.
[0476] Input: summarized character data
[0477] The server uses a chart generation library (e.g. D3.js or Google Chart API) to display the data as text boxes or diagrams.
[0478] Output: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0479] Step 6: Submit and download results
[0480] The server provides the generated image file of the graphical representation to the user.
[0481] Input: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0482] A user uses a terminal to download and view image files from a web interface.
[0483] Output: Image file downloaded to the user
[0484] (Application example 1)
[0485] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0486] To improve the efficiency and accuracy of communication between workers and robots in factories, a means is needed to quickly understand voice instructions and transmit them to the robot as accurate work instructions. It is also necessary to improve work efficiency by visually monitoring progress. Currently, there is a lack of systems that meet these requirements, making it difficult to provide an efficient work environment.
[0487] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0488] In this invention, the server includes means for a user to upload an audio file, means for the server to receive and store the audio file, means for the server to generate text from the audio file using speech recognition, means for summarizing the generated text, means for graphically displaying the summary, means for providing the summarized graphical display to the user, means for extracting instructions from the audio file and transmitting them to the robot as work instructions, and means for providing a dashboard for visually monitoring the instructions. This enables the voice instructions of workers in a factory to be efficiently and accurately transmitted to robots, and progress to be monitored in real time.
[0489] An "audio file" is digital data in which a user has recorded their voice.
[0490] "Upload" is the act of a user sending data from their own terminal to a server.
[0491] A "server" is a computer system that provides various services over a network.
[0492] "Speech recognition" is a technology that analyzes voice data and converts the content into text.
[0493] "Text" is a digital string of characters generated by speech recognition.
[0494] A "summary" refers to text data that has been shortened and only important information has been extracted.
[0495] A "graphical display" is a digital representation that includes pictures and text boxes to visually display the summarized text.
[0496] "User" refers to the entity that uses the system.
[0497] A "work instruction" is a command that indicates the specific actions or tasks that a robot should perform.
[0498] A "robot" is a mechanical device that operates autonomously or semi-autonomously.
[0499] A "dashboard" is an interface that visually displays system information and monitors progress and key indicators.
[0500] To implement the present invention, the following system configuration and algorithm are used: The system mainly consists of a user terminal, a server, and a storage device.
[0501] Detailed system configuration
[0502] 1. Upload and save audio files
[0503] Users select and upload audio files to the server using a web interface or a dedicated smartphone application. The audio files are stored as digital data and received and saved by the server. This allows the user's work instructions to be stored on the server as audio files.
[0504] 2. Speech Recognition and Text Generation
[0505] The server loads the received audio file from the storage device and converts the audio data into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API or IBM Watson Speech to Text). In this step, the instructions in the audio file are extracted as text data.
[0506] 3. Summary Generation
[0507] The generated text is then summarized using a natural language processing model (e.g., BART or T5) using Hugging Face's transformers library. The summarization model shortens the text data and outputs a shortened version that extracts only the important information. This process eliminates redundant information and summarizes the key instructions.
[0508] 4. Generating the Graphical Display
[0509] The server uses a diagram generation library such as matplotlib to create a graphical representation of the summarized text. The summary is displayed in a visually understandable format and generated as an image file. This image file allows the user to intuitively understand the work instructions.
[0510] 5. Transferring work instructions to robots and providing dashboards
[0511] The instructions in the audio file are extracted as text, a summary is generated, and the text is sent to the robot as work instructions. A dashboard is also provided for visually monitoring the instructions and progress. This dashboard is used to check the progress of work in real time, enabling efficient work management within the factory.
[0512] Specific examples
[0513] For example, a user might use their smartphone to voice instructions such as, "After completing the bearing replacement on Line 1, please adjust the machine on Line 2." This voice file is then uploaded to the server, which converts the voice data into text and further summarizes this text to "Line 1: Bearing replacement completed, Line 2: Machine adjustment." The summarized text is then displayed graphically and provided as an image file. The robot receives the summarized work instructions, and real-time progress is displayed on a dashboard.
[0514] Example prompt sentence:
[0515] "Please input operator voice instructions: After completing the bearing replacement on Line 1, please adjust the machine on Line 2."
[0516] In this way, the present invention can significantly improve work efficiency within a factory.
[0517] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0518] Step 1:
[0519] The user uploads the audio file via their smartphone or web interface. Specifically, the user launches the application, selects the recorded audio file (e.g., "Work Instructions_audio.wav"), and presses the "Upload" button. At this time, the audio file is sent to the server.
[0520] Input: Audio file
[0521] Output: Audio file uploaded to the server
[0522] Step 2:
[0523] The server receives the audio file and saves it to a storage device. The server receives the audio file through an HTTP request and saves it in a specified folder in the file system for further processing.
[0524] Input: Uploaded audio file
[0525] Output: Audio file saved on your storage device
[0526] Step 3:
[0527] The server loads the audio file from the storage device and converts the audio into text using a highly accurate speech recognition library. The server then reads the audio file and converts the audio data into text using the Google Speech-to-Text API or IBM Watson Speech to Text. During this process, the instructions in the audio are extracted as a string.
[0528] Input: Audio file saved on your storage device
[0529] Output: Converted text data
[0530] Step 4:
[0531] The server summarizes the generated text using a natural language processing model. The server uses Hugging Face's transformers library to preprocess the text data before inputting it into the summarization model. Summarization is performed using BART or T5 models, and important instructions are extracted without redundant information.
[0532] Input: Converted text data
[0533] Output: Summarized text data
[0534] Step 5:
[0535] The server creates a graphical display to visualize the summarized text. The server uses a chart generation library such as matplotlib to create an image to graphically represent the summarized text. This image is then provided to a dashboard or other display medium.
[0536] Input: Summarized text data
[0537] Output: Image file of the graphical display
[0538] Step 6:
[0539] The server sends the summarized instructions to the robot as work instructions, and the server sends messages to the robot via the network, and the robot performs the work based on these instructions.
[0540] Input: Summarized text data
[0541] Output: Instruction message to the robot
[0542] Step 7:
[0543] The server provides a dashboard for visually monitoring the instructions and progress of work. The server updates data in real time through the dashboard application, allowing users to monitor the progress. This allows users to check the progress of work at a glance.
[0544] Input: Feedback data from the robot
[0545] Output: Real-time updated dashboard
[0546] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0547] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0548] System Overview
[0549] 1. Upload and save audio files:
[0550] The user selects an audio file from their device and uploads it to the server. The server receives the audio file sent by the user and saves it in the specified storage.
[0551] 2. Speech Recognition and Text Generation:
[0552] The server loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library.
[0553] 3. Emotion recognition:
[0554] Apart from the text data recognized from the audio file, the server also analyzes the user's emotions using an emotion engine, which uses multiple audio features such as intonation, stress, and speed as input.
[0555] 4. Summary generation:
[0556] The converted text is then fed into a summary generation algorithm to generate a summary, which uses a natural language processing model to extract only the important parts.
[0557] 5. Emotion-based regulation:
[0558] After the summary is generated, the server adjusts the summary content and graphical display format based on the perceived user emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[0559] 6. Generate the graphical display:
[0560] The server creates a graphical display based on the adjusted summary, using text boxes and diagrams for visual comprehension, and saves this display as an image file.
[0561] 7. Providing Results:
[0562] The server provides the user with a graphical representation stored as an image file, which the user can download and view through their terminal.
[0563] Program processing overview
[0564] Uploading and saving audio files
[0565] The user selects an audio file from the device's web interface and uploads it to the server, which then stores the received audio file in its storage.
[0566] Speech Recognition and Text Generation
[0567] The server loads the saved audio files and performs speech recognition to generate text that contains the full content of the meeting.
[0568] emotion recognition
[0569] The server uses an emotion engine to recognize the user's emotions from the audio file. The emotion engine analyzes the voice characteristics and infers the user's emotional state, for example, stress or satisfaction based on changes in speech rate and volume.
[0570] Summary Generation
[0571] The generated text is then summarized using a natural language processing model, which turns long texts into short, important summaries.
[0572] Emotion-Based Adjustment
[0573] The server adjusts the summary content based on the user's recognized emotions, for example, by making the summary more concise if the user is in a hurry.
[0574] Generating a Graphical Display
[0575] The server creates a graphical display from the tailored summary, which may include text boxes and diagrams to present the information in a visually easy-to-understand format.
[0576] Providing results
[0577] Finally, the server saves the generated graphical display as an image file and provides the path to the file to the user, who can then download the image file to their terminal and view its contents.
[0578] Specific examples
[0579] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0580] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0581] The processing flow will be explained below.
[0582] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The purpose of this invention is to provide an efficient understanding of meetings and information that adapts to the user's emotions.
[0583] Processing step details
[0584] Step 1:
[0585] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0586] Step 2:
[0587] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0588] Step 3:
[0589] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0590] Step 4:
[0591] The server analyzes the tone, speed, volume, etc. of the voice in the audio file and uses an emotion engine to recognize the user's emotion. This recognition uses an emotion analysis algorithm to evaluate the user's emotional state (e.g., stress, joy, fatigue, etc.).
[0592] Step 5:
[0593] The server inputs the converted text into a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is then retrieved.
[0594] Step 6:
[0595] The server adjusts the summary content and the format of the graphical display based on the user's perceived emotions, for example, if the user is tired, the summary text will be more concise and displayed in a format that is easy to understand visually.
[0596] Step 7:
[0597] The server then creates a graphical display based on the emotion-adjusted summary, including text boxes and diagrams that present the information in a visually accessible format.
[0598] Step 8:
[0599] The server saves the generated graphical representation image file (e.g. "summary_visual.png") and obtains its path.
[0600] Step 9:
[0601] The server sends a response containing the path of the image file of the generated graphical representation to the user terminal, and the user receives the response and uses the path of the image file to download and view the graphical representation.
[0602] Specific examples
[0603] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0604] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0605] Example 2
[0606] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0607] In conventional meeting recording systems, simply converting audio data into text and summarizing it was difficult to fully support users' understanding of the meeting content. Furthermore, since information was not provided according to the user's emotional state, there were issues with differences in how information was received and the level of understanding.
[0608] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving and saving an audio file, means for generating text from the audio file using speech recognition technology, means for summarizing the generated text using a natural language processing model, means for using an emotion recognition engine to analyze the user's emotional state, means for adjusting the summary content based on the recognized user emotion, means for graphically displaying the summary content in a visually easy-to-understand format, and means for providing the generated graphical display content to the user. This allows the user to grasp the meeting content efficiently and in an easy-to-understand format, and makes it possible to provide information adapted to the user's emotions.
[0609] An "audio file" is a digital file containing audio data that can be uploaded by a user.
[0610] A "server" is a computer system that stores received audio files and performs speech recognition, natural language processing, emotion recognition, and graphical display.
[0611] "Voice recognition technology" is a technology for analyzing voice data from an audio file and converting it into text data.
[0612] A "natural language processing model" is an algorithm or machine learning model for analyzing generated text and performing summarization or other language processing.
[0613] An "emotion recognition engine" is a technology for analyzing and estimating a user's emotional state from voice data.
[0614] A "summary" is a text that extracts important information from the original text data and summarizes it in a short, concise format.
[0615] A "graphical display" is a display format that includes graphic elements to display text or data in a visually understandable format.
[0616] A "visualization library" is a program library for displaying text and data graphically.
[0617] This invention relates to a system that allows users to upload audio files to a server and then performs speech recognition, emotion recognition, summary generation, and graphical display of the content. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0618] First, the user selects an audio file through the device's web interface and uploads it to the server. The uploaded audio file is received by the server and saved in the specified storage. For example, let's assume that the user uploads an audio file called "meeting_audio.wav."
[0619] The server then loads the saved audio file and converts it into text using speech recognition technology, such as a reliable speech recognition library like the Google Speech-to-Text API. The audio file, "In today's meeting, project progress was reviewed and new tasks were assigned," is converted into text using this technology.
[0620] Furthermore, the server uses an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the audio data extracted from the audio file and recognize the user's emotions. The emotion recognition engine analyzes speech characteristics such as speed, stress, and intonation to determine the user's emotional state. Through this process, it can be recognized that the user is tired, for example.
[0621] The generated text is then fed into a natural language processing model (e.g., the OpenAI GPT-3 model) to generate a summary. The generated text, "In today's meeting, project progress was reviewed and new tasks were assigned," is summarized by the summary generation algorithm as "Progress review and new tasks assigned."
[0622] Based on the user's emotions, the server adjusts the summary. For example, if the server detects that the user is tired, it will make the summary more concise and easier to understand. The adjusted summary is then presented as "progress and new tasks."
[0623] Based on the tailored summary, the server creates a graphical display, using text boxes and diagrams to arrange the information in a visually understandable way, and saves the display as an image file.
[0624] Finally, the server provides the generated graphical display to the user, who can download the image file through his / her terminal and visually confirm the important points of the meeting.
[0625] Specific examples
[0626] For example, if a user uploads a meeting recording file called "meeting_audio.wav," the system works as follows: First, the server receives the audio file and saves it in storage. Then, the Google Speech-to-Text API is used to convert the audio into text, generating the following: "In today's meeting, project progress was reviewed and new tasks were assigned." At the same time, an emotion recognition engine recognizes the user's fatigue, and a summary generation algorithm summarizes the text as "Progress review and new tasks assigned." Finally, the summary is adjusted and presented to the user as a graphical representation of "progress and new tasks."
[0627] Prompt Sentence Examples
[0628] "Perform speech recognition on the meeting recording file 'meeting_audio.wav' and generate a graphical representation of the results of the summary and sentiment analysis."
[0629] This system allows users to grasp the contents of a meeting efficiently and in an easy-to-understand format, and provides information that adapts to the user's emotions.
[0630] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0631] Step 1: Upload and save your audio file
[0632] Specific operation: The user selects an audio file using their own device and uploads it to the server via a web interface. Specifically, the user opens a browser and accesses the system's upload page. They click the file selection button, select "meeting_audio.wav" from the device's file system, and click the upload button to send the audio file to the server. The server receives the file and saves it in a specific directory.
[0633] Input: An audio file selected and uploaded by the user (e.g. "meeting_audio.wav").
[0634] Output: Audio file saved on the server.
[0635] Step 2: Speech recognition and text generation
[0636] Specific operation: The server loads the audio file from storage. Then, it calls a speech recognition library (e.g., Google Speech-to-Text API) to convert the audio data into text. This text contains the entire content of the meeting. The converted text is stored in a temporary variable. For example, the audio "In today's meeting, the project progress was reviewed and new tasks were assigned." is converted to text.
[0637] Input: The audio file saved on the server in step 1.
[0638] Output: Text data generated by the server using speech recognition technology.
[0639] Step 3: Emotion Recognition
[0640] Specific operation: The server inputs the voice data extracted from the audio file into an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the user's emotions. The emotion recognition engine analyzes the speech rate, stress, intonation, etc. to determine the user's emotional state. In this process, it may recognize, for example, that the user is tired.
[0641] Input: The audio file saved on the server in step 1.
[0642] Output: Data about the user's emotional state (e.g., fatigue, satisfaction).
[0643] Step 4: Summary generation
[0644] How it works: The server inputs the generated text into a natural language processing model (e.g., OpenAI GPT-3 model) to create a summary. This converts long text into a short summary that includes only the key points. For example, "At today's meeting, project progress was reviewed and new tasks were assigned." is converted into the summary "Progress review and new tasks assigned."
[0645] Input: The text data generated in step 2.
[0646] Output: Summarized text data.
[0647] Step 5: Emotional Adjustment
[0648] Specific behavior: The server adjusts the summary content based on the user's recognized emotions. For example, if the user is tired, the summary text will be more concise and easy to understand. After this adjustment, the summary will be provided as "progress and new tasks."
[0649] Input: The user's emotional state recognized in step 3, and the summary text generated in step 4.
[0650] Output: A summary text adjusted according to the user's sentiment.
[0651] Step 6: Generate the graphical display
[0652] Specific operation: The server creates a graphical display based on the adjusted summary. For example, it arranges the information in a visually easy-to-understand manner using text boxes and diagrams. The generated graphical display is saved as an image file.
[0653] Input: Summary text adjusted in step 5.
[0654] Output: An image file that graphically displays the summary in a visually easy-to-understand format.
[0655] Step 7: Delivering results
[0656] Specific operation: The server provides the generated graphical display to the user. Specifically, the server generates a path to the image file of the graphical display and notifies the user of this path or provides a direct download link. The user clicks the notified link to download the image file to their device and check its contents.
[0657] Input: The image file of the graphical display generated in step 6.
[0658] Output: An image file of the graphical representation that is presented to the user.
[0659] (Application example 2)
[0660] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0661] In order to improve employee customer service skills and customer satisfaction, it is essential to accurately understand the content of conversations between employees and customers and provide feedback. However, current methods require manual analysis of conversation content and emotion recognition, which is time-consuming and labor-intensive. In addition, there are few tools available for providing appropriate feedback based on emotions. This makes it difficult to efficiently support the improvement of employee customer service skills.
[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to upload an audio file; means for the server to receive and store the audio file; means for the server to generate text from the audio file using speech recognition; means for summarizing the generated text; means for graphically displaying the summary; means for providing the summarized graphical display content to the user; means for the server to recognize emotions from the audio file; means for adjusting the summary content based on the emotion recognition; and means for graphically displaying the adjusted summary content. This makes it possible to efficiently analyze the content of a conversation between an employee and a customer and provide appropriate feedback based on emotions.
[0663] "User" refers to a person who uses this system.
[0664] "Audio file" refers to a data file that digitally records sound.
[0665] "Server" refers to a computing device for receiving and processing audio files.
[0666] "Speech recognition" refers to the technology of converting the speech contained in an audio file into text.
[0667] "Text" refers to character string data generated by speech recognition.
[0668] A "summary" refers to a concise summary that extracts only the important parts of the generated text.
[0669] "Graphical display" means a visual representation of the summarized content, provided in the form of an image or chart.
[0670] "Emotion recognition" refers to the technology of analyzing a person's emotional state from the audio features contained in an audio file.
[0671] "Adjustment" refers to changing or adapting the summary content based on the results of emotion recognition.
[0672] The present invention relates to a system for recording and analyzing conversations between store employees while they are serving customers, with the aim of improving the customer service skills of store employees. Specific embodiments of the system are described below.
[0673] This system allows users (employees) to upload audio files while serving customers, and provides the content with speech recognition, emotion recognition, summary generation, and graphical display.
[0674] First, the user uploads the audio file of the customer service session from their own device (such as a smartphone) to the server. The server then stores the received audio file in the specified storage.
[0675] The server then loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library (e.g., the SpeechRecognition library).
[0676] Apart from the converted text data, the server also performs emotion recognition on the audio file using an emotion engine, which uses multiple audio features as input, such as intonation, stress, and speed, using a generative AI model (e.g., Transformers emotion recognition model).
[0677] The converted text is then fed into a summary generation algorithm to generate a summary that extracts only the important parts. A natural language processing model (e.g., Transformers' summary generation model) is used for summary generation.
[0678] After the summary is generated, the server adjusts the summary content and graphical display format based on the recognized emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[0679] Finally, the server creates a graphical representation of the adjusted summary, using text boxes and diagrams for easy visual interpretation, and saves the representation as an image file that users can download to their devices to view the content.
[0680] As a concrete example, suppose a user uploads a recording file of a customer service session called "meeting_audio.wav." After receiving and saving this file, the server performs speech recognition to convert it into text. The converted text is "Today's conversation confirmed the customer's request and explained additional options." If this text is input into a summary generation algorithm, a summary of "Confirmation of customer request and explanation of additional options" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server further simplifies the summary and presents it as a graphical representation such as "Customer request and options." The user can view this graphical representation and easily understand the key points of the conversation.
[0681] An example prompt is, "Please upload an audio file of your conversation. We will convert the audio into text and perform sentiment analysis. We will also summarize the conversation and provide a graphical representation."
[0682] This system allows employees to efficiently analyze conversations during customer service and provide appropriate feedback based on customer sentiment.
[0683] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0684] Step 1:
[0685] The user uploads the audio file recorded during customer service from a device (such as a smartphone or PC) to the server. The audio file (e.g., "meeting_audio.wav") is used as input. The audio file is saved on the server as output. Specifically, the user selects the file from the web interface and clicks the upload button.
[0686] Step 2:
[0687] The server saves the received audio file in the specified storage. At this time, the audio file received from the user is used as input. The path to the audio file saved in the storage is obtained as output. Specifically, the server analyzes the HTTP request and writes the audio file to the file system.
[0688] Step 3:
[0689] The server loads the saved audio file and performs speech recognition, using the saved audio file as input and generating text data as output. Specifically, the server uses a speech recognition library to analyze the audio file and convert it into text.
[0690] Step 4:
[0691] The server inputs the generated text data into an emotion recognition engine to recognize emotions. The generated text data is used as input, and the emotion analysis results are obtained as output. Specifically, the server applies an emotion recognition AI model to the text data and analyzes the main emotional state.
[0692] Step 5:
[0693] The server inputs the generated text data into a summary generation algorithm to create a summary. The generated text data is used as input, and a summary text is generated as output. Specifically, the server uses a natural language processing model to extract important parts from the long text.
[0694] Step 6:
[0695] The server adjusts the summary content based on the emotion recognition results. The inputs are the emotion analysis results and the summary text. The output is the adjusted summary text. Specifically, the server simplifies and emphasizes the text according to the emotional state based on rules.
[0696] Step 7:
[0697] The server generates a graphical display based on the adjusted summary text. The adjusted summary text is used as input. A visual graphical display (e.g., an image file) is generated as output. Specifically, the server uses WordCloud or Matplotlib to visually represent the words in the summary text.
[0698] Step 8:
[0699] The server saves the generated graphical representation as an image file and provides the path to the file to the user. At this time, the generated graphical representation is used as input. As output, a URL of the image file that the user can access is obtained. Specifically, the server saves the image in the file system and notifies the user of the path.
[0700] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0701] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0702] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0703] [Third embodiment]
[0704] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0705] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0706] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0707] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0708] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0709] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0710] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0711] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0712] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0713] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0714] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0715] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0716] To implement this invention, a system consisting of three main components is utilized: a user terminal, a server, and a storage device. The present invention involves a process in which a user uploads an audio file, the server recognizes and summarizes the content, and then displays the summary graphically.
[0717] System Overview
[0718] 1. Upload your audio file
[0719] Users select audio files from their devices and upload them to the server, typically via a web interface.
[0720] The server receives the audio file and stores it on a storage device.
[0721] 2. Speech Recognition and Summarization
[0722] The server loads the audio file stored on the storage device and performs speech recognition to convert it into text, preferably using a high-precision speech recognition library.
[0723] The converted text is input into a summary generation model to generate a summary, which uses a natural language processing algorithm.
[0724] 3. Generating the Graphical Display
[0725] The server generates a graphical display from the summarized text, which can be in the form of diagrams and text boxes in a visually friendly format.
[0726] The graphical representation is saved as an image file and presented to the user.
[0727] Program processing overview
[0728] Uploading and saving audio files
[0729] The user uses the device's web interface to select and upload audio files to the server, which receives the uploaded audio files and stores them on a designated storage device.
[0730] Speech Recognition and Text Generation
[0731] The server loads the audio file from the storage device and uses a speech recognition library to convert the speech into text, which reflects the content of the meeting.
[0732] Generate a summary
[0733] The generated text is then summarized using a summary generation algorithm, which reconstructs long texts into short, meaningful summaries.
[0734] Generating a Graphical Display
[0735] The server creates a graphical display from the summarized text using visually friendly formats (e.g., text boxes and diagrams) and stores the generated image files.
[0736] Providing results
[0737] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files using a terminal.
[0738] Specific examples
[0739] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and stores this file, it performs speech recognition and converts it into text. Suppose the converted text is "Today's meeting reviewed project progress and assigned new tasks." This text is summarized to generate a summary: "Review progress and assigned new tasks." A graphical representation is created based on this summary and provided to the user as an image file.
[0740] This system allows users to easily and quickly grasp the important points of a meeting, enabling them to smoothly proceed with decision-making and next actions.
[0741] The processing flow will be explained below.
[0742] Step 1:
[0743] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0744] Step 2:
[0745] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0746] Step 3:
[0747] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0748] Step 4:
[0749] The converted text is input to a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is retrieved.
[0750] Step 5:
[0751] Configure the server to create a graphical display of the generated summary text, visualizing the summary content in a visually understandable format (e.g., text boxes).
[0752] Step 6:
[0753] The server saves the graphical representation as an image file (e.g. "summary_visual.png") and retrieves the path of the saved image file.
[0754] Step 7:
[0755] The server generates a response including a path to the image file for providing the generated image file of the graphical representation to the user, and sends the response to the user terminal.
[0756] Step 8:
[0757] The user receives the response from the server on the device, and uses the image file path included in the response to download and view the generated graphical display.
[0758] Example 1
[0759] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0760] In today's business environment, it is important to efficiently process, summarize, and share audio recordings of meetings and conferences. However, conventional methods require converting speech to text, summarizing it, and providing it in a visually understandable format, which is time-consuming and inefficient. Furthermore, there is a lack of systems that can consistently perform highly accurate speech recognition, generate summaries, and provide intuitively understandable visualizations, making it difficult for users to easily grasp the content of meetings.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0762] In this invention, the server includes means for a user to upload voice data, means for the server to receive and store the voice data, means for the server to generate character data from the voice data using voice recognition technology, means for summarizing the generated character data using a summary generation algorithm, means for generating a graphical representation of the summarized character data, and means for providing the summarized graphical representation to the user. This makes it possible to efficiently generate a summary from voice data and further provide the summary in an intuitively easy-to-understand form.
[0763] "User" means a person or entity that uses the system to upload audio data and is responsible for receiving the results.
[0764] "Audio data" refers to digital files in audio format that users upload to the system, including records of meetings, conferences, etc.
[0765] A "server" is a computer system that receives, stores, analyzes, and provides results from audio data.
[0766] "Speech recognition technology" is a technology for converting voice data into text data, analyzing the content of the voice, and outputting it in text format.
[0767] "Text data" is text-format data generated by speech recognition technology, and is a recording of the contents of speech.
[0768] A "summary generation algorithm" is an algorithm for shortening long text data and extracting only the main points to create a concise summary.
[0769] "Graphical display" refers to converting summarized text data into a visually understandable format and displaying it in the form of charts, tables, text boxes, etc.
[0770] A "generative AI model" is an artificial intelligence model used to perform complex tasks such as natural language processing and summary generation.
[0771] A "prompt" is an instruction to be input to a generative AI model, which serves as a guideline for the model to generate appropriate output.
[0772] "Storage Device" means a storage device used by the server to store audio data and other data.
[0773] This invention is implemented using a system that mainly consists of a user terminal, a server, and a storage device. The system shows a series of processes in which a user uploads voice data, the server performs speech recognition and summarization of the data, and provides a graphical display of the results.
[0774] Uploading audio data
[0775] Users select audio data from their devices and upload it to the server via a web interface. The server receives the data and stores it on a designated storage device. The storage device can be a standard cloud storage device (e.g., an Amazon S3 bucket).
[0776] Speech Recognition and Text Generation
[0777] The server loads the saved audio data and converts it into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API). This text data is a recorded text of the meeting or discussion.
[0778] Text summary generation
[0779] The server inputs the generated text data into a natural language processing model (e.g., Hugging Face's BART or T5 model) and runs a summary generation algorithm. The generated text data is shortened to generate a summary that includes only the key points.
[0780] Generating a Graphical Display
[0781] The server creates a visually easy-to-understand graphical representation based on the summary content. This representation uses a chart generation library (e.g., D3.js or Google Chart API) to display the summary in an easy-to-understand format, such as a text box or diagram. The generated graphical representation is saved to the storage device as an image file (e.g., PNG format).
[0782] Results provided and downloaded
[0783] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files through a web interface using their terminal.
[0784] Specific examples
[0785] For example, consider a case where a user uploads a meeting recording file called "meeting_audio.wav." This file is received by the server and saved to a storage device. The server then converts this audio file into text using a speech recognition library, resulting in the text "Today's meeting reviewed project progress and assigned new tasks." This text is then summarized using a natural language processing model to generate a summary: "Reviewed progress and assigned new tasks." A visually friendly graphical representation is then generated from this summary and served to the user as an image file.
[0786] Prompt Sentence Examples
[0787] A user has uploaded a meeting recording file, "meeting_audio.wav." Convert the contents of this audio file into text and generate a summary. Then, display the summary graphically.
[0788] Audio: Today's meeting reviewed project progress and assigned new tasks.
[0789] Summary: Check progress and assign new tasks
[0790] The generated image file of the graphical display will be provided to the user, so please name the file "summary_graphic.png".
[0791] This invention allows users to quickly and efficiently grasp important information about meetings and conferences, which can lead to decision-making and next actions.
[0792] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0793] Step 1: Select and upload audio data
[0794] The user accesses the web interface from their device and selects the audio data (e.g., "meeting_audio.wav").
[0795] Input: Audio data file
[0796] When the user clicks the "Upload" button, the terminal sends the audio data to the server.
[0797] Output: Audio data file sent to the server
[0798] Step 2: Receiving and storing audio data
[0799] The server receives the voice data sent from the user's terminal.
[0800] Input: Audio data file sent from the user's device
[0801] The server stores the received audio data in a designated storage device.
[0802] Output: An audio data file saved on your storage device (e.g. "meeting_audio.wav")
[0803] Step 3: Loading audio data and recognizing it
[0804] The server loads the audio data from the storage device where it is stored.
[0805] Input: Audio data file saved on a storage device
[0806] The server converts the audio data into text data using a speech recognition library such as the Google Speech-to-Text API.
[0807] Output: Converted text data (e.g. "In today's meeting, project progress was reviewed and new tasks were assigned.")
[0808] Step 4: Text summary generation
[0809] The server inputs the generated text data into a natural language processing library (e.g., Hugging Face's BART model).
[0810] Input: Generated character data
[0811] The server executes a summary generation algorithm to summarize the text data.
[0812] Output: Summary text data (e.g. "Check progress and assign new tasks")
[0813] Step 5: Generate the graphical display
[0814] The server creates a visually easy-to-understand graphical display based on the summarized character data.
[0815] Input: summarized character data
[0816] The server uses a chart generation library (e.g. D3.js or Google Chart API) to display the data as text boxes or diagrams.
[0817] Output: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0818] Step 6: Submit and download results
[0819] The server provides the generated image file of the graphical representation to the user.
[0820] Input: An image file containing the graphical representation (e.g. "summary_graphic.png")
[0821] A user uses a terminal to download and view image files from a web interface.
[0822] Output: Image file downloaded to the user
[0823] (Application example 1)
[0824] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0825] To improve the efficiency and accuracy of communication between workers and robots in factories, a means is needed to quickly understand voice instructions and transmit them to the robot as accurate work instructions. It is also necessary to improve work efficiency by visually monitoring progress. Currently, there is a lack of systems that meet these requirements, making it difficult to provide an efficient work environment.
[0826] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0827] In this invention, the server includes means for a user to upload an audio file, means for the server to receive and store the audio file, means for the server to generate text from the audio file using speech recognition, means for summarizing the generated text, means for graphically displaying the summary, means for providing the summarized graphical display to the user, means for extracting instructions from the audio file and transmitting them to the robot as work instructions, and means for providing a dashboard for visually monitoring the instructions. This enables the voice instructions of workers in a factory to be efficiently and accurately transmitted to robots, and progress to be monitored in real time.
[0828] An "audio file" is digital data in which a user has recorded their voice.
[0829] "Upload" is the act of a user sending data from their own terminal to a server.
[0830] A "server" is a computer system that provides various services over a network.
[0831] "Speech recognition" is a technology that analyzes voice data and converts the content into text.
[0832] "Text" is a digital string of characters generated by speech recognition.
[0833] A "summary" refers to text data that has been shortened and only important information has been extracted.
[0834] A "graphical display" is a digital representation that includes pictures and text boxes to visually display the summarized text.
[0835] "User" refers to the entity that uses the system.
[0836] A "work instruction" is a command that indicates the specific actions or tasks that a robot should perform.
[0837] A "robot" is a mechanical device that operates autonomously or semi-autonomously.
[0838] A "dashboard" is an interface that visually displays system information and monitors progress and key indicators.
[0839] To implement the present invention, the following system configuration and algorithm are used: The system mainly consists of a user terminal, a server, and a storage device.
[0840] Detailed system configuration
[0841] 1. Upload and save audio files
[0842] Users select and upload audio files to the server using a web interface or a dedicated smartphone application. The audio files are stored as digital data and received and saved by the server. This allows the user's work instructions to be stored on the server as audio files.
[0843] 2. Speech Recognition and Text Generation
[0844] The server loads the received audio file from the storage device and converts the audio data into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API or IBM Watson Speech to Text). In this step, the instructions in the audio file are extracted as text data.
[0845] 3. Summary Generation
[0846] The generated text is then summarized using a natural language processing model (e.g., BART or T5) using Hugging Face's transformers library. The summarization model shortens the text data and outputs a shortened version that extracts only the important information. This process eliminates redundant information and summarizes the key instructions.
[0847] 4. Generating the Graphical Display
[0848] The server uses a diagram generation library such as matplotlib to create a graphical representation of the summarized text. The summary is displayed in a visually understandable format and generated as an image file. This image file allows the user to intuitively understand the work instructions.
[0849] 5. Transferring work instructions to robots and providing dashboards
[0850] The instructions in the audio file are extracted as text, a summary is generated, and the text is sent to the robot as work instructions. A dashboard is also provided for visually monitoring the instructions and progress. This dashboard is used to check the progress of work in real time, enabling efficient work management within the factory.
[0851] Specific examples
[0852] For example, a user might use their smartphone to voice instructions such as, "After completing the bearing replacement on Line 1, please adjust the machine on Line 2." This voice file is then uploaded to the server, which converts the voice data into text and further summarizes this text to "Line 1: Bearing replacement completed, Line 2: Machine adjustment." The summarized text is then displayed graphically and provided as an image file. The robot receives the summarized work instructions, and real-time progress is displayed on a dashboard.
[0853] Example prompt sentence:
[0854] "Please input operator voice instructions: After completing the bearing replacement on Line 1, please adjust the machine on Line 2."
[0855] In this way, the present invention can significantly improve work efficiency within a factory.
[0856] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0857] Step 1:
[0858] The user uploads the audio file via their smartphone or web interface. Specifically, the user launches the application, selects the recorded audio file (e.g., "Work Instructions_audio.wav"), and presses the "Upload" button. At this time, the audio file is sent to the server.
[0859] Input: Audio file
[0860] Output: Audio file uploaded to the server
[0861] Step 2:
[0862] The server receives the audio file and saves it to a storage device. The server receives the audio file through an HTTP request and saves it in a specified folder in the file system for further processing.
[0863] Input: Uploaded audio file
[0864] Output: Audio file saved on your storage device
[0865] Step 3:
[0866] The server loads the audio file from the storage device and converts the audio into text using a highly accurate speech recognition library. The server then reads the audio file and converts the audio data into text using the Google Speech-to-Text API or IBM Watson Speech to Text. During this process, the instructions in the audio are extracted as a string.
[0867] Input: Audio file saved on your storage device
[0868] Output: Converted text data
[0869] Step 4:
[0870] The server summarizes the generated text using a natural language processing model. The server uses Hugging Face's transformers library to preprocess the text data before inputting it into the summarization model. Summarization is performed using BART or T5 models, and important instructions are extracted without redundant information.
[0871] Input: Converted text data
[0872] Output: Summarized text data
[0873] Step 5:
[0874] The server creates a graphical display to visualize the summarized text. The server uses a chart generation library such as matplotlib to create an image to graphically represent the summarized text. This image is then provided to a dashboard or other display medium.
[0875] Input: Summarized text data
[0876] Output: Image file of the graphical display
[0877] Step 6:
[0878] The server sends the summarized instructions to the robot as work instructions, and the server sends messages to the robot via the network, and the robot performs the work based on these instructions.
[0879] Input: Summarized text data
[0880] Output: Instruction message to the robot
[0881] Step 7:
[0882] The server provides a dashboard for visually monitoring the instructions and progress of work. The server updates data in real time through the dashboard application, allowing users to monitor the progress. This allows users to check the progress of work at a glance.
[0883] Input: Feedback data from the robot
[0884] Output: Real-time updated dashboard
[0885] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0886] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0887] System Overview
[0888] 1. Upload and save audio files:
[0889] The user selects an audio file from their device and uploads it to the server. The server receives the audio file sent by the user and saves it in the specified storage.
[0890] 2. Speech Recognition and Text Generation:
[0891] The server loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library.
[0892] 3. Emotion recognition:
[0893] Apart from the text data recognized from the audio file, the server also analyzes the user's emotions using an emotion engine, which uses multiple audio features such as intonation, stress, and speed as input.
[0894] 4. Summary generation:
[0895] The converted text is then fed into a summary generation algorithm to generate a summary, which uses a natural language processing model to extract only the important parts.
[0896] 5. Emotion-based regulation:
[0897] After the summary is generated, the server adjusts the summary content and graphical display format based on the perceived user emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[0898] 6. Generate the graphical display:
[0899] The server creates a graphical display based on the adjusted summary, using text boxes and diagrams for visual comprehension, and saves this display as an image file.
[0900] 7. Providing Results:
[0901] The server provides the user with a graphical representation stored as an image file, which the user can download and view through their terminal.
[0902] Program processing overview
[0903] Uploading and saving audio files
[0904] The user selects an audio file from the device's web interface and uploads it to the server, which then stores the received audio file in its storage.
[0905] Speech Recognition and Text Generation
[0906] The server loads the saved audio files and performs speech recognition to generate text that contains the full content of the meeting.
[0907] emotion recognition
[0908] The server uses an emotion engine to recognize the user's emotions from the audio file. The emotion engine analyzes the voice characteristics and infers the user's emotional state, for example, stress or satisfaction based on changes in speech rate and volume.
[0909] Summary Generation
[0910] The generated text is then summarized using a natural language processing model, which turns long texts into short, important summaries.
[0911] Emotion-Based Adjustment
[0912] The server adjusts the summary content based on the user's recognized emotions, for example, by making the summary more concise if the user is in a hurry.
[0913] Generating a Graphical Display
[0914] The server creates a graphical display from the tailored summary, which may include text boxes and diagrams to present the information in a visually easy-to-understand format.
[0915] Providing results
[0916] Finally, the server saves the generated graphical display as an image file and provides the path to the file to the user, who can then download the image file to their terminal and view its contents.
[0917] Specific examples
[0918] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0919] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0920] The processing flow will be explained below.
[0921] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The purpose of this invention is to provide an efficient understanding of meetings and information that adapts to the user's emotions.
[0922] Processing step details
[0923] Step 1:
[0924] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[0925] Step 2:
[0926] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[0927] Step 3:
[0928] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[0929] Step 4:
[0930] The server analyzes the tone, speed, volume, etc. of the voice in the audio file and uses an emotion engine to recognize the user's emotion. This recognition uses an emotion analysis algorithm to evaluate the user's emotional state (e.g., stress, joy, fatigue, etc.).
[0931] Step 5:
[0932] The server inputs the converted text into a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is then retrieved.
[0933] Step 6:
[0934] The server adjusts the summary content and the format of the graphical display based on the user's perceived emotions, for example, if the user is tired, the summary text will be more concise and displayed in a format that is easy to understand visually.
[0935] Step 7:
[0936] The server then creates a graphical display based on the emotion-adjusted summary, including text boxes and diagrams that present the information in a visually accessible format.
[0937] Step 8:
[0938] The server saves the generated graphical image file (e.g. "summary_visual.png") and obtains its path.
[0939] Step 9:
[0940] The server sends a response containing the path of the image file of the generated graphical representation to the user terminal, and the user receives the response and uses the path of the image file to download and view the graphical representation.
[0941] Specific examples
[0942] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After receiving and saving this file, the server performs speech recognition to convert it into text. The converted text is "Today's meeting reviewed project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Review progress and assigned new tasks" will be generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[0943] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[0944] Example 2
[0945] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0946] In conventional meeting recording systems, simply converting audio data into text and summarizing it was difficult to fully support users' understanding of the meeting content. Furthermore, since information was not provided according to the user's emotional state, there were issues with differences in how information was received and the level of understanding.
[0947] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving and saving an audio file, means for generating text from the audio file using speech recognition technology, means for summarizing the generated text using a natural language processing model, means for using an emotion recognition engine to analyze the user's emotional state, means for adjusting the summary content based on the recognized user emotion, means for graphically displaying the summary content in a visually easy-to-understand format, and means for providing the generated graphical display content to the user. This allows the user to grasp the meeting content efficiently and in an easy-to-understand format, and makes it possible to provide information adapted to the user's emotions.
[0948] An "audio file" is a digital file containing audio data that can be uploaded by a user.
[0949] A "server" is a computer system that stores received audio files and performs speech recognition, natural language processing, emotion recognition, and graphical display.
[0950] "Voice recognition technology" is a technology for analyzing voice data from an audio file and converting it into text data.
[0951] A "natural language processing model" is an algorithm or machine learning model for analyzing generated text and performing summarization or other language processing.
[0952] An "emotion recognition engine" is a technology for analyzing and estimating a user's emotional state from voice data.
[0953] A "summary" is a text that extracts important information from the original text data and summarizes it in a short, concise format.
[0954] A "graphical display" is a display format that includes graphic elements to display text or data in a visually understandable format.
[0955] A "visualization library" is a program library for displaying text and data graphically.
[0956] This invention relates to a system that allows users to upload audio files to a server and then performs speech recognition, emotion recognition, summary generation, and graphical display of the content. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[0957] First, the user selects an audio file through the device's web interface and uploads it to the server. The uploaded audio file is received by the server and saved in the specified storage. For example, let's assume that the user uploads an audio file called "meeting_audio.wav."
[0958] The server then loads the saved audio file and converts it into text using speech recognition technology, such as a reliable speech recognition library like the Google Speech-to-Text API. The audio file, "In today's meeting, project progress was reviewed and new tasks were assigned," is converted into text using this technology.
[0959] Furthermore, the server uses an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the audio data extracted from the audio file and recognize the user's emotions. The emotion recognition engine analyzes speech characteristics such as speed, stress, and intonation to determine the user's emotional state. Through this process, it can be recognized that the user is tired, for example.
[0960] The generated text is then fed into a natural language processing model (e.g., the OpenAI GPT-3 model) to generate a summary. The generated text, "In today's meeting, project progress was reviewed and new tasks were assigned," is summarized by the summary generation algorithm as "Progress review and new tasks assigned."
[0961] Based on the user's emotions, the server adjusts the summary. For example, if the server detects that the user is tired, it will make the summary more concise and easier to understand. The adjusted summary is then presented as "progress and new tasks."
[0962] Based on the tailored summary, the server creates a graphical display, using text boxes and diagrams to arrange the information in a visually understandable way, and saves the display as an image file.
[0963] Finally, the server provides the generated graphical display to the user, who can download the image file through his / her terminal and visually confirm the important points of the meeting.
[0964] Specific examples
[0965] For example, if a user uploads a meeting recording file called "meeting_audio.wav," the system works as follows: First, the server receives the audio file and saves it in storage. Then, the Google Speech-to-Text API is used to convert the audio into text, generating the following: "In today's meeting, project progress was reviewed and new tasks were assigned." At the same time, an emotion recognition engine recognizes the user's fatigue, and a summary generation algorithm summarizes the text as "Progress review and new tasks assigned." Finally, the summary is adjusted and presented to the user as a graphical representation of "progress and new tasks."
[0966] Prompt Sentence Examples
[0967] "Perform speech recognition on the meeting recording file 'meeting_audio.wav' and generate a graphical representation of the results of the summary and sentiment analysis."
[0968] This system allows users to grasp the contents of a meeting efficiently and in an easy-to-understand format, and provides information that adapts to the user's emotions.
[0969] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0970] Step 1: Upload and save your audio file
[0971] Specific operation: The user selects an audio file using their own device and uploads it to the server via a web interface. Specifically, the user opens a browser and accesses the system's upload page. They click the file selection button, select "meeting_audio.wav" from the device's file system, and click the upload button to send the audio file to the server. The server receives the file and saves it in a specific directory.
[0972] Input: An audio file selected and uploaded by the user (e.g. "meeting_audio.wav").
[0973] Output: Audio file saved on the server.
[0974] Step 2: Speech recognition and text generation
[0975] Specific operation: The server loads the audio file from storage. Then, it calls a speech recognition library (e.g., Google Speech-to-Text API) to convert the audio data into text. This text contains the entire content of the meeting. The converted text is stored in a temporary variable. For example, the audio "In today's meeting, the project progress was reviewed and new tasks were assigned." is converted to text.
[0976] Input: The audio file saved on the server in step 1.
[0977] Output: Text data generated by the server using speech recognition technology.
[0978] Step 3: Emotion Recognition
[0979] Specific operation: The server inputs the voice data extracted from the audio file into an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the user's emotions. The emotion recognition engine analyzes the speech rate, stress, intonation, etc. to determine the user's emotional state. In this process, it may recognize, for example, that the user is tired.
[0980] Input: The audio file saved on the server in step 1.
[0981] Output: Data about the user's emotional state (e.g., fatigue, satisfaction).
[0982] Step 4: Summary generation
[0983] How it works: The server inputs the generated text into a natural language processing model (e.g., OpenAI GPT-3 model) to create a summary. This converts long text into a short summary that includes only the key points. For example, "At today's meeting, project progress was reviewed and new tasks were assigned." is converted into the summary "Progress review and new tasks assigned."
[0984] Input: The text data generated in step 2.
[0985] Output: Summarized text data.
[0986] Step 5: Emotional Adjustment
[0987] Specific behavior: The server adjusts the summary content based on the user's recognized emotions. For example, if the user is tired, the summary text will be more concise and easy to understand. After this adjustment, the summary will be provided as "progress and new tasks."
[0988] Input: The user's emotional state recognized in step 3, and the summary text generated in step 4.
[0989] Output: A summary text adjusted according to the user's sentiment.
[0990] Step 6: Generate the graphical display
[0991] Specific operation: The server creates a graphical display based on the adjusted summary. For example, it arranges the information in a visually easy-to-understand manner using text boxes and diagrams. The generated graphical display is saved as an image file.
[0992] Input: Summary text adjusted in step 5.
[0993] Output: An image file that graphically displays the summary in a visually easy-to-understand format.
[0994] Step 7: Delivering results
[0995] Specific operation: The server provides the generated graphical display to the user. Specifically, the server generates a path to the image file of the graphical display and notifies the user of this path or provides a direct download link. The user clicks the notified link to download the image file to their device and check its contents.
[0996] Input: The image file of the graphical display generated in step 6.
[0997] Output: An image file of the graphical representation that is presented to the user.
[0998] (Application example 2)
[0999] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1000] In order to improve employee customer service skills and customer satisfaction, it is essential to accurately understand the content of conversations between employees and customers and provide feedback. However, current methods require manual analysis of conversation content and emotion recognition, which is time-consuming and labor-intensive. In addition, there are few tools available for providing appropriate feedback based on emotions. This makes it difficult to efficiently support the improvement of employee customer service skills.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to upload an audio file; means for the server to receive and store the audio file; means for the server to generate text from the audio file using speech recognition; means for summarizing the generated text; means for graphically displaying the summary; means for providing the summarized graphical display content to the user; means for the server to recognize emotions from the audio file; means for adjusting the summary content based on the emotion recognition; and means for graphically displaying the adjusted summary content. This makes it possible to efficiently analyze the content of a conversation between an employee and a customer and provide appropriate feedback based on emotions.
[1002] "User" refers to a person who uses this system.
[1003] "Audio file" refers to a data file that digitally records sound.
[1004] "Server" refers to a computing device for receiving and processing audio files.
[1005] "Speech recognition" refers to the technology of converting the speech contained in an audio file into text.
[1006] "Text" refers to character string data generated by speech recognition.
[1007] A "summary" refers to a concise summary that extracts only the important parts of the generated text.
[1008] "Graphical display" means a visual representation of the summarized content, provided in the form of an image or chart.
[1009] "Emotion recognition" refers to the technology of analyzing a person's emotional state from the audio features contained in an audio file.
[1010] "Adjustment" refers to changing or adapting the summary content based on the results of emotion recognition.
[1011] The present invention relates to a system for recording and analyzing conversations between store employees while they are serving customers, with the aim of improving the customer service skills of store employees. Specific embodiments of the system are described below.
[1012] This system allows users (employees) to upload audio files while serving customers, and provides the content with speech recognition, emotion recognition, summary generation, and graphical display.
[1013] First, the user uploads the audio file of the customer service from their own device (such as a smartphone) to the server. The server then stores the received audio file in the specified storage.
[1014] The server then loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library (e.g., the SpeechRecognition library).
[1015] Apart from the converted text data, the server also performs emotion recognition on the audio file using an emotion engine, which uses multiple audio features as input, such as intonation, stress, and speed, using a generative AI model (e.g., Transformers emotion recognition model).
[1016] The converted text is then fed into a summary generation algorithm to generate a summary that extracts only the important parts. A natural language processing model (e.g., Transformers' summary generation model) is used for summary generation.
[1017] After the summary is generated, the server adjusts the summary content and graphical display format based on the recognized emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[1018] Finally, the server creates a graphical representation of the adjusted summary, using text boxes and diagrams for easy visual interpretation, and saves the representation as an image file that users can download to their devices to view the content.
[1019] As a concrete example, suppose a user uploads a recording file of a customer service session called "meeting_audio.wav." After receiving and saving this file, the server performs speech recognition to convert it into text. The converted text is "Today's conversation confirmed the customer's request and explained additional options." If this text is input into a summary generation algorithm, a summary of "Confirmation of customer request and explanation of additional options" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server further simplifies the summary and presents it as a graphical representation such as "Customer request and options." The user can view this graphical representation and easily understand the key points of the conversation.
[1020] An example prompt is, "Please upload an audio file of your conversation. We will convert the audio into text and perform sentiment analysis. We will also summarize the conversation and provide a graphical representation."
[1021] This system allows employees to efficiently analyze conversations during customer service and provide appropriate feedback based on customer sentiment.
[1022] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1023] Step 1:
[1024] The user uploads the audio file recorded during customer service from a device (such as a smartphone or PC) to the server. The audio file (e.g., "meeting_audio.wav") is used as input. The audio file is saved on the server as output. Specifically, the user selects the file from the web interface and clicks the upload button.
[1025] Step 2:
[1026] The server saves the received audio file in the specified storage. At this time, the audio file received from the user is used as input. The path to the audio file saved in the storage is obtained as output. Specifically, the server analyzes the HTTP request and writes the audio file to the file system.
[1027] Step 3:
[1028] The server loads the saved audio file and performs speech recognition, using the saved audio file as input and generating text data as output. Specifically, the server uses a speech recognition library to analyze the audio file and convert it into text.
[1029] Step 4:
[1030] The server inputs the generated text data into an emotion recognition engine to recognize emotions. The generated text data is used as input, and the emotion analysis results are obtained as output. Specifically, the server applies an emotion recognition AI model to the text data and analyzes the main emotional state.
[1031] Step 5:
[1032] The server inputs the generated text data into a summary generation algorithm to create a summary. The generated text data is used as input, and a summary text is generated as output. Specifically, the server uses a natural language processing model to extract important parts from the long text.
[1033] Step 6:
[1034] The server adjusts the summary content based on the emotion recognition results. The inputs are the emotion analysis results and the summary text. The output is the adjusted summary text. Specifically, the server simplifies and emphasizes the text according to the emotional state based on rules.
[1035] Step 7:
[1036] The server generates a graphical display based on the adjusted summary text. The adjusted summary text is used as input. A visual graphical display (e.g., an image file) is generated as output. Specifically, the server uses WordCloud or Matplotlib to visually represent the words in the summary text.
[1037] Step 8:
[1038] The server saves the generated graphical representation as an image file and provides the path to the file to the user. At this time, the generated graphical representation is used as input. As output, a URL of the image file that the user can access is obtained. Specifically, the server saves the image in the file system and notifies the user of the path.
[1039] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1040] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1041] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1042] [Fourth embodiment]
[1043] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1044] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1045] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1046] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1047] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1048] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1049] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1050] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1051] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1052] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1053] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1054] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1055] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1056] To implement this invention, a system consisting of three main components is utilized: a user terminal, a server, and a storage device. The present invention involves a process in which a user uploads an audio file, the server recognizes and summarizes the content, and then displays the summary graphically.
[1057] System Overview
[1058] 1. Upload your audio file
[1059] Users select audio files from their devices and upload them to the server, typically via a web interface.
[1060] The server receives the audio file and stores it on a storage device.
[1061] 2. Speech Recognition and Summarization
[1062] The server loads the audio file stored on the storage device and performs speech recognition to convert it into text, preferably using a high-precision speech recognition library.
[1063] The converted text is input into a summary generation model to generate a summary, which uses a natural language processing algorithm.
[1064] 3. Generating the Graphical Display
[1065] The server generates a graphical display from the summarized text, which can be in the form of diagrams and text boxes in a visually friendly format.
[1066] The graphical representation is saved as an image file and presented to the user.
[1067] Program processing overview
[1068] Uploading and saving audio files
[1069] The user uses the device's web interface to select and upload audio files to the server, which receives the uploaded audio files and stores them on a designated storage device.
[1070] Speech Recognition and Text Generation
[1071] The server loads the audio file from the storage device and uses a speech recognition library to convert the speech into text, which reflects the content of the meeting.
[1072] Generate a summary
[1073] The generated text is then summarized using a summary generation algorithm, which reconstructs long texts into short, meaningful summaries.
[1074] Generating a Graphical Display
[1075] The server creates a graphical display from the summarized text using visually friendly formats (e.g., text boxes and diagrams) and stores the generated image files.
[1076] Providing results
[1077] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files using a terminal.
[1078] Specific examples
[1079] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and stores this file, it performs speech recognition and converts it into text. Suppose the converted text is "Today's meeting reviewed project progress and assigned new tasks." This text is summarized to generate a summary: "Review progress and assigned new tasks." A graphical representation is created based on this summary and provided to the user as an image file.
[1080] This system allows users to easily and quickly grasp the important points of a meeting, enabling them to smoothly proceed with decision-making and next actions.
[1081] The processing flow will be explained below.
[1082] Step 1:
[1083] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[1084] Step 2:
[1085] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[1086] Step 3:
[1087] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[1088] Step 4:
[1089] The converted text is input to a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is retrieved.
[1090] Step 5:
[1091] Configure the server to create a graphical display of the generated summary text, visualizing the summary content in a visually understandable format (e.g., text boxes).
[1092] Step 6:
[1093] The server saves the graphical representation as an image file (e.g. "summary_visual.png") and retrieves the path of the saved image file.
[1094] Step 7:
[1095] The server generates a response including a path to the image file for providing the generated image file of the graphical representation to the user, and sends the response to the user terminal.
[1096] Step 8:
[1097] The user receives the response from the server on the device, and uses the image file path included in the response to download and view the generated graphical display.
[1098] Example 1
[1099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1100] In today's business environment, it is important to efficiently process, summarize, and share audio recordings of meetings and conferences. However, conventional methods require converting speech to text, summarizing it, and providing it in a visually understandable format, which is time-consuming and inefficient. Furthermore, there is a lack of systems that can consistently perform highly accurate speech recognition, generate summaries, and provide intuitively understandable visualizations, making it difficult for users to easily grasp the content of meetings.
[1101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1102] In this invention, the server includes means for a user to upload voice data, means for the server to receive and store the voice data, means for the server to generate character data from the voice data using voice recognition technology, means for summarizing the generated character data using a summary generation algorithm, means for generating a graphical representation of the summarized character data, and means for providing the summarized graphical representation to the user. This makes it possible to efficiently generate a summary from voice data and further provide the summary in an intuitively easy-to-understand form.
[1103] "User" means a person or entity that uses the system to upload audio data and is responsible for receiving the results.
[1104] "Audio data" refers to digital files in audio format that users upload to the system, including records of meetings, conferences, etc.
[1105] A "server" is a computer system that receives, stores, analyzes, and provides results from audio data.
[1106] "Speech recognition technology" is a technology for converting voice data into text data, analyzing the content of the voice, and outputting it in text format.
[1107] "Text data" is text-format data generated by speech recognition technology, and is a recording of the contents of speech.
[1108] A "summary generation algorithm" is an algorithm for shortening long text data and extracting only the main points to create a concise summary.
[1109] "Graphical display" refers to converting summarized text data into a visually understandable format and displaying it in the form of charts, tables, text boxes, etc.
[1110] A "generative AI model" is an artificial intelligence model used to perform complex tasks such as natural language processing and summary generation.
[1111] A "prompt" is an instruction to be input to a generative AI model, which serves as a guideline for the model to generate appropriate output.
[1112] "Storage Device" means a storage device used by the server to store audio data and other data.
[1113] This invention is implemented using a system that mainly consists of a user terminal, a server, and a storage device. The system shows a series of processes in which a user uploads voice data, the server performs speech recognition and summarization of the data, and provides a graphical display of the results.
[1114] Uploading audio data
[1115] Users select audio data from their devices and upload it to the server via a web interface. The server receives the data and stores it on a designated storage device. The storage device can be a standard cloud storage device (e.g., an Amazon S3 bucket).
[1116] Speech Recognition and Text Generation
[1117] The server loads the saved audio data and converts it into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API). This text data is a recorded text of the meeting or discussion.
[1118] Text summary generation
[1119] The server inputs the generated text data into a natural language processing model (e.g., Hugging Face's BART or T5 model) and runs a summary generation algorithm. The generated text data is shortened to generate a summary that includes only the key points.
[1120] Generating a Graphical Display
[1121] The server creates a visually easy-to-understand graphical representation based on the summary content. This representation uses a chart generation library (e.g., D3.js or Google Chart API) to display the summary in an easy-to-understand format, such as a text box or diagram. The generated graphical representation is saved to the storage device as an image file (e.g., PNG format).
[1122] Results provided and downloaded
[1123] The server provides the generated image files of the graphical representation to the user, who can then download and view the image files through a web interface using their terminal.
[1124] Specific examples
[1125] For example, consider a case where a user uploads a meeting recording file called "meeting_audio.wav." This file is received by the server and saved to a storage device. The server then converts this audio file into text using a speech recognition library, resulting in the text "Today's meeting reviewed project progress and assigned new tasks." This text is then summarized using a natural language processing model to generate a summary: "Reviewed progress and assigned new tasks." A visually friendly graphical representation is then generated from this summary and served to the user as an image file.
[1126] Prompt Sentence Examples
[1127] A user has uploaded a meeting recording file, "meeting_audio.wav." Convert the contents of this audio file into text and generate a summary. Then, display the summary graphically.
[1128] Audio: Today's meeting reviewed project progress and assigned new tasks.
[1129] Summary: Check progress and assign new tasks
[1130] The generated image file of the graphical display will be provided to the user, so please name the file "summary_graphic.png".
[1131] This invention allows users to quickly and efficiently grasp important information about meetings and conferences, which can lead to decision-making and next actions.
[1132] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1133] Step 1: Select and upload audio data
[1134] The user accesses the web interface from their device and selects the audio data (e.g., "meeting_audio.wav").
[1135] Input: Audio data file
[1136] When the user clicks the "Upload" button, the terminal sends the audio data to the server.
[1137] Output: Audio data file sent to the server
[1138] Step 2: Receiving and storing audio data
[1139] The server receives the voice data sent from the user's terminal.
[1140] Input: Audio data file sent from the user's device
[1141] The server stores the received audio data in a designated storage device.
[1142] Output: An audio data file saved on your storage device (e.g. "meeting_audio.wav")
[1143] Step 3: Loading audio data and recognizing it
[1144] The server loads the audio data from the storage device where it is stored.
[1145] Input: Audio data file saved on a storage device
[1146] The server converts the audio data into text data using a speech recognition library such as the Google Speech-to-Text API.
[1147] Output: Converted text data (e.g. "In today's meeting, project progress was reviewed and new tasks were assigned.")
[1148] Step 4: Text summary generation
[1149] The server inputs the generated text data into a natural language processing library (e.g., Hugging Face's BART model).
[1150] Input: Generated character data
[1151] The server executes a summary generation algorithm to summarize the text data.
[1152] Output: Summary text data (e.g. "Check progress and assign new tasks")
[1153] Step 5: Generate the graphical display
[1154] The server creates a visually easy-to-understand graphical display based on the summarized character data.
[1155] Input: summarized character data
[1156] The server uses a chart generation library (e.g. D3.js or Google Chart API) to display the data as text boxes or diagrams.
[1157] Output: An image file containing the graphical representation (e.g. "summary_graphic.png")
[1158] Step 6: Submit and download results
[1159] The server provides the generated image file of the graphical representation to the user.
[1160] Input: An image file containing the graphical representation (e.g. "summary_graphic.png")
[1161] A user uses a terminal to download and view image files from a web interface.
[1162] Output: Image file downloaded to the user
[1163] (Application example 1)
[1164] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1165] To improve the efficiency and accuracy of communication between workers and robots in factories, a means is needed to quickly understand voice instructions and transmit them to the robot as accurate work instructions. It is also necessary to improve work efficiency by visually monitoring progress. Currently, there is a lack of systems that meet these requirements, making it difficult to provide an efficient work environment.
[1166] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1167] In this invention, the server includes means for a user to upload an audio file, means for the server to receive and store the audio file, means for the server to generate text from the audio file using speech recognition, means for summarizing the generated text, means for graphically displaying the summary, means for providing the summarized graphical display to the user, means for extracting instructions from the audio file and transmitting them to the robot as work instructions, and means for providing a dashboard for visually monitoring the instructions. This enables the voice instructions of workers in a factory to be efficiently and accurately transmitted to robots, and progress to be monitored in real time.
[1168] An "audio file" is digital data in which a user has recorded their voice.
[1169] "Upload" is the act of a user sending data from their own terminal to a server.
[1170] A "server" is a computer system that provides various services over a network.
[1171] "Speech recognition" is a technology that analyzes voice data and converts the content into text.
[1172] "Text" is a digital string of characters generated by speech recognition.
[1173] A "summary" refers to text data that has been shortened and only important information has been extracted.
[1174] A "graphical display" is a digital representation that includes pictures and text boxes to visually display the summarized text.
[1175] "User" refers to the entity that uses the system.
[1176] A "work instruction" is a command that indicates the specific actions or tasks that a robot should perform.
[1177] A "robot" is a mechanical device that operates autonomously or semi-autonomously.
[1178] A "dashboard" is an interface that visually displays system information and monitors progress and key indicators.
[1179] To implement the present invention, the following system configuration and algorithm are used: The system mainly consists of a user terminal, a server, and a storage device.
[1180] Detailed system configuration
[1181] 1. Upload and save audio files
[1182] Users select and upload audio files to the server using a web interface or a dedicated smartphone application. The audio files are stored as digital data and received and saved by the server. This allows the user's work instructions to be stored on the server as audio files.
[1183] 2. Speech Recognition and Text Generation
[1184] The server loads the received audio file from the storage device and converts the audio data into text using a high-precision speech recognition library (e.g., Google Speech-to-Text API or IBM Watson Speech to Text). In this step, the instructions in the audio file are extracted as text data.
[1185] 3. Summary Generation
[1186] The generated text is then summarized using a natural language processing model (e.g., BART or T5) using Hugging Face's transformers library. The summarization model shortens the text data and outputs a shortened version that extracts only the important information. This process eliminates redundant information and summarizes the key instructions.
[1187] 4. Generating the Graphical Display
[1188] The server uses a diagram generation library such as matplotlib to create a graphical representation of the summarized text. The summary is displayed in a visually understandable format and generated as an image file. This image file allows the user to intuitively understand the work instructions.
[1189] 5. Transferring work instructions to robots and providing dashboards
[1190] The instructions in the audio file are extracted as text, a summary is generated, and the text is sent to the robot as work instructions. A dashboard is also provided for visually monitoring the instructions and progress. This dashboard is used to check the progress of work in real time, enabling efficient work management within the factory.
[1191] Specific examples
[1192] For example, a user might use their smartphone to voice instructions such as, "After completing the bearing replacement on Line 1, please adjust the machine on Line 2." This voice file is then uploaded to the server, which converts the voice data into text and further summarizes this text to "Line 1: Bearing replacement completed, Line 2: Machine adjustment." The summarized text is then displayed graphically and provided as an image file. The robot receives the summarized work instructions, and real-time progress is displayed on a dashboard.
[1193] Example prompt sentence:
[1194] "Please input operator voice instructions: After completing the bearing replacement on Line 1, please adjust the machine on Line 2."
[1195] In this way, the present invention can significantly improve work efficiency within a factory.
[1196] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1197] Step 1:
[1198] The user uploads the audio file via their smartphone or web interface. Specifically, the user launches the application, selects the recorded audio file (e.g., "Work Instructions_audio.wav"), and presses the "Upload" button. At this time, the audio file is sent to the server.
[1199] Input: Audio file
[1200] Output: Audio file uploaded to the server
[1201] Step 2:
[1202] The server receives the audio file and saves it to a storage device. The server receives the audio file through an HTTP request and saves it in a specified folder in the file system for further processing.
[1203] Input: Uploaded audio file
[1204] Output: Audio file saved on your storage device
[1205] Step 3:
[1206] The server loads the audio file from the storage device and converts the audio into text using a highly accurate speech recognition library. The server then reads the audio file and converts the audio data into text using the Google Speech-to-Text API or IBM Watson Speech to Text. During this process, the instructions in the audio are extracted as a string.
[1207] Input: Audio file saved on your storage device
[1208] Output: Converted text data
[1209] Step 4:
[1210] The server summarizes the generated text using a natural language processing model. The server uses Hugging Face's transformers library to preprocess the text data before inputting it into the summarization model. Summarization is performed using BART or T5 models, and important instructions are extracted without redundant information.
[1211] Input: Converted text data
[1212] Output: Summarized text data
[1213] Step 5:
[1214] The server creates a graphical display to visualize the summarized text. The server uses a chart generation library such as matplotlib to create an image to graphically represent the summarized text. This image is then provided to a dashboard or other display medium.
[1215] Input: Summarized text data
[1216] Output: Image file of the graphical display
[1217] Step 6:
[1218] The server sends the summarized instructions to the robot as work instructions, and the server sends messages to the robot via the network, and the robot performs the work based on these instructions.
[1219] Input: Summarized text data
[1220] Output: Instruction message to the robot
[1221] Step 7:
[1222] The server provides a dashboard for visually monitoring the instructions and progress of work. The server updates data in real time through the dashboard application, allowing users to monitor the progress. This allows users to check the progress of work at a glance.
[1223] Input: Feedback data from the robot
[1224] Output: Real-time updated dashboard
[1225] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1226] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[1227] System Overview
[1228] 1. Upload and save audio files:
[1229] The user selects an audio file from their device and uploads it to the server. The server receives the audio file sent by the user and saves it in the specified storage.
[1230] 2. Speech Recognition and Text Generation:
[1231] The server loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library.
[1232] 3. Emotion recognition:
[1233] Apart from the text data recognized from the audio file, the server also analyzes the user's emotions using an emotion engine, which uses multiple audio features such as intonation, stress, and speed as input.
[1234] 4. Summary generation:
[1235] The converted text is then fed into a summary generation algorithm to generate a summary, which uses a natural language processing model to extract only the important parts.
[1236] 5. Emotion-based regulation:
[1237] After the summary is generated, the server adjusts the summary content and graphical display format based on the perceived user emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[1238] 6. Generate the graphical display:
[1239] The server creates a graphical display based on the adjusted summary, using text boxes and diagrams for visual comprehension, and saves this display as an image file.
[1240] 7. Providing Results:
[1241] The server provides the user with a graphical representation stored as an image file, which the user can download and view through their terminal.
[1242] Program processing overview
[1243] Uploading and saving audio files
[1244] The user selects an audio file from the device's web interface and uploads it to the server, which then stores the received audio file in its storage.
[1245] Speech Recognition and Text Generation
[1246] The server loads the saved audio files and performs speech recognition to generate text that contains the full content of the meeting.
[1247] emotion recognition
[1248] The server uses an emotion engine to recognize the user's emotions from the audio file. The emotion engine analyzes the voice characteristics and infers the user's emotional state, for example, stress or satisfaction based on changes in speech rate and volume.
[1249] Summary Generation
[1250] The generated text is then summarized using a natural language processing model, which turns long texts into short, important summaries.
[1251] Emotion-Based Adjustment
[1252] The server adjusts the summary content based on the user's recognized emotions, for example, by making the summary more concise if the user is in a hurry.
[1253] Generating a Graphical Display
[1254] The server creates a graphical display from the tailored summary, which may include text boxes and diagrams to present the information in a visually easy-to-understand format.
[1255] Providing results
[1256] Finally, the server saves the generated graphical display as an image file and provides the path to the file to the user, who can then download the image file to their terminal and view its contents.
[1257] Specific examples
[1258] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[1259] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[1260] The processing flow will be explained below.
[1261] This invention relates to a system that allows users to upload audio files to a server, and performs speech recognition, summary generation, graphical display, and emotion recognition on the content of the files. The purpose of this invention is to provide an efficient understanding of meetings and information that adapts to the user's emotions.
[1262] Processing step details
[1263] Step 1:
[1264] The user opens the web interface of the device, selects an audio file (e.g., "meeting_audio.wav"), and clicks the "Upload" button to upload the selected audio file to the server.
[1265] Step 2:
[1266] The server receives the audio file sent by the user and saves it in the specified storage device (e.g., "upload_folder" directory).
[1267] Step 3:
[1268] The server loads the saved audio file from the storage device, converts the audio file to text using a speech recognition library (e.g., Google Speech Recognition), and retrieves the converted text.
[1269] Step 4:
[1270] The server analyzes the tone, speed, volume, etc. of the voice in the audio file and uses an emotion engine to recognize the user's emotion. This recognition uses an emotion analysis algorithm to evaluate the user's emotional state (e.g., stress, joy, fatigue, etc.).
[1271] Step 5:
[1272] The server inputs the converted text into a summary generation algorithm (e.g., a natural language processing model) to generate a summary. The generated summary text is then retrieved.
[1273] Step 6:
[1274] The server adjusts the summary content and the format of the graphical display based on the user's perceived emotions, for example, if the user is tired, the summary text will be more concise and displayed in a format that is easy to understand visually.
[1275] Step 7:
[1276] The server then creates a graphical display based on the emotion-adjusted summary, including text boxes and diagrams that present the information in a visually accessible format.
[1277] Step 8:
[1278] The server saves the generated graphical representation image file (e.g. "summary_visual.png") and obtains its path.
[1279] Step 9:
[1280] The server sends a response containing the path of the image file of the generated graphical representation to the user terminal, and the user receives the response and uses the path of the image file to download and view the graphical representation.
[1281] Specific examples
[1282] For example, suppose a user uploads a meeting recording file called "meeting_audio.wav." After the server receives and saves this file, it performs speech recognition and converts it into text. The converted text is "Today's meeting reviewed the project progress and assigned new tasks." If this text is input into a summary generation algorithm, the summary "Reviewing progress and assigning new tasks" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server will further simplify the summary and provide it as a graphical representation, such as "progress and new tasks." The user can then browse this graphical representation to easily understand the key points of the meeting.
[1283] This system allows users to understand the content of meetings in an efficient and easy-to-understand format, and supports better decision-making by providing emotion-based information.
[1284] Example 2
[1285] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1286] In conventional meeting recording systems, simply converting audio data into text and summarizing it was difficult to fully support users' understanding of the meeting content. Furthermore, since information was not provided according to the user's emotional state, there were issues with differences in how information was received and the level of understanding.
[1287] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving and saving an audio file, means for generating text from the audio file using speech recognition technology, means for summarizing the generated text using a natural language processing model, means for using an emotion recognition engine to analyze the user's emotional state, means for adjusting the summary content based on the recognized user emotion, means for graphically displaying the summary content in a visually easy-to-understand format, and means for providing the generated graphical display content to the user. This allows the user to grasp the meeting content efficiently and in an easy-to-understand format, and makes it possible to provide information adapted to the user's emotions.
[1288] An "audio file" is a digital file containing audio data that can be uploaded by a user.
[1289] A "server" is a computer system that stores received audio files and performs speech recognition, natural language processing, emotion recognition, and graphical display.
[1290] "Voice recognition technology" is a technology for analyzing voice data from an audio file and converting it into text data.
[1291] A "natural language processing model" is an algorithm or machine learning model for analyzing generated text and performing summarization or other language processing.
[1292] An "emotion recognition engine" is a technology for analyzing and estimating a user's emotional state from voice data.
[1293] A "summary" is a text that extracts important information from the original text data and summarizes it in a short, concise format.
[1294] A "graphical display" is a display format that includes graphic elements to display text or data in a visually understandable format.
[1295] A "visualization library" is a program library for displaying text and data graphically.
[1296] This invention relates to a system that allows users to upload audio files to a server and then performs speech recognition, emotion recognition, summary generation, and graphical display of the content. The system aims to provide efficient understanding of meetings and information adapted to the user's emotions.
[1297] First, the user selects an audio file through the device's web interface and uploads it to the server. The uploaded audio file is received by the server and saved in the specified storage. For example, let's assume that the user uploads an audio file called "meeting_audio.wav."
[1298] The server then loads the saved audio file and converts it into text using speech recognition technology, such as a reliable speech recognition library like the Google Speech-to-Text API. The audio file, "In today's meeting, project progress was reviewed and new tasks were assigned," is converted into text using this technology.
[1299] Furthermore, the server uses an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the audio data extracted from the audio file and recognize the user's emotions. The emotion recognition engine analyzes speech characteristics such as speed, stress, and intonation to determine the user's emotional state. Through this process, it can be recognized that the user is tired, for example.
[1300] The generated text is then fed into a natural language processing model (e.g., the OpenAI GPT-3 model) to generate a summary. The generated text, "In today's meeting, project progress was reviewed and new tasks were assigned," is summarized by the summary generation algorithm as "Progress review and new tasks assigned."
[1301] Based on the user's emotions, the server adjusts the summary. For example, if the server detects that the user is tired, it will make the summary more concise and easier to understand. The adjusted summary is then presented as "progress and new tasks."
[1302] Based on the tailored summary, the server creates a graphical display, using text boxes and diagrams to arrange the information in a visually understandable way, and saves the display as an image file.
[1303] Finally, the server provides the generated graphical display to the user, who can download the image file through his / her terminal and visually confirm the important points of the meeting.
[1304] Specific examples
[1305] For example, if a user uploads a meeting recording file called "meeting_audio.wav," the system works as follows: First, the server receives the audio file and saves it in storage. Then, the Google Speech-to-Text API is used to convert the audio into text, generating the following: "In today's meeting, project progress was reviewed and new tasks were assigned." At the same time, an emotion recognition engine recognizes the user's fatigue, and a summary generation algorithm summarizes the text as "Progress review and new tasks assigned." Finally, the summary is adjusted and presented to the user as a graphical representation of "progress and new tasks."
[1306] Prompt Sentence Examples
[1307] "Perform speech recognition on the meeting recording file 'meeting_audio.wav' and generate a graphical representation of the results of the summary and sentiment analysis."
[1308] This system allows users to grasp the contents of a meeting efficiently and in an easy-to-understand format, and provides information that adapts to the user's emotions.
[1309] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1310] Step 1: Upload and save your audio file
[1311] Specific operation: The user selects an audio file using their own device and uploads it to the server via a web interface. Specifically, the user opens a browser and accesses the system's upload page. They click the file selection button, select "meeting_audio.wav" from the device's file system, and click the upload button to send the audio file to the server. The server receives the file and saves it in a specific directory.
[1312] Input: An audio file selected and uploaded by the user (e.g. "meeting_audio.wav").
[1313] Output: Audio file saved on the server.
[1314] Step 2: Speech recognition and text generation
[1315] Specific operation: The server loads the audio file from storage. Then, it calls a speech recognition library (e.g., Google Speech-to-Text API) to convert the audio data into text. This text contains the entire content of the meeting. The converted text is stored in a temporary variable. For example, the audio "In today's meeting, the project progress was reviewed and new tasks were assigned." is converted to text.
[1316] Input: The audio file saved on the server in step 1.
[1317] Output: Text data generated by the server using speech recognition technology.
[1318] Step 3: Emotion Recognition
[1319] Specific operation: The server inputs the voice data extracted from the audio file into an emotion recognition engine (e.g., Microsoft Azure Emotion API) to analyze the user's emotions. The emotion recognition engine analyzes the speech rate, stress, intonation, etc. to determine the user's emotional state. In this process, it may recognize, for example, that the user is tired.
[1320] Input: The audio file saved on the server in step 1.
[1321] Output: Data about the user's emotional state (e.g., fatigue, satisfaction).
[1322] Step 4: Summary generation
[1323] How it works: The server inputs the generated text into a natural language processing model (e.g., OpenAI GPT-3 model) to create a summary. This converts long text into a short summary that includes only the key points. For example, "At today's meeting, project progress was reviewed and new tasks were assigned." is converted into the summary "Progress review and new tasks assigned."
[1324] Input: The text data generated in step 2.
[1325] Output: Summarized text data.
[1326] Step 5: Emotional Adjustment
[1327] Specific behavior: The server adjusts the summary content based on the user's recognized emotions. For example, if the user is tired, the summary text will be more concise and easy to understand. After this adjustment, the summary will be provided as "progress and new tasks."
[1328] Input: The user's emotional state recognized in step 3, and the summary text generated in step 4.
[1329] Output: A summary text adjusted according to the user's sentiment.
[1330] Step 6: Generate the graphical display
[1331] Specific operation: The server creates a graphical display based on the adjusted summary. For example, it arranges the information in a visually easy-to-understand manner using text boxes and diagrams. The generated graphical display is saved as an image file.
[1332] Input: Summary text adjusted in step 5.
[1333] Output: An image file that graphically displays the summary in a visually easy-to-understand format.
[1334] Step 7: Delivering results
[1335] Specific operation: The server provides the generated graphical display to the user. Specifically, the server generates a path to the image file of the graphical display and notifies the user of this path or provides a direct download link. The user clicks the notified link to download the image file to their device and check its contents.
[1336] Input: The image file of the graphical display generated in step 6.
[1337] Output: An image file of the graphical representation that is presented to the user.
[1338] (Application example 2)
[1339] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1340] In order to improve employee customer service skills and customer satisfaction, it is essential to accurately understand the content of conversations between employees and customers and provide feedback. However, current methods require manual analysis of conversation content and emotion recognition, which is time-consuming and labor-intensive. In addition, there are few tools available for providing appropriate feedback based on emotions. This makes it difficult to efficiently support the improvement of employee customer service skills.
[1341] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for a user to upload an audio file; means for the server to receive and store the audio file; means for the server to generate text from the audio file using speech recognition; means for summarizing the generated text; means for graphically displaying the summary; means for providing the summarized graphical display content to the user; means for the server to recognize emotions from the audio file; means for adjusting the summary content based on the emotion recognition; and means for graphically displaying the adjusted summary content. This makes it possible to efficiently analyze the content of a conversation between an employee and a customer and provide appropriate feedback based on emotions.
[1342] "User" refers to a person who uses this system.
[1343] "Audio file" refers to a data file that digitally records sound.
[1344] "Server" refers to a computing device for receiving and processing audio files.
[1345] "Speech recognition" refers to the technology of converting the speech contained in an audio file into text.
[1346] "Text" refers to character string data generated by speech recognition.
[1347] A "summary" refers to a concise summary that extracts only the important parts of the generated text.
[1348] "Graphical display" means a visual representation of the summarized content, provided in the form of an image or chart.
[1349] "Emotion recognition" refers to the technology of analyzing a person's emotional state from the audio features contained in an audio file.
[1350] "Adjustment" refers to changing or adapting the summary content based on the results of emotion recognition.
[1351] The present invention relates to a system for recording and analyzing conversations between store employees while they are serving customers, with the aim of improving the customer service skills of store employees. Specific embodiments of the system are described below.
[1352] This system allows users (employees) to upload audio files while serving customers, and provides the content with speech recognition, emotion recognition, summary generation, and graphical display.
[1353] First, the user uploads the audio file of the customer service from their own device (such as a smartphone) to the server. The server then stores the received audio file in the specified storage.
[1354] The server then loads the saved audio file and performs speech recognition to convert the speech to text using a reliable speech recognition library (e.g., the SpeechRecognition library).
[1355] Apart from the converted text data, the server also performs emotion recognition on the audio file using an emotion engine, which uses multiple audio features as input, such as intonation, stress, and speed, using a generative AI model (e.g., Transformers emotion recognition model).
[1356] The converted text is then fed into a summary generation algorithm to generate a summary that extracts only the important parts. A natural language processing model (e.g., Transformers' summary generation model) is used for summary generation.
[1357] After the summary is generated, the server adjusts the summary content and graphical display format based on the recognized emotion, for example, if the user is tired, the information is presented in a more concise and easy-to-understand format.
[1358] Finally, the server creates a graphical representation of the adjusted summary, using text boxes and diagrams for easy visual interpretation, and saves the representation as an image file that users can download to their devices to view the content.
[1359] As a concrete example, suppose a user uploads a recording file of a customer service session called "meeting_audio.wav." After receiving and saving this file, the server performs speech recognition to convert it into text. The converted text is "Today's conversation confirmed the customer's request and explained additional options." If this text is input into a summary generation algorithm, a summary of "Confirmation of customer request and explanation of additional options" is generated. At the same time, if the emotion engine recognizes the user's fatigue, the server further simplifies the summary and presents it as a graphical representation such as "Customer request and options." The user can view this graphical representation and easily understand the key points of the conversation.
[1360] An example prompt is, "Please upload an audio file of your conversation. We will convert the audio into text and perform sentiment analysis. We will also summarize the conversation and provide a graphical representation."
[1361] This system allows employees to efficiently analyze conversations during customer service and provide appropriate feedback based on customer sentiment.
[1362] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1363] Step 1:
[1364] The user uploads the audio file recorded during customer service from a device (such as a smartphone or PC) to the server. The audio file (e.g., "meeting_audio.wav") is used as input. The audio file is saved on the server as output. Specifically, the user selects the file from the web interface and clicks the upload button.
[1365] Step 2:
[1366] The server saves the received audio file in the specified storage. At this time, the audio file received from the user is used as input. The path to the audio file saved in the storage is obtained as output. Specifically, the server analyzes the HTTP request and writes the audio file to the file system.
[1367] Step 3:
[1368] The server loads the saved audio file and performs speech recognition, using the saved audio file as input and generating text data as output. Specifically, the server uses a speech recognition library to analyze the audio file and convert it into text.
[1369] Step 4:
[1370] The server inputs the generated text data into an emotion recognition engine to recognize emotions. The generated text data is used as input, and the emotion analysis results are obtained as output. Specifically, the server applies an emotion recognition AI model to the text data and analyzes the main emotional state.
[1371] Step 5:
[1372] The server inputs the generated text data into a summary generation algorithm to create a summary. The generated text data is used as input, and a summary text is generated as output. Specifically, the server uses a natural language processing model to extract important parts from the long text.
[1373] Step 6:
[1374] The server adjusts the summary content based on the emotion recognition results. The inputs are the emotion analysis results and the summary text. The output is the adjusted summary text. Specifically, the server simplifies and emphasizes the text according to the emotional state based on rules.
[1375] Step 7:
[1376] The server generates a graphical display based on the adjusted summary text. The adjusted summary text is used as input. A visual graphical display (e.g., an image file) is generated as output. Specifically, the server uses WordCloud or Matplotlib to visually represent the words in the summary text.
[1377] Step 8:
[1378] The server saves the generated graphical representation as an image file and provides the path to the file to the user. At this time, the generated graphical representation is used as input. As output, a URL of the image file that the user can access is obtained. Specifically, the server saves the image in the file system and notifies the user of the path.
[1379] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1380] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1381] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1382] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1383] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1384] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1385] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1386] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1387] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1388] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1389] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1390] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1391] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1392] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1393] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1394] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1395] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1396] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1397] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1398] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1399] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1400] The following is further disclosed regarding the above embodiment.
[1401] (Claim 1)
[1402] a means for a user to upload an audio file;
[1403] a means for the server to receive and store the audio files;
[1404] means for the server to generate text from the audio file using speech recognition;
[1405] a means for summarizing the generated text;
[1406] a means for graphically displaying the summary content;
[1407] means for providing a summarized graphical display to a user;
[1408] A system including:
[1409] (Claim 2)
[1410] 10. The system of claim 1, wherein the server summarizes the generated text using a natural language processing model.
[1411] (Claim 3)
[1412] 10. The system of claim 1, wherein the server uses a chart generation library to create a graphical display for visualizing the summary content.
[1413] "Example 1"
[1414] (Claim 1)
[1415] a means for a user to upload audio data;
[1416] a means for the server to receive and store the audio data;
[1417] A means for the server to generate character data from voice data using voice recognition technology;
[1418] A means for summarizing the generated character data using a summary generation algorithm;
[1419] means for generating a graphical representation of the summarized text data;
[1420] means for providing a summarized graphical display to a user;
[1421] A system including:
[1422] (Claim 2)
[1423] 2. The system of claim 1, wherein the server summarizes the generated character data using a natural language processing model.
[1424] (Claim 3)
[1425] 10. The system of claim 1, wherein the server uses a diagram generation library to visually display the summarized content.
[1426] "Application Example 1"
[1427] (Claim 1)
[1428] a means for a user to upload an audio file;
[1429] a means for the server to receive and store the audio files;
[1430] means for the server to generate text from the audio file using speech recognition;
[1431] a means for summarizing the generated text;
[1432] a means for graphically displaying the summary content;
[1433] means for providing a summarized graphical display to a user;
[1434] A means for extracting instructions from the voice file and transmitting the instructions to the robot as work instructions;
[1435] a means for providing a dashboard for visually monitoring instructions;
[1436] A system including:
[1437] (Claim 2)
[1438] 10. The system of claim 1, wherein the server summarizes the generated text using a natural language processing model.
[1439] (Claim 3)
[1440] 10. The system of claim 1, wherein the server uses a chart generation library to create a graphical display for visualizing the summary content.
[1441] "Example 2: Combining Emotion Engines"
[1442] (Claim 1)
[1443] a means for a user to upload an audio file;
[1444] a means for the server to receive and store the audio files;
[1445] a means for the server to generate text from the audio file using speech recognition technology;
[1446] a means for the server to summarize the generated text using a natural language processing model;
[1447] means for the server to use an emotion recognition engine to analyze the user's emotional state;
[1448] means for the server to adjust the summary content based on the recognized user sentiment;
[1449] a means for the server to graphically display the summary content in a visually easy-to-understand format;
[1450] means for providing the generated graphical display to a user;
[1451] A system including:
[1452] (Claim 2)
[1453] 10. The system of claim 1, wherein the server uses an emotion recognition engine to infer the user's emotional state by analyzing voice features.
[1454] (Claim 3)
[1455] 10. The system of claim 1, wherein the server uses a visualization library to generate a graphical display based on the summary content.
[1456] "Application example 2 when combining emotion engines"
[1457] (Claim 1)
[1458] a means for a user to upload an audio file;
[1459] a means for the server to receive and store the audio files;
[1460] means for the server to generate text from the audio file using speech recognition;
[1461] a means for summarizing the generated text;
[1462] a means for graphically displaying the summary content;
[1463] means for providing a summarized graphical display to a user;
[1464] a means for the server to recognize emotions from the audio file;
[1465] means for adjusting summary content based on emotion recognition;
[1466] means for graphically displaying the adjusted summary content;
[1467] A system including:
[1468] (Claim 2)
[1469] 10. The system of claim 1, wherein the server summarizes the generated text using a natural language processing model.
[1470] (Claim 3)
[1471] 10. The system of claim 1, wherein the server uses a chart generation library to create a graphical display for visualizing the summary content. [Explanation of symbols]
[1472] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for a user to upload an audio file; a means for the server to receive and store the audio files; means for the server to generate text from the audio file using speech recognition; a means for summarizing the generated text; a means for graphically displaying the summary content; means for providing a summarized graphical display to a user; A system including:
2. The system of claim 1 , wherein the server summarizes the generated text using a natural language processing model.
3. 2. The system of claim 1, wherein the server uses a chart generation library to create a graphical display for visualizing the summary content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A